Release notes#
The full changelog is reproduced below, newest first; the format follows
Keep a Changelog and versions are YYYY.MINOR.PATCH (a calendar year,
then a minor number that increments with each release of that year). Start with the upgrade notes if
you are moving an existing codebase forward.
Upgrading from 2026.1.0#
The optimization layer’s axis objects became classes. Everything here raises rather than failing silently, and the constructor spellings you already use mostly still work.
Geometries moved and are constructed, not fetched. backend.optimizers.MANIFOLD_OPS becomes
backend.geometry.ManifoldGeometryOps(), COREWISE_OPS becomes CorewiseGeometryOps(), and
shared_geometry_ops(base, groups) becomes base.with_sharing(sharing, shape) (ragged; the uniform twins take
with_sharing(sharing) – they carry their shape already). On the uniform side,
uniform_geometry_ops(name, x0_data, sharing) becomes
UniformManifoldGeometryOps.from_point(x0_data, sharing) (or the Corewise twin). The frontend
geometry singletons – MANIFOLD, COREWISE, UNIFORM_MANIFOLD, UNIFORM_COREWISE,
shared_manifold(...) – are untouched, so ordinary fitting code needs no change.
A custom sampling kind is a subclass now. Building one by copying an existing kind and swapping a
function – dataclasses.replace(APPLY, forward=...) – no longer type-checks. This is deliberate: that
construction copied the kind’s identity along with it, so the variant claimed to be APPLY and jit
handed back APPLY’s compiled program. Subclass ApplyKind (or the relevant base) and override the
operations you need; keep your parameters as dataclass fields and you get correct jit cache
behaviour for free.
A custom backend geometry needs three more members. The 2026.0.0 notes below show a geometry
object with frame / project / retract / inner / precompute = None. The optimizers now also call
stack_shape(x_cores) (returns the frame stack C, used by the stacked-point guard) and
base_point(frame) (the point the frame is attached to, used by the regularizer), and precompute must
be a callable returning None rather than None itself. Adding the three reproduces the old
behaviour exactly. The full surface is the Geometry protocol in backend/optimizers.py.
The six uniform kind builders are gone. uniform_apply_kind(x0_data) and its five siblings become
UniformApplyKind.from_point(x0_data) and theirs; the uniform_sampling_kind(name, x0_data, weight) /
uniform_derivatives_kind(name, x0_data, order, weight, chunk_size) dispatchers are unchanged and are
the easier route. The four ragged constructors (probe_kind, the three *_derivatives_kind) keep
both their positional and their weight= keyword spellings.
fitting.UniformGaussNewtonModel is gone. One GaussNewtonModel now serves both representations,
and the *_model factories still dispatch on x, so fitting.apply_model(UNIFORM_MANIFOLD, ux, ...)
returns what it always did – just of the merged type. If you were checking
isinstance(m, UniformGaussNewtonModel), check the frame instead:
isinstance(m.frame, t3toolbox.UT3Frame).
Stacked points raise in the optimizers. If you were passing a stacked x0 to newton_cg,
gradient_descent, mc_sgd or adam, you were getting a TypeError from inside the loop; you now get
a NotImplementedError at the entry that says what to do. Fit the stack elements separately, or build
the local model and drive your own loop – that path supports stacked points and returns a shape-C
objective. A regularizer on a stacked point raises for a separate reason (see the changelog).
Upgrading from 2026.0.0#
Three changes in this release can break existing code. All three are mechanical, and each raises a clear error rather than failing silently.
The named contraction functions are gone. backend.contractions used to export ~104 functions
whose names spelled their own subscripts (WCa_Caib_WCi_to_WCb and relatives). They are replaced by
one interpreter, contract, and the migration is a direct rewrite — the function name is the
subscripts string:
from t3toolbox.backend.contractions import contract
# before: WCa_Caib_WCi_to_WCb(mu, G, xi, n_probe, n_frame)
# after:
contract('WCa,Caib,WCi->WCb', mu, G, xi, len_W=n_probe, len_C=n_frame)
A trailing n_probe / n_frame argument becomes the keyword len_W= / len_C=, and you only need
to supply one when the subscripts alone cannot pin the split — the error message names exactly which.
The results are numerically identical (each named function was checked against its contract call
before removal). See grouped_contractions.md.
A custom GeometryOps must accept aux. The backend geometry protocol gained an optional
precompute(frame) slot, and project / retract now take a third argument. If you implement your
own geometry, accept and ignore it:
def project(frame, variations, aux=None): ...
def retract(frame, variations, aux=None): ...
precompute = None # or a callable frame -> whatever project/retract need
The slot exists so a geometry with expensive per-frame setup pays for it once per local model instead
of once per Hessian matvec; the built-in geometries pass None. Rationale:
contributor/precompute_and_caching.md.
A custom Regularizer must accept aux. Same shape of change, for the same reason:
gradient(geom, frame, aux=None), hessian(geom, frame, p, aux=None) and
quadratic(geom, frame, p, aux=None) now receive the geometry aux, so a regularizer composed with a
shared-factor geometry reuses the frame companion instead of rebuilding it per matvec. Accept and
ignore the parameter if you do not need it.
One behavior change worth knowing about, which is not a signature break: use_jit=True now moves
your inputs onto jax rather than silently running eager when they are numpy. Runs that previously
reported “jit” timings while executing eagerly will now genuinely compile, return jax-backed
results (float32 unless you enable jax_enable_x64), and raise if jax is not installed. If you were
relying on the old behavior you were not getting jit; see
fitting_and_optimization.md §4.5.
[2026.2.0] — 2026-08-25#
Changed — breaking#
The optimization layer’s axis objects are classes, not records of closures. A geometry and a sampling kind are now frozen dataclasses whose parameters are fields; their behaviour is methods. This is what makes them compare and hash by value, which matters because both ride as jax
aux_dataand therefore form part of the compilation cache key. Rationale, measurements and the rejected alternatives:contributor/parameters_not_closures.md.New
backend/geometry.pywithManifoldGeometryOps/CorewiseGeometryOpsand theirUniform*twins (the uniform ones carry the fixed rank they were built at; usefrom_point). Removed:backend.optimizers.GeometryOps,MANIFOLD_OPS,COREWISE_OPS,shared_geometry_ops, andbackend.uniform_fitting.uniform_{manifold,corewise,geometry}_ops.Sharing is a
groupsfield, not a wrapper — one class and one code path for shared and unshared. The frontendshared_manifold/shared_corewiseconstructors are unchanged.SamplingKindis a class hierarchy (ApplyKind,EntriesKind,ProbeKind, the three*DerivativesKind, and theirUniform*twins). The*_kindconstructor spellings still work. Removed: theidentityfield and its__eq__/__hash__, and the six uniform builder functionsuniform_{apply,entries,probe}_kind/uniform_{apply,entries,probe}_derivatives_kind(useUniform*Kind.from_point(x0_data, **parameters), or the unchangeduniform_sampling_kind/uniform_derivatives_kinddispatchers). A custom kind is now a subclass; deriving one withdataclasses.replaceno longer type-checks, and that is the point — see below.ScalarOutputKindandProbeOutputKindare public, so a new operator can inherit the‖·‖²reductions.fitting.UniformGaussNewtonModelis removed. OneGaussNewtonModelserves both representations; the factories still dispatch onx. It is deliberately not aliased.optimizer_displaygates its error table onkind.has_block_sumsqrather than on ablock_sumsqfield beingNone.block_sumsqis optional; the flag is declared by whoever implements it.A custom backend geometry needs three more members:
stack_shape(x_cores),base_point(frame), and aprecomputethat is a callable returningNonerather thanNoneitself. TheGeometryprotocol inbackend/optimizers.pylists the full surface.
Optimizing a STACKED point raises
NotImplementedError. It never worked –newton_cgandgradient_descentfailed onfloat()of a shape-Cobjective,mc_sgdon a broadcast, andadamonce its loss logging first fired – but it failed obscurely, from deep inside the loop. Building the local model on a stacked point is supported and still works, so rolling your own loop over a stacked problem is the route. Stacked optimization is possible future work.
Fixed#
The items below came out of the 2026-08-22 whole-library review (dev/review_2026-08-22/).
The safe-mode trace detector is the committed-array probe alone (review H6-12). The old primary detector,
jax.core.trace_state_clean, does not exist in jax 0.10 – every check on jax inputs paid a raise/catch before falling back – and its docstring promised any-transform detection the probe never delivered. Behavior under jit is unchanged (checks skip); undergrad/vmapwith concrete operands the checks run, as they always effectively did – harmless and now documented. The batching docs also record why a stackedUniformTuckerTensorTraincannot bevmap-ed at all (masks are static aux; use the nativeCstack).The uniform same-frame guard compares real (masked) content only (review H5-5). Two frames identical up to the don’t-care padding now count as the SAME frame – the tangent space depends only on the real content – so
UT3Tangentarithmetic /stack_tangentsbetween a frame and a re-padded copy of it works instead of raising “different tangent spaces”.safety.frames_equalitself stays representation-agnostic; its contract is now explicit: callers pass real content, and the uniform call sites mask their supercores first. (This also makesUT3Frame’s documented “frames_equal on the masked supercores” true.)The review’s oracle sweeps are permanent tests (Phase D):
tests/test_oracle_sweep.pyruns the O1 sampling/tangent sweep (dense einsum + exact combinatorial jet oracles + adjoint identities, ragged and uniform incl. varying-rank stacks) and the O2 fitting sweep (polynomial-interpolation oracle vsGaussNewtonModel; optimizer trajectories vs hand recomputation, ragged == uniform) from verbatim copies of the review scripts intests/oracle_sweeps/. Two tiers: always-on (full O1 + full O2 models + a pairwise-covering optimizer subset) andT3TOOLBOX_SLOW_TESTS=1for the full optimizer matrix – a required release-gate step, so nothing the review matrix covered stops being enforced.Small ergonomics from the review’s E list (R9-10, R7-11, R3-7):
chunk_sizenow threads through both frontends’probe_corewise_derivatives_transpose(and the uniform backend wrapper) – it existed only on the ragged backend and the tangent-level transposes;chunk_size='auto'resolution reads supercore shapes without pulling the arrays to host, and the fallback chunk size is the singlebackend.sampling_derivatives.DEFAULT_CHUNK_SIZEconstant; the minimal-rank recurrence inbackend.ranksuses the dispatched backend throughout (three straynp.minimumcalls made it jit-hostile).jit cache-key hygiene on the value-hashed aux family (reviews H1-5, H1-6, R2-7, H2-8, R7-13). The mask holders’
hashand==now derive from one content key (they could disagree on dtype and silently zip-truncate on length); a jax-array field in aValueHashedFieldsaux raises a directive TypeError (device arrays belong in leaves); the aux objects (mask holders, uniform geometries and kinds) store defensive read-only copies of their mask arrays, so mutating a caller’s mask in place can no longer stale a cached cache key – in-place writes to stored masks now raise;shared(...)geometries key their identity (and jit aux) on the canonical partition, so label spellings of the same partition share one compiled program, and aSharedGeometrysubclass now survives the jit round-trip as itself. Thestack(unstack(x))-under-jit findings were verified already fixed by the review’s static-axesmoveaxisand are pinned by a regression test.Obscure errors on malformed inputs became structural guards (the review’s E-cluster sweep, 2026-08-24: R2-8, R2-12, R2-13, R3-6, H3-7, R8-6, R8-9, R1-15, R10-6, H3-8, R9-9). Str leaves no longer recurse the tree helpers forever;
tree_ziprejects structure mismatches instead of silently truncating;truncated_svd(min_rank > max_rank)raises (min used to silently win);t3_inner_producttakes list/tuple mixes andt3_sum_stackan ndarray axis; uniform*is scalar-only with a clear TypeError, uniformstackvalidates the static layout, and uniformminimal_ranksstays host numpy on jax-backed trains; the ragged tensor-weight ops check rank/stack consistency (one-metric broadcast still allowed);COREWISE.retractrejects a detectable manifold-frame tangent;SharedGeometrynames a wrong-layer point/frame. One regression test per finding.The uniform frame at a numerically rank-deficient point loses no tangent directions (review S1b – the last open review item). Every kept-basis SVD in the uniform orthogonalization and T3-SVD sweeps now goes through the new mask-aware
backend.linalg.pad_safe_svd, so thesigma ~= 0completion columns stay OFF the padding instead of landing in padded slots that the masks then erase. Before the fix, a zero-paddedresizewarm start – the rank-continuation restart – produced a uniform frame whose masked real block was genuinely rank-deficient (60/60 trials), silently losing exactly the tangent directions continuation needs to grow into. AffectsUT3Frame.from_ut3/ut3_orthogonal_representations, the per-coreut3_*_orthogonalize_*ops, andut3svd/ut3_rank_adjustment_sweep/ut3svd_orthogonal_representations(unshared and SF-T3 shared). Regression tests:TestPadSafeFrameintests/test_uniform_frame_variations_format.py.Uniform
+,-,squash_tailsandsum_stackread padding garbage.ut3_squash_tailssummed the leading/trailing TT bond over every slot, including the padded ones the equivalence contract declares don’t-care – and a corewise gradient step (adam(UNIFORM_COREWISE, …),UNIFORM_COREWISE.retract) leaves exactly those slots nonzero, sox_fit + ywas silently wrong whileto_dense/norm(which mask on entry) were right. It now masks on entry.t3mwith boundary TT ranksr0, rd != 1.inplace_fused(the default) never SVD’d the last bond, so an unsquashed trailing bond survived asrd_x · rd_yand escapedmax_tt_ranks;swapcontracted the two operands’ trailing bonds against each other (silently wrong when equal, a raw einsum error when not) and ignored a per-positionmax_tt_ranksgiven as a numpy array. Everyt3mmethod now canonicalizes both operands’ boundary bonds to 1 on entry – the rule, and why the exact*/t3_multdeliberately do not, is int3m_methods.md.entries_ambient_transposereturned zero factors for a negative index. Its one-hots were built witharange(N) == index, which matches nothing for-1; the forwardentriesand the tangent/corewise transposes wrap numpy-style, so the adjoint identity silently failed. All four now agree (eye(N)[index]). Documented alongside: an out-of-range index raises under numpy but clamps under jax (gather semantics).The documented singular-value-metric recipe applied the σ’s in the wrong gauge.
docs/weighting.mdpairedT3FrameWeights.from_t3weights(T3Weights.from_t3svd(x))with a frame fromt3_orthogonal_representations(x), which re-SVDs the already-orthonormal Tucker factors of a T3-SVD result – a degenerate spectrum, so the frame’s Tucker basis comes back rotated by an arbitrary orthogonal matrix relative to the singular basis, and per-coordinate σ-weights then weight the wrong directions (unstacked too; the planning measurement said “stacked only”). Newt3svd_orthogonal_representations(x, **t3svd_kwargs)(+ theut3svd_twin, both exported at the root) returns(frame, variations, T3Weights)from ONE SVD with the frame builtalready_left_orthogonal=True, so its Tucker basis is the singular basis; the recipe and thefrom_t3svddocstring now say so.The shared-factor companion was wrong on a frame converted between layers.
fv_shared_frame_datarecomputed the centers by re-running the SVD sweep that built the frame, which reproduces the construction’s sign choices only when it is the same SVD on the same arrays; on aUT3Frame.to_t3frame()leaf – the routeUniformManifoldGeometry.project_ambientrecommends for dense gradients – the tied projection was silently 30% off (the group spectrum, being gauge-invariant, looked right). The centers are now the zipper of the stored chains (H_i = L_i Z_{i+1}, no SVD), exact and gauge-consistent for a frame built by either layer; the uniform twin masks first, and the “frames packed from ragged are not guaranteed” caveat is gone.jax.gradthrough uniformnorm()/inner()/ut3_weighted_norm/ut3_weighted_innerwas all-NaN on any train with rank slack – the default orthogonalized path differentiated through the SVD, whose JVP has1/(σᵢ²−σⱼ²)terms and a padded train has several exactly-zero σ’s. New backend twinsut3_norm/ut3_inner(closing a ledger item) keep the precise orthogonalized VALUE and carry acustom_jvpwith the exact multilinear derivative (2⟨T, dT⟩, via the zipper with the orthogonalized side held fixed – no SVD in the derivative path); the frontend and the weighted ops route through them.backend.optimizers.Problem.objective(x, data=…)/.local_model(x, data=…)ignoreddatawhensamplewas omitted, silently scoring the training data;sample=alone crashed with a bareTypeError. The pair now goes together or not at all (a structural error otherwise).The regularized objective was wrong at a raw point on the ragged manifold.
ManifoldGeometryOps.point_norm_sqread‖last TT core‖², exact only for a left-orthogonal point (measured 3 where the truth was 1400 on arandnpoint). Backend-only: the optimizers evaluate the regularizer through the frame or at retraction outputs, so no optimizer result orstatsrow was affected. It now left-orthogonalizes first (noWfactor, negligible beside the misfit it sits next to), matching the uniform twin.UT3Variations.sum_stack(axis=-1)/UT3Tangent.sum_tangents(axis=-1)summed over the MODE axis (the raw1 + axissent-1to array axis 0), silently whenK == dand with an obscure shape error otherwise. Negative axes now count from the last stack axis, as in numpy and the ragged twin; out of range raises.A structural tangent mismatch was caught only by the numerical same-frame guard, which is skipped under
safety.unsafe()/ a jax trace – soa + bfor tangents at frames of different rank structure broadcast silently there. The core-shape comparison now runs unconditionally (structural problems always error), inT3Tangentand the fitting model’s trial-tangent guard.30 structural checks were written as
assert(x * ndarrayshape, stack-axis ranges, core-tuple lengths,corewiseoperand types,dense_probevector shapes, …): a bareAssertionErrornormally, and no check at all underpython -O, where a wrong-shapedx * ndarraybroadcast silently. All areValueError/TypeErrorwith a message now.The gauge precondition compared an absolute residual against the relative tolerance.
gauge_residualwasmax|UᵀV|, which scales with the tangent, whileMANIFOLD.inner/normthreshold it atrtol(1e-9 numpy, 1e-5 jax): a correctly gauged tangent of norm ≳1e8 failed in numpy, and on jax’s default float32 one of norm ≳50 – the documented jax-eager safe-mode use – while a tiny ungauged tangent passed. The residual is now relative (each gram divided by its variation core’s norm, per stack element), on both layers.The tangent/corewise
apply/entriestransposes rejected a bare Python float residual (the natural unstacked case; the ambient twin’s own doctest passes1.7).d = 1on the ragged manifold:MANIFOLD.project_ambient(frame, <T3>),project_ambient(…, method='t3svd')andMANIFOLD.transportraised anIndexError(an empty TT chain has no zipper);dense_proberaised on a one-mode tensor with a probe stack. A one-mode T3 is a vector; these now degenerate to that case like everything else.UNIFORM_MANIFOLD.transport/project_ambientcrashed on aK-stacked tangent or gradient (the frame’sC-stack gauge masks were handed toK+Cvariations and failed in a reshape); the ragged twins acceptedK. The masks are now broadcast overK(newufv_masking.ufv_variation_masks_over_stack, the same rule the sampling transposes already used).stack()was not jittable (TuckerTensorTrain.stack,T3Frame.stack,T3Tangent.stack_tangents, the uniform twins):moveaxiswas given anarangesource, which is a tracer underjit.adamon a ragged manifold with a non-minimal-rankx0crashed at step 2 (its moment trees were allocated atx0’s core shapes; the first retraction drops the redundant rank). A ragged manifoldx0is now reduced to minimal ranks on entry, mirroring the uniform path’suniform_minimal; a minimal start is untouched.A slack-padded
x0crashed every uniform manifold optimizer at step 2.UniformTuckerTensorTrain.from_t3(x, n=, r=)(the documented “force a larger pad”) passes the minimal-rank gate, bututv_retractreturned its point at the max rankut3svdkept rather than at the frame’s padded dims (its docstring promised the latter), so the loop-invariant masks no longer fit. The retraction now re-pads (newbackend.ut3_operations.ut3_pad_ranks).newton_cg(verbose=True, val_data=…)withoutval_samplefailed deep inside the kind, andval_samplealone was silently ignored; both now raise at the entry (make_newton_displaytoo).uniform_least_squares_problemaccepted any non-'manifold'geometry string as corewise (so'Manifold'silently fit on the wrong geometry), and built aProblemfor a derivative kind withoutorder=that failed on first use. Geometry and kind strings are validated (case-insensitively, see below) andorderis required up front.Structural validation gaps and wrong messages (each a deep or misleading failure before):
check_fv_pairaccepted a frame and variations with different numbers of cores (four ops then returned partial sums);T3Frame.validate’s Tucker-rank message printed stack dimensions on a stacked frame;entries()with a wrong-length list index raisedAttributeErrorinstead of the documentedValueError;share(rtol=…)on a stacked T3 failed insidetruncated_svd; uniform+/inneron operands with different padded widthsNraised raw numpy errors; the sixGaussNewtonModelfactories did no validation of the sample / residual against the point;set_default_safety(None, None)was accepted and then every check crashed (it now means unsafe by default, the one script-level way to get it; a loneNoneis rejected);has_numerically_minimal_ranks()raised on a stacked train (now a per-element bool array like the other checkers).
Changed#
==andhashare now explicit-only on every runtime array-carrying class (review R1-13 extended; ruled 2026-08-24).TuckerTensorTrain,T3Weights,T3Frame,T3Variations,T3FrameWeights,T3Tangentand their six uniform twins refuse==with a TypeError naming the named checks (x == xstays True via the identity fast path) and are unhashable – the equality of a floating-point tensor representation is not a well-posed operator question, so the user says which equality they mean:a.allclose(b[, rtol=, atol=])for what the objects REPRESENT (numerical, per stack element; gauge/padding-invariant for trains; the BASE POINT for frames, withsafety.frames_equalanswering the same-FRAME question),a.corewise_equal(b)for the bitwise representation,a is bfor identity. Previously==raised numpy’s ambiguous-truth error (T3Frame: silently compared identity – a vestige of the pre-safety-redesign aux_data era, per its own history). Value-based hash/eq remains exactly where jit needs it: the aux_data family (mask holders, geometries, kinds). Rule + rationale:backend.common.ExplicitEquality. The pre-existingallclosemethods (T3Tangent/UT3Tangent/UT3Frame) are unified into the family: defaultrtolis now the ambient safety default (same value on numpy float64), the reference is the symmetricmax(||a||, ||b||)(was||other||), and the tangent twins now CHECK the same-frame precondition they used to assume.Citing T3Toolbox means citing the software, not the T4S paper.
CITATION.cff’spreferred-citationused to redirect citations to the research preprint the library grew out of; the library is a general-purpose T3 toolkit and is cited as software. The paper moved to the CFFreferenceslist (cite it when building on its research contributions); README / docs wording aligned.The six
utv_*_transpose_from_sweephooks now defaultsum_over_probes=False, matching their ragged twins (review H2-4). Each says it “shares” the ragged function, but the uniform defaults summed the probe stackWwhile ragged kept it – so a raw-.datauser porting ragged split-seam code to uniform silently got a different result shape. No shipped path changes (both frontends already defaultedFalse; the fitting kinds passTrueexplicitly). Breaking only for raw-backend callers that relied on the old uniform default: passsum_over_probes=Trueexplicitly.A failed Armijo line search stops
newton_cg/gradient_descentinstead of silently accepting an ascent (review R7-7). Exhausting every halving (40 / 50) used to accept the last trial anyway – measured objective increases accepted at the precision floor, ~2×40 objective evaluations per stagnant iteration – and the recordedalphawas one halving past the step actually taken, so that row’sdelta_f/rho/step_reldescribed a step never tried. Now: the defaulton_line_search_failure='stop'rejects the step and terminates;'accept'opts into the old take-the-step behavior as a deliberate escape mode (it can bump a stagnant iterate loose);NewtonInfo/stats['history']gainls_failed,gradient_descent’s stats gain'line_search_failed', andalphanow records the last trial actually evaluated.The ragged backend geometries’
inner/point_norm_sqare per-element over the frame stack, matching the uniform twins and the frontend (review H3-5).ManifoldGeometryOps.inner,CorewiseGeometryOps.inner/point_norm_sq, andt3_left_orthogonal_norm_sqused to collapse the stackCto a scalar whileUniformManifoldGeometryOps.innerandMANIFOLD.innerreturned shapeC– a silent cross-layer asymmetry for a raw-.datauser composing a stacked loop. Unstacked behavior is unchanged (a 0-d scalar); shipped optimizer paths reject stacked points up front and are unaffected.The uniform manifold frame is gauge-equivalent to ragged, no longer bit-identical. The old bit-equality was an accident of both layers feeding the same matrices to the same LAPACK kernel; the pad-safe sweep (previous section) computes the same subspaces and singular values through a different – still fully deterministic – factorization route, so gauge-carrying cores agree with ragged only up to the SVD’s inherent gauge freedom (column signs, rotations within equal-
sigmaclusters, the choice of null-space completion). Every gauge-invariant quantity – objectives, gradient norms, optimizer trajectories, represented tensors – still matches ragged exactly. The contract statement for gauge-carrying operations is now indocs/uniform_equivalence_contract.md.jax requested but not installed: run on numpy and warn, never raise.
TuckerTensorTrain.randn(use_jax=True)/UniformTuckerTensorTrain.randn(use_jax=True)used to die with a bareNameError,to_jax()/load(use_jax=True)silently returned numpy, and the optimizers’use_jit=Trueraised. The policy is now one rule (backend.common.jax_or_warn): the request runs on numpy with a one-timeRuntimeWarningnaming thet3toolbox[jax]extra – a project developed without jax on one machine and deployed with it on another runs unchanged on both. (use_jit=Truewith jax present still auto-converts to jax, as in 2026.1.0.)Geometry and kind strings are case-insensitive on the frontend optimizers and
uniform_least_squares_problem('Manifold','PROBE_DERIVATIVES', …).adamon a manifold geometry warns: its per-coordinate moments live in the gauge-dependent core coordinates, so its iterates depend on the backend / representation (measured ~1e-2);COREWISEis the intended geometry,mc_sgd/newton_cgthe gauge-invariant manifold optimizers.A per-mode residual weight may have more rows than modes. The extra rows are ignored (intended: one weight rides through a continuation scheme that adds modes), now on both layers – the frontend used to reject it while the backend truncated silently. Fewer rows than modes is a structural error on both.
Fixed (continued)#
d = 1on the uniform layer. The uniform boundary-bond squash duplicated the single core, which took down every squashing op (+,-,norm,inner,t3svd,rank_adjustment_sweep,UT3Frame.from_ut3,is_left_orthogonal). A one-mode T3 now degenerates to the vector case on both layers, as intended.A regularized fit of a stacked point silently mis-weighted, and now raises. The data misfit keeps the frame stack
C(one value per element) while every regularizer scalar collapses it, soobjective = misfit + ρadded the whole-stack regularization total to each element – inflating the effectiveλby about|C|, unevenly. The regularizer gradient was per-element correct, which is what made it easy to miss. No test combinedstack_shapewith a regularizer.A derived sampling kind silently reused its parent’s compiled program.
dataclasses.replacecopied the hand-maintainedidentitytuple unchanged, sodc.replace(APPLY, forward=<other math>)compared equal toAPPLY; underjitit returnedAPPLY’s answer (measured: 115.302888 where eager gave 28.825722). Unrepresentable now — a variant is a subclass, hence a distinct type.
Documentation#
rank_adjustment_sweep’s orthogonality is conditional, and the docs said so unconditionally: the sweep re-SVDs the TT cores and bonds and preserves Tucker orthonormality but never creates it, so its result is right-/left-orthogonal in the full sense only when the input’s Tucker cores were already orthonormal (anyt3svdresult) – the documented precondition, now stated at every site. The one real gap behind “compose both directions for guaranteed minimal ranks” – a Tucker rank above its mode size (n_i > N_i), which no TT-side step can reduce – is closed by a Tucker up-SVD at the start of the sweep.A sweep of doc ↔ code mismatches found by the 2026-08-22 review: the README / user-guide core diagram and frame-core listings (
dcores, notd+1;up_tucker_cores), the Tucker-rank tuple, the uniform mask form (canonical prefix vs gappy working form), the registered-pytree list (the uniform classes were missing), the paper section cited for probing, stale “slice”/”deferred” module docstrings on the uniform layer, theUT3FrameMasksvalue-hash comment,chunking.md’s'auto'default (it is the frontend’s),sharing.md’s “closure state”, sixdev/regularization_design.mdpaths (now indev/archive/), thet3m_methods.mdboundary-bond sentence, theweighting.mdrecipes and “same operations” claim, theGeometryOpsname in the project guide,backend.geometryandcorewisein the API reference, and thesave/loaddoctests, which now write into a temporary directory instead of the working directory; and, found by the same review:truncated_svd’s documented truncation rule (it is the tail-Frobenius TT-SVD rule, not a per-singular-value cut), theGeometryprotocol’s member list, theshared(…).framecontract row and the rank-deficient-frame retraction caveat innumerical_contracts.md, the corewise transpose’s description intransposes.md(a pullback, not a projection), the λ-vs-drawscaling note, the chunking memory caveats (two-chunk and kept-path regimes), theresizecall in the uniform continuation loop, the retired identity-guard wording onT3Tangent, andt3_norm(use_orthogonalization=False), which orthogonalized anyway (it now honors the flag; stack-shape mismatches in+/innerareValueErrors, as documented).
Performance#
The inner CG compiles once per fitting run instead of once per Newton iteration (measured 1 → 0 compiles per iteration on the uniform
probe_derivativespath, warm)._cg_solve(local_model, rhs, tol, maxiter, use_jit)dispatches to one function compiled over(local_model, rhs, tol, maxiter), so the cache key is the model’s pytree treedef — value-based, thanks to the change above.mc_sgd/adamper-step kernels compile once per shape signature, process-wide, rather than once per optimizer call: the kernels are module-level functions and the jit wrapper is memoized on them, replacing a per-call closure that discarded its own cache.A user-defined kind or geometry now gets this for free, provided its parameters are dataclass fields.
[2026.1.0] — 2026-08-20#
Added#
The explicit equality checks:
allclose(other, rtol=, atol=)andcorewise_equal(other)on all twelve runtime classes (see the==entry under Changed),corewise.corewise_equal(the bitwise tree comparison underneath), andsafety.comparison_rtol(the ambient jax-aware default tolerance for pure comparisons; mode-agnostic, likeframes_equal). Trainallclosecompares the represented TENSORS bynorm(a - b)– the difference train through the orthogonalized norm, numerically stable exactly in thea ~= boptimization-residual case.Scalar-left multiplication and scalar division on the point/tangent classes (review R1-12 / H4-8):
TuckerTensorTraingains__rmul__(2.0 * xused to raise whilex * 2.0worked), andTuckerTensorTrain/UniformTuckerTensorTrain/T3Tangentall gain scalar-only__truediv__(x / 2). The ragged weighted frontendst3_weighted_norm/t3_weighted_innergain theuse_orthogonalizationkwarg their backend and uniform twins always had (review H2-6), andcheck_fw_pairis exported as the docs already said (review R10-7). Per-axis uniformsum_stack(axis=)is deferred-and-documented (review H2-7;contributor/deferred_and_rejected.md).backend.linalg.pad_safe_svd– the mask-aware SVD of a zero-padded matrix: the firstmin(n, m)triplets are exactly the economy SVD of the unpadded real block, bitwise zero on the padding, with no rank tolerance anywhere and one jit compile across mask patterns (masks are runtime data). Batched, backend-agnostic, symmetric in(n, m)(wide real blocks need no caller-side transpose), deterministic (fixed Haar sketch per width). The algorithm is Method D (“sketch-project”) of the pad-safe SVD design packet,dev/review_2026-08-22/repros/S1b/packet/: pads are permuted to the trailing QR pivots (Householder reflectors then never touch a padded row), one small augmented SVD pins the surplus, and the right factor is rebuilt fromA^T U = V diag(ss). The uniform sweeps use it wherever an SVD produces kept basis columns. The complete derivation note – problem, contract, algorithm, measurements, and the alternatives considered – isdocs/pad_safe_svd.tex(+pdf).Shared Tucker factors (SF-T3) — optimize over Tucker tensor trains whose Tucker factors are constrained equal within user-specified groups of modes (the SF-ETT decomposition of Molozhavenko & Rakhuba (2026), generalized from one trailing block to an arbitrary partition; the partition is always user-provided). A
sharingspec is one hashable group label per mode, e.g.(0, 0, 1); a shared T3 is an ordinaryTuckerTensorTrainwhose group factors are equal (redundant storage — a compute-not-memory feature). Ragged layer and its full uniform mirror:The shared geometries —
shared(MANIFOLD, sharing)with shorthandsshared_manifold(sharing)/shared_corewise(sharing), exported from the package root: drop-in geometry wrappers for every optimizer and fitting model. Projections post-pass onto the tied tangent subspace in the base geometry’s own metric (manifold: the tilted least-squares projection through a per-frame SVD companion; corewise: the per-group mean), and the manifold retraction goes through a tied doubled-rank embedding plus the groupedt3svd, so every iterate stays exactly tied (one array per group).Grouped
t3svd(sharing=)/rank_adjustment_sweep(..., sharing=)— the paper-faithful two-phase truncation (TT rounding, then simultaneous group Tucker SVDs on concatenated matricizations), reporting one spectrums_gper group;sharing=Noneor an all-singleton partition is the existing sweep exactly.TuckerTensorTrain.share(sharing, ...)— the quasi-optimal shared initializer (exact common-span rewrite + grouped truncation), and the checker methodhas_shared_tucker_factors(sharing, rtol=)(per-stack-element bool).Shared rank bookkeeping —
get_minimal_ranks(..., sharing=)/backend.ranks.compute_minimal_ranks(sharing=)with the group ceilingn_g <= min(N_g, sum_i min(N_g, rL_i*rR_i))(per-mode ceilings ADD across a group, so a shared rank may exceed an individual mode’srL_i*rR_i— the unshared reduction would clip it and untie the group),manifold.manifold_dim(s, sharing=)(one Stiefel term per group; validated against dense tied-tangent ranks), andframe_has_minimal_ranks(..., sharing=).Shared rank continuation —
continuation_ranks(sharing=)/backend.ranks.compute_continuation_ranks(sharing=): a group’s Tucker edges are ONE edge — oneκ_g = s_g[0]/s_g[-1]in the conditioning pool, one growth decision applied group-wide, onemax_growcandidate — with the shared useless-rank removal (κ_gis never worse than the group’s worst per-mode condition number, and can be far better on complementary spectra). Plusresize(..., sharing=)for the zero-padded warm start: the group factor is padded once (one array per group), the represented tensor unchanged. A freshly padded restart carries exactly-zero new spectrum levels (the tied Tucker channel is gated); the escape runs through the untied TT-variation channel within the first Newton steps — which is why full shared rank is a diagnostic, never an enforced precondition.Uniform mirror, grouped truncation family —
UniformTuckerTensorTrain.t3svd(sharing=)/rank_adjustment_sweep(..., sharing=)/has_shared_tucker_factors(sharing, rtol=), backendut3svd(sharing=)(the two-phase grouped sweep in scan/supercore form: TT-bond rounding scan with the Tucker steps skipped, centers collected by the polymorphic right sweep, per-group SVDs on statically-gathered concatenations — mask-only truncation, ONE group rank mask at every group mode),ut3_rank_adjustment_sweep(sharing=), the grouped host recurrencecompute_raw_sweep_ranks(sharing=)(verified == the ragged grouped output ranks over randomized structures/caps), the masked checkersut3_sharing_residual/ut3_tucker_factors_shared, and a sharing-awareuniform_minimal— required: the per-mode reduction silently unties a shared uniform start (it can clip a group rank to unequal per-mode values). All verified under the uniform equivalence contract (== the ragged grouped ops on real parts, per stack element, varying-rank stacks included), with exact output-mask assertions, garbage robustness, and jit-clean dispatch.Uniform mirror, tied tangent machinery — the uniform companion
ufv_shared_frame_data(the identical polymorphic derivation on the frame’s stored supercores — deliberately NOT re-masked: the companion’s exactness rests on reproducing the construction’s own sweep on the same arrays, and the padded rows of eachS_i^Tvanish because completion rows are orthogonal to the centers’ row space), the tied post-passesufv_share_tucker_variations/ufv_share_tucker_variations_corewise(mask-and-delegate to the polymorphic ragged solves), andshared_data=threading throughutv_orthogonal_gauge_projection/utv_to_ut3(the TIED doubled-rank embedding:Udotat every group mode, the companion’s centers replacing the down cores, the variation block rebuilt at the up width) /utv_retract(tied embedding + the groupedut3svd). Verified gauge-invariantly against the ragged twins (dense tangents/points at machine precision; outputs exactly tied).Uniform mirror, shared geometries + fitting —
shared(UNIFORM_MANIFOLD, sharing)/shared(UNIFORM_COREWISE, sharing): theSharedGeometrywrapper now takes the uniform singletons too (uniform points/frames/tangents in and out; same value-hashed identity), the backend factories takesharing=(uniform_geometry_opsand both singles — the closures capture the partition beside the masks and populateprecomputewith the uniform companion),uniform_least_squares_problem(sharing=)with a shared-minimal gate, and the frontend fitting models carry the companion as ageometry_auxleaf (UniformGaussNewtonModel), so the packed compile-once path holds for shared fits (one trace across rebuilt same-rank models). All four optimizers run shared on the uniform layer; deterministic trajectories match the ragged shared runs, and every iterate stays exactly tied.Docs: the user page
docs/sharing.md(the format, the grouped truncation, what the group spectrum is, the two geometries, rank machinery, batching, uniform, scope — including sharing ≠ symmetry); the design recorddocs/contributor/sharing_internals.md; a shared section in the CI-doctested getting-started tour; the sharing section ofdocs/rank_continuation.md; and the TIED precondition rows indocs/numerical_contracts.md.Weights × sharing — the combination composes within the existing framework (absorbing keeps a tied T3 tied iff the group Tucker weight vectors are equal; TT-bond weights never touch the factors;
from_t3svd(x, sharing=…)builds group-equal weights; the weight algebra preserves group-equality), so the only addition is the non-enforcing compatibility checker:T3Weights.has_shared_tucker_weights(sharing, rtol=)and theUT3Weightstwin (masked content), witht3_tucker_weights_sharing_residual/t3_tucker_weights_shared(+ut3_*) inbackend.sharing. Nothing gates — absorbing group-unequal weights legitimately unties.Example:
examples/fit_shared_factors_jetted_probes.py— a groupwise-symmetric five-mode target (two Hilbert tensors coupled by a random matrix; two sharing groups of different mode sizes) fit from noisy probe-derivative jets, running the SAME rank-continuation fit shared vs unshared: tying the factors the target’s symmetry justifies reaches ~35% lower true error with ~37% fewer parameters before overfitting.Entering the format is the geometry’s job.
shared(...).frame(x)ties an untiedxfirst, silently, by the per-group mean — so an untied initial guess is a non-event for the optimizers, and the slow drift a long run of low-precision first-order steps can produce is absorbed at the next frame. An already-tied point is a bitwise fixed point, so the ordinary path is unchanged. The uniform layer gets the same route without a round trip through ragged:ut3_tie_tucker_factors, the twin oft3_tie_tucker_factors(garbage-transparent, so it needs no masking; masks and TT cores untouched). The shared corewise retraction ties the sum rather than aliasing one mode’s copy of it, which makes it total:mean_i(U_i + V_i) = mean_i(U_i) + mean_i(V_i), so an untied tangent (always handled) and an untied base point both land on the shared set with nothing silently discarded. The TIED-tangent precondition is now one backend checker —fv_tied_variations_residual/ufv_tied_variations_residual, a single global Frobenius ratio per stack element — replacing two hand-rolled formulas that disagreed between the ragged and uniform paths and collapsed the whole stack into one scalar (so one untied stack element could hide behind many tied ones).Backend surface in
backend.sharing(validate_sharing,t3_sharing_residual,t3_tucker_factors_shared,t3_tie_tucker_factors,SharedFrameData+fv_shared_frame_data, the tied post-passes) andbackend.t3_svd.t3_share_tucker_factors. Safe mode checks tied factors at shared entry points; full shared rank is a diagnostic, never a precondition (rank-continuation restarts legitimately sit below it).
The grouped-einsum interpreter
backend.contractions.contract— one general entry point for every grouped contraction:contract('WCa,Caib,WCi->WCb', *operands, len_W=...). Standard einsum strings where an UPPERCASE letter is a group of zero or more axes (Wprobe stack,Ktangent stack,Cframe stack, …); group sizes are solved exactly from the operand ndims (identifiability is decided from the subscripts alone — a call site either always needs alen_<G>supplement or never does, and the error names precisely what is missing); groups expand into fresh single-axis letters and one ordinary einsum runs on the operands as given. No reshape ever happens, so every sub-axis of every group is shardable (compiler-verified across the whole library vocabulary) and fusing two groups is inexpressible. numpy keeps the greedy pairwise BLAS path (computed on the grouped string); jax gets a single fused einsum.Weighting (edge weights) — diagonal weights on the internal edges of a T3, as a lightweight data format plus
absorbinto cores, in both the ragged and uniform representations. Two classes per layer, because a tensor and a tangent have genuinely different edges:T3Weights/UT3Weightsweight a Tucker tensor train as a tensor (tucker[d],tt[d+1]— exactly the shapet3svdreturns, so the singular values are the canonical weight object).T3FrameWeights/UT3FrameWeightsare a metric on a tangent’s coordinates (up/down/left/right, eachlen=d) — the Grasedyck–Kramer preconditioner — absorbed into the variation cores with the frame left orthonormal, so they areO(ranks).Operations on all four:
absorb_weights,weighted_norm/weighted_inner,reciprocal/sqrt, andconcatenate/kronecker(the+/⊙duality: ranks add / multiply). Constructorsfrom_t3svd/from_ut3svdandfrom_t3weights/from_ut3weights, plus ragged↔uniform conversions for both weight types.The frontend free functions carry the family prefix —
t3_absorb_weights,ut3_absorb_weights,fv_absorb_weights,ufv_absorb_weights(+t3_/ut3_weighted_norm/_weighted_inner) — and the whole surface is exported from the package root. Docs:docs/weighting.md; design records:docs/contributor/weighted_internals.md.
Regularization on the fitting objective —
regularizer=on every optimizer (gradient_descent/mc_sgd/adam/newton_cg) and on the six model factories, minimizing½‖ω⊙(S(x) − y)‖² + ρ(x). Shipped implementation:IdentityRegularizer(λ)=½λ‖x‖²— a Hilbert–Schmidt ridge onMANIFOLD, weight decay onCOREWISE— composing with every optimizer, sampling kind, geometry and representation (ragged and uniform), withλauto-scaled bybatch/nin the minibatch steps so the stochastic objective is an unbiased estimate of the full one. Extensible through the smallRegularizerprotocol inbackend/regularization.py. Docs:docs/fitting_and_optimization.md§4.9; exampleexamples/fit_hilbert_regularized.py(denoising, with λ chosen by held-out validation).Per-mode residual weighting — the fitting layer’s residual weight
ωgeneralizes from a per-order vector to anω[mode, order]matrix, owned by the sampling kind:probe_model(weight=)takes a length-dper-mode vector andprobe_derivatives_model(weight=)a(d, order+1)matrix (probe is the only kind with a per-mode axis, soapply/entriesstay per-order). A bare vector is still read as per-order, so existing calls are unchanged. Docs:docs/fitting_and_optimization.md§4.6; exampleexamples/fit_per_mode_weight_probes.py(inverse-noise weighting of a noisier mode).Newton-CG diagnostics —
newton_cg(..., verbose=True)prints a per-iteration block (objective and gradient norms, CG statistics, line search, ρ, wall time) plus a per-(mode, order)relative-error table, with an optional held-out validation column viaval_sample=/val_data=; the same records are returned instats['diagnostics']andstats['history'], and a regularized run splits the objective asobj = misfit + reg. Backend-owned (backend.optimizer_display.make_newton_display+ acallback=hook), so a raw-.datauser gets the identical display; ragged and uniform. Exampleexamples/fit_probe_display.py.Newton-CG warm-start controls —
g0norm_newton/g0norm_cgpin the reference‖g₀‖that the Newton stopping test and the CG forcing term are measured against (the computed initial norm is misleadingly small after a rank-continuation warm start;g0norm_newtonalso feeds CG unlessg0norm_cgis given), andcg_forcing_power(default0.5) trades CG iterations per Newton step against the number of Newton steps.NewtonInfo.g0normreports the effective reference. Docs:docs/fitting_and_optimization.md§5.chunk_sizefor the probe-derivative transpose — the𝒥ᵀvariation assemblies are computed in chunks over the probe stackW, trading recompute for peak memory:chunk_size(default100,None= dense) threads through the whole transpose chain and bothT3Tangent/UT3Tangent.probe_derivatives_transpose, and the optimizers acceptchunk_size='auto'(the default), which sizes it fromx0’s shapes. The sizing helpersbackend.sampling_derivatives.estimate_chunk_size/max_chunk_size_withinare eager (outsidejit), measure the per-row cost by lowering the real kernel rather than estimating it analytically, and take ann_shards=for per-device sizing and astack_shape=for a batch of base points (the frame stackCmultiplies the assembly but not an absolute byte budget, so omitting it on a stacked frame would makemax_chunk_size_withinreturn a chunk up toprod(C)times too large). Docs:docs/chunking.md.Recurrence/scan jets are now the standard derivative path — the pushthrough, combine and variation-assembly jets are computed by affine two-term recurrences and order-scans rather than by contracting the dense binomial
trstensor. The lean forms own the canonical names (compute_mu_jets,compute_eta_jets,compute_{sigma,tau}_jets,compute_deta_jets, the tilde twins, andassemble_{tt,tucker}_variation_jets); the dense forms remain public as the*_trsreference twins (numerically equal, and the test oracle). Underjiton the uniform layer this is a large memory win at unchanged accuracy — measured at rank 128,W=32000: ~14–28× less XLA temporary for theetajets and 168 GB → 2.63 GB (64×) for the TT-core variation assembly.
Changed#
BREAKING: the backend
optimizers.GeometryOpsprotocol gains an optionalprecomputeslot (frame -> geometry aux,Nonefor the existing geometries), andproject/retracttake a third argument ((frame, variations, aux=None)).Problem.local_modelbuilds the aux once per Newton step (aLocalModel.geom_auxleaf field besidesweep) and passes it toproject/retract; the frontendGaussNewtonModelmirrors it as ageometry_auxleaf. Migration for customGeometryOps: accept (and ignore)aux=Noneinproject/retract; the shared geometries use the slot to compute their per-frame SVD companion once per local model instead of once per CG matvec.The
Regularizerprotocol threads the geometry aux:gradient/hessian/quadraticgain anaux=Noneparameter (the per-frame geometry companion, e.g. the SF-T3SharedFrameData), and every model/LocalModelcall site passes its stored companion — closing the one seam where a regularized shared fit rebuilt the companion per CG matvec. CustomRegularizerimplementations: accept (and may ignore)aux=None. Design record:docs/contributor/precompute_and_caching.md(the precompute/caching principle and audit).backend/common.pygainsprefix_mask(the boolean prefix indicator shared by every uniform prefix structure) and now hostsrequire_concrete_masks, which moved frombackend/ut3_masking.py— it is infrastructure for the uniform mask-representation contract, not part of any one object family.use_jit=Truenow converts its inputs to jax instead of silently running eager. Previously the flag was accepted but ignored unlessx0/sample/datawere already jax arrays, so a “jit” run could be an eager one (and a meaningless benchmark). Requesting jit is now taken as opting into jax-world precision:x0,problem.sampleandproblem.dataare moved onto jax, the result comes back jax-backed (float32 unlessjax_enable_x64is set), and the call raises if jax is not installed.use_jitis also now an explicit keyword onnewton_cg/mc_sgd/adamrather than arriving through**kwargs. Docs:docs/fitting_and_optimization.md§4.5.
Fixed#
A
GaussNewtonModelbuilt from a parameterized sampling kind recompiled on every rebuild. The model carries itsSamplingKindas jax pytreeaux_data, and jax keys its compilation cache on the aux; the derivative kinds and the weighted probe kind are built per model out of fresh closures, which under dataclass field equality never compare equal. So the documented “roll your own optimizer” pattern — rebuild the model at each outer point,jita function of it — paid a full recompile every step (measured: 3 traces for 3 rebuilds, against 1 for theapply/entries/probesingletons).SamplingKindnow compares on a valueidentity(name, order, residual weight,chunk_size) rather than on its closures, so a rebuilt kind with the same parameters is the same cache key while a genuinely different one still gets its own compilation. The uniform model already did this through its value-hashed aux; the ragged twin now matches, and both are pinned by regression tests.backend.fv_conversionsignoredsquash_tails=False— the parameter was shadowed inside the function, so the tails were squashed regardless of what the caller asked for.The uniform frame/variation entry points (
ufv_apply_frame_masks,ufv_apply_variations_masks,ut3frame_to_t3frame,ut3variations_to_t3variations) accepted traced masks and failed deep inside jax with aTracerArrayConversionError; they now callrequire_concrete_masksand raise the actionable message, like the plain-layer entry points already did.
Removed#
BREAKING: the ~104 named contraction functions in
backend.contractions(WCa_Caib_WCi_to_WCb-style), replaced by thecontractinterpreter above. Migration is mechanical: the function name is the subscripts string (X_Y_to_Z(a, b)→contract('X,Y->Z', a, b)), and the trailingn_probe/n_framearguments becomelen_W=/len_C=keywords. Numerically identical (each named function was verified equal to itscontractcall over an empty/single/multi-axis block-shape matrix before removal).The old parked weighted layer (
weighted_tucker_tensor_train.py,backend/wt3_operations.py, thewt3_prefix, and the brokenabsorb_weights_into_tangent_cores), superseded by the above.
[2026.0.0] — 2026-07-13#
The first public release — the initial public surface:
The Tucker tensor train (T3) format — arithmetic with dense-tensor semantics, orthogonalization, minimal ranks, T3-SVD, save/load, batching on every operation.
The three sampling operations (
entries/apply/probe), their symmetric directional derivatives, and the ambient/corewise/tangent transposes.The fixed-rank T3 manifold: orthogonal frame + gauged variations, tangent vectors, gauge projections, retraction, and the
MANIFOLD/COREWISEgeometries.Least-squares fitting from any sampling operation or its derivatives (Gauss-Newton models) with four optimizers (
gradient_descent,mc_sgd,adam,newton_cg).The uniform layer: supercores + boolean rank masks mirroring the whole stack, for
jax.lax.scanvectorization and compile-oncejit(optimizers included).NumPy / JAX backends with dispatch inferred from input array types; safe mode for numerical-precondition checking.