AbstractPhil's picture
references+attribution section, inline citations
4fdc332 verified
|
Raw
History Blame Contribute Delete
18.4 kB

Deep Arm β€” Depth Extension for captionbert-8192-v2-B

A technical companion to the week-in-review article. Ships as deep-arm/TECHNICAL.md inside the AbstractPhil/captionbert-8192-v2-B repository.

This document describes deep-1: a ~5M-parameter depth-extension module ("deep arm") that extends the encoder's usable conditioning depth from 256 tokens to 2048+ while leaving the base encoder untouched. It covers the measurement that motivated it, the construction, the preregistered gate and its result, a zero-parameter finding about positional initialization, the measured cost of the engaged mode, and the variant catalogue this work opens.


1. Overview

captionbert-8192-v2-B is a 512-dimensional caption encoder whose position table carries 8192 rows, of which only rows 0–255 were ever trained (the training corpus averaged 48 tokens per caption). The deep arm is a switchable long-prompt mode layered over that frozen base:

  • Disengaged, the model is bit-identical to stock β€” the added adapters are gated wrappers initialized nearly closed, and position rows 0–255 are frozen throughout, so short-input behavior is preserved by construction (and verified at benchmark, section 6).
  • Engaged, the model carries attribute binding β€” the semantic property long prompts actually need β€” at every measured depth from 10 to 2288 tokens, where the stock encoder's binding collapses to chance near depth 480.

The arm adds 5M trainable parameters (4.06M re-initialized deep position rows + ~0.92M gated relay adapters) against a frozen trunk, trained in 8k steps overnight on a single consumer GPU by distilling token-state behavior from Longformer-base-4096 across a learned cross-dimensional bridge. Its preregistered depth gate passed, and a labeled harder secondary passed as well.


2. Motivation β€” the gauge must match the consumer

2.1 Similarity plateaus; binding dies

Downstream consumers of this encoder (diffusion cross-attention among them) do not read a pooled sentence vector β€” they read per-token states. So depth claims have to be made on a gauge those consumers care about.

Under a wide-spectrum depth instrument (15 depths out to 8000 tokens, 138 measurement rows, 10 models), the encoder's state-similarity curve looks healthy at depth: a degraded-but-stable plateau around .77 with no cliff. Token recoverability also holds up, at .61–.66 through depth 1024 β€” the strongest of the models tested. But the semantic gauge tells a different story: attribute binding collapses to roughly chance (~1.0) near token depth 480 and never returns. For reference on the same instrument, t5-base holds binding 1.9–3.0 through 2048 (relative position encodings are the real long-context contribution there), and bert-base reads 3.6 within its native 512.

The binding gauge, briefly: a contrast score read from the token states at the two noun positions of a relational caption ("a red cube on a blue sphere") against its attribute-swapped counterpart. A score near 1.0 means the states cannot distinguish the correct assignment from the swapped one β€” chance. Higher is stronger binding. bind1/bind2 denote the first and second noun of the relation.

The law this minted: depth claims require spectrum sampling plus the semantic gauge; similarity curves lie. A "stable plateau to 1500" read from sparse sampling on the wrong gauge had stood for exactly one day before re-instrumentation overturned it.

2.2 The failure decomposes into two mechanical parts

An offset-curve instrument placed the binding payload at window offsets 2–240 inside the trained positional range, embedded in realistic caption context. Binding is alive (1.4–2.6) at every offset from 2 through 192 and dies only at offset 240 β€” the truncation edge, where a bidirectional encoder's second noun loses its right context. The depth cliff therefore resolves into exactly two defects:

  1. Untrained position rows past 256 β€” an initialization problem.
  2. Right-context starvation in the last ~16 pre-edge tokens β€” a windowing problem.

Context interference is refuted. Both defects are mechanical, and each fix below targets one or both; each fix's success is itself a confirmation of the decomposition.

A zero-training wrapper (V0, section 7) already validates the story: encoding in 256-token windows with a ~16-token edge margin (stride ~224–240, clause-aligned splits) recovers shallow-grade binding at any prompt depth, with cross-window anaphora as the only loss. The deep arm is the native fix β€” no wrapper, single forward pass, trunk untouched.


3. Design

Seven load-bearing elements:

  1. Frozen trunk. captionbert-8192-v2-B loaded through the certified loader. Disengaged mode is bit-identical stock by construction: adapters are sigma-gated wrappers born nearly closed, and position rows 0–255 are never updated.
  2. Mod-256 tiled deep position rows (~4.06M trainable). Rows 256+ are re-initialized by tiling the trained 0–255 table (plus small noise), and made trainable under a masked-gradient hook so only the deep rows learn.
  3. Gated relay adapters (MslRelay, ~0.92M trainable). Small certified adapter modules wrapped around each frozen block, sigma-gated.
  4. Longformer targets. Teacher token states from Longformer-base-4096 on long documents, span-resampled onto the encoder's token grid, then mapped 768β†’512.
  5. Cross-dimensional Procrustes bridge. A whitened-Procrustes map into the encoder's real 512-d token frame, fit at cos .5013 / retrieval R@1 .849, with a paired shuffled-null control at floor. This exceeds the earlier 768-d proxy-frame fit (.35). Stage 0 aborts by construction if the fit lands below cos .20.
  6. Loss. Per-token (1 βˆ’ cosine) to the mapped teacher targets on valid positions, with 2Γ— weight past position 256. Optimizer: Adam, weight decay 0.
  7. The gate is the instrument. The binding-at-depth battery runs inside the eval loop, so the preregistered criterion is measured continuously rather than once at the end.

Total: ~5M trainable parameters over a frozen trunk; 8k training steps; overnight on a single consumer GPU.

Gates cleared before any training compute

  • Teacher check. Longformer-base-4096's deep binding measures 2.1–3.5 at depths 480–3800 β€” the strongest measured in this battery β€” and it strengthens with depth. A teacher without the target property would make the distillation pointless; this one has it.
  • Bridge check. The 512-d Procrustes fit above (.5013 / .849) cleared its .20 abort threshold with margin.

4. Preregistered gate and result

Preregistered bar (set before training): deep binding at depths {1024, 2048} must reach at least 0.8Γ— the shallow value.

Final battery (bind1 / bind2):

Depth bind1 bind2
10 2.69 2.27
480 1.63 3.50
1024 2.66 3.96
1248 3.18 3.80
2048 4.83 13.14
2288 1.63 1.58

All cells above 1.5, no inversions. The preregistered gate passed.

Two discipline notes belong next to that result:

  • The "too lucky" rule. The preregistered gate passed at initialization, before any training, because of the tiling effect described in section 5. Rather than declaring a quiet victory, harder off-phase depths (1248 and 2288 β€” depths not aligned to the 256-token tile) were added to the battery as a labeled post-hoc secondary. They passed as well, which is what makes the headline pass credible.
  • Open flag, owed before any strength claim. The aligned bind2 value of 13.1 at depth 2048 requires an absolute-distance decomposition before it can be read as strength β€” it may be inflated by aligned-tile geometry. Until that decomposition is published, read it as "well above threshold," not as a magnitude claim.

Consequence. The encoder's "≀256 tokens" conditioning constraint is repealed: captionbert-B + deep arm binds natively through 2048+ tokens, trunk untouched, at a cost of ~5M trainable parameters.


5. The zero-parameter finding, and what training actually bought

5.1 Tiling alone restores aligned-depth binding

Re-initializing rows 256+ by tiling the trained 0–255 table β€” zero parameters trained β€” restores binding of 2.7–3.3 at tile-aligned depths (1024, 2048: depth mod 256 β‰ˆ 0), versus ~1.0 (chance) under random row initialization. The failure does not disappear; it becomes periodic: tile-edge offsets stay at chance (depth 480, which lands at offset ~224 within its tile, reads 1.16 at init) β€” the section 2.2 edge pathology repeating every 256 positions.

This re-scopes the training problem. The arm's job is not to learn depth from scratch; it is boundary-smoothing and cross-window integration β€” repairing the edges the tiling leaves behind. The tiling law is free to any model with a partially trained positional table.

5.2 Contribution ledger

The result is reported per component, because the contributions are layered and non-competing β€” a single scoreboard number would erase the mechanism:

  • Tiled init (0 params): aligned-depth binding 1.0 β†’ 2.7–3.3, at initialization.
  • Deep position rows (~4.06M params): edge repair arriving as a near-to-far wave β€” depth 480's bind2 climbs .81 β†’ 3.50 by step 4.5k; depth 1248 passes through a transient inversion on its way from .72 β†’ 3.80; depth 2288 crosses 1.5 only in the final 500 of 8k steps.
  • Relay adapters (~0.92M params): retroactive-binding amplification wherever gradient reached β€” bind1 moved first at every depth. The adapter gates finished .35 β†’ .51, opening monotonically with depth β€” the third independent architecture in this research line to show the depth-gradient regularity (its third substrate).
  • Longformer supervision through the .5013-cos bridge: shallow behavior improved under it β€” depth-10 binding 1.59/2.01 β†’ 2.69/2.27. Depth was bought without taxing the trained regime.

6. Honest costs

The engaged mode is not free, and both numbers ship together.

  • Engaged-mode short-input cost: on the certified 8-task benchmark, engaged mode scores .5655 versus stock .6031 β€” a βˆ’.0376 mean difference β€” with token self-cosine up +.288: the Longformer-mapped frame's anisotropy pulls shallow states toward a cone.
  • Disengaged mode: bit-identical stock. Because the adapters are sigma-gated wrappers and rows 0–255 are frozen, disengaging the arm reproduces stock behavior exactly β€” guaranteed by construction and verified at benchmark.

The arm is therefore carded as a length-conditional mode: engage it for long prompts, leave it off for short ones. The capability trade is a mode switch, not a regression. Closing the βˆ’.0376 gap natively is the named refinement path (V2, below).


7. Variant catalogue β€” V0 through V6

# Name Status What it is / why it exists
V0 Chunk wrapper Validated, zero training 256-token windows with a ~16-token edge margin (stride ~224–240), clause-aligned; shallow-grade binding at any depth; window placement is part of the contract (an edge-placed payload falsely reads "no help")
V1 deep-1 as shipped Trained, gate PASS Tiled deep rows + relay adapters under Longformer supervision; the baseline variant, described in this document
V2 Joint-objective revision Named, unbuilt Distillation plus a shallow-capability anchor, to close the .5655 engaged-mode gap; the refinement path recorded on the model card
V3 Token-fidelity masked-span arm Named, unbuilt A span-reconstruction objective on the existing trunks β€” the targeted fix for the encoder's binding gap and the conditioning line's shelf upgrade
V4 Interpolation-init deep rows Named, unbuilt Interpolation initialization as an ablation against tiling (which is free and strong)
V5 Anchored-trunk deep arm Named, unbuilt β€” queue head The V1 recipe applied to the AlephLM-0 anchored trunks β€” the variant that extends that family's inference depth. Two free reads: whether signed dispatch shapes depth repair differently than a dense trunk, and whether the a3 trunk's early-grid diversity advantage survives depth extension. Requires normalized-gauge grid reads and context layer-norm in any consumer
V6 RoPE-native next trunk Design-level A related hybrid with a RoPE trunk measures ~.98 positional integrity, flat to depth 8000 β€” evidence that rotary positions remove the position-row problem at the root for the next pretrain, rather than retrofitting this one

Queue by leverage: V5 β†’ V2 β†’ V3 β†’ V6. All of V2/V3/V5 are single-consumer-GPU overnight runs at the deep arm's ~5M-parameter scale.


8. Laws established or confirmed on this line

  1. The spectrum-sampling law. Depth claims require spectrum sampling plus the semantic gauge (binding); state-similarity curves lie. A plausible "stable plateau to 1500" survived exactly one day before the right instrument overturned it.
  2. The two-part depth decomposition. The long-context failure is exactly (a) untrained position rows and (b) pre-edge right-context starvation; context interference is refuted. Fixes become mechanical, and their success confirms the theory.
  3. The tiling law. Tiling a trained positional table into untrained rows restores aligned-depth semantics at zero cost and makes the residual failure periodic β€” re-scoping depth training as boundary-smoothing. Free to any model with a partially trained positional table.
  4. Window placement is part of the wrapper's contract. An edge-placed payload falsely reads "no help"; the naive test's false negative would have discarded the cheapest fix in the catalogue.
  5. The "too lucky" rule. A preregistered gate that passes at initialization triggers labeled harder secondaries rather than a quiet victory. deep-1's pass is credible because off-phase depths were added β€” and passed.
  6. Honest dual-mode accounting. A module that buys depth at a measured shallow cost ships with both numbers (βˆ’.0376 here) and a construction-guaranteed disengaged mode.
  7. The depth-gradient law, third substrate. Adapter gates opening monotonically with depth (.35 β†’ .51) reappeared here β€” the third independent architecture in this line to show it.
  8. Contribution ledgers, not leaderboards. The result is reported as per-component contributions (init / rows / relays / supervision) because the components are layered and non-competing; a single scoreboard number would have erased the mechanism.

Two further measurement laws established in the same coverage window β€” amplitude-vs-direction in grid-rank reads, and the invariance of binding under per-token normalization β€” belong to the adjacent conditioning-viability line and are covered in the accompanying week-in-review article.


9. Provenance

Research and writing: AbstractPhil. Coverage window: 2026-08-04 through 2026-08-06. Shipped artifacts, in this repository under deep-arm/: the trained arm weights (deep1_arm_s0.pt), the full training and battery record (deep1_results.json), and the engaged-mode benchmark (deep_arm_bench.json), alongside the model-card section covering construction, the before/after battery, the bridge fit, the honest cost, toggle semantics, and usage. Every number in this document is inherited from the registered measurement record for that window.


References and attribution

Models

  • Longformer β€” Beltagy et al., arXiv:2004.05150. The distillation teacher (longformer-base-4096), source of the deep token-state targets.
  • T5 β€” Raffel et al., arXiv:1910.10683. Depth-battery comparator (t5-base), whose relative position encodings hold binding through 2048.
  • BERT β€” Devlin et al., arXiv:1810.04805. Depth-battery comparator (bert-base), read within its native 512.
  • CLIP β€” Radford et al., arXiv:2103.00020. Lineage of the conditioning-consumer framing: downstream consumers read per-token encoder states through cross-attention rather than a pooled vector.

Data

  • MTEB β€” arXiv:2210.07316. Source of the suites behind the certified 8-task short-input benchmark used for the engaged-mode cost.

Methods

  • Procrustes analysis β€” SchΓΆnemann (1966), doi:10.1007/BF02291478. The cross-dimensional bridge: a whitened-Procrustes map from the teacher's 768-d frame into the encoder's 512-d token frame.
  • Knowledge distillation β€” Hinton et al., arXiv:1503.02531. The training scheme's lineage; here, per-token state distillation from the teacher across the bridge.
  • LoRA β€” Hu et al., arXiv:2106.09685. Lineage of the small gated relay adapters trained over the frozen trunk.
  • RoPE β€” Su et al., arXiv:2104.09864. Rotary position embedding; the V6 design direction that removes the position-row problem at the root.

Infrastructure

  • PyTorch β€” pytorch.org. Training and evaluation stack for the arm and its batteries.

Related artifacts: the trunk this arm extends, AbstractPhil/captionbert-8192-v2-B (this repository; the arm ships under deep-arm/) Β· the anchored trunks named as the V5 target, AbstractPhil/alephlm-0 Β· the prior installment of this series, geometric-memory-ft4.