# Deep Arm — Depth Extension for captionbert-8192-v2-B **A technical companion to the week-in-review article. Ships as `deep-arm/TECHNICAL.md` inside the [`AbstractPhil/captionbert-8192-v2-B`](https://huggingface.co/AbstractPhil/captionbert-8192-v2-B) repository.** This document describes **deep-1**: a ~5M-parameter depth-extension module ("deep arm") that extends the encoder's usable conditioning depth from 256 tokens to 2048+ while leaving the base encoder untouched. It covers the measurement that motivated it, the construction, the preregistered gate and its result, a zero-parameter finding about positional initialization, the measured cost of the engaged mode, and the variant catalogue this work opens. --- ## 1. Overview captionbert-8192-v2-B is a 512-dimensional caption encoder whose position table carries 8192 rows, of which only rows 0–255 were ever trained (the training corpus averaged 48 tokens per caption). The deep arm is a **switchable long-prompt mode** layered over that frozen base: - **Disengaged**, the model is bit-identical to stock — the added adapters are gated wrappers initialized nearly closed, and position rows 0–255 are frozen throughout, so short-input behavior is preserved by construction (and verified at benchmark, section 6). - **Engaged**, the model carries attribute binding — the semantic property long prompts actually need — at every measured depth from 10 to 2288 tokens, where the stock encoder's binding collapses to chance near depth 480. The arm adds ~5M trainable parameters (~4.06M re-initialized deep position rows + ~0.92M gated relay adapters) against a frozen trunk, trained in 8k steps overnight on a single consumer GPU by [distilling](https://arxiv.org/abs/1503.02531) token-state behavior from [Longformer-base-4096](https://arxiv.org/abs/2004.05150) across a learned cross-dimensional bridge. Its preregistered depth gate passed, and a labeled harder secondary passed as well. --- ## 2. Motivation — the gauge must match the consumer ### 2.1 Similarity plateaus; binding dies Downstream consumers of this encoder (diffusion cross-attention among them) do not read a pooled sentence vector — they read **per-token states**. So depth claims have to be made on a gauge those consumers care about. Under a wide-spectrum depth instrument (15 depths out to 8000 tokens, 138 measurement rows, 10 models), the encoder's state-similarity curve looks healthy at depth: a degraded-but-stable plateau around **.77** with no cliff. Token recoverability also holds up, at **.61–.66 through depth 1024** — the strongest of the models tested. But the semantic gauge tells a different story: **attribute binding collapses to roughly chance (~1.0) near token depth 480 and never returns.** For reference on the same instrument, [t5-base](https://arxiv.org/abs/1910.10683) holds binding 1.9–3.0 through 2048 (relative position encodings are the real long-context contribution there), and [bert-base](https://arxiv.org/abs/1810.04805) reads 3.6 within its native 512. *The binding gauge, briefly:* a contrast score read from the token states at the two noun positions of a relational caption ("a red cube on a blue sphere") against its attribute-swapped counterpart. A score near 1.0 means the states cannot distinguish the correct assignment from the swapped one — chance. Higher is stronger binding. `bind1`/`bind2` denote the first and second noun of the relation. The law this minted: **depth claims require spectrum sampling plus the semantic gauge; similarity curves lie.** A "stable plateau to 1500" read from sparse sampling on the wrong gauge had stood for exactly one day before re-instrumentation overturned it. ### 2.2 The failure decomposes into two mechanical parts An offset-curve instrument placed the binding payload at window offsets 2–240 inside the trained positional range, embedded in realistic caption context. Binding is alive (**1.4–2.6**) at every offset from 2 through 192 and dies only at offset 240 — the truncation edge, where a bidirectional encoder's second noun loses its right context. The depth cliff therefore resolves into exactly two defects: 1. **Untrained position rows past 256** — an initialization problem. 2. **Right-context starvation in the last ~16 pre-edge tokens** — a windowing problem. Context interference is refuted. Both defects are mechanical, and each fix below targets one or both; each fix's success is itself a confirmation of the decomposition. A zero-training wrapper (V0, section 7) already validates the story: encoding in 256-token windows with a ~16-token edge margin (stride ~224–240, clause-aligned splits) recovers shallow-grade binding at any prompt depth, with cross-window anaphora as the only loss. The deep arm is the **native** fix — no wrapper, single forward pass, trunk untouched. --- ## 3. Design Seven load-bearing elements: 1. **Frozen trunk.** captionbert-8192-v2-B loaded through the certified loader. Disengaged mode is bit-identical stock by construction: adapters are sigma-gated wrappers born nearly closed, and position rows 0–255 are never updated. 2. **Mod-256 tiled deep position rows (~4.06M trainable).** Rows 256+ are re-initialized by tiling the trained 0–255 table (plus small noise), and made trainable under a masked-gradient hook so only the deep rows learn. 3. **Gated relay adapters (MslRelay, ~0.92M trainable).** Small certified adapter modules wrapped around each frozen block, sigma-gated. 4. **Longformer targets.** Teacher token states from Longformer-base-4096 on long documents, span-resampled onto the encoder's token grid, then mapped 768→512. 5. **Cross-dimensional [Procrustes](https://doi.org/10.1007/BF02291478) bridge.** A whitened-Procrustes map into the encoder's real 512-d token frame, fit at **cos .5013 / retrieval R@1 .849**, with a paired shuffled-null control at floor. This exceeds the earlier 768-d proxy-frame fit (.35). Stage 0 aborts by construction if the fit lands below cos .20. 6. **Loss.** Per-token (1 − cosine) to the mapped teacher targets on valid positions, with 2× weight past position 256. Optimizer: Adam, weight decay 0. 7. **The gate is the instrument.** The binding-at-depth battery runs inside the eval loop, so the preregistered criterion is measured continuously rather than once at the end. Total: **~5M trainable parameters over a frozen trunk**; 8k training steps; overnight on a single consumer GPU. ### Gates cleared before any training compute - **Teacher check.** Longformer-base-4096's deep binding measures **2.1–3.5 at depths 480–3800** — the strongest measured in this battery — and it strengthens with depth. A teacher without the target property would make the distillation pointless; this one has it. - **Bridge check.** The 512-d Procrustes fit above (.5013 / .849) cleared its .20 abort threshold with margin. --- ## 4. Preregistered gate and result **Preregistered bar (set before training):** deep binding at depths {1024, 2048} must reach at least **0.8× the shallow value**. **Final battery (bind1 / bind2):** | Depth | bind1 | bind2 | |---|---|---| | 10 | 2.69 | 2.27 | | 480 | 1.63 | 3.50 | | 1024 | 2.66 | 3.96 | | 1248 | 3.18 | 3.80 | | 2048 | 4.83 | 13.14 | | 2288 | 1.63 | 1.58 | **All cells above 1.5, no inversions. The preregistered gate passed.** Two discipline notes belong next to that result: - **The "too lucky" rule.** The preregistered gate passed *at initialization*, before any training, because of the tiling effect described in section 5. Rather than declaring a quiet victory, harder off-phase depths (1248 and 2288 — depths not aligned to the 256-token tile) were added to the battery as a **labeled post-hoc secondary**. They passed as well, which is what makes the headline pass credible. - **Open flag, owed before any strength claim.** The aligned bind2 value of **13.1 at depth 2048** requires an absolute-distance decomposition before it can be read as strength — it may be inflated by aligned-tile geometry. Until that decomposition is published, read it as "well above threshold," not as a magnitude claim. **Consequence.** The encoder's "≤256 tokens" conditioning constraint is repealed: captionbert-B + deep arm binds natively through 2048+ tokens, trunk untouched, at a cost of ~5M trainable parameters. --- ## 5. The zero-parameter finding, and what training actually bought ### 5.1 Tiling alone restores aligned-depth binding Re-initializing rows 256+ by tiling the trained 0–255 table — **zero parameters trained** — restores binding of **2.7–3.3 at tile-aligned depths** (1024, 2048: depth mod 256 ≈ 0), versus ~1.0 (chance) under random row initialization. The failure does not disappear; it becomes **periodic**: tile-edge offsets stay at chance (depth 480, which lands at offset ~224 within its tile, reads 1.16 at init) — the section 2.2 edge pathology repeating every 256 positions. This re-scopes the training problem. The arm's job is not to learn depth from scratch; it is **boundary-smoothing and cross-window integration** — repairing the edges the tiling leaves behind. The tiling law is free to any model with a partially trained positional table. ### 5.2 Contribution ledger The result is reported per component, because the contributions are layered and non-competing — a single scoreboard number would erase the mechanism: - **Tiled init (0 params):** aligned-depth binding 1.0 → 2.7–3.3, at initialization. - **Deep position rows (~4.06M params):** edge repair arriving as a **near-to-far wave** — depth 480's bind2 climbs .81 → 3.50 by step 4.5k; depth 1248 passes through a transient inversion on its way from .72 → 3.80; depth 2288 crosses 1.5 only in the final 500 of 8k steps. - **Relay adapters (~0.92M params):** retroactive-binding amplification wherever gradient reached — bind1 moved first at every depth. The adapter gates finished **.35 → .51, opening monotonically with depth** — the third independent architecture in [this research line](https://huggingface.co/blog/AbstractPhil/geometric-memory-ft4) to show the depth-gradient regularity (its third substrate). - **Longformer supervision through the .5013-cos bridge:** shallow behavior **improved** under it — depth-10 binding 1.59/2.01 → 2.69/2.27. Depth was bought without taxing the trained regime. --- ## 6. Honest costs The engaged mode is not free, and both numbers ship together. - **Engaged-mode short-input cost:** on the certified 8-task benchmark, engaged mode scores **.5655 versus stock .6031 — a −.0376 mean difference** — with token self-cosine up +.288: the Longformer-mapped frame's anisotropy pulls shallow states toward a cone. - **Disengaged mode: bit-identical stock.** Because the adapters are sigma-gated wrappers and rows 0–255 are frozen, disengaging the arm reproduces stock behavior exactly — guaranteed by construction and verified at benchmark. The arm is therefore carded as a **length-conditional mode**: engage it for long prompts, leave it off for short ones. The capability trade is a mode switch, not a regression. Closing the −.0376 gap natively is the named refinement path (V2, below). --- ## 7. Variant catalogue — V0 through V6 | # | Name | Status | What it is / why it exists | |---|---|---|---| | V0 | Chunk wrapper | Validated, zero training | 256-token windows with a ~16-token edge margin (stride ~224–240), clause-aligned; shallow-grade binding at any depth; window placement is part of the contract (an edge-placed payload falsely reads "no help") | | V1 | deep-1 as shipped | Trained, gate PASS | Tiled deep rows + relay adapters under Longformer supervision; the baseline variant, described in this document | | V2 | Joint-objective revision | Named, unbuilt | Distillation plus a shallow-capability anchor, to close the .5655 engaged-mode gap; the refinement path recorded on the model card | | V3 | Token-fidelity masked-span arm | Named, unbuilt | A span-reconstruction objective on the existing trunks — the targeted fix for the encoder's binding gap and the conditioning line's shelf upgrade | | V4 | Interpolation-init deep rows | Named, unbuilt | Interpolation initialization as an ablation against tiling (which is free and strong) | | V5 | Anchored-trunk deep arm | Named, unbuilt — queue head | The V1 recipe applied to the [AlephLM-0](https://huggingface.co/AbstractPhil/alephlm-0) anchored trunks — the variant that extends that family's inference depth. Two free reads: whether signed dispatch shapes depth repair differently than a dense trunk, and whether the a3 trunk's early-grid diversity advantage survives depth extension. Requires normalized-gauge grid reads and context layer-norm in any consumer | | V6 | [RoPE](https://arxiv.org/abs/2104.09864)-native next trunk | Design-level | A related hybrid with a RoPE trunk measures ~.98 positional integrity, flat to depth 8000 — evidence that rotary positions remove the position-row problem at the root for the *next* pretrain, rather than retrofitting this one | **Queue by leverage:** V5 → V2 → V3 → V6. All of V2/V3/V5 are single-consumer-GPU overnight runs at the deep arm's ~5M-parameter scale. --- ## 8. Laws established or confirmed on this line 1. **The spectrum-sampling law.** Depth claims require spectrum sampling plus the semantic gauge (binding); state-similarity curves lie. A plausible "stable plateau to 1500" survived exactly one day before the right instrument overturned it. 2. **The two-part depth decomposition.** The long-context failure is exactly (a) untrained position rows and (b) pre-edge right-context starvation; context interference is refuted. Fixes become mechanical, and their success confirms the theory. 3. **The tiling law.** Tiling a trained positional table into untrained rows restores aligned-depth semantics at zero cost and makes the residual failure periodic — re-scoping depth training as boundary-smoothing. Free to any model with a partially trained positional table. 4. **Window placement is part of the wrapper's contract.** An edge-placed payload falsely reads "no help"; the naive test's false negative would have discarded the cheapest fix in the catalogue. 5. **The "too lucky" rule.** A preregistered gate that passes at initialization triggers labeled harder secondaries rather than a quiet victory. deep-1's pass is credible because off-phase depths were added — and passed. 6. **Honest dual-mode accounting.** A module that buys depth at a measured shallow cost ships with both numbers (−.0376 here) and a construction-guaranteed disengaged mode. 7. **The depth-gradient law, third substrate.** Adapter gates opening monotonically with depth (.35 → .51) reappeared here — the third independent architecture in this line to show it. 8. **Contribution ledgers, not leaderboards.** The result is reported as per-component contributions (init / rows / relays / supervision) because the components are layered and non-competing; a single scoreboard number would have erased the mechanism. Two further measurement laws established in the same coverage window — amplitude-vs-direction in grid-rank reads, and the invariance of binding under per-token normalization — belong to the adjacent conditioning-viability line and are covered in the accompanying week-in-review article. --- ## 9. Provenance Research and writing: **AbstractPhil**. Coverage window: **2026-08-04 through 2026-08-06**. Shipped artifacts, in this repository under `deep-arm/`: the trained arm weights (`deep1_arm_s0.pt`), the full training and battery record (`deep1_results.json`), and the engaged-mode benchmark (`deep_arm_bench.json`), alongside the model-card section covering construction, the before/after battery, the bridge fit, the honest cost, toggle semantics, and usage. Every number in this document is inherited from the registered measurement record for that window. --- ## References and attribution **Models** - Longformer — Beltagy et al., [arXiv:2004.05150](https://arxiv.org/abs/2004.05150). The distillation teacher (longformer-base-4096), source of the deep token-state targets. - T5 — Raffel et al., [arXiv:1910.10683](https://arxiv.org/abs/1910.10683). Depth-battery comparator (t5-base), whose relative position encodings hold binding through 2048. - BERT — Devlin et al., [arXiv:1810.04805](https://arxiv.org/abs/1810.04805). Depth-battery comparator (bert-base), read within its native 512. - CLIP — Radford et al., [arXiv:2103.00020](https://arxiv.org/abs/2103.00020). Lineage of the conditioning-consumer framing: downstream consumers read per-token encoder states through cross-attention rather than a pooled vector. **Data** - MTEB — [arXiv:2210.07316](https://arxiv.org/abs/2210.07316). Source of the suites behind the certified 8-task short-input benchmark used for the engaged-mode cost. **Methods** - Procrustes analysis — Schönemann (1966), [doi:10.1007/BF02291478](https://doi.org/10.1007/BF02291478). The cross-dimensional bridge: a whitened-Procrustes map from the teacher's 768-d frame into the encoder's 512-d token frame. - Knowledge distillation — Hinton et al., [arXiv:1503.02531](https://arxiv.org/abs/1503.02531). The training scheme's lineage; here, per-token state distillation from the teacher across the bridge. - LoRA — Hu et al., [arXiv:2106.09685](https://arxiv.org/abs/2106.09685). Lineage of the small gated relay adapters trained over the frozen trunk. - RoPE — Su et al., [arXiv:2104.09864](https://arxiv.org/abs/2104.09864). Rotary position embedding; the V6 design direction that removes the position-row problem at the root. **Infrastructure** - PyTorch — [pytorch.org](https://pytorch.org). Training and evaluation stack for the arm and its batteries. **Related artifacts:** the trunk this arm extends, [AbstractPhil/captionbert-8192-v2-B](https://huggingface.co/AbstractPhil/captionbert-8192-v2-B) (this repository; the arm ships under `deep-arm/`) · the anchored trunks named as the V5 target, [AbstractPhil/alephlm-0](https://huggingface.co/AbstractPhil/alephlm-0) · the prior installment of this series, [geometric-memory-ft4](https://huggingface.co/blog/AbstractPhil/geometric-memory-ft4).