I think the concerns about attribution and untested accuracy claims are reasonable. A few technical details from the blog also deserve consideration, though:
The 512 local tokens are included in the 2,048-token budget. The example specifies 512 local + 1,536 remotely selected tokens. With 128-token blocks, that means 12 remote blocks plus the local window, rather than 16 remote blocks with another 512 tokens added. This makes the proposed attention budget internally consistent.
Using the same
k = 2048as DSA does not make the mechanisms equivalent. DSA scores individual tokens with its lightning indexer. BGA proposes scoring block representations before retrieving their full-resolution tokens. At 1,048,576 tokens and 128 tokens per block, there are 8,192 block candidates. That is a plausible way to reduce routing work, although fewer candidates do not automatically imply a proportional speedup or equally accurate selection.Banaxi's distinction about the recurrent state is technically meaningful. NSA attends over a growing collection of compressed block representations. BGA proposes an additional small state updated recurrently from its previous contents and active information. That idea was already present in the original blog version. It is a conceptual difference, although the proposal still needs to specify how that state contributes to the output. It also does not make the whole system constant-memory, since older exact blocks remain searchable.
The 256× versus 512× distinction has a valid mathematical basis. Under the blog's assumptions, causal attention averages roughly
n/2accessible positions across a sequence, while a decoding step near its end sees approximatelyn. Comparing those withkgivesn/(2k)andn/k. These are estimates for the main attention interactions, not measured training or generation speedups. The blog explicitly acknowledges other costs, so that qualification deserves recognition.Selecting contiguous blocks has a sound hardware motivation. NSA itself discusses more efficient memory access and Tensor Core utilization with blockwise selection. This supports the engineering rationale behind BGA's routing choice, even though the idea is established and BGA's actual implementation would still need benchmarking.
The current blog also acknowledges NSA, MoBA, Landmark Attention, Routing Transformer, and Infini-attention. That is a useful correction to the original omission.
None of these points establishes that BGA is novel or more accurate than DSA. They do make the proposal more nuanced than a block-size change alone. A useful next step would be a concrete implementation and a controlled comparison, including an ablation with and without the recurrent state, measuring quality, memory usage, and actual runtime.