Modern linear attention variants are complex, and it is not easy to see what they are designed to achieve. a state kernel performs the only inter-chunk scan, producing the state entering each chunk and resolving its delta errors. the resulting matrices can be used to construct the causal akkakkakk interaction matrix, which is used for decoding if only one new token is available at the time. ...