Sequence modelling · State-space models
My contribution
I implemented a diagonal state-space layer (S4D) from scratch and compared it with a causal Transformer in a shared backbone across context lengths 64 to 1024, on Apple MPS and on a Tesla T4.
Verified resultsResearch implementation · one seed, fixed training budget
S4D reached lower validation loss at every tested context length. On a Tesla T4 at context 1024 it trained at about 272k tokens/s against 67k for attention, with 1,349 against 6,455 MiB peak allocated memory (one seed, fixed training budget).
Problem
Deep state-space models share their core recurrence, x′(t) = Ax(t) + Bu(t), with classical control theory. This study asks how a diagonal state-space layer compares with causal self-attention in accuracy, throughput and memory as the context length grows, when everything else in the model and the training is held the same.
Approach
A from-scratch PyTorch implementation of S4D (Gu, Gupta, Goel and Ré, 2022) and a from-scratch causal Transformer share one backbone with a swappable mixer: embeddings, MLP, normalisation and weight tying are identical. Both are trained on character-level Tiny Shakespeare with the same optimiser, schedule, batch size, seed, depth and width, at context lengths 64 to 1024, for 1,200 steps (400 at context 1024).
Verification
The training-time convolutional kernel and the step-by-step recurrence of the S4D layer are tested to compute the same function, at initialisation and after an optimiser step, with a measured maximum difference of about 9 × 10⁻⁸ to 2.4 × 10⁻⁶ in float32.
Result
S4D reached lower final validation loss than attention at every tested context length, on both backends. On the Tesla T4, attention had the higher end-to-end throughput at context 64, the two were close at 128 (204.8k against 199.9k tokens/s), and S4D was higher from 256 upwards. At context 1024, S4D trained at about 272.6k tokens/s against 67.1k for attention, and peak allocated CUDA memory was 1,349 MiB against 6,455 MiB.
Two hardware datasets
The original sweep ran on Apple MPS and is kept as its own dataset: there, at context 1024, S4D trained at about 172.9k tokens/s against 38.2k for attention, and at context 64 it was slower (about 105k against 138k). The Tesla T4 sweep repeats the same settings on CUDA. Final validation losses on the two backends differ by at most 0.0088. Each throughput comparison is within one device; rerunning MPS configurations changed throughput by about 25% between runs, so no CUDA-versus-MPS speed-up is inferred.
Limitations
One dataset, one seed and one run per configuration, with no error bars. Neither model is trained to convergence, there was no per-architecture hyperparameter tuning, and the parameter counts differ by about 12%. The attention mixer is a plain implementation without a fused kernel, so the throughput and memory ratios are specific to this implementation and device. The result does not show that S4D is better than Transformers in general.