Adebanji Adelowo
← Projects

State-space model vs. Transformer

Sequence modelling · State-space models

My contribution

I implemented a diagonal state-space layer (S4D) from scratch and compared it with a causal Transformer in a shared backbone across context lengths 64 to 1024, on Apple MPS and on a Tesla T4.

Verified resultsResearch implementation · one seed, fixed training budget

S4D reached lower validation loss at every tested context length. On a Tesla T4 at context 1024 it trained at about 272k tokens/s against 67k for attention, with 1,349 against 6,455 MiB peak allocated memory (one seed, fixed training budget).

State-space modelsS4DCausal attentionPyTorch

Problem

Deep state-space models share their core recurrence, x′(t) = Ax(t) + Bu(t), with classical control theory. This study asks how a diagonal state-space layer compares with causal self-attention in accuracy, throughput and memory as the context length grows, when everything else in the model and the training is held the same.

Approach

A from-scratch PyTorch implementation of S4D (Gu, Gupta, Goel and Ré, 2022) and a from-scratch causal Transformer share one backbone with a swappable mixer: embeddings, MLP, normalisation and weight tying are identical. Both are trained on character-level Tiny Shakespeare with the same optimiser, schedule, batch size, seed, depth and width, at context lengths 64 to 1024, for 1,200 steps (400 at context 1024).

Verification

The training-time convolutional kernel and the step-by-step recurrence of the S4D layer are tested to compute the same function, at initialisation and after an optimiser step, with a measured maximum difference of about 9 × 10⁻⁸ to 2.4 × 10⁻⁶ in float32.

Result

S4D reached lower final validation loss than attention at every tested context length, on both backends. On the Tesla T4, attention had the higher end-to-end throughput at context 64, the two were close at 128 (204.8k against 199.9k tokens/s), and S4D was higher from 256 upwards. At context 1024, S4D trained at about 272.6k tokens/s against 67.1k for attention, and peak allocated CUDA memory was 1,349 MiB against 6,455 MiB.

272.6k vs. 67.1kTraining tokens/s, S4D vs. attention, context 1024, same Tesla T4
1,349 vs. 6,455 MiBPeak allocated CUDA memory, context 1024
≤ 0.0088Largest CUDA vs. MPS difference in final validation loss
Final validation loss, training throughput and peak CUDA memory against context length for attention and S4D mixers on one Tesla T4
Tesla T4 sweep, seed 42, one run per configuration: final validation loss, training throughput (end-to-end and step-only) and peak CUDA memory against context length. Context 1024 uses 400 steps, the others 1,200. Select figure to enlarge.

Two hardware datasets

The original sweep ran on Apple MPS and is kept as its own dataset: there, at context 1024, S4D trained at about 172.9k tokens/s against 38.2k for attention, and at context 64 it was slower (about 105k against 138k). The Tesla T4 sweep repeats the same settings on CUDA. Final validation losses on the two backends differ by at most 0.0088. Each throughput comparison is within one device; rerunning MPS configurations changed throughput by about 25% between runs, so no CUDA-versus-MPS speed-up is inferred.

Limitations

One dataset, one seed and one run per configuration, with no error bars. Neither model is trained to convergence, there was no per-architecture hyperparameter tuning, and the parameter counts differ by about 12%. The attention mixer is a plain implementation without a fused kernel, so the throughput and memory ratios are specific to this implementation and device. The result does not show that S4D is better than Transformers in general.