[home]   [coding projects]   [research projects]   [research interests]


Spectral Attention

Notes on spectral relative bias, FFT token mixing, and persistent oscillatory state.

This project explores two related ideas. The first keeps ordinary causal self-attention but gives it a learnable relative-position bias initialized from a spectral basis. The second removes quadratic attention from the mixer and carries a bank of damped oscillatory modes from chunk to chunk.

The code is experimental on purpose: it is a place to ask whether spectral structure can supply useful long-range organization without turning the whole model into an opaque new architecture.

Contents

1. Spectral attention bias
2. Where the bias enters
3. Streaming mixer
4. Running the experiments
5. What is compared
6. Source map

1. Spectral relative-position bias

The relative-position initialization begins with ordinates \(\gamma_m\) of nontrivial zeta zeros. The code normalizes them by the first ordinate and evaluates a cosine basis against logarithmic distance:

\[ b_0(d)=\frac1M\sum_{m=1}^M \cos\!\left(\frac{\gamma_m}{\gamma_1}\log(1+d)\right). \]
How the spectral distance profile enters causal self-attention.
How the spectral distance profile enters causal self-attention.
model/spectral_attention.py
gammas = torch.tensor(
    ZETA_GAMMAS[:n_modes],
    dtype=torch.float32
)
gammas = gammas / gammas[0]

dist = torch.arange(
    self.max_seq_len,
    dtype=torch.float32
)
log_dist = torch.log1p(dist)

basis = torch.cos(
    gammas.view(n_modes, 1)
    * log_dist.view(1, self.max_seq_len)
)

init = basis.mean(dim=0)
init = init - init.mean()
init = init / init.std().clamp_min(1e-6)
init = 0.02 * init

This is only an initialization. The result is copied into a learnable relative-bias table for every attention head, so training is free to move away from the initial spectral profile.


2. Where the bias actually enters attention

The model still computes ordinary query-key compatibility. A learned token salience term and the spectral relative bias are added before the causal mask and softmax:

model/spectral_attention.py
scale = math.sqrt(self.head_dim) * torch.exp(
    self.temperature
).view(1, self.n_heads, 1, 1)

att = (q @ k.transpose(-2, -1)) / scale

salience = torch.tanh(
    self.salience(x)
).transpose(1, 2)

bias, mask = self.spectral_bias(seq_len)

att = att + salience + bias
att = att.masked_fill(
    ~mask.view(1, 1, seq_len, seq_len),
    float("-inf")
)

att = torch.softmax(att, dim=-1)
out = att @ v

This makes the experiment easy to interpret. The content-based term \(QK^ op\) is still present; the spectral object is a structured additive preference over relative distances.


3. Streaming spectral mixer

The second branch is more radical. Instead of materializing a full \(T\times T\) attention matrix, a chunk is mixed in three ways: a spectral convolution within the chunk, a carried complex state from earlier chunks, and a short depthwise local convolution. A learned gate controls the combined output.

Within-chunk spectral convolution and persistent real/imaginary mode state.
Within-chunk spectral convolution and persistent real/imaginary mode state.

The within-chunk path forms a learned damped-cosine kernel and applies it by FFT convolution:

model/streaming_spectral_mixer.py
kernel = self._kernel(
    seq_len, v.device, v.dtype
).transpose(0, 1).unsqueeze(0)

v_t = v.transpose(1, 2)

v_f = torch.fft.rfft(
    v_t, n=n_fft, dim=-1
)
k_f = torch.fft.rfft(
    kernel, n=n_fft, dim=-1
)

y = torch.fft.irfft(
    v_f * k_f,
    n=n_fft,
    dim=-1
)[..., :seq_len]

Across chunks, the state is represented by real and imaginary mode amplitudes. Each mode has a learned decay and frequency. Old state is rotated and damped forward by the chunk length, while the current chunk contributes new real/imaginary components. The result behaves like a bank of persistent damped oscillators rather than a cache of old token vectors.


4. Running the experiments

The standard attention comparison runs the Transformer baseline and spectral engine:

shell
bash scripts/run_all.sh

There is also a small FineWeb smoke test for the streaming model. It uses stream length 11,520, chunk length 720, two blocks, and a 10-million-token training target:

shell
bash scripts/streaming_spectral_fineweb_smoke.sh

The larger script uses the same architecture with a one-billion-token training target. The scripts are intentionally separate because the streaming experiment is a different computational regime from the ordinary sequence-length sweep.


5. What is compared

The ordinary training harness reports train loss, validation loss, validation perplexity, validation token accuracy, and tokens per second. The spectral script uses much longer configured sequence lengths than the Transformer baseline script, while the streaming implementation also contains an explicit forward-FLOP estimator so compute scaling can be inspected rather than inferred only from wall-clock time.


6. Source map

PathPurpose
model/spectral_attention.pyLearnable spectral relative-position bias added to causal attention.
model/spectral_mixer.pyNon-streaming spectral mixer model.
model/streaming_spectral_mixer.pyFFT within-chunk mixing plus persistent complex mode state.
train_streaming_spectral_fineweb.pyStreaming FineWeb training loop.
scripts/spectral.shSequence-length sweep for the spectral engine.
scripts/streaming_spectral_fineweb_*.shSmoke and larger streaming runs.

These pages are selective technical manuals: enough source to expose the mechanism, not a mirror of the entire repository.
Last updated: September 14, 2026.