[home] [coding projects] [research projects] [research interests]
A software manual for the encoder-only / decoder-only / encoder-decoder forecasting experiments.
This codebase is a controlled architecture experiment for multivariate time-series forecasting. The research question is deliberately narrow: if the training procedure, datasets, embedding, forecast horizon, and model family are held as fixed as possible, what changes when the encoder-decoder structure itself is removed?
The implementation supports the Transformer, Informer, Autoformer, and FEDformer families. A single --stack argument selects encoder, decoder, or both. The associated study reports that decoder-only variants frequently matched or outperformed the full architectures, which is why the code is organized around making this ablation easy to repeat rather than around introducing an entirely new forecasting model.
| 1. Architecture 2. The stack switch 3. Invariant memory experiment |
4. Running an ablation 5. Outputs and comparisons 6. Where the code lives |
The diagrams below are the three cases implemented by the same forecasting code. In the full case, the encoder produces a representation that the decoder can cross-attend to. In the decoder-only case that cross-attention source is absent. In the encoder-only case the encoded sequence is projected directly to the forecasting channels.

![]() | ![]() |
The point of the ablation is that the stack choice sits inside the same experiment machinery. The data loaders, forecasting losses, prediction lengths, optimization loop, and most model hyperparameters do not need separate code paths for each architecture. The architecture changes; the rest of the experiment can remain comparable.
The most important source excerpt is almost comically small. The model records the stack mode and uses it to decide which pieces should be constructed.
self.stack = getattr(configs, "stack", "both")
assert self.stack in ["encoder", "decoder", "both"]
self.use_encoder = self.stack in ["encoder", "both"]
self.use_decoder = self.stack in ["decoder", "both"]
That decision then changes the forward pass. The full model gives the decoder an encoder output; the decoder-only model passes no cross-attention source at all.
if self.stack == "both":
dec_out = self.decoder(
dec_out,
enc_out,
x_mask=dec_self_mask,
cross_mask=dec_enc_mask
)
else:
dec_out = self.decoder(
dec_out,
None,
x_mask=dec_self_mask,
cross_mask=None
)
For an encoder-only run, the encoder output is sent through a linear projection and the last pred_len positions are returned. This means the ablation is not implemented as three unrelated programs; it is one model family whose information path is changed by one structural flag.
There is also an optional experiment called Invariant Memory. This is separate from the basic encoder/decoder ablation. The encoder hidden states are projected from dimension \(d_{\mathrm{model}}\) to a smaller dimension \(k\), and a least-squares linear evolution operator \(A\) is fitted so that consecutive projected states approximately satisfy
Instead of storing the whole fitted matrix as a token, the code summarizes it using traces of powers,
and maps those invariant summaries back into the model dimension.
Z0 = z0.transpose(1, 2)
Z1 = z1.transpose(1, 2)
ZZt = Z0 @ Z0.transpose(-1, -2)
reg = 1e-4 * torch.eye(self.k, device=h.device)
A = (Z1 @ Z0.transpose(-1, -2)) @ torch.linalg.inv(ZZt + reg)
traces = []
Ap = A
for _ in range(self.m):
traces.append(torch.diagonal(Ap, dim1=-2, dim2=-1).sum(-1))
Ap = Ap @ A
M = torch.stack(traces, dim=-1)
token = self.to_token(M)
When memory is enabled, that token is appended to the encoded sequence before the decoder sees it. The idea is therefore not “remember every hidden state”; it is “fit a local linear dynamical summary, compress it to basis-invariant trace statistics, and expose that summary as one additional token.”
The repository already contains dataset scripts for ETTh1, ETTh2, ETTm1, ETTm2, ILI, and Weather. The convenience runner executes the dataset scripts in sequence:
bash scripts/run_all.sh
For a single experiment, the key argument is --stack. The test script uses the same forecasting input length and evaluates several prediction horizons. For example, the core of a single decoder-only run is conceptually:
python main.py --seed 2025 --data ETTh1 --model Transformer --stack decoder --root_path ./data/datasets/ --data_path ETTh1.csv --seq_len 720 --label_len 96 --pred_len 96 --learning_rate 1e-4 --d_model 512 --n_heads 8 --e_layers 2 --d_layers 1 --d_ff 2048 --dropout 0.05 --itr 1 --train_epochs 10
Changing only --stack decoder to --stack encoder or --stack both gives the architectural comparison. The scripts repeat this idea over multiple datasets and forecast horizons.
Training is performed with the shared experiment class in exp/exp_main.py. Checkpoints are stored under a setting name containing the model, stack, dataset, sequence length, forecast length, and major architectural hyperparameters. Testing uses the same setting identifier, so the experiment directory itself records which ablation produced a result.
The important scientific comparison is not merely final error. The codebase is useful because it makes it possible to ask whether a simpler information path reaches the same forecasting quality under the same data and training setup.
| Path | Purpose |
|---|---|
| models/Transformer.py | Encoder / decoder / both structural switch and optional memory injection. |
| models/Informer.py | Informer version of the same ablation idea. |
| models/Autoformer.py | Autoformer version. |
| models/FEDformer.py | FEDformer version. |
| layers/InvariantMemory.py | Least-squares evolution operator and trace-invariant memory token. |
| exp/exp_main.py | Shared train / validation / test loop. |
| scripts/ | Dataset-specific ablation runs. |
Rethinking the Encoder-Decoder Structure for Transformer-Based Time Series Forecasting
These pages are selective technical manuals: enough source to expose the mechanism, not a mirror of the entire repository.
Last updated: September 14, 2026.