Abstract
The problem
Data that unfolds over time, like heartbeats, factory sensor readings, or the motion signals from a phone in your pocket, has patterns at several scales at once. A walking signal has tiny wiggles from each footstep, repeating strides lasting a second or two, and an overall character that says “this person is climbing stairs.” Most methods that learn from this kind of data without human labels focus on the big picture, asking whether two whole recordings look alike. That leaves the middle scale, the recurring motifs where much of the meaningful structure lives, to be learned only by accident or squeezed away.
Our approach
We teach the model at all three scales separately. To make this possible, we propose the Progressive Memory Transformer (PMT), a transformer that keeps a small, writable memory as it reads a signal window by window. The memory gives the model a dedicated place to store mid-range patterns, and the training can check and shape what goes into it directly.

Figure 1: PMT reads a signal window by window and learns representations at three scales (low, mid, and high). Each level carries a memory state forward in time, and the final states at the mid and high levels form the readout.
Results
We tested PMT on seven standard classification benchmarks, where the model had to recognize categories such as human activities, engine faults, or electrical devices after seeing labels for only 1–5% of the examples. PMT had the highest average accuracy and won 11 of 14 settings. On the human activity recognition dataset, PMT trained with 1% of the labels beat competing methods that got 5%.
We also designed a new test of memory. We hid a short synthetic blip early in a signal and asked whether the model’s summary at the very end still “remembered” it, even though the model had never been trained to look for such blips. PMT detected the hidden blip almost perfectly (0.98 on a scale where 0.5 is guessing). It clearly beat other memory-based designs, and it far outperformed a popular method that pools information away, which scored 0.65.
For forecasting electricity use and electrical transformer data, PMT made smaller errors than a comparable method, especially when predicting far into the future. Longer forecasts are where a memory that carries context across windows should help most.

Figure 2: How the PMT sees activities. The top-left panel shows a smartphone motion recording stitched together from three separate recordings of a person walking downstairs, walking on flat ground, and walking upstairs. To a human eye, these three activities produce very similar-looking signals. Each square below is a heat map comparing how the model describes each moment of the stitched signal, with brighter colors meaning more similar descriptions. The middle row compares the model’s moment-to-moment features, and the bottom row compares its memory within the model. The dashed lines mark where the pieces were joined. In the highlighted left column, where the stitched signal is compared with itself, the similarity stays inside each piece and does not spill across the joins. Both the moment-to-moment features and the memory keep the three kinds of walking apart, even though the model was never told which activity was which. The other columns compare the stitched signal with untouched recordings of each of the three activities.
Citation
@inproceedings{stangeland2026,
title = {Progressive Memory Transformer: Memory-Aware Attention for Time-Series},
author = {Stangeland, Tord Sture and Köhler, Andreas and Mæland, Steffen and Ramírez Rivera, Adín},
booktitle = {Advances in Neural Information Processing Systems},
volume = {39},
year = {2026}
}