Latent Oscillator Thompson Sampling accepted at ICAIF
I am very happy to share that my paper Latent Oscillator Thompson Sampling for Recurrent Nonstationary Bandits has been accepted at ICAIF 2026, the ACM International Conference on AI in Finance, in Milan 🎉
This post walks through the problem, the idea, and the part I find most interesting: a simple test that tells you in advance whether modelling cycles will pay off at all.
Table of contents
- Why does this problem exist? 🤔
- The idea: give every arm an oscillator 🌀
- What can be proved 📐
- What I found 📊
- When does a periodic prior help? 🔍
- Conclusion 🎯
Why does this problem exist? 🤔
A multi-armed bandit is the simplest model of learning by doing. At every round you pick one option (an “arm”), you see the reward of that option only, and you try to collect as much reward as possible. You never see what the other arms would have paid.
Most bandit methods assume that each arm’s reward is either stationary or slowly drifting. They track it with sample means, discounting or sliding windows. That is often right, but it is a blunt tool when an arm’s reward rises and falls in a recurring pattern.
This happens in many places:
- Finance. A strategy that prioritises one security at a time by its trading activity sees a security look weak during a quiet spell and strong when activity returns. The instrument did not change; its liquidity state did.
- Energy. Power demand rises and falls through the day, and differently in each zone. Allocating a limited forecasting or bidding budget across zones means following those cycles as they shift.
- Transport. Ridership peaks at different hours at different stations, and the peaks move with seasons and events.
- Advertising, traffic allocation and staffing. A placement, a server or a team performs best at certain times of day or week, and the pattern drifts.
In all of them, averaging across the cycle throws the pattern away, and hard-coding a fixed weekday rule breaks as soon as the pattern shifts.
The regime I study sits in between: the reward is not just drift, but it is not a rigid calendar effect either. The cycles have a phase, an amplitude and a period, and all three drift over time. Because the best arm keeps changing, the right yardstick is dynamic regret: at every round, how much worse was the chosen arm than the best arm at that round?
Here is the benchmark from the paper, running live in your browser: five arms, 720 rounds, and three policies that each see only the arm they pull.
The idea: give every arm an oscillator 🌀
Latent Oscillator Thompson Sampling (LOTS) builds the recurrence into what it believes about each arm. Every arm carries a hidden state:
- a drifting baseline, for slow changes in level;
- a 2D oscillator whose state rotates at an angular frequency, so its length is the amplitude and its angle is the phase;
- that frequency itself, which drifts within a band, so the period can change online.
The expected reward is the baseline plus the oscillator’s projection on one axis. Because frequencies drift, the model is nonlinear, so LOTS tracks it with a lightweight particle filter: 128 guesses of the hidden state per arm, reweighted every time the arm is pulled.
Then it acts by Thompson sampling: it draws one plausible state for each arm, predicts its next reward, and pulls the arm with the best draw. Only the pulled arm gets an observation. Every other arm keeps moving forward through its predicted dynamics alone, so LOTS knows roughly where in its cycle an arm is even when it has not looked at it for a while.
LOTS also adapts online. When the observed rewards surprise it, LOTS temporarily loosens its dynamics and lowers its damping, so it can re-lock when a cycle shifts.
What can be proved 📐
I want to be precise here, because the theory covers less than the method does.
- A general reduction. For the full model, dynamic regret is bounded by the one-step error in predicting each arm’s latent score. Good prediction is enough for low regret.
- An explicit bound for a simplified core. With fixed frequencies, a stable baseline and exact Kalman inference instead of particles, I prove an explicit dynamic-regret bound that depends on how long each arm has gone unobserved, on damping, and on the number of oscillators.
Both are tracking guarantees, not sublinear-regret results, and neither separates Thompson sampling from greedy play. Frequency drift and the particle approximation remain outside the theorem, so the experiments are the main evidence for the method as implemented.
What I found 📊
The benchmark has six synthetic families: stable cycles, drifting phase, drifting period, cycles that appear and disappear, cycles that collapse halfway, and a control with no recurrence at all. Hyperparameters were tuned once on 12 development seeds, frozen, and evaluated on 80 held-out seeds, with every method seeing the same rewards.
Final cumulative dynamic regret, lower is better; the thin line is the 95% interval and ★ marks the lowest.
Final cumulative dynamic regret over 80 held-out seeds, from the paper. HarmonicLinTS uses fixed Fourier features, SeasonalTS a fixed set of phase states, D-UCB and SW-UCB discount or window old rewards, and RW-KalmanTS tracks each arm as a random walk. LOTS-Gate adds a random-walk fallback to LOTS.
LOTS has the lowest regret in five of the six families, and every win holds up under paired tests with Holm correction across families. The margin over the strongest competitor, D-UCB, is largest on stable and drifting-phase cycles (about 12 points). On drifting periods, both fixed-dictionary methods are structurally disadvantaged, and a generic random-walk posterior is not enough either: what helps is tracking the frequency itself.
The one loss is the control with no recurrence, where LOTS is far behind the drift trackers. That is the price of a prior that expects cycles. LOTS-Gate, which adds a random-walk fallback, repairs that case (81.5 → 50.1) but loses the advantage everywhere else, so I report it as an ablation and leave robust model selection open.
Two more checks I like. Giving the fixed-Fourier baseline the true period of each arm as side information makes it beat LOTS when the period is near-constant, but LOTS still wins wherever the frequency, presence or regime changes online. And doubling arms and rounds (10 arms, 1,440 rounds) keeps LOTS 23–30% below the strongest baseline in every recurring family.
Real data. I also replayed four real panels as bandits, with no tuning on any of them: MTA subway hourly ridership by station, hourly power demand by client, and daily and 15-minute trading volume across ten liquid ETFs. Each panel holds every arm’s true reward, but each method only sees the arm it picks. These are partial-feedback replays, not trading backtests.
Differential seasonal R² 0.61, cycle rigidity 0.94.
Final cumulative dynamic regret, lower is better; the thin line is the 95% interval and ★ marks the lowest.
Real-data replays, 80 episodes each, from the paper. The last row of each panel is an ablation, not a competitor: LOTS with its oscillators removed, keeping the same filter, drifting baseline and adaptation.
LOTS is best on power demand and on daily ETF volume, and runner-up on the other two. On MTA, the fixed-Fourier baseline wins because every station’s period is exactly 24 hours, right on its grid, so it pays nothing to estimate it. On intraday volume, D-UCB wins. The ablation row is honest about the daily ETF win too: removing the oscillators makes LOTS better there, so that win comes from its slowly drifting baseline, not from the cycles.
When does a periodic prior help? 🔍
The intraday result puzzled me at first. Intraday volume has a famous U-shaped daily pattern, and in this panel that cycle explains 48% of the raw variance. So why does modelling cycles not help?
Because a bandit ranks arms. Write each arm’s reward as a part common to all arms plus a part specific to it. At every round the common part shifts every arm by the same amount, so it never changes which arm is best, and it cancels in the regret. The intraday cycle is almost identical across tickers (mean pairwise correlation 0.95): it is a market-wide factor, and no prior can exploit it.
What should matter is the cycle left after comparing arms: the seasonal R² of the panel after subtracting, at each round, the average over arms. I call it the differential seasonal R². Slide the shared share of the cycle below to see what it measures.
Two experiments back this up, after the fact and then by design:
- The real panels. The differential R² is large where a periodic method wins (0.68 on MTA, 0.61 on power demand) and collapses where drift trackers do best (0.09 intraday, 0.00 daily). Removing LOTS’s oscillators costs 10.8 on MTA and 3.6 on power demand, changes nothing on intraday volume, and helps by 4.5 on daily volume: the same order. Raw seasonal R² would have predicted a clear gain on intraday volume, and there is none.
- A controlled sweep. I formulated the statistic after seeing the four panels, so I tested it by intervention. In a synthetic family where only the shared share of the cycle varies, the raw seasonal share stays at about 0.95 throughout, while LOTS’s advantage decays monotonically from +13.9 to +1.2.
The statistic predicts how much a periodic prior can help, not who wins. And low differential R² looks pervasive in intraday markets: on six further public constructions (ETFs by liquidity, regional ETFs, realised-range volatility, hourly crypto volume, hourly FX realised range) it stays below 0.08, even where raw profiles peak at different hours. The model fits cross-sectional problems where cycles differ by arm, such as power demand across zones, or execution policies that peak at different points of the trading session.
Conclusion 🎯
LOTS treats a recurring pattern as something to infer, not to assume: every arm carries an oscillator whose phase, amplitude and period are tracked online. When cycles drift, appear, disappear or fade, that pays off, with the lowest regret in every recurring synthetic setting, at double the scale, on a harder family never used for tuning, and on real power-demand data.
The result I find most useful goes beyond the method. A cycle helps a bandit only if it differs across arms, because whatever all arms share cancels when you compare them. The differential seasonal R² measures exactly that, offline and before deployment, and it predicts how much a periodic model can gain. Where it is high, as with power demand across zones or ridership across stations, modelling the cycle is worth it. Where it is close to zero, as with intraday market volume, a simple drift tracker is enough.
Next steps include a theory that covers drifting frequencies, and hybrids that switch smoothly between recurring and non-recurring regimes.