Better Data, Better Models: What 184 Experiments Changed for Cascade

To build the best decoder for Cascade, we needed to optimize across streaming, covariates, context and the training distribution. Thanks to compute credits from @ Targoncompute , we were able to conduct a series of experiments across each of these areas - whose results help us understand how to complement the data generated by the subnet with a decoder designed to maximise its performance. 

Cascade is running two searches in parallel. Bringing the results of those two together is what makes us unique. 

On the subnet, miners compete to find better synthetic pretraining data. They submit programs that generate time-series data, not trained models. We hold the model and training setup fixed, train a fresh model on each submission, and ask which data produces the better forecaster.

Alongside that, we run a separate model-side research programme. Starting from Toto 2.0, we are testing what should change in the decoder - the part of the model that turns historical inputs into future forecasts - to design a forecasting model we can eventually deploy on live, real-world data.

Keeping those searches separate is deliberate. If the model and the data changed at the same time, we would not know what caused an improvement. The subnet isolates the data question; the research programme isolates the model question.

Across the two searches, we wanted five practical questions answered:

  • Can we make the decoder stream continuously without its memory growing forever?
  • Can it make better use of long histories, rather than simply extending dense attention?
  • Can we get more learning signal from the same training sequence?
  • Can it use real-world predictive context - covariates such as weather, promotions or related series - while understanding what each input is for?
  • Does Cascade ultimately want one winning synthetic generator, or a diverse mixture of them?

To find out, we trained 184 models across 24 experiment groups over roughly nine days. The point was not to collect benchmark numbers. It was to decide what should actually change in the decoder, what should stay, which ideas were not worth carrying forward, and what the subnet's generator population was producing.

The headline results

Before getting into the detail, the experiments pointed to five practical conclusions:

  • Generator mixtures may change what Cascade is searching for. A mixture of four strong, diverse generators improved forecasting by about 3.8%, compared with 2.25% for the best individual generator, and beat it in all nine paired comparisons. The competition may be building a valuable portfolio, rather than searching for one ultimate winner.
  • Windowed attention offers a credible path to streaming. At half of Toto’s normal context, forecasting performance was effectively unchanged.
  • Denser supervision can extract more value from the same data. It improved forecasting by roughly 2.1% to 2.6%, without adding parameters or inference cost.
  • Role-aware covariates help the decoder use real-world predictive context. The model learned to benefit from a supporting signal when it understood that signal’s role.
  • Longer context will require more than a positional trick. More sophisticated positional approaches produced no meaningful improvement, while attention sinks made performance worse.

The rest of the article explains how we reached those conclusions, beginning with the model-side experiments before returning to what the generator-mixture result means for the subnet.

Starting from Toto 2.0

Cascade currently uses Datadog's Toto 2.0 as its fixed backbone. Toto 2.0 is a decoder-only time-series transformer. In practical terms, decoder-only means one causal transformer stack turns past observations into future forecasts, rather than using a separate encoder to first represent the history and a decoder to read from it.

Toto groups observations into patches, then alternates attention over time and across variables. Its standard context is 4,096 time steps and it produces probabilistic forecasts rather than a single point estimate.

It is also a natural starting point for Cascade because its own pretraining recipe already leans heavily on synthetic data. Toto 2.0 uses a 57.5% synthetic pretraining mix, with no public time-series datasets in pretraining.

That matters to our thesis. Synthetic data is not a small side ingredient in modern forecasting models. It can be a major part of the training distribution. Cascade asks what happens when that distribution becomes something an open network can continuously compete to improve.

At the same time, benchmark strength is not the same as being the decoder we want to deploy. That is what the model-side experiments were for.

1. Can windowed attention make the decoder stream?

Toto's time attention is full causal attention. In plain English, every point can look back at everything that came before it.

That is reasonable for a fixed benchmark window. It is less attractive for a model watching a live stream, because the history it may need to keep around can grow as the stream keeps running.

We tested windowed attention. Instead of letting the model look back over its entire available history, windowed attention restricts it to the most recent N time steps.

At 512 steps, forecasting got noticeably worse. At 1,024 steps, most of the performance returned. At 2,048 steps, half of Toto's normal 4,096-step context, performance was effectively unchanged within the resolution of our experiments.

That gives us something much more useful than a small benchmark gain: a credible path to bounded-memory streaming. A model could keep receiving new data without needing to preserve an ever-growing attention history.

Attention sinks: a useful failure

We also tested attention sinks, a technique used in streaming LLMs. The idea is to permanently keep a few early tokens because language models can learn to put disproportionate attention on them even after the rest of the old context is discarded.

It transferred badly. At the smallest window, adding attention sinks made forecasting dramatically worse than using the bounded window on its own.

That result matters because it stops us importing LLM techniques simply because they solve a problem with the same name. A time-series model needs its own evidence.

What this changed for us: Windowed attention is now part of the direction we want to carry forward. Attention sinks are not. The design goal is not just lower memory use; it is a forecaster that can sit on a live stream and keep operating predictably over time.

2. Can we extend context just by changing positional encoding?

We also tested the opposite direction. If we can safely shorten attention for streaming, can we simply stretch the decoder beyond its trained 4,096-step context when more history is available?

This becomes partly a positional-encoding problem. Positional encoding is how a transformer knows where observations sit in a sequence and how far apart they are in time.

Removing positional information made medium and long-horizon forecasting roughly one third worse. So position is clearly important.

We then tried several more sophisticated positional approaches, including alternative positional geometries and xPos, which changes how attention behaves as observations get further apart. None produced a meaningful improvement in our experiments.

That is a null result, but it changed the research direction. We started with the reasonable hypothesis that longer context might mainly require a better positional trick. The experiments did not support that.

We are now more interested in how the decoder manages history: what needs to stay at full resolution, what can be compressed or summarized, and what distant information should be brought back only when it is useful.

If that works, a long-lived forecasting model could keep access to important seasonal cycles, earlier regime changes or past incidents without paying full attention cost for every old observation. That could matter for continuous systems where useful history stretches far beyond one fixed input window.

So while long-context handling remains open research for us, these experiments have narrowed the search space. We now have a clearer idea of what not to spend compute on next.

What this changed for us: Long-context handling remains open research. What we learned is that blindly extending dense attention is probably not the most useful next step. The streaming result and the positional nulls push us toward smarter context management rather than simply a bigger window.

3. Can denser supervision learn more from the same sequence?

One of the best results came from one of the simplest changes.

During training, a forecasting model does not necessarily need to be supervised at every possible prediction point. We tested denser supervision: asking the model to make, and learn from, more predictions inside each training sequence.

The model did not get larger. We added no new parameters, no extra inference cost and no additional training-point budget.

Forecasting improved by roughly 2.1% to 2.6%, depending on the horizon. It was one of the strongest gains in the whole programme.

That is useful beyond the number itself. It says there was still learning signal sitting unused inside the sequences we were already paying to train on.

What this changed for us: Denser supervision is a change we expect to keep. Before spending more compute on a larger decoder, we should first make sure the decoder extracts as much learning signal as possible from each sequence it already sees.

4. Can role-aware covariates make forecasts use real-world context?

A lot of forecasting benchmarks simplify the task to one question: here is a series, what happens next? Real forecasting rarely works like that.

Retail demand can depend on a promotion. Electricity demand can depend on weather. Infrastructure load can depend on a product launch or incident. Another time series can itself move before the target and provide an early signal.

These extra signals are usually called covariates.

Toto 2.0 is multivariate, so it can process several series together. But the decoder is effectively role-blind: it does not inherently know that one channel is the target while another channel is there because it may help predict that target.

We built role-aware covariates into the decoder and ran a simple diagnostic. We gave the model a signal that moved ahead of the target by a known number of steps. If the mechanism worked, the model should learn to use that lead.

It did. When the supporting signal aligned closely with the future target, forecast error fell sharply. As the lead became less useful, the benefit decayed. A control model given the same underlying signal without the role-aware mechanism got essentially no benefit.

That makes this more than an interface feature. The mechanism changed what the decoder was capable of learning from the same input information.

What this changed for us: Role-aware covariates move the decoder toward the kind of forecasts people actually need. The deployment goal is a model that can use known future information and supporting signals explicitly, rather than treating every channel as if it means the same thing.

What these experiments add up to

The important outcome is not five isolated experiment results. They are starting to define the decoder we want Cascade to build on.

The direction today is:

  • Toto 2.0 as the strong base architecture;
  • windowed time attention for bounded-memory streaming;
  • role-aware covariates for real-world predictive context;
  • denser supervision to extract more learning signal from every sequence; and
  • new work on smarter context management rather than assuming longer dense attention will solve long-history forecasting.

We would not describe this as a finished new architecture yet. Some pieces have clean experimental support and some are still active research. But it is now a much more concrete model direction than simply saying we want to improve Toto 2.0.

The practical target is a Toto-derived decoder that can run continuously, use the information around a forecast, and eventually make better use of long histories without paying for every old observation at full attention cost.

Data: Cascade’s other half

While we research the model independently, the subnet continues to search over the training distribution.

Keeping those two search problems separate is deliberate. If miners could change the model, optimiser, architecture and data at the same time, a better result would be difficult to interpret. We would not know whether the improvement came from the data or the model.

So Phase 1 keeps the model byte-identical and asks one clean question: which generator produces training data that results in a better forecaster?

One of the most important experiments in this programme was testing what happens when we take several strong generators produced by that competition and combine them.

5. Is the best training distribution one generator or a mixture?

A competition naturally makes you think the end product is a single best generator. Find the winner, use its data, and move on.

Our results suggest that may be the wrong mental model.

Diversity: the key design choice

These were not simply four high-scoring generators. We selected four strong generators that were meaningfully different from one another. Every model received the same total number of training points, so a mixture could not win simply by seeing more data.

The best single generator improved average forecasting performance by about 2.25%. A uniform mixture of four diverse generators improved it by about 3.8%, and the four-generator mix beat the best single generator in all nine paired comparisons across seeds and forecast horizons.

We also changed the mixture weights. Giving the four generators equal weight versus making one generator dominant produced essentially the same result.

That points to a different interpretation of what the subnet may be producing. The value may not be one perfect synthetic-data algorithm. It may be a portfolio of different useful generating processes that cover more of the training distribution together than any one can cover alone.

Which means the best portfolio may not be the same as the generators ranked highest one by one. A generator can be useful because it contributes regimes or behaviours the others do not cover. That gives us a reason to preserve useful diversity rather than let the population collapse onto one narrow family of synthetic data.

What this changed for us: This changes how we think about selection pressure on the subnet. We still need generators to be individually strong, but useful diversity may itself be valuable. A field of miners that discovers different ways to win can be more useful than a field that converges on copies of the same recipe.

What about real data?

We also tested public real-world training data. Adding either 20% or 50% real data to the synthetic mixture did not produce a change large enough for us to confidently separate from noise. Training on the public real-world data alone performed about 4.45% worse than the control in this experiment.

We do not interpret that as "real data is bad." The comparison contains important confounds. Cascade generators have already survived competitive selection. The synthetic path also used augmentation that the real-data path did not, and generators can keep producing fresh series while a real corpus is finite.

We have follow-up work aimed at separating those effects. The supported conclusion today is narrower: competition-surviving synthetic generators produced very useful training data, and several diverse generators together were stronger than the best one alone.

Two search problems, deliberately distinguished

This is the clearest way we now think about Cascade.

Today, we deliberately distinguish between these two search problems so the results remain interpretable. Later, the goal is to bring them back together: a more capable decoder trained on a better distribution discovered through competition.

The data

  • synthetic generators
  • competition
  • strong, diverse generators
  • generator mixtures
  • better training distribution

The model

  • Toto 2.0
  • streaming attention
  • role-aware covariates
  • denser supervision
  • smarter context handling
  • larger model scales

Why the improvements need to compound

Cascade already has a mechanism for compounding progress on the data side. When a king survives enough rounds, its best checkpoint can be promoted and become the starting point for the next generation of competition.

That means a successful generation can become the foundation for the next one instead of the network repeatedly returning to the same starting point.

Promotion acts like a ratchet: once a generation earns a better starting point, the next competition begins from there rather than slipping back to the original model.

The intended loop is simple:

  • better generators -> better training data
  • better training data -> better model
  • better model -> promoted checkpoint
  • promoted checkpoint -> a new round of competition on top of it

Eventually we want the model-side improvements from this research programme to join that loop too.

What 184 experiments changed for us

We started this programme expecting some sophisticated architectural ideas to be the biggest wins. Many were not.

Attention sinks, borrowed from LLMs, made the model worse. More complicated positional approaches did not meaningfully improve the result. Simply making things more complex was often less useful than expected.

The experiments that did change our direction were more concrete:

  • windowed attention gave us a credible path to streaming;
  • denser supervision showed we can get more learning from the same sequence;
  • role-aware covariates gave the decoder access to information real forecasts actually depend on; and
  • diverse competition-surviving generators worked better together than the best generator did alone.

The important part is not that every hypothesis worked. It is that the failures and the wins now tell us more clearly what to build next.

Long-context handling remains open research for us. We have not yet shown that every model-side improvement survives all the way up the size ladder. And we are still separating how much of the synthetic-data result comes from generation itself versus the selection pressure of the competition.

Those are the next questions, not details to hide.

Cascade's thesis is not that we already have the perfect time-series model. It is that we can build a system that keeps discovering better ingredients for one, while using controlled research to decide which model changes are actually worth carrying forward.

Better data. A better decoder. Together, they cascade.