Building frontier systems for forecasting. axion.eternis.ai

Pinned Tweet
1/ We found that reading internal model signals gave more reliable confidence estimates than asking the model how confident it was! With @GoodfireAI, we cut Eternis-Forecaster 8B’s calibration error by more than half in our tests. This continues to apply for much larger models, having strong implications on decision making. Below on what we found—and why it matters. Paper (arxiv.org/abs/2607.08046)
3
18
47
2,704
Eternis retweeted
Models know when they’re reward hacking. But they still do it a ton - in 50-96% of rollouts we studied! We built activation monitors that detect the behavior behind the Hugging Face hack in real time. This can help us stop hacks now - and train future models that don’t cheat. 🧵
30
99
842
147,320
1/ We found that reading internal model signals gave more reliable confidence estimates than asking the model how confident it was! With @GoodfireAI, we cut Eternis-Forecaster 8B’s calibration error by more than half in our tests. This continues to apply for much larger models, having strong implications on decision making. Below on what we found—and why it matters. Paper (arxiv.org/abs/2607.08046)
3
18
47
2,704
10/ The ambition is to bring that depth of analysis to almost every consequential decision. For each action we might take: what could follow, how likely is it, who benefits, and what could go wrong? We want people to weigh those possibilities against the futures they want.
1
10
121
11/ The future will never be solvable. However we can act today to make the futures we want more likely. At Eternis (eternis.ai), we’re building forecasting systems and coordination tools for new markets and systems where humans and autonomous agents make decisions. Reach out if you're interested in working on coordination systems for the future. We're looking for people with a love for philosophy, market design, work around governance and creating new models for forecasting.
10
124
New insights on using model internals for much more reliable forecasting will soon be live on Axion, our multi-agent system for reasoning through high-stakes forecasts and decision-making. Exciting results from our partnership with @GoodfireAI!
Can LLMs predict the next World Cup champion? Goodfire partnered with @EternisAI to improve how LLM forecasters use available evidence and manage uncertainty. We found models were overconfident in their predictions – but probes significantly improved calibration. (1/6)
3
2
15
1,566
Language models as forecasters can be accurate while being badly calibrated and overconfident. Their chain-of-thought can also omit the evidence that actually changed a forecast. In this new work, we find that internal activations give us a much more reliable signal! 1/6
Can LLMs predict the next World Cup champion? Goodfire partnered with @EternisAI to improve how LLM forecasters use available evidence and manage uncertainty. We found models were overconfident in their predictions – but probes significantly improved calibration. (1/6)
2
4
11
1,459
Eternis retweeted
Can LLMs predict the next World Cup champion? Goodfire partnered with @EternisAI to improve how LLM forecasters use available evidence and manage uncertainty. We found models were overconfident in their predictions – but probes significantly improved calibration. (1/6)
11
29
277
26,757
Eternis retweeted
the greatest experiment in human history turns 250 years old 🇺🇸 deeply grateful to the opportunities this country has given me and for the people that have fought to maintain its freedom
79
76
2,482
131,760
1/ Forecasting models We trained an 8B-parameter model that surpasses all published baselines on open-ended forecasting, including models 10–15x larger. Post-training a small model to reason about uncertainty the way a good forecaster does turns out to give a significant portion of the gains!
7
11
64
6,338
6/ At 40% of training, EF-8B already matches or exceeds every published baseline on the OpenForesight benchmark in both accuracy and Brier score.
3
15
834
7/ With these results, we’re now scaling up post-training to match the best reasoning models on general-purpose forecasting at a fraction of the cost. We are also creating internal evaluations for what makes a "good forecaster". This will be critical in a world where models reinforce beliefs and what humans and agents alike coordinate over is heavily influenced by these The goal: cheap, continually learning world models that maintain forecasts across every human-relevant question. If this is exciting, reach out to us! Full post: eternis.ai/blog/towards-sota…
12
782