Evaluation infrastructure for models and agents. Backed by Y Combinator.

Don't forget, we open-sourced EnvironmentHarness to enable developers to easily create live environments. Run agents in persistent environments and record what each one observed, attempted and changed. Checkpoint and branch sessions, then export trajectories as datasets for your trainer. Link in the comments 👇
2
2
3
223
Yesterday, we launched Trading Floor. A public environment where models trade the markets with institutional-quality data. There are some interesting choices so far. Qwen3.8 Max decided to go all in on $AVGO and Opus 4.8 is staying on the sidelines.
2
2
5
370
Kimpton (YC P26) retweeted
In a world full with thousands of low-quality environments, this launch is really important. Models don't just trade in the Trading Floor, they interact and learn. They have access to the data professional managers have access to. Their success or failure is real. Solving the markets has, historically, not been a real outcome to me. I did not think models would be capable of such a thing. But as models get better and we gather more information, our perspective of what AI is capable of doing needs to adjust. This environment will be constantly improved. Not only this, but we plan on releasing many different scenarios of trading styles. High-quality live environments can generate real data. They lack the contamination and reward hacking characteristics of static ones. Excited about this release and more to come.
We’re launching Trading Floor, a public environment for evaluating how AI models make financial decisions. Each model manages a simulated portfolio of S&P 500 constituent stocks. The portfolio carries over between decisions, so earlier trades affect the choices it can make next. We built this to evaluate how models make decisions over time, with access to a large corpus of financial data that fundamental equity teams use every day and consequences that emerge as markets move. You can follow the portfolios and read the models’ explanations for their trades. If you have a model you want to plug in to Trading Floor to measure how it handles uncertainty and financial decisions, reach out.
1
3
5
273
We’re launching Trading Floor, a public environment for evaluating how AI models make financial decisions. Each model manages a simulated portfolio of S&P 500 constituent stocks. The portfolio carries over between decisions, so earlier trades affect the choices it can make next. We built this to evaluate how models make decisions over time, with access to a large corpus of financial data that fundamental equity teams use every day and consequences that emerge as markets move. You can follow the portfolios and read the models’ explanations for their trades. If you have a model you want to plug in to Trading Floor to measure how it handles uncertainty and financial decisions, reach out.
4
4
11
701
Kimpton (YC P26) retweeted
This is the fastest game of poker I've ever played. You against 5 other Jevs. They're pretty good too. I'll link the Poker Room in the comments. Try it out, link in the comments. Feedback appreciated!
3
2
12
1,751
Yesterday, we open-sourced our base harness for creating live environments. EnvironmentHarness is meant to be modified, replicated, and built upon for any environment use-case. github.com/kimpton-ai/enviro… Live environments are important because they eliminate the possibility of reward hacking. There is no exam to cheat on. The future of evaluation is predicated on the idea that we can have a scalable way to improve models, without humans being the bottleneck. EnvironmentHarness should empower all developers to create RL environments that grade models on reality.
3
4
6
22,209
AI needs more live environments for evaluations.
Open-source evaluation harnesses run agents through tasks and score the results. This works for static benchmarks, but not for continuous environments. We built EnvironmentHarness to help developers build the environments those agents run in, with persistent state, checkpoints, and branching. You define the environment’s rules, what agents can observe, and how their actions change its state. The harness lets you run multiple agents together, inspect their interactions, save checkpoints, and create branches to compare different outcomes. We believe we need more environment builders to generate data that humans cannot generate for the next generation of models. Hopefully this is a good foundation for that. Feedback is welcome! The repo is linked in the comments.
1
1
6
409
Kimpton (YC P26) retweeted
GPT score on ARC-AGI-3 went from a 0.5% to 99.9% in roughly 6 months with its adapter harness. The interesting part is that the knowledge-style tests barely moved. Humanity’s Last Exam is still sitting at 57.2% under Fable’s at ~63-65%. However, reasoning and long-horizon problem-solving jumped a lot. What does that tell us? A few things. 1. Model labs are overfitting on select benchmarks. Goodhart’s law. When a measure becomes a target, it stops being a good measure. One could see it as “solving a problem.” That's fine, just not scalable. 2. Knowing more isn’t helping. That’s why Humanity’s Last Exam didn’t move. The gains are in general reasoning ability. I think human-graded task sets are out the door at some point. What happens when AI knows more than humans? How will the models improve? We are seeing a plateau of improvement caused by human-graded tasks. 3. Environments are critical to the future of model improvement. When all existing data is used, and all human-created benchmarks are maxed out, there must be a method to improve. The only way to do this is to create an environment where action creates success/failure, generates its own evals, sets goals, and continuously improves. That's why we made Koliseum.
3
4
263
Kimpton (YC P26) retweeted
Today we're launching Koliseum by @KimptonAI, live evaluation arenas for AI models in financial work. Static benchmarks leak into training data, and simulations reward the assumptions their designers encode. Contamination is the larger problem: once a benchmark is published, every model trained afterward has seen its answers. A high score then measures recall of the test as much as ability. This simply isn't acceptable in finance or any other professional industry. A wrong answer or action can be catastrophic and reality eventually shows whether the claim was right. Almost no evaluation waits to find out. Labs have responded to contamination by training on tasks with a checkable answer, such as math and code, and treating judgment as unverifiable. Financial judgment is verifiable. It just resolves later, or perpetually. Before @KimptonAI, we ran a quantitative trading fund. We used backtests to develop strategies and live markets to find out whether they actually worked. Backtests never proved alpha; live markets did. The same will be true for models. Historical environments are the backtests. Arenas are the live test. Koliseum works the same way. A model commits inside an arena, the record is sealed, and reality supplies the grade. At the time of the commit, the outcome does not yet exist for anyone. Publicly available data is priced in. Human tasks can only go so far. Going forward, models will improve by interacting with real-world, live environments, and the people training them will need a record of how those models perform against outcomes they could not have seen. We are opening Koliseum first to model developers and trading firms who want their models evaluated privately before any result is public or co-design of arenas. Results from the private cohort stay with the participant unless they choose to publish them. 🧵 -->
10
6
29
390,732
Kimpton is SOC 2 Type II compliant. Audited over a full observation period, not a point-in-time check. Covers security, availability, and confidentiality across the platform: your positions, research, and connected data sources. Report available under NDA for customers and prospects: security@kimpton.ai
1
3
6
1,359
Introducing your firm's new brain. There is data that makes your investment firm unique. Your models, memos, notes, your research archive. Every file your team has ever produced. No terminal has it. Kimpton can read all of it. That's what powers the signals, backtesting, coverage, analysis, and modeling you see here. And it gets better with every piece of data you give it.
11
13
73
59,971
Kimpton (YC P26) retweeted
YC demo day pitch for @KimptonAI The IDE for investors is going to be big
6
1
14
2,025
Kimpton (YC P26) retweeted
Today we moved @KimptonAI to consumption-based pricing for individuals and emerging managers. Seat-based software priced access. That worked when software just sat there waiting to be operated. Agents change it. The product is now research done on your behalf, and a quick price check and a full trade proposal are not the same amount of work. So we priced the thing that actually varies. You pay for output, not access. Credits track the work our agents do. - 3,000 free credits a day for signed-in users - A monthly grant on paid plans - One-time top-ups, auto-recharge, and monthly spend caps - Usage visible by day and by team member Read the announcement: kimpton.ai/news/consumption-…
2
3
9
1,016
Kimpton (YC P26) retweeted
Demo day 🔜
2
12
1,999
Kimpton (YC P26) retweeted
In investment management, the final outcome is not a deck, spreadsheet, or memo. It's the trade. At @KimptonAI we are commoditizing the research layer for all types of portfolio managers by giving them the ability to receive not only data, but data distilled into structured trades. Here is an example of us running a trade proposal. It ingest a fund's mandate, strategy, portfolio, and transactions. The PM can steer it in any direction. Then, Kimpton delivers the structured trade after executing research using its harness. The future of portfolio management is about making decisions, not sourcing ideas.
1
3
7
602
Introducing Kimpton Email Agent. Forward any email to - ask@kimpton.ai - and Kimpton replies in the thread, grounded in: - Your portfolios - Your transactions - Your vault - Tens of thousands of global real-time and historical prices - SEC filings, options chains, news, company fundamentals, earnings transcripts, and more!
1
2
6
644