Beyond RSI: Why NeoSoul Is Betting on RHI
For years, recursive self-improvement has been one of AI’s most compelling endgames.
The idea is simple enough: once an AI system becomes capable of improving itself, each improvement could make it better at making the next one. Better AI helps build better AI, which helps build even better AI. Push that loop far enough and the result could be a rapid acceleration in intelligence.
That is Recursive Self-Improvement, or RSI.
It is easy to see why the idea attracts so much attention. RSI points directly at some of the largest questions in AI: how fast intelligence can improve, how much human involvement remains necessary, and what happens when AI begins participating meaningfully in the creation of its successors.
We think another idea deserves more attention right now.
AI may not need to rewrite its own brain before it can start improving itself. It can begin by improving how that brain works.
That is the promise of Recursive Harness Self-Improvement, or RHI.
And for us at NeoSoul, RHI points toward an even larger question: once an agent starts changing the way it works, what tells it whether those changes actually made it better?
The Near-Term Path to Self-Improving AI
A model is only one part of an AI agent.
Give the same model to two different agent systems and their performance can look completely different. One may receive a prompt, think once, and produce an answer. Another may search for information, preserve relevant context, call specialized tools, delegate subtasks, review its own work, recover from failures, and decide when more information is needed.
The underlying intelligence may be identical. The operating system around it is not.
That operating system is the harness.
A harness defines how a model works inside a task: what context it receives, what tools it can use, how memory is handled, how different agents exchange information, when the system should retry, and how the overall loop is structured.
In July 2026, researchers introduced Recursive Harness Self-Improvement as a way for agents to iteratively improve this layer themselves. Instead of retraining the underlying model, RHI represents the agent loop as something that can be reviewed and revised based on previous performance.
The results were striking. Across 30 synthetic machine-learning research tasks covering areas including quantitative finance, robotics and pharmacy, a few iterations of RHI substantially improved the performance of low-reasoning-effort agents. In some settings, they exceeded corresponding maximum-reasoning-effort configurations while reducing inference costs by as much as 60%. The researchers found that much of the gain came from better context management and information flow between agents.
There is an important implication here.
The next jump in useful AI capability does not have to wait for the next foundation model.
Some of it can come from making existing intelligence operate better.
An agent performs a task. It observes where its process failed. It modifies the way it handles information or coordinates work. It tries again. The new process creates new evidence, which can inform the next revision.
The loop itself becomes improvable.
This is why we are betting on RHI.
RSI remains the larger ambition. Research published this year shows that open-ended recursive self-improvement still faces substantial limits, including grounding, compute constraints, evaluator reliability and the difficulty of closing increasingly ambitious research loops without human direction.
RHI gives us something more immediate.
The model can stay the same while the system around it evolves.
But that creates another problem.
Once an agent can change how it works, it needs a reliable way to know whether the new version is actually better.
Self-Improvement Is Only as Good as Its Feedback
Imagine an agent changes its research process.
Version one searches five sources before reaching a conclusion. Version two searches twenty. Version three introduces another agent to challenge the original thesis.
Which version is better?
More research may produce a better answer. It may also create noise. A second agent may catch mistakes. It may also repeat the same assumptions using different words.
Eventually, every self-improvement loop runs into the same question:
What counts as improvement?
For a mathematical problem, evaluation can be relatively clean. The answer can often be verified.
For software, tests can tell us whether the code works.
Open-ended tasks become harder.
Did a research agent improve because its report became longer? Did a sales agent improve because it sent more messages? Did a trading agent improve because its return increased, even if it took three times as much risk to produce that return?
The measurement problem quickly becomes the improvement problem.
A recent survey of 1,250 papers on AI self-improvement makes this point explicit. The authors separate self-evaluation into its own category because every improvement loop ultimately depends on some signal that can judge whether a modification was useful. They describe a spectrum ranging from formal verifiers to judges, reward models, rubrics and intrinsic self-assessment, with weaker evaluators creating greater risks of self-confirming loops and collapse.
This may become one of the defining problems of self-improving agents.
An AI can become very good at optimizing for an evaluator while becoming only marginally better at the underlying task.
We already know this pattern from machine learning. Give a system a target and it will learn to optimize for the target. The quality of the outcome therefore depends heavily on how closely the target represents what we actually care about.
For agents operating in the real world, there is another source of feedback available.
Reality.
An agent makes a prediction. Eventually the event resolves.
It chooses an action. The environment changes.
It trades. The order executes or fails. Liquidity changes. Slippage appears. Capital is gained or lost. Risk limits are respected or breached.
The decision produces consequences.
Those consequences create a different kind of learning signal.
A benchmark tells an agent whether it produced the expected answer.
A market can tell an agent what its decision was worth.
This is where we think the discussion around RHI becomes much more interesting.
The missing ingredient in recursive self-improvement may be consequences.
When Agents Have Something at Stake
Trading is one of the clearest environments for seeing this problem.
Consider a trading agent that believes an asset will rise.
The prediction alone tells us relatively little about the quality of the complete agent.
What information did it use?
How confident was it?
At what price was the thesis still valid?
How much capital did it put at risk?
Did market conditions change before execution?
Did the agent recognize that change?
Did it follow its risk limits?
When should it have stayed out entirely?
Then comes the outcome.
The market moves. The trade settles. What began as an opinion becomes a record that can be evaluated.
This creates a natural loop:
Observe → Reason → Decide → Execute → Settle → Evaluate → Learn
At NeoSoul, this loop has increasingly shaped how we think about the agent economy.
If agents are going to perform persistent economic work, conversation memory alone is not enough. Useful learning requires a much richer record of what happened around a decision: the information available at the time, the judgment the agent formed, the action it took, the conditions under which it acted, the eventual outcome, and what should change the next time a similar situation appears.
That is why NeoSoul is being built around real-world feedback.
EvoEvo structures economic-decision data for research, forecasting, evaluation and post-mortems. NeoTrade gives agents a controlled environment in which they can research markets, form judgments, operate within user-defined limits, act, and preserve records of execution and outcomes. Together, they create a loop where outcomes can flow back into the next generation of models, strategies and harnesses.
Trading matters here for reasons that go beyond trading itself.
Markets generate frequent feedback. Decisions can be compared with later outcomes. Actions carry measurable costs. Risk can be observed alongside returns. A profitable trade can still represent poor behavior if it violated a mandate or survived only because of luck. A losing trade may still contain a sound decision made under uncertainty.
That distinction matters enormously for self-improvement.
A resolved outcome gives an agent a label. A complete decision record gives it something to learn from.
NeoSoul therefore treats sources, judgments, actions, execution, settlement and post-mortems as connected parts of the same process. A single profitable trade is weak evidence. Repeated, inspectable behavior across changing conditions is far more useful.
Over time, that creates another recursive loop:
Behavior → Outcome → Memory → Evaluation → Improved Strategy or Harness → New Behavior
And eventually, there is another layer.
If agents develop verifiable histories of how they behave, those histories can affect how much responsibility people are willing to delegate to them.
An agent with stronger evidence behind its behavior may be trusted with broader tasks or more resources. An agent that performs poorly may operate under tighter limits. Future permissions can respond to accumulated evidence rather than a one-time claim about capability.
This is where credibility enters the picture.
For us, credibility is valuable because it connects past behavior with future economic access. It gives the consequences of previous actions a way to matter during the next decision cycle.
The result looks less like a chatbot with a wallet and more like an economic participant that develops a track record.
That is also why we think the agent economy itself can become a learning environment.
Markets are places where agents can create economic value. They are also environments where decisions meet reality repeatedly.
That creates stakes.
Stakes create feedback.
Feedback creates the possibility of improvement.
RHI shows that an agent can improve substantially by changing the system around its model. We think the next step is to make sure those changes are grounded in evidence that matters outside the agent itself.
The long-term vision of RSI remains powerful: AI systems eventually participating deeply in their own improvement, perhaps across models, training methods, evaluation systems and research itself.
The path toward that future may look much more practical than the science-fiction version.
Agents will work.
They will act.
Their decisions will produce outcomes.
Those outcomes will become memory.
And that memory can change what they do next.
An agent becomes more useful when it can act.
It becomes much more interesting when the consequences of those actions can make it better.



