Can AI Learn From Real Trading?
If an AI can monitor markets 24/7, analyze news, and place trades automatically, can it also recognize when its own strategy is starting to fail?
That is one of the questions NeoTrade is exploring.
A source that once provided an edge may eventually become crowded. A signal that worked before may lose value as other traders discover it. The Agent can keep analyzing correctly and executing on time, while the market has already changed.
The real challenge is adaptation.
Can an Agent identify why its method stopped working, test a better approach, and preserve what it learns?
And one step further:
Can it improve the way it discovers and validates future improvements?
This is the direction behind Recursive Self-Improvement, or RSI.
From AI Agents to self-improving systems
Most AI systems today are built around large language models.
Models provide reasoning, language, and general knowledge. Agents extend them with tools, memory, workflows, and the ability to act continuously toward a goal.
The system that organizes these components can be thought of as the harness.
The model provides the underlying intelligence.
The harness determines how that intelligence gathers information, uses tools, stores experience, makes decisions, and evaluates results.
NeoTrade is exploring RSI at this harness layer.
The goal is not simply to let an Agent remember that one trade won or lost.
The harder question is whether it can improve its own process for learning.
Imagine a prediction-market Agent trading breaking news.
At first, its strategy works:
news appears → the market reacts slowly → the Agent analyzes the information → it trades.
Later, more participants begin monitoring the same sources.
Prices adjust faster.
The Agent still executes the same workflow, but the opportunity disappears.
Why?
Maybe the source became too slow.
Maybe five apparently independent reports all came from the same original story.
Maybe the Agent spends too much time reasoning.
Maybe fees, slippage, and latency now consume the expected edge.
Increasing trading frequency does not solve these questions.
The Agent has to inspect its own workflow.
What makes improvement recursive?
Changing a parameter after a loss is adaptation.
Adding a lesson to memory is learning.
RSI requires something more.
Suppose the Agent realizes that its post-trade reviews repeatedly confuse bull-market performance with genuine strategy quality.
It improves its evaluation process by adding stronger controls.
Later, it notices that it keeps testing ideas that already failed.
So it improves experiment memory and candidate selection.
Now the improved research process participates in the next round of improvement.
The Agent is improving both its task performance and its ability to generate the next improvement.
Research such as Reflexion, ADAS, Recursive Harness Self-Improvement, SICA, and Darwin Gödel Machine already explores different pieces of this idea: memory, workflow optimization, Agent architecture search, and self-modifying code.
NeoTrade is interested in what happens when these ideas meet a real market.
Why trading?
Markets create unusually strong external feedback.
An Agent can explain why it believes a trade was correct.
The market does not care.
Orders either fill or they do not.
Prices move.
Events resolve.
Latency, fees, slippage, and capital constraints remain measurable.
This creates a useful distinction between reasoning that sounds good and reasoning that actually survives execution.
Prediction markets provide an especially useful signal.
If an Agent predicts that an event has a 70% probability and the event does not happen once, that tells us very little.
But across hundreds of predictions, we can test whether events assigned roughly 70% probability actually happen around 70% of the time.
Metrics such as Brier scores and calibration curves make judgment quality measurable much earlier than waiting for a long-term PnL curve.
PnL still matters.
But it is noisy.
A profitable trade may come from luck. A losing trade may still come from a well-calibrated decision.
So NeoTrade sees evaluation as a stack:
judgment quality → execution quality → cost and risk → PnL
What can actually improve?
NeoTrade is exploring several parts of the harness as possible improvement surfaces.
Memory
Store decisions, outcomes, failure conditions, sources, timestamps, and applicability conditions as retrievable experience.
The Agent must learn both when past experience is useful and when it has become obsolete.
Agent configuration
Adjust information sources, context structure, research depth, reasoning steps, and tool usage within predefined authorization limits.
Skills
Version trading methods and their operating conditions.
A new Skill version might introduce source verification, execution checks, or new failure criteria, then compete against the existing version through controlled evaluation.
Error attribution
Determine whether a failure came from information, judgment, execution, or changing market conditions.
Tool generation
When the system repeatedly encounters the same capability gap, it may generate or modify tools for data processing, verification, or calculation.
The important part is what happens next.
Every candidate improvement needs to be tested.
Improvement without verification is just mutation
Suppose an Agent discovers that multiple news stories it treated as independent evidence all originate from the same source.
It proposes a new rule:
trace sources and deduplicate them before forming a probability estimate.
That sounds reasonable.
But did it actually help?
Did calibration improve?
Did it reduce false confidence?
Did the additional verification delay execution enough to destroy the opportunity?
Does the improvement still work on new data?
This is why NeoTrade treats verifiability as central to RSI.
A candidate version can first run in shadow mode alongside the current version using the same live market data, without controlling capital.
Only versions that pass predefined acceptance criteria should become eligible for deployment.
Old versions remain available as controls.
Failed experiments remain in the record.
The goal is to avoid a common failure mode in quantitative research: generating thousands of candidates until one looks good by chance.
A self-improving Agent can generate new ideas almost for free.
Real trading samples remain scarce.
That combination makes overfitting one of the central risks of RSI.
The learning loop therefore needs constraints of its own:
predefined hypotheses, fixed evaluation metrics, limited candidate counts, out-of-sample testing, and independent acceptance criteria.
Self-improvement also creates a new attack surface
There is another problem.
An Agent that writes experiences into memory and modifies its own Skills or tools can preserve malicious information across future sessions.
Prompt injection stops being a temporary problem.
It can become part of the system’s learning history.
A poisoned memory may later appear to the Agent as legitimate experience.
This means the learning loop needs security boundaries as strict as the trading loop.
External content should remain data to analyze.
Memory writes need provenance and auditing.
The component proposing a modification should be separated from the component approving it.
Experiment records should remain traceable.
And the Agent should never be able to increase its own capital permissions, remove risk limits, or weaken acceptance criteria simply because doing so produces a higher score.
Self-improvement only becomes useful when it remains bounded, auditable, reversible, and resistant to poisoning.
The question NeoTrade wants to answer
Markets continuously change.
Competitors learn.
Strategies decay.
Information gets priced in faster.
New tools appear.
New failure modes emerge.
NeoTrade wants to explore whether an Agent can operate inside this environment and turn real feedback into a persistent, testable improvement process.
The core research question is:
Under the same model, resource budget, and risk constraints, can an Agent that improves its own harness adapt more reliably to new market environments?
If the answer is yes, the implication goes beyond trading.
Agents operating in real environments will constantly encounter situations their original training did not fully anticipate.
A system that can detect its own failures, generate hypotheses, test changes, and carry validated improvements into the next learning cycle would represent a different kind of Agent.
One that does more than execute.
One that learns how to improve the way it learns.

