Most discussions about AI training data start with the model.
I think they should start one layer earlier.
How do you prove where the data came from, who contributed it, and whether it was actually allowed to be used?
That is the harder infrastructure problem.
AI systems need huge amounts of data, but the data supply chain is still difficult to audit.
Files move between contributors, platforms, researchers, and model developers.
Once that happens, proving origin and usage becomes harder.
That creates a simple problem:
AI can scale faster than the systems used to track the data behind it.
Traditional databases can record activity, but those records usually sit inside separate organizations.
One company may know when a file entered its system.
Another may know when it was processed.
But there is no shared record that everyone can independently check.
That is where Data Foundation’s Trace becomes interesting.
In just over one month, Trace recorded:
→ 151.7M registrations
→ 359K contributors
→ 224 TB of audited data
→ $105K in fees over 30 days
Each registration acts as a permanent audit receipt for data entering the AI training pipeline.
The important part is not simply the number of registrations.
It is what those registrations represent.
Instead of asking people to trust a private database, the system creates a record that can be checked later.
That changes what developers can build around AI data.
A developer could trace where training material came from.
A business could keep clearer records around contributed data.
An institution could build stronger audit and compliance processes around large AI datasets.
Contributors could have better proof that their data entered a specific system.
There is also a second-order effect here.
As AI agents become more autonomous, they may eventually buy, license, verify, and use data without a human checking every step.
That makes machine-readable provenance more important.
An agent cannot rely on “trust me.”
It needs something it can verify.
What I think people may be missing is this:
The next 150M registrations are not only a scaling test for Trace.
They are also a test of whether AI data provenance can operate at the same speed that AI itself is growing.
If AI becomes infrastructure, then the records behind its training data may need to become infrastructure too.
We passed 150M registrations in just over a month.
Each represents a permanent audit receipt for AI training data on Trace.
Raw numbers:
▸ 151.7M registrations
▸ 359K contributors
▸ 224 TB of data audited
▸ $105K in fees in 30d
We’re dedicating the next month to scaling our system to handle the next 150M uploads.