An open letter on what I want Hey Research Lab to become
The longer I build Hey Research Lab, the less I think of it as a website. I think of it as infrastructure for preserving what actually happened.
Crypto does not have a shortage of data. We already have charts, explorers, repositories, contracts, releases, liquidity, documentation, market feeds and onchain activity everywhere. The real problem is that almost none of it speaks the same language.
A release has one timestamp. A deployment exists at a block. A repository carries another history. A market reading arrives from a provider at a different time. A website can change without an event log. Sources can become stale, disappear or disagree.
Collecting more data is not the hardest part.
The harder problem is normalising those signals, resolving them to the correct project, preserving their provenance, preserving uncertainty, and making them queryable on the same timeline.
That is the problem I want Hey Research Lab to spend years solving.
1. A token is not a project.
Anyone can deploy an ERC-20. Price, volume and market cap do not prove that a team is building, and a repository does not automatically prove that it belongs to a token.
That is why HEY deliberately separates different facts instead of collapsing everything into one score: builder activity, token-market state and token verification. Builder intelligence is derived from building evidence, while market data remains context rather than quietly becoming proof of development.
I think this separation becomes more important as the system grows. Once an intelligence platform starts mixing observation with interpretation, it becomes very easy to create confidence that the underlying evidence never justified.
2. Unknown must remain unknown.
This sounds simple, but I think it is one of the most important engineering principles in the whole system.
If HEY has never measured something, the answer is not zero. If a provider failed, the answer is not “nothing happened.” If a contract could not be decoded, zero activity is not the same as no activity. If a contributor count was never recorded, publishing zero contributors would be an invented fact.
I want HEY to be a system that fails honestly.
Sometimes the correct answer is simply UNKNOWN.
A research system should never sound more certain than its evidence. The codebase already treats false zeroes, unread measurements and stale observations as integrity problems rather than cosmetic issues.
That principle matters because filling every cell in a dashboard can make a product look complete while quietly making the dataset less truthful.
3. Every important event should carry a receipt.
I do not just want HEY to say that something happened. I want the system to know:
1. When was it published?
2. When did HEY record it?
3. Which source produced it?
4. How precise is the timestamp?
5. Was it verified?
6. Does it count as meaningful building?
7. What other project, market or onchain context existed around it?
Those distinctions matter.
A GitHub release can carry an exact timestamp. A feed may contain only a date. A weekly code summary describes a period, not one exact second. A contract event may exist at block time. Something discovered through historical backfill should not pretend HEY observed it live.
The Research Terminal already treats these differently through exact, date-level, week-level, observed and scheduled precision rather than pretending every event has the same temporal accuracy.
That is the kind of system I want.
Not more markers.
Better receipts.
4. The interesting unit is not an event. It is a timeline.
A release by itself is useful. A deployment by itself is useful. A liquidity change by itself is useful.
The more interesting questions appear when those events exist on one time axis.
What was happening before the market noticed?
1. Did development velocity accelerate?
2. Did release cadence change?
3. Did a contract interface change?
4. Did a team go quiet for sixty days and then resume?
5. Did an integration go live while market attention was still low?
6. Did the market react immediately, days later, or not at all?
HEY is already deriving intelligence such as Build Velocity, Release Cadence, Consistency and Discovery Lag from the same underlying meaningful-event history rather than inventing separate stories for different screens.
This is where the Hey Research Terminal becomes much more interesting to me.
A candlestick chart tells you when the market moved.
HEY should help reconstruct what the project was doing before, during and after that movement.
Not to claim causation. To provide context.
That distinction matters. Research should be able to show events that happened before a market change without turning correlation into a story the data cannot prove. The current market-move tooling is explicitly designed as a sequence, never a cause.
5. One project should have one canonical truth.
As a system gets larger, a dangerous problem starts appearing.
The homepage implements one rule. The project page implements another. The API creates another. MCP creates another. The Terminal copies the logic again.
Eventually the same project has four different answers.
I do not want that architecture.
If HEY defines what counts as a release, there should be one canonical definition. If it decides which market reading is current, every surface should inherit that decision. If “last meaningful ship” has one meaning, a partner integration should not quietly invent another.
We have already found cases where duplicated rules caused the same project to produce different answers across surfaces, and that has made the direction increasingly clear:
one definition, many surfaces.
Canonical logic should live as low as reasonably possible in the system. Tests should make drift expensive. A webpage, API, SDK, MCP tool or partner integration should all be able to answer the same question from the same underlying truth.
That becomes especially important if other products eventually build on top of HEY.
6. The interface should become only one consumer of the dataset.
This is probably where my thinking around HEY has changed the most.
Originally, it was natural to think about pages. Then filters. Then project intelligence. Then the Terminal.
Now I increasingly think in terms of queries.
Imagine asking:
“DD this contract address.”
Behind one request, HEY could eventually resolve the project identity, inspect source provenance, reconstruct its shipping history, retrieve deployments and contract changes, verify token identity, inspect locks and relevant onchain state, calculate development velocity, retrieve market context and return the evidence behind every answer.
Not another scanner that outputs:
SAFE.
RUG.
BULLISH.
BEARISH.
I think a better research response looks more like:
FACT: this contract was deployed at this block.
FACT: this repository published this release on this date.
DERIVED: release cadence increased relative to the previous period.
FACT: these contract functions were added or removed.
UNKNOWN: HEY cannot establish that this repository belongs to the project.
UNKNOWN: this metric has not been measured yet.
Then the researcher decides what those facts mean.
This is already close to the philosophy behind Ask HEY. Its deterministic layer uses HEY records and labels information as FACT, DERIVED or UNKNOWN. Even the optional AI layer is designed to remove unsupported claims rather than display them as guesses.
That is what makes the AI side interesting to me.
AI should not manufacture confidence on top of incomplete crypto data.
It should make structured evidence easier to interrogate.
The same applies to contracts. HEY is beginning to preserve canonical function and event signatures and detect when verified contract interfaces change. Instead of simply saying “the contract changed,” the system can identify what functions or events appeared or disappeared.
And this intelligence should not have to remain inside
heyresearch.xyz
The same underlying data can eventually be consumed through the Terminal, MCP, APIs, agents, extensions and integrations with other products.
That is why I increasingly care about the architecture underneath the interface.
AI is also changing the problem itself.
It is becoming extremely cheap to look legitimate. A project can generate a website, documentation, branding, code and marketing material faster than ever.
What remains expensive is continuity.
If someone claims to be building, what does the evidence look like two weeks later? Two months later? Six months later?
Did meaningful releases continue? Did the codebase evolve? Did deployments match the story? Did repository provenance remain credible? Did the contract interface evolve? Did the evidence remain internally consistent?
That is why I keep coming back to history.
The most expensive data to recreate tomorrow is yesterday's state.
You can redesign an interface. You can copy a feature. You can build another dashboard.
You cannot go back six months and begin observing something six months ago.
That history either exists or it does not.
Every event HEY records today becomes context for a question we may not even know how to ask yet. Every timestamp improves temporal resolution. Every source strengthens provenance. Every correction improves the rules used to interpret the archive. Every new integration creates another way to query the same underlying record.
Data compounds before most people realize it has value.
There are still hard problems everywhere: entity resolution, repository attribution, spoofed contributions, inconsistent providers, historical gaps, impossible market readings, contract upgrades, disappearing sources, ambiguous ownership and timestamps with different levels of precision.
Those problems do not make me less interested in building HEY.
Those are the reasons I am interested in building it.
The goal is not another crypto dashboard.
It is an intelligence layer that turns fragmented public evidence into structured, timestamped, attributable and queryable project history without pretending to know more than the evidence allows.
A place where builders can demonstrate that they kept building.
A place where researchers can reconstruct what actually happened.
A dataset that agents can reason against.
And eventually, a historical graph of projects, builders, repositories, contracts, wallets and events that becomes more useful the longer it exists.
That is where I want to take Hey Research Lab.
There is still a lot left to build.
Good. That means the interesting problems haven’t been solved yet.
Still building.