In the spare minutes of weekend dad mode, I've been exploring code parsing and indexing for Verity.md. Which is something we decided not to do early on.
A few years ago, if you wanted to work on or modify code, you needed to index it somehow. Vector databases, embeddings, creating indexes of your code were all then important to help you search efficiently. Famously, Cursor did this.
But then Claude Code and Codex came along and started having great performance with just greping code.
@bcherny wrote "agentic search generally works better", and is simpler, with fewer problems around security, privacy, staleness and reliability. Search in Claude code is glob, grep and file reads. Codex uses ripgrep.
In Verity we've been taking this same approach. You certainly don't want to be like google antigravity introducing plan mode when the leaders are removing it.
However, in Verity (which is a peer reviewer for your agents) we have a specific set of restrictions:
- We review code without repo access, we get what we get from what the agent worked on or saw or did.
- We have a maximum time for each review. Reviews need to be done in under a min (or 3 min in a git commit/push).
So these things force us to be smarter in how we look at code. With our new formal verification feature, the boundaries of functions that we're modeling are now more sensitive. In our agent that reviews code in multiple turns, less search means also more time for code reasoning. Specifically I'm hoping that better understanding of code will let us catch bugs that span different functions, build better models, make our review agents more intelligent and ultimately support more programming languages well.
So I've been implementing a code map with tree-sitter to get exact functions and index system to find effects and function calls. I hope this returns good results.
(image is generated by claude)