Making agentic workflows hallucination and poison free powered by our verifiable databases and cryptography tools.

Switzerland
Pinned Tweet
SourceryKit in early adoption: - detects 100% of agent tool-call errors. - improves accuracy by up to 50% in some models. - works across single-agent and multi-agent workflows. This is why we are building runtime verification for production agents.
1
2
6
216
1/7 Even the best agents get tool calls wrong. What helps them recover? Our previous study showed why a convincing answer isn’t enough. We took 60 failed runs and tried three recovery methods on them. Here’s what we learned.
1
4
7
195
6/7 One hard task followed a chain from a museum to its architect, his birthplace, a library and the nearest subway. Guided recovery found evidence for all five expected claims. The final answer still omitted them. Finding the evidence isn’t the same as using it.
1
6
7/7 Of 32 guided runs that recovered complete evidence, only 11 also passed the final-answer check. The next challenge: getting that evidence into the answer. These were selected failures, weighted toward missing tool evidence. provably.ai/blogs/the-trace-…
7
Provably retweeted
anthropic.com/research/multi… This is fantastic read. 1. agents struggle to coordinate with each other especially for common resources. 2. they don't know how to detect when they are being lied to. And as usual Anthropic wants to solve it through more learning..
2
3
51
We spent 9 weeks testing agent reliability during software tool use. Nearly 60% of claims accepted as correct could not be backed by evidence from recorded tool calls. The answer looked right. The evidence wasn't. 🧵
2
5
6
130
We believe agents should capture evidence natively as they work. SourceryKit captures tool calls in a state of the art verifiable database, allowing agents to evaluate themselves and downstream agents to independently verify their answers.
1
12
SourceryKit detected 100% of the covered deterministic grounding errors in this study. Next, we will test how much accuracy agents can recover through answer healing and targeted retries. Read the study: provably.ai/blogs/The-Agent-…
9
Provably retweeted
AI is going collaborative. In the next 6 months all we are gonna talk about is verifiability (for safety, security and accuracy), protocols and crazy new use cases.
3
3
45
Provably retweeted
Replying to @hot_town @jack
The main thing in AI 2.0 - which Buzz is triggering with collaborative AI - is verifiability. Bolting on & releasing shortly. Then whenever another AI does a task it, it can prove it really did what claimed it did. It saves you needing access to the “actual” Terminal / end point.
2
5
133
Provably retweeted
Three months ago, we decided to work on one of the really big problems in AI: Agent reliability in tool calling.
1
5
13
6,899
SourceryKit in early adoption: - detects 100% of agent tool-call errors. - improves accuracy by up to 50% in some models. - works across single-agent and multi-agent workflows. This is why we are building runtime verification for production agents.
1
2
6
216
Our view: The future is agents that can verify their own work, catch errors, and heal while the workflow is still running. As agents get smarter, this feedback loop will help them adapt and continuously learn from new workloads, powered by verifiable data from each other. It is time to take the training wheels off.
1
29
SourceryKit is available now. It is a Python SDK for runtime verification of agent tool calls. You can verify the tool response while the workflow is still running, before the error spreads. Get started in minutes at provably.ai/ with pip install sourcerykit Check our launch blog: provably.ai/blogs/Runtime-Ve…
31