Had a great time presenting at Builders & Brews Amsterdam organised by @nebiusai & @tavilyai at the @jetbrains office.
For the people in the back: verification is the next scaling axis. Powering next generation monitoring inside @moyai_ai
Looking to finally build something people want!
Customerbase is going live. Claim me as your customer with reverse lead forms. Sign me up instead for free.
For those who hate and say product hunt is already doing this. True
Didn’t the architecture role not die when most enterprises should have sunsetted their mainframes?
If history rhymes solution is prop we need to be more like architects instead of coders.
Making a case against spec-driven development ...
Here you see that the spec adherence is least differentiating criteria for finding agent failures. Checking output (maybe together with execution feedback) and errors in the agent trace works better.
Here we show scoring distribution over four verifiers on Terminal Bench 2.0
The model doesn't seem to influence the score imbalance to the highest and lowest score granularities. Context optimisation can be used to achieve more nuance verification results uniformly over models.
Presenting during @nebiusai and @tavilyai Builder & Brews event in Amsterdam.
Verification scaling the software factory. Using @deepseek_ai V4.1 flash and self-verification to hit SOTA, CI/CD green and money to spare.
Join now!
luma.com/tavily-7u0e
We ran four verifiers over terminal-bench@2.0, DeepSeek family of V4 Pro/Flash and GLM 5.2/5.3-flash.
Interesting to see is the difference in in verification tokens generated and the observed wall time. Overall verification results are relatively uniform across these four models
Just asked Fable if we can use game theory and byzantium proofs to safeguard against P(Doom), simple:
P(Doom) = Σ over coalitions C of P(attempt by C) × P(capability | attempt) × P(all safeguards fail | capable attempt)
Running GLM-5.3-flash on Nebius and Together AI shows very high overall Pearson correlation but small average score differences. Verifier optimization is necessary to create more sensitivity from the model.
AgentCon Amsterdam: Beach Drinks by boxd × Moyai
Join the teams of boxd and Moyai for beach drinks at Strandzuid, Amsterdam. Thursday 17 September, 19:00 to 22:00. It is a few minutes on foot from the conference.
Register here now: luma.com/z6jgvk0k
We are trying out new models and inference providers for our LLM verifiers. Recently we started looking int GLM-5.3-flash on Nebius and comparing to Together.
Graph shows Pearson correlations (r) on terminal-bench 2.0