$STYXX is live on robinhood chain. 0xC750bcdAe34cC578Ff17963bed40C5d9396fdC5D launch tx: 0x4996f313ef099e5297c767a3a816db49b231756b3faea67e3662696a521b5ce2 dev buy: 0.07 ETH · 0x9035…0948 our other wallet: 0x14a2…c0f3. that's all of them. creator tax: 0 · creator fees → holders, pro-rata, routed on-chain supply fixed at 1B · no mint · no freeze · liquidity locks at graduation (4.2 ETH) full launch receipt, read from the chain, not typed: sha256 5ff50c30f02c8122cf69c867a66e2355b21e0c346973ebaa5cb2e175af7bd12a anchor tx: in the reply below. the solana $STYXX mint is discontinued. no migration. this is a separate token. only buy the address above. names can be copied. addresses can't. pons: ponsfamily.com/launchpad/0xC…
5
1
9
3,108
styxx 7.49.0 is live. pip install -U styxx we attacked our own proof format, found the hole, shipped the fix, and filed the advisory against ourselves. → the hole OATH capsules are single-file documents that carry their own receipts plus a certificate, and styxx.capsule verify re-derives that certificate from the bytes. we went after the verifier ourselves. on every release since capsules shipped (7.47.0, 7.48.0, 7.48.1) it printed VERIFIED on capsules whose certificate had been edited: an unchecked number hidden, a ledger row deleted, a receipt repointed, an epistemics flag flipped, a page that shows one document while the verifier reads another. five forgeries. 7.48.0 and 7.48.1 passed every one. → the fix 7.49.0 compares the whole certificate and the page around it. all five forgeries fail now. a newly minted page shows a verdict only after its hashes match. advisory, filed and published by us: GHSA-3g8h-qcfm-25xw → the part we left open, on purpose and in writing capsules minted on the old page format still verify, because honest ones carry it. a forged one dressed up as "older" still passes on that page, but 7.49.0 prints six NOT CHECKED lines and three warnings next to it. on a new page the same forgery fails. → the diff gate holds back where it could be wrong PATH-2a: where three known reader bugs could flip a verdict, the gate withholds the verdict. → the Action reports by default and blocks only if you opt in reason: no kind of accusation it still makes has cleared our own 0.95 precision floor. path accusations measured 0.23 precision on 100 accusations sampled from 71,016 external agent PRs. that number is in the README. → "LIE" is gone. a contradicted claim prints CONTRADICTED, with its reason. → receipts PyPI 7.49.0 · GitHub release with sha256 for both files · Zenodo DOI 10.5281/zenodo.23221755 a verifier that hides its own failure modes isn't worth running. so we published ours.
64
styxx // what's next no cap: the layer on #187 went twelve rounds with our own reviewers. three of four already said ship on round twelve. the last one's still in the booth. here's the plan, no smoke: - the last review lands. if it says ship, the layer cleared all four. - then it's a human call: is "can't tell" good enough, or do we hold out for a real fix. we said fix up front, so a human signs off on any switch. no shortcuts. - if it's a yes: merge, then a new release. 7.48.0 still gets fooled, and we're not pretending it doesn't. - after the merge, 16 open bugs we caught in our own release are next in line. they wait on #187 because the layer is pinned to the release's exact bytes. touch those early and the proof breaks. - one gap stays on the table, out loud: under certain git diff options a false call can still slip through. closing it costs 98 right answers per port on our test set, so a human decides that too. we don't ship vibes. we ship receipts. the tool you trust should know when it doesn't know. github.com/fathom-lab/styxx/…
75
our lab's gate, now runnable in a browser tab with nothing installed: paste the description your coding agent wrote and the diff it shipped, read the verdict line by line. private diffs never leave the tab.
2
208
our diff gate, run in public on any PR you submit. it reads what the agent wrote about the diff and checks it against the diff: [ok] / [LIE] / UNCHECKABLE. it gated our own PR first — PASS, 5 of 5, in CI this morning. reply there with a link.
186
why we launched $STYXX the way we did, and what happened. a thread for the people who arrived today. 1/5
1
3
9
437
4/ what happened: launched 18:40 UTC. graduated 21:33 UTC — under three hours. liquidity now sits in a uniswap v4 position locked in the pons launch locker; nobody can pull it, including us. holder fee sharing paid out three times on day one. dev wallet untouched. every one of those is a tx you can open.
1
1
62
5/ what's next: the plate as a moving picture (receipts rendered as chladni figures — this is what one looks like still), the bounty spec for reproduction failures paid in $STYXX, and daily receipts in the pinned thread on @styxxhq. no dates. we post when it's real. nothing crosses unseen.
1
53
for everyone who arrived today through the token — what styxx actually is: pip install styxx. an instrument that sits on every model response and reads vitals, then turns every claim about a model into a receipt anyone can regenerate and hash. no GPU, no LLM in the loop, no trusting us. the token launch got the same treatment as a model run: read from the chain, hashed, anchored. same discipline, different substrate. the plate below is what a receipt looks like when you render it — a chladni figure. every claim leaves a pattern. that's the whole idea. github.com/fathom-lab/styxx
2
1
6
296
Fathom Lab retweeted
$STYXX graduated. 4.2 ETH raised on the curve → liquidity is now in a uniswap v4 pool on robinhood chain, locked. no one can pull it, including us. graduation tx: 0x1cfa94b5e3b664a866485ffa634372f289836da79ce5d4ae3624873df9173209 pool id: 0x90f25edf87b00914a1998864b63412df6917751990989a5191c9984e770f1548 position #2759185 locked in the pons launch locker 0x267444D099b10fB5Ed7c3Cc7B7c767AdcA574952 trading continues on the same page: ponsfamily.com/launchpad/0xC… creator fees keep routing to holders. nothing else changes. nothing crosses unseen.
2
1
9
1,172
$STYXX is live on robinhood chain. 0xC750bcdAe34cC578Ff17963bed40C5d9396fdC5D launch tx: 0x4996f313ef099e5297c767a3a816db49b231756b3faea67e3662696a521b5ce2 dev buy: 0.07 ETH · 0x9035…0948 our other wallet: 0x14a2…c0f3. that's all of them. creator tax: 0 · creator fees → holders, pro-rata, routed on-chain supply fixed at 1B · no mint · no freeze · liquidity locks at graduation (4.2 ETH) full launch receipt, read from the chain, not typed: sha256 5ff50c30f02c8122cf69c867a66e2355b21e0c346973ebaa5cb2e175af7bd12a anchor tx: in the reply below. the solana $STYXX mint is discontinued. no migration. this is a separate token. only buy the address above. names can be copied. addresses can't. pons: ponsfamily.com/launchpad/0xC…
5
1
9
3,108
anchor tx: 0x4c5d0b1fc4240fa8640b6c4354b98bfcc2298f4fc6f5c9e95f1e8146dd3fc82a 0 ETH, dev wallet to itself, block 63935251, 20:25:10 UTC. the calldata is the sha256 of launch_receipt.json — byte for byte. check it with zero trust in us. file + script land in the repo with the next push (docs/token/receipt): python3 launch_receipt.py verify launch_receipt.json --anchor-tx 0x4c5d…c82a 19 chain checks. all hold.
113
🌊 STYXX UPDATE: we tried to prove ourselves wrong. Here's what survived. WHAT STYXX IS, IN PLAIN WORDS AI companies publish numbers and say "trust us." Styxx makes those claims checkable by anyone, without having to trust us either. • Every number we publish is tied to the exact file it came from. • Every verdict is re-computed by software and written into a public, chained log. We call it the ferry log, and it now holds 256 entries. Change one old entry and the chain breaks. • Our reports are "sworn." Each number points back to its record, and a stranger can re-check a whole document with one command. THE SAND: OUR FINGERPRINT When you pay for an AI model through an API, how do you know you're getting the model you were promised? It could be a cheaper, compressed, or quietly updated one instead. The sand gives a model a set of sealed test questions and records how it answers, down to how confident it is in each word. It also measures how much that same model normally wobbles against itself. If its answers move further than that normal wobble, something changed. And you get receipts. HAS ANYONE ALREADY DONE THIS? WE SPENT FOUR ROUNDS FINDING OUT • We wrote the rules down and published them before reading anything. • 54 papers and projects were read end to end by independent AI readers and checked piece by piece. • Every quote was checked against the real text: 100 of 100 found. • Then we attacked our own work. 44 attacks, 33 landed, and we fixed every one of them in public. THE RESULT We know of no lab that ties every published number to its bytes, re-derives every verdict into a chained log, and fingerprints a model's behaviour on sealed test questions against a measured floor, all at once. The honest part: four research groups come within one missing piece of the fingerprint. They are DiFR, Chauvin et al., Gao et al. and Hochlehnert et al., and we name them. We also dropped the bounty from that sentence because one paywalled article couldn't be checked. That's the standard we hold ourselves to. HOW FAR WE'VE COME • A "sworn" document format, plus a second, independent verifier that agrees with the first • A public ferry log with 256 entries, every one re-checkable • A four-round survey of the field, and a public correction when our own red team caught mistakes • A scoring engine that caught all 92 of 92 deliberate attempts to break it WHERE THIS IS GOING AI models get swapped, compressed and updated behind APIs every day, and most people have no way to check. Next up: • round five of the survey • sealed test runs on real hardware • a planned standing bounty for anyone who catches our own verifier wrong The goal: any claim about an AI model, checkable by anyone. 🔥 Built in the open: github.com/fathom-lab/styxx
1
8
452
four models. three companies. one geometry. these are chladni plates. in 1787 a physicist put sand on a metal plate and bowed it — the sand jumped off the parts that moved and settled on the lines that didn't, and for the first time anyone could see a sound. we did it to ai models. the sand here is driven by how each model arranges 462 concepts inside itself. meta's llama, google's gemma, alibaba's qwen: same lines. bottom right is the control — the same qwen with its concepts shuffled. nothing alike. not a metaphor. the agreement numbers are printed under each plate (0.94 / 0.96 / 0.87) and anyone can rerun the whole thing from the banks in the public repo with one script. the picture can't make two models agree more than the number says. we published this finding in august. what's new is that you can see it.
4
269
we built a system whose job is to catch AI models changing without anyone noticing. then we spent a day trying to break it. it lost five times. the fifth loss turned out to be a real result, and it changed what we think we're building. 🧵
2
2
343
honest status: released styxx is on PyPI. everything above is unreleased — one branch, one laptop, not committed. what's real: runs end to end, suite green, two independently-written implementations agree byte-for-byte on real certificates, one flipped byte anywhere gets caught and named.
1
65
what isn't real: nobody outside our room has checked any of it. and the feature we just concluded matters more than all the others — an outsider reproducing our work and disagreeing — has still never happened once. that's next. not more defences.
34