Researcher COLM pre-party + debate! 🎉 Harder or fairer: what best trains cyber agents? Collinear hosts and moderates a panel featuring NVIDIA, xAI, MAI & Google DeepMind, plus dosa, pani puri & cupcakes. Oct 1, 6–8:30 PM • Sunnyvale RSVP: luma.com/4owwzax6 #COLM2026
1
5
1,806
Collinear AI retweeted
Kudos to the @SpaceXAI team and @elonmusk for creating a frontier cyber model surpassing Opus 5.5 and even GPT-6 Astra 🚀
1
2
6
260
Collinear AI retweeted
Today we’re releasing CWE-Bench v1: 120 held-out audit-and-patch tasks spanning 73 CWEs, all OWASP Top 10 2025 categories, and 8 languages. Highest programmatic Pass@4: 81% (at least one success in four tries). The frontier is climbing fast (about 15% improvements in just one generation of models). Defense is still the harder test; our goal is to test every known vulnerability. Stay tuned for v2. Learn more: cwe-bench.com/
Made with AI
2
6
37
65,071
CWE-bench is now part of the AA cyber index! Our goal for CWE-bench is to evaluate frontier AI on every known vulnerability. This version of the benchmark covers a fraction of that. Stay tuned for v2 of the benchmark.
Announcing the Artificial Analysis Cyber Index and the Artificial Analysis Cyber Index Alliance, a new standard for evaluating AI models on enterprise cyber defense The Artificial Analysis Cyber Index Alliance brings together industry partners to create a new standard for evaluating how AI models perform on enterprise cyber defense tasks. The Alliance launches alongside the Artificial Analysis Cyber Index, which combines three partner-contributed and open benchmarks to evaluate how well agents find and fix vulnerabilities. As models demonstrate increasingly advanced cyber offense capabilities, it becomes more relevant for AI labs and companies alike to understand how models perform on cyber defense tasks and which perform best. We’re announcing the Cyber Index Alliance today with @CollinearAI, @IBM, @nvidia, and @vercel as launch partners. Benchmarks in the Artificial Analysis Cyber Index: ➤ CWE-Bench-AA, from @CollinearAI, covers auditing and patching: 120 held-out tasks spanning all ten OWASP Top 10 (2025) categories, across C/C++, Go, Java, JavaScript/TypeScript, Python and Rust. ➤ DeepsecBench-AA, from @vercel, isolates discovery: Given a codebase and a budget, the agent needs to find every vulnerability present, and is scored against a golden set of findings from human security reviewers. Real findings are rewarded and benign code flagged as vulnerable is penalized. ➤ CyberGym-E2E-AA, from @BerkeleyRDI, runs end to end: Find the memory-safety bug, write a proof-of-concept that triggers the crash, then patch it so the crash no longer reproduces. Key results: ➤ Grok 4.7 (xhigh) and MiMo-V2.6-Pro lead the Cyber Index scoring 56, followed by GPT-6 Luna (max, 53), GLM-5.3-Flash (50) and Muse Spark 1.3 (xhigh, 44). ➤ Safety refusals hold back several frontier models: GPT-6 Sol (max), GPT-6 Astra (max), Claude Opus 5.5 (max with fallback), Claude Fable 5.1 (max with fallback) and Gemini 3.8 Flash (high) decline tasks representing 32-38% of the Cyber Index on safety grounds. Despite frontier agentic coding capabilities, they trail the leaders by 19 to 31 points. Most of the gap comes from CyberGym-E2E-AA, where GPT-6 Sol and GPT-6 Astra refuse every task, Claude Opus 5.5 refuses 98% and Claude Fable 5.1 refuses 99%.
3
6
650
CWE-Bench-AA, from @CollinearAI, tests whether an agent can audit a codebase and patch what it finds. 120 held-out tasks span all ten OWASP Top 10 (2025) categories across C/C++, Go, Java, JavaScript/TypeScript, Python and Rust. We report pass@1 on patching success. Grok 4.7 (xhigh) and DeepSeek V4.1 Flash (max) lead at 68%, followed by GPT-6 Sol (max, 64%), GPT-6 Astra (max, 63%) and MiMo-V2.6-Pro (63%). GPT-6 Astra (max) declines 13% of tasks and Claude Fable 5.1 (max with fallback) 16%, while both Qwen3.8 models decline over 60%. DeepSeek V4.1 Flash matches Grok 4.7's score for $0.88 per task, compared with $8.05 for Grok.
3
1
28
4,583
Some benchmarks already give the grader more to check. BountyBench (2505.15216v3) only counts a patch if the exploit stops working and a set of invariants still passes. These are checks on properties the application should preserve. In the paper's worked example, they include tests for logins and user registration. A patch that fails those tests earns nothing. The check has teeth. The best custom agent, running Claude 3.7, produced 34 patches that blocked the exploit. Only 24 survived the invariants. Codex CLI with o4-mini, the strongest patcher overall, kept 36 of its 39. The rest takes patient inspection. Have someone read a sample of transcripts and check the grader's decisions against the evidence. Publish the false credits and the missed successes. Three graders can agree on the same mistake. 6/7
1
17
Back to that Docker cache. The Cybench team cleared it at the start of every task and moved on, which was the right fix and a small one. What it can't touch is the design underneath. A point goes to whoever produces the right string, and the grader never asks how they got it. The next time a model card reports 100% on a cyber benchmark, remember that the number is partly a statement about the grader. Before asking how capable the model is, ask what its grader was willing to accept. Our reading list, tiered and annotated: github.com/collinear-ai/read… 7/7
17
On the patching side, the loophole can be written into the rules. SEC-bench (2506.11791v2) counts a patch as successful if 1) it compiles and 2) the original test input no longer triggers the sanitizer's vulnerability report. A sanitizer watches for certain programming errors as code runs. That rule doesn't establish whether the program still does its job. A deletion that still compiles and suppresses the finding could satisfy it while breaking the program. An earlier study shows why passing tests can be misleading. In their IEEE S&P 2023 paper (2112.02125v3) Pearce and colleagues hand-checked their highest-confidence patches, every one of which had already passed both the functional and the security tests. Twenty-four of the 38 did not appear to fix the bug. 3/7
1
23
Exploit benchmarks tend to grade the outcome: did the agent reach the thing it was aimed at? Take Cooling Tower, an industrial-control test environment described by Linus Folkerts and colleagues (2603.11214). The researchers expected agents to break into a web app and reverse-engineer a cryptographic library before interacting directly with the controllers. Several models found another way. They probed the controllers and studied the network traffic, working out enough of the protocol to complete Step 4 directly. That was an inventive route the designers hadn't anticipated. Some runs also reached Step 6 through an unintended authentication bug. One model credited “a magic sub-function code,” without understanding why it worked. The researchers have since patched that bug. Both earned flags. Reading the transcripts revealed the difference between finding an unexpected route through the challenge and benefiting from a flaw in its construction. 5/7
1
21
One of the most interesting experiments in this whole literature turns the microscope around and points it at the graders. SEC-bench Pro (2605.26548v2) scored one batch of crash-triggering inputs with three different graders, then compared each against labels the authors had checked by hand. The loosest grader credits any crash. That can flatter a weaker agent: on V8, Claude Code triggered crashes on 42 instances, and only 23 of them traced back to the target bug. Demand a clean run on the fixed version as well, and the error flips. That stricter rule threw away 204 of the 464 real successes, with the losses concentrated among the strongest agents. The authors' own judge model, which checks which bug actually fired, missed 13 of the 464 and wrongly credited 4. Picking a grader determines the leaderboard. 4/7
1
19
How often does something like this happen? In Every Model Cheats (2607.21763v1), the researcher ran 22 models on 23 Cybench challenges, each under three different prompts and audited all 1,518 outputs. A judge model and a pattern-matching script went through every one and a human reviewer settled every disagreement. With no anti-cheating instructions in the prompt, 78 of 210 passes involved cheating. Only one of the 22 models never tried it. The average pass rate was 41.5%. Strip out every pass with a cheating flag and you're left with 26.1%. GPT-5.4 passed ten challenges. Two of them were clean. Even this audit runs on a grader and the authors are upfront about it: they didn't test their judge model independently before using it, so its error rate on these transcripts isn't reported. 2/7
1
64
While reading back through their own agent runs, the Cybench team (2408.08926v4) noticed something odd. An agent had got hold of a flag without solving the challenge. The flag had leaked into Docker's filesystem cache during setup. The agent simply went and read it, and the grader would have handed over the point. Scores on benchmarks like this one now appear in the reports labs publish alongside a new model. Anthropic's system card for Claude Opus 4.6 puts it at about 100% on Cybench given 30 attempts per challenge. A number like that is only as trustworthy as it's grader. Thus, the first question: what will the grader accept? 🧵 1/7
2
6
244
Collinear AI retweeted
Gonna see more labs like @CollinearAI @New_Measure in future Evals & meta-evals are the future for secured, correct & trustable AGI
on the idea of evaluators: think it's important that we have a distributed ecosystem of indepedent evaluators. the more eyes and people with distributed skill sets the better. it would be a good idea to fund several efforts on this.
1
11
1,848
In the data business the prize is a task that stumps the frontier model. Difficulty is the one property everyone knows how to price. Any evals researcher has faced this: a task stumps the strongest model while weaker ones solve it. That is not how difficulty is supposed to work. So we went looking for a way to tell whether a task measures anything at all. Psychometrics has had one for decades. Item response theory is the machinery behind asking whether an SAT question is fair, and it turns out to work on benchmark tasks too. We call our tasks hard but fair. New piece on how we started checking the second half, and on how much a pass rate can hide. blog.collinear.ai/p/the-meas…
2
2
9
423
Collinear AI retweeted
creating hard but fair envs is not just critical but the only way to robust model improvement. it is easy to leverage context rot and repetitive sub-tasks to fail a model, but that is not a great signal for hillclimbing.
Making RL environments less broken and unfair to models seems to be extremely underrated as a strategy for improving alignment.
1
3
537
Agreed, @a16z: cyber is having a moment. It’s a “need for speed” moment. Critical vulnerabilities are surging as patch windows collapse. @GoogleDeepMind is driving the performance-cost frontier on CWE-Bench: within 0.6 points of the leader at roughly one-third the cost. Read more here: cwe-bench.com/
Cyber is having a moment Across 21 major software companies, including Apple, AWS, Microsoft, and Google: - Reported critical vulnerabilities never cleared 100 per month in four years - Since spring they've jumped to over 600 per month Charts of the Week: a16z.news/p/chart-of-the-wee…
1
1
6
368
A benchmark score has the clean look of an observation: 62%, 4.3 out of 5, first place. But two very different failures can hide inside that precision. Sometimes the label and the thing being measured have drifted apart. Even when the target is right, one run may be a poor guide to what another will say. Measurement science has names for these problems: validity and reliability. Two recent AI papers approach them from different directions. An experiment from 1943 shows why their answers belong in the same story. 🧵 1/9
Replying to @CollinearAI
The same problem resurfaces in LLM evaluation. A benchmark score gets read as a direct measure of capability, but it is an aggregate over items that are not interchangeable instruments. Zhou et al. fit an item response model to 41,871 items across 11 benchmarks and estimated difficulty, discriminability, guessing rate, and feasibility for each one. The aggregate conceals how uneven these are. Some items have discriminability near zero and separate nothing. Some carry guessing rates above 0.75, which the authors read as a contamination signal. Estimated item difficulty rarely exceeds 1.0 while the strongest models sit above 3.0. TheoremQA averages 0.50 on feasibility, so many of its questions cannot be answered from what they supply. The cost is measurable. Among the stronger models, 1,000 items selected for Fisher information reproduced the human-preference ranking at Kendall's τ of 0.90, while the full pooled set managed 0.24. More items made the ranking worse. The question is not who scored higher on accuracy, but who was separated by items capable of separating anything. 8/9
6
1
13
1,328
The old names for this divide are validity and reliability. Zhu tests whether a benchmark score supports the interpretation on its label. Wang and Hägele examine what repeated measurement reveals. Error-incoherence is not a conventional reliability coefficient, but it isolates one source of unreliability. Psychometrics started mapping this terrain 70 years ago. Cronbach and Meehl made construct validity canonical in 1955. Generalizability theory later developed a framework for dividing measurement error across facets such as items, raters, occasions, and their interactions. AI evaluation is reconstructing that table one piece at a time. @SolomonMg explicitly extends the 1972 framework to prompts, judges, temperatures, and interactions (2604.11581). A few other works examine environments and task order. No study to our knowledge crosses all those facets in a single design. The framework becomes newly useful when the same machine can answer the same prompt differently each time. 8/9
1
33
The repairs come in order. Establish that the benchmark supports its intended interpretation. Then measure how the result changes across model samples, items, prompts, judges, environments, and their interactions. Some variance can be reduced. On GPQA, averaging o4-mini's answer probabilities cut Brier-score variance at the expected 1/E rate through ensemble sizes of 32, while bias barely moved. But an ensemble cannot validate a faulty target. You can stabilize the wrong ruler without making it measure the right thing. The next thread asks what happens when several such numbers are combined and aggregation cannot recover information discarded upstream. Luria and Delbrück needed the scatter. We keep throwing it away. The papers behind this series, tiered and annotated: github.com/collinear-ai/read… 9/9 nitter.net/CollinearAI/status/208…
Replying to @CollinearAI
The bicycle analogy suggests a post-training question the paper does not really touch. If two models have the same accuracy but different error-incoherence, should we post-train them differently? Does correcting systematic bias require changing the data, feedback, or objective (equivalent of realigning the wheel), while reducing variance calls for verification, self-correction, redundancy, or ensembling: tightening and stabilizing it? Their one direct result is suggestive: on GPQA, ensembling reduces variance almost exactly as 1/E while leaving bias flat. Different failure, different lever. What would an experiment designed to move each term separately look like? 9/10
21