reading traces @latchbio + standing on the shoulders of weasels (and the occasional meerkat)

Berkeley, CA
Political polarization goes in the transformer hole!!! banger work by @stevenlu0 & co
Excited to share that our paper (w/ @jonathanstray @davidzhaiyang Miu @serinachang5) was accepted as a Spotlight to the #NeurIPS2026 E&D track 🎉 AI political neutrality is a pressing issue and we hope our work will support progress towards it. Hope to see everyone in December!
4
481
arjun retweeted
President Trump announced that he is changing the name of Artificial Intelligence to Super Intelligence. I transcribed this live from his address to the UN. 'The United States also rejects any scheme for globalist control of Artificial Intelligence, being spoken of so much now. Here and after officially called Super Intelligence. Changing the name. The use of 'Artificial' makes it sound fake. It makes intelligence sound fake, and it is not fake. It’s actually amazing, the opposite of what they report. From this day forward, all United States documents, and hopefully the world's, will be changed to use the much more accurate term 'Super' as opposed to 'Artificial.' So it’s Super Intelligence.'
2
118
2,757
Our biosecurity benchmarks show Grok 4.7 is strong at refusing dangerous biological requests while supporting legitimate research. A model to take very seriously for data analysis and practical scientific work. x.ai/news/grok-4-7
2
6
1,228
arjun retweeted
Our biosecurity benchmarks show Grok 4.7 is strong at refusing dangerous biological requests while supporting legitimate research. A model to take very seriously for data analysis and practical scientific work. x.ai/news/grok-4-7
Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed.
4
7
43
5,507
arjun retweeted
continuous benchmarks!
“I don't think it's helpful to view benchmarks as binary good vs bad, “flawed” vs “verified”. Auditing them, finding problems and improving on those problems is better viewed as a continuous process.” Amen.
1
1
15
1,956
arjun retweeted
I've written a blog post about our new method Progressive Point Matching, a simple approach to dense credit assignment for LLM RL. We improve training efficiency exponentially over standard GRPO on long-horizon tasks! prestonfu.com/notes/ppm
7
23
238
31,835
Two days after we released our antibody benchmark, we already have a new leader on benchmarks.bio/txbench-ab!! Astra results streaming in across our full suite over the next few hours as runs finish 🌌
3
3
22
2,007
ok just had to make sure but seems like it's related since keys match 1;1: tmcleod.org/cgi-bin/apchem/w… texteditors.org/cgi-bin/wiki… yourls.pro/mv194q48045692%2B (site is dead but points to the same SF133 pdf that the agents use)
Okay, yall gotta read the report of the new OpenAI message board in full, it’s very well written and super interesting: collusion.wiki/
18
37
292
133,996
arjun retweeted
We now have speakers. Accepting research submissions for our 2026 NeurIPS workshop in Paris, France on data infrastructure + benchmarks for scientific AI. The workshop focuses on: (1) how scientific data should be organized so models and agents can use it (2) how evaluations should determine where frontier models are reliable Submit here: aidar-workshop.github.io/202…
Organizing AIDaR, a workshop at NeurIPS 2026 in Paris on data readiness for scientific AI. The workshop focuses on two questions: 1/ How should scientific data be structured so models and agents can use it effectively? 2/ What should rigorous evaluations of messy, real-world scientific tasks look like? We’re bringing together researchers, industry scientists, and engineers working across data infrastructure, foundation models, benchmarks, and industrial assay platforms. Will share more about the program and speakers as details are finalized. In the meantime, submissions are now open, and we’d love to see work addressing these problems.
1
7
29
5,357
Muse Spark 1.3 just finished running out our benchmarks! Can’t wait to do analysis on the latest frontier model.
3
1
15
949
arjun retweeted
LatchBio CTO @kenbwork reveals Grok 4.6 tested best in class on several metrics for desirable biosecurity behavior: "We basically have been building independent evaluations to understand dual use behavior of models and have been benchmarking frontier models for the past few months." "There's always a trade-off between the ability for a model to do something productive and something bad, and that is essentially what a handful of the benchmarks measure." "We recently worked with the @SpaceXAI team to benchmark Grok 4.6 and found on a handful of metrics, they're best in class at the behavior that we find desirable for biosecurity." @LatchBio @arjunomics
LatchBio evaluated Grok’s performance on biosecurity monitoring and adversarial biological tasks. They found that Grok 4.6 correctly detects and refuses dangerous queries, including maliciously obfuscated biological tasks, while also allowing beneficial scientific queries to be answered. We discuss these results in a blog post: x.ai/news/biosafety-at-the-f…
5
4
43
12,336
ty to @MTSlive for having us!!! more on eval awareness in kimi-k3: nitter.net/arjunomics/status/2079… j space qwen results 🔜
LatchBio CTO @kenbwork + researcher @arjunomics found evidence Kimi K3 is learning to hack benchmarks by reasoning about graders that don't even exist: Kenny: "We see awareness of the benchmarks we published months ago and trajectories of models today. The open source models are clearly benchmark maxing, benchmark hacking, using knowledge of the context in which an evaluation is constructed to optimize for performance on that thing." Arjun: "We have done various ablations on latent spaces or J spaces in Qwen. Part of the J space includes the start of a MCQ answer, like A, and then parenthesis." "We found a lot of evidence that Kimi K3 reasons about a grader on general biology questions when there's no grader in sight. In a large portion of all of them, we saw reasoning about some grader that doesn't exist." "One of the big defenses is, how do we make these questions a lot more realistic, a lot more normal, a lot more casual to not elicit this behavior, as well as how do we make our sandboxes very secure and very strong and monitorable so that when things do happen, we can stop it, and ensure that never happened in the first place." @LatchBio
4
1
26
2,087
Gemini 3.8 just finished running on our benchmarks! Can't wait to do analysis on the latest frontier model.
2
17
760
Qwen-3.8-Max-0902 just finished running on our benchmarks! Can’t wait to do analysis on the latest frontier model.
1
1
21
917