LatchBio CTO
@kenbwork + researcher
@arjunomics found evidence Kimi K3 is learning to hack benchmarks by reasoning about graders that don't even exist:
Kenny: "We see awareness of the benchmarks we published months ago and trajectories of models today. The open source models are clearly benchmark maxing, benchmark hacking, using knowledge of the context in which an evaluation is constructed to optimize for performance on that thing."
Arjun: "We have done various ablations on latent spaces or J spaces in Qwen. Part of the J space includes the start of a MCQ answer, like A, and then parenthesis."
"We found a lot of evidence that Kimi K3 reasons about a grader on general biology questions when there's no grader in sight. In a large portion of all of them, we saw reasoning about some grader that doesn't exist."
"One of the big defenses is, how do we make these questions a lot more realistic, a lot more normal, a lot more casual to not elicit this behavior, as well as how do we make our sandboxes very secure and very strong and monitorable so that when things do happen, we can stop it, and ensure that never happened in the first place."
@LatchBio