I think today is the first time I've ever gotten an actually good hypothesis out of a model on an out of distribution research project. There was what I thought was a bug or sign of a problem and mentioned my hypothesis, but the model voluntarily countered with it's own, better hypothesis.
Also, it (in advance) determined a problem in a mechanism and placed a constraint on the weights to mitigate it. It also makes pretty decent hypothesis on what's wrong, runs isolations and ablations, etc in a way that isn't slop.
This has basically never happened before, this project in particular is very OOD and I don't think I've seen anything even remotely like it in public or arxiv and yet it is able to make meaningful contributions to the implementation.
I would say that we are like 10% of the way there to me being able to describe very high level ideas to Claude and it implementing it flawlessly, with all the requisite attention to detail, nitty gritty, etc. handled by the model.
It's not quite "there" yet, it still needed a lot of hand holding and I needed to tell it where to look for issues, but it's very clearly the first time LLMs don't feel like a stochastic parrot on ML research tasks. Maybe if I ran it on max it would be even better.
Sparks of RSI