Our AI alignment competition with
@AureliusAligned is live.
Join us and help steer society towards a safer, healthier AI future.
- Participants steer a real LLM toward a target concept using a single learned direction in its internal representation space, based on research regarding the Golden Gate Claude, opened up to the whole network
- Each round, participants submit their direction and it's scored on how strongly it steers the model; the best steer takes the round.
- You are free to engage in a range of solutions, from classic diff-of-means to sparse autoencoders, so participants can bring whatever interpretability method they trust.
It is a hands-on contribution to AI interpretability, with no retraining or prompt tricks, This is all about the model’s internals.