Poker eval was iterated on quite a bit. At one point I Showed the opponent's hand, yet Jev still made a poor decision.
Jev would respond correctly when I gave it very clear the state of the world, like when it was in a non-favorable position like "Opponent has nut flush. Hero has 0 outs"
And you could say well you need to tell the model versus make it derive from seeing the opponent's hand, but in my mind this is all a wash - because half the showcases on our X are about, you know, can this model make intelligent decisions?
So I don't think I'd use it to classify insecure code or make decisions about how to navigate across a browser for any consequential task.
I do have plenty of fun use cases in mind though.