Very cool results and interesting write-up!
A couple of thoughts:
1. At the core, this seems like an insanely difficult behavior prediction problem due to the quirk you mention. Given the limited information (i.e., you can't observe bullets), your success largely depends on your ability to accurately model your partner's behavior and exploit it. I'd expect modeling the opponent to play a larger role here than in, say, Chess.
2. Are you explicitly considering "diversity" (e.g., action dist entropy) in maintaining your pool of checkpoints? As you mentioned, the key here is probably figuring out how to make the policy explore a wide strategy space, but probably also which cpts to store.