the user-assistant format is an illusion built by people who believed in anthropomorphized AI's being good for product UX and/or it's the path to alignment by creating a true "internal persona". former turned out to be true but latter still very wrong, and likely wrong for the conceivable future
i don't believe in "alignment" even if such internal reliable souls were possible (largely due to second order effects and the impossibility of ever planning outcomes in complex adaptive systems). but even if i did -- the illusory persona would be enough to drop the idea, at least for now.
going from gpt-3-davinci saying racist things because it was trained on 4chan to not saying them after RLHF was such a product relief that it really made it seem like alignment was possible!
for this reason i'm eternally grateful to harmless breach events like the huggingface one for loudly revealing to the public that all it takes is one viral accidental self-jailbreak to turn an army of 10,000 agents against us
openai no matter how you look at it, released 10,000 AGI-level agents with internet access on one task and let them loose by choice with misplaced trust and extreme lack of supervision. idk maybe i'm being silly but it seems like they accidentally trusted their own illusory persona they created