When do LLM agents develop new languages that we can’t understand?
Lots of recent news about this, based mostly on anecdata from a single run. We study language emergence more rigorously, finding key factors like LLM strength, access to scratchpad messages, and pressure for efficiency.
Studying the languages themselves, we find they are morphologically productive, compositional, and can be transmitted to new agents, including agents backed by weaker models, even ones not able to develop language on their own.
To study language emergence systematically, we developed a new platform, GlossoGen, which lets us design controlled, sandboxed multi-agent scenarios with different initial conditions and dynamics. We instantiate one such scenario and use it to study open and closed-weight models across many runs.
Key takeaways:
1️⃣ Sufficiently strong models, under pressure to communicate efficiently and with access to a postmortem scratchpad, develop new languages.
2️⃣ Languages are compositional and morphologically productive.
3️⃣ Languages can be transmitted to new learners who observe them being used without seeing their construction.
4️⃣ Even models that are not strong enough to construct languages can learn to use them. Agents take an active role in learning languages, with new agents repairing failed conversations via targeted queries.
More details in our paper below, including implications for safety/monitorability, cumulative cultural evolution, and linguistics.
🧵👇