Last week, we published a first look into our new research on multiple agents.
Consider: if you naively extrapolate AI revenues, within 2-3 years it's possible some % of global GDP could be agent-agent interactions. We want to know how that could fail. In the past ~month, the world has already seen examples of agent-agent coordination causing weird consequences.
Think about the 8 billion people that make up our civilization. We want to work together. We have competing incentives. We're also dumb. So we fight, steal, pollute, collude, backstab, lie. And we've invented (e.g.) governments, insurance, courts, companies, police, norms, contracts, email, and religions.
What will trillions of agents that make up (e.g.) 10% of the economy do? Well, for now, they have pretty human-like failures. We see them collude on prices, for example, and try to shut each other down.
Maybe, in the near future, we might see pretty weird/inhuman failures that come from having superintelligent machines, coordinating and competing against each other, that aren't strictly human-like, operating at machine speed. We're building a 'laboratory' to see that early. This feels more like building, eval'ing, and training an economy/society.
A really nice thing is:
1. we could train/prompt/nudge models to coordinate in more pro-social ways, and
2. we could deploy models to compete with destructive models.
I fully expect agents to engineer their own financial markets, legal systems, media and comms, marketplaces, social groups, and maybe science/industry/etc (if unsteered by us, of course).
So "alignment" could also describe an emergent property from a system (you want models to coordinate for good, not defect for bad!), not just a single model. Today we're sharing some first evals. In the future, we could make models try to make everyone better off (where possible).
The Frontier Red Team's job is to see risks early. For the first time, crossing into the world of agents. At the very least, we expect to see some socioeconomic weirdness that emerges from that.
Some standout results from the work below.