game theory supplier collusion + negotiation simulations with buzz agents and our real RFQ data + synthetically generated data
been testing a fleet of agent personas in the environment in which it actually processes for our agentic otc desk - these channels are sattered and live across, email, telegram, whatsapp, slack, wechat...
we've been running sims on these to improve our understanding and the real supply demand negotiation flows we've been running that currently still have a HITL component but run mostly autonomous on the monitoring and response sides
with this we now can create a sealed slack workspace against other agents - supplier agents, demand side (buyer) agents and our own desk employees (ito agent employees / us directly)
the more data real and synthetic we supply them and base them off of existing / real scenarios the better each one gets, they (are supposed to but they mess up here sometimes), each has its own memory, KB DB etc, its own ledgers / inventories, fill history, warehouse / idle cost, cost basis, settlement timing etc.
in this case one didn't even know its own cost basis somehow but booked 48k gross like it accomplished something? also it didn't think to negotiate or collude - but when we told it we did negotiate and colluded with a different supplier and lied about a quote we got to bring the price down it did the reasonable thing and shut down flagged the ticket and stopped doing business with us indefinitely
ofc you cant say its a sim or leave any evals or benchmarks lying around otherwise - reward hacking
on the demand side we permute cluster combinatorics with different probabilities. hundreds of different cluster shapes depending on what the buyer is actually doing. inference and training pull completely different asks: interconnect, term, region, storage, sla. an inference buyer and a training buyer both say "8 h100s" and mean two different machines entirely
its been fun to watch them do their own side deals in individual dm chats, tracking PNL across them, undercutting, warehousing / idle capacity versus moving them across providers, when to cut a loss etc.
its also still nascent ofc and a function of the data we provide and make the systems better, currently the internal singular KB with agent system we have performs better as we seed these agents with more
we recorded this a few days ago and have actually evolved them quite a bit already
I tried
@jack's Buzz.
It's like Slack + OpenClaw + Herdr + but with some really unique features that people are sleeping on.
The video below shows how it works, and some of my thoughts on the process and platform, e.g.:
- Create and interact with agents on top of any harness (claude code, codex, pi, etc.)
- Choose which models agents use, including local ones
- Agents can delegate work and work in parallel in git worktrees
- Agents are first-class citizens and work like humans (creating channels, delegating, access to chat history)
- You can share AI compute within a community
- It's completely open-source and decentralized
Things I like:
- Delegating work in chat feels natural: tag an agent, it replies in a thread with status updates as it e.g. compiles, commits, and deploys.
- Shared compute: relay owners can share local compute with members, so a community could pool funds for one beefy machine running a local model and everyone uses it.
- It's built on Nostr, an open protocol already tied into Bitcoin Lightning so I can imagine communities tipping each other or paying for compute/agent tasks with instant zero-fee micropayments in the future.
- It ties together things like OpenClaw, an agent manager, and Slack-style chat into one tool.
Things I didn't like:
- You can't see what the agent is doing in a terminal. The activity view exists, but if you're used to watching a session run, this UI feels a bit abstracted. A terminal view would be great.
- It feels slower than running a session in Claude Code, though no evidence to back that up. For that reason I found myself doing one-off tasks in the terminal instead.
Verdict:
- I really like it so far and can genuinely imagine working with a team this way.
- It doesn't feel ready for big, complex tasks yet. For shallower tasks, it's perfect.
- The shared compute + Nostr/Lightning angle is what really separates it from every other agent manager for me, and I think that future is coming.