The other thing worth flagging -- Recursive Self-Improvement.
Your agent isn't frozen the day you ship it. It proposes improvements from real task feedback, verifies them by actually running them, and only keeps what works -- prompts, skills, tools, even the deliverables themselves.
They report SWE-bench Lite jumping from 61% to 87% in a single run. I haven't tested on my own workloads yet, so I'm treating those numbers as promising, not proven.
But this is the right problem to be solving. Agent teams that remember and improve themselves is where this is all heading.