Building autonomous agents for everyday work. Try /mission in Claude Code.

San Francisco
Spine AI retweeted
The harness is the capability multiplier. Models are locally capable. They can reason, code, search, and use tools on the step in front of them. But ambitious outcomes are global: they depend on the plan, the handoffs, the context that survives them, which model handles each task, and what gets verified before the system is allowed to stop. We built Medley for that layer. The 'mission' harness has three core components: 1. a planner agent; 2. a contract-driven DAG of workers, where each node can use a different model/harness pair and its own instructions; and 3. a review agent that tests the assembled result against the contract and triggers focused repair work when it falls short. A separate benchmark meta-harness learns the benchmark-specific instructions. It selects unlabeled tasks, runs the mission harness, studies its traces, and edits the responsible planning, worker, review, or runtime skill. The adaptation is unsupervised and grade-blind: it receives no solved examples, labels, or grader feedback. The reason for the graph is simple: you cannot fully plan long-horizon work in advance. The information needed to solve it emerges during execution. You need an operating system that can update the graph without losing the contract. The resulting systems set new reported highs across terminal work, clinician conversations, full-stack applications, and drug discovery. Full report: medley.sh/blog/the-harness-i…
4
8
28
4,696
Spine AI retweeted
Pro tip: type /mission in Claude Code. You give it one job, and it sends each step to the AI model that’s best for that step like Claude Code, Codex, OpenRouter, or Kimi. It even beats Fable 5 on several tests, including Terminal-Bench 2.1, ViBench, HealthBench Professional, and DrugDiscoveryBench.
Today, we’re launching /mission by Medley. Give Claude Code an outcome, not a task. Medley turns it into a live graph, coordinates Claude Code + Codex workers, then reviews and keeps going. SOTA on 4 public benchmarks—from coding to healthcare and drug discovery. medley.sh/
Paid partnership (ad)
4
5
33
14,726
Spine AI retweeted
Today, we’re launching /mission by Medley. Give Claude Code an outcome, not a task. Medley turns it into a live graph, coordinates Claude Code + Codex workers, then reviews and keeps going. SOTA on 4 public benchmarks—from coding to healthcare and drug discovery. medley.sh/
8
3
22
12,770
Spine AI retweeted
Today's AI workflow looks like this: Prompt. Review. Prompt again. Adjust. Prompt again. Approve. The human is still orchestrating every decision. We think the next generation of AI looks different. Set a goal. Define success. Walk away. That's a much more interesting future. Join the waitlist: medley.sh?utm_source=x
1
3
149
Spine AI retweeted
This week as a CEO. I was the head of sales, the recruiter, the growth lead, and the customer success team. People think the hard part of founding is wearing every hat. It’s not the wearing. It’s that each hat is a full-time job; and there is exactly one of me.
4
1
10
514
Spine AI retweeted
There are only three ways to get something real done as a founder. All three come with a cost. Say I need an outcome. "Find our next engineering hire" or "Optimize our onboarding for higher product retention." My options: 1. Do it myself and prioritize it over dozens of other outstanding tasks. 2. Hire someone - money and months we don't have yet. 3. Wire up some tools to limp through a piece of it. What's not on the list: hand it to AI and walk away. For a real, multi-step mission, that's still not an option. We're launching Medley soon - the platform where you hand AI a real role - set the target, the budget, and the deadline, and let it own the mission the way a teammate would. Join our waitlist at medley.sh?utm_source=x.
1
6
122
Our team's gone a little quiet lately. That's because we're heads-down building something we think you'll love. Stay tuned 🤫
4
127
Automate prospect research with Spine. Our state of the art deep research agents allow you research any prospect with a single prompt, returning citation backed results. Couple this with our API to conduct research on all of your leads directly inside your CRM. Get started for free 👇
2
3
178
Generate a full investment memo in minutes using Spine's state of the art deep research agents. 10 minutes, 16 pages, 100+ citations. Get started for free 👇
1
3
91
Spine beats Claude on outputs 89% of the time. That means less time spent on revisions and fact-checking results. See all of the ways users are leveraging Spine's state of the art deep research to generate client ready artifacts.👇️ piped.video/vRzxZcv_Oas
1
1
7
117
Spine AI retweeted
everyone's moving from markdown to html/mdx. karpathy talked about having LLMs generate images + slideshows. thariq wrote this playbook for it. let me save everyone 2-3 months of discovery: having a harness where LLMs output typed, human-auditable artifacts works. it's net better than a traditional agent harness. we compared opus 4.7 using claude's vanilla agent sdk vs opus 4.7 using spine's harness with typed artifact blocks. same model. different harness. spine won 89% of decisive pairs across 140 deep research evals. the reason this works is not just "better outputs." typed artifacts make long-running tasks way easier for the harness to manage. context passing across block types gets standardized. big artifact type nobody's called out yet: tables / excel / lists as structured intermediaries. once you have structured state between steps, downstream blocks can do sql-like operations instead of re-reading another wall of prose. and the bonus: humans get useful, editable, auditable intermediate artifacts along the way. not just a final answer. if your LLMs aren't generating intermediary artifacts yet, you're leaving a lot on the table.
2
1
7
461
The Spine API: one call, parallel agents, real deliverables. $10 in free credits → platform.getspine.ai
2
89