The ledger for your AI software factory. Turn time and tokens into outcomes.

San Francisco
Pinned Tweet
Mainstream discourse about agent effectiveness tends to focus on model selection. We looked at complete agent trajectories across 103 engineering teams and found that prompt clarity, environment readiness, and quality stewardship are heavily correlated with the right outcomes. Read the report: span.app/research/agent-effe…
1
6
911
.@trygenius_ai had doubled engineering output with coding tools in six months. Then Span surfaced a bottleneck that directly impacted their velocity: a cross-team code review backlog that was slowing the whole pipeline. The result? A 25% drop in time to first review.
1
4
532
Teams with stronger quality stewardship complete 39% fewer review cycles per PR. Every extra cycle is time and tokens spent re-reviewing code instead of shipping the work that moves the business forward. Span helps teams point time and tokens to what matters. Read our full Q3 report: span.app/research/agent-effe…
3
92
Span retweeted
Human review is declining. AI code review scaled more than 6x since Jan 2025. Defect rates did not meaningfully move. AI reviewer bots now produce 28% of PR review comments across our dataset, up from 5% in January 2025. Today, 96% of companies use at least one AI code review tool, with the median using two. That growth is happening as review demand outpaces human capacity. Over the last nine months, merged PR volume increased 53%, while the reviewer pool grew only 10%. At the same time, the share of PRs merging with no human review rose from 13% to 19%. The obvious question is whether that changes software quality. So far, we do not see evidence in either direction for equally sized PRs. PRs merged with no human review were about as likely to be associated with a bug within seven days as comparable human-reviewed PRs. Mostly AI-written PRs were also about as likely to be associated with a bug as mostly human-written PRs of the same size. The one exception is very large changes: in the largest tenth of PRs, mostly AI-written changes were associated with bugs about twice as often as comparable human-written ones. We cannot yet tell whether that reflects the AI itself or the kind of large change engineers still write by hand. My interpretation: AI review may be helping teams absorb growing code volume without an obvious decline in near-term quality. But we also do not yet see evidence that more automated review reduces defects, and this analysis does not capture the broader benefits traditionally associated with human review. One factor had a much stronger relationship with bugs: PR size. Larger PRs had a higher probability of being associated with a bug. I suspect practices such as smaller changes, stronger testing, type safety, and focused bug bashes remain your most important quality levers.
2
5
248
Everyone is talking about graph engineering. In simple terms, it means designing how several specialized AI agents hand work off to each other, instead of making one agent do everything in a single loop. This feels very similar to designing how a team works. Roles and responsibilities, handoffs between functions and teammates. Once you've designed the graph, you implement different sub-agents to do different "jobs" / parts of the workflow. Writing a PRD or a tech spec probably requires frontier intelligence (Opus 5, etc.). Once the tech spec is broken down into granular tickets, a smaller, cheaper model probably works (Sonnet 5, or Haiku 4.5 for the really mechanical stuff). Time for review? If it requires security review, you probably have a dedicated security review agent again using frontier intelligence, or even better, a fine-tuned security focused model. The analogy to team design is almost one to one. The PM does the PRD, the Tech Lead does the tech spec, the L1 new grad implements simple tickets, the security team does sensitive security review. If this is how software is developed in the future, you can pin specific model versions to different sub-agents / nodes in the graph. If you want to, you can run the same task through different "versions" of the graph with different models (A/B test) and eval which hits the cost/quality tradeoff you're comfortable with. Won't this be more powerful than using a model router that has no knowledge of your graph?
2
1
4
186
The AI coding debate is focused too much on model selection - just see all the model router launches this week - Ramp, Cursor, etc. We studied 103 engineering teams and found big levers are also on the dev platform and human side: • +1 pt prompt clarity → ~27% lower token cost per merged AI line • +1 pt environment readiness → ~88% higher turn yield • +1 pt quality stewardship → ~39% fewer review cycles Prompts, environments, and stewardship shape every task and are things you can coach your team to improve today. Full report + a practical field guide below. Would love to hear what matches or contradicts your experience. span.app/research/agent-effe…
1
2
4
174
Two people joined @Span_App last week and by the end of their first week, they'd both had something that they worked on ship to production. They were a little unsettled by it. Both came from 300+ person companies, where the norm is three to six months before anything ships. Here's why I like to operate this way: spend a month building before it's in front of anyone, and you might pivot before you ever learn whether it had traction. That's a month you can't get back. Ship a rough version instead. If there's traction, expand it. If not, you've lost days, not months. A big company can afford slow bets, a startup can't.
1
4
146
Span retweeted
What does our SVP of Engineering @IcchaSethi have to say about maturity frameworks, evaluation processes, and more? 🤔 Find out on the latest @Span_app episode here: bit.ly/4gFvTvR
5
12
447
Span retweeted
Fable produces fewer defects for just $0.08 more per shipped AI line than Opus. That’s what we observed with Fable 5 versus Opus 4.7 even though Fable handled a more complex task mix (94% vs. 83%). Composer was also a clear cost outlier on the Pareto curve. Span makes this analysis possible by connecting agent trajectories to merged code and downstream defects. Early results but still interesting. What is one defect worth to you and your customers? We’ll repeat this for GPT-5.6 Sol and control for confounders for Composer if there's interest.
4
18
133
197,062
Span retweeted
Excited to share that I’ve joined the @Span_App design team to help teams move faster with less friction. Working at an earlier stage company has always been something that I wanted to do. And this moment felt like the right time to take the leap. Thanks to @erondu for spinning the block so we can finally work together. 🔥
6
1
18
816
Tl;dr: We deleted our product homepage. At a previous org, we had a KPI to increase key actions, so we threw a bunch of stuff onto the product homepage to drive it. It technically worked, and the number went up enough to call it a win. At @Span_App, I didn't want to chase an internal number like that. What we actually cared about was shortening the time it took new users to hit an "aha" moment, the point where the product finally clicks and they get why it matters. So we did the opposite of loading up the homepage - we deleted it entirely. New users used to land on a generic dashboard, think "what is this product," and leave before they ever reached the parts that mattered. We dropped them straight onto the page they already use most, with three prompts pointed at the questions people actually come to us with: what should I pay attention to this week, how is AI changing our output, catch me up on X. Then the numbers we weren't chasing moved anyway. Key actions per active user rose 208%, day-7 return climbed 213%, and key actions per week went up 61%. At the previous company we chased the metric and nudged it. At Span, @erondu and I focused on the experience, ignored the metric, and ultimately increased the number 3x.
3
14
234
Quick demo of @Span_App's new product, Src! Customers are already using Src to: - Understand ROI of AI tokens on projects and types of work - Run scenario and capacity planning - Kick off cost capitalization cycles - ... and more!
5
13
878
Which teams are using AI most effectively? Why did defect rates increase this month? Engineering leaders currently spend hours digging through a dozen dashboards to answer those questions. They shouldn't have to. Src, Span's AI agent, makes your entire engineering system instantly queryable with automatic reports and a full audit trail that lets you verify the answers. See Src in action: span.app/platform/src
1
5
8
1,119
Different parts of a product, and different kinds of companies, are meant to move at different paces. @TownAI co-founder and CEO @jgreze stopped by our Stack Trace podcast to discuss how to match shipping velocity with what you're building.
2
2
11
2,878
Stop talking about tokens and come waste some the old-fashioned way. Span is hosting Tokenmaxxing Arcade Night on July 1st at Thriller Social Club in SF, just 10 minutes away from @aiDotEngineer World's Fair. From 6-9pm, spend tokens all you want. We won't tell your CFO.
2
2
16
344,723