co-founder @columntax (acq). prev @waymo @pioneerdotapp @google. see my popular tweets using this side project i made: toptweets.by/@michaelrbock

San Francisco, CA
1/ After 5 years, I’m proud to share that @ColumnTax has found a new home. We’ve been acquired by @AiwynAI. I couldn’t be more sure this is the right move for our business, tech, and team. These pics are the moments we started & sold the company:
17
6
82
14,335
pretty cool: Dario predicted AI solving taxes in ~2024 we've been tracking progress on that since 2025 with TaxCalcBench: github.com/column-tax/tax-ca…
4
14
2,915
14 months ago I built v1 of TaxCalcBench (eval to test AI model's ability to calculate tax returns) based on the >1M returns we filed @ColumnTax Now every tax startup (there were none in 2021 when I started Column Tax, but many many now) using it as a standard:
4
1
33
2,886
congrats @nedwize on the results!
1
6
191
Claude Opus 5.5 is the #2 model in the world at filing taxes according to TaxCalcBench. Unfortunately, it's also 10x more expensive than the new #1 model: GPT-6 Sol.
8
5
111
7,600
But Anthropic made the bigger generational leap: Opus 5 → 5.5: 42% → 62% (+10 returns) GPT-5.6 Sol → GPT-6 Sol: 62% → 64% (+1 return) And without web search, Opus 5.5 wins: 38% vs. 26%.
1
6
240
if you have any friends that are lawyers, have any of them ever talked about using Harvey?
2
2
744
hedge fund aside, it's kinda wild to sit in 2026 and read situational awareness (written in 2024) and it reads like a history book
10
658
Last week was pretty wild for AI: there were four major AI models launched in four days. All four are already on TaxCalcBench, including GPT-6 Astra. But the surprising result: Meta & Gemini are back in the race. Last week’s release calendar: - Tuesday: Anthropic released Claude Fable 5.1 - Wednesday: Google released Gemini 3.8 Flash - Wednesday: Meta released Muse Spark 1.3 - Friday: OpenAI announced GPT-6 Astra All four releases make some version of the same promise: better coding, better agents, and better long-horizon knowledge work. So we tested all of the new models on 50 hyper-realistic federal and state tax returns: - Gemini 3.8 Flash w/ web search: 52.5% (21 of 40 completed returns) - Claude Fable 5.1 w/ web search: 46% (23 of 50) - Meta Muse Spark 1.3 w/ web search: 26% (13 of 50) - GPT-6 Astra w/ web search: 60% (20 of 50) The headline is Gemini. Compared to 3.7 Flash, 3.8 is almost double as good at tax return calculation. The upshot of all these releases? Release announcements tell you what a model is designed to do. But they don’t tell you how it performs on your task, with your tools, at your reasoning budget. The model release cycle is now faster than most companies’ evaluation cycle. That is a problem. For tax filing, we built TaxCalcBench so we can test new models within days, inspect every failure, and share the full results publicly. Whatever domain you work in, you need the same thing. Full results: github.com/column-tax/tax-ca…
3
10
1,647
Not sure if this will be useful unless you're a CTO, Head of Engineering, or maybe Engineering Mgr, but figured I'd share: At some point (this has probably already happened), you won't be able to individually lead every project yourself. At that point, you'll want to make sure that your DRIs (you have DRIs for every project, right??) are executing well, and step in to help when things are off-track. This is honestly the most-fun part of the job: you're a manager now so you don't get to code or solve fun technical problems, but you do get to scratch the problem solving itch here: what's wrong with the project and how do you get it back on track? In any case, here's the process I used, which worked really well - hope it's helpful: michaelrbock.com/execution/
16
676
I just advised a startup (seed stage) to help them get acquired. They raised about $5m, had a team of ~6, and got acquired by a hot AI company (can't say who yet, but you know them). The things that came up most often: - Storytelling for getting acquired is different than storytelling for raising. When raising VC, you have to tell a "perfect" story about how you'll become huge. When getting acquired, it's ok to have "spikes" - some things that are bad are OK - as long as your spikes fit well into the acquirer. Tell a story about 1+1 = 3. - Momentum is everything. Keep telling & reinforcing the story every day/week all the way through close. - Related: a million things can kill the deal (legal, finance, an exec wakes up one day and changes their mind), so momentum around your combined story is the thing that can help pull your deal through those inevitable issues that will come up. - Legal issues _will_ come up. Expect 2-3 months & hundreds of thousand of dollars of legal work. Be as buttoned up as you can be with your contracts/docs/etc. through the life of your company. - Your negotiation leverage is based on your storytelling ability. Does your potential acquirer believe they can't survive without you? - Most acquisitions only get done if there is a huge need from the acquirer. The metaphor we use is: "is there a hole in the lobby? does the CEO need to walk around it every day on the way into their office?" - your startup should fix that hole. - Most companies can only acquirer 1 (maybe 2) companies a year. Why should your company be that 1 big bet? Good luck!
5
1
99
17,538
This is something pretty interesting I'm watching right now: the race between U.S. vs. China for the best open model. Meta says it will release an open-weight version of Muse Spark 1.2 soon. China's Moonshot AI has already released the weights for Kimi K3. So we put the two models through TaxCalcBench: 50 hyper-realistic U.S. federal and state tax returns. Who won? It depends on what you measure. Without web search: - Meta Muse Spark 1.2 calculated 4 of 50 returns perfectly. Kimi K3 calculated 3. - Under lenient scoring, Kimi won: 6 correct returns to Meta's 4. - By-line, Kimi won by more than 10 percentage points: 65.49% to 55.35%. Meta produced one more perfect return. Kimi got much more of the tax forms right overall. Then we gave Muse Spark web search. It jumped to: - 8 of 50 perfect returns - 11 of 50 under lenient scoring - 68.28% line-level accuracy That last number is only 2.79 points above Kimi, despite Kimi running without web search. The upshot? There isn't one winner in the U.S. vs. China open-model race. Plus, Kimi K3's weights are available today. Meta has announced Muse Spark 1.2's open-weight release, but it hasn't happened yet. Model comparisons need more than one score. The winner can change with the metric, the tools, and even what you mean by “open.” All 150 model outputs behind these three rows, plus their evaluation reports, are public in TaxCalcBench. Full results and model releases: github.com/column-tax/tax-ca…
5
656
Before starting @ColumnTax, I had never hired anyone. By the time we were acquired by @AiwynAI, I had hired 50 amazing folks. I used the same algorithm for every hire:
2
19
1,641
Before I made my first-ever hires, I was lucky enough to learn this 4-step algorithm that I've personally used for every hire I've made since. I also forced every hiring manager who worked for me to use it too.
1
3
136
This blog post includes the algorithm, examples, and more on every part of the hiring lifecycle in the age of AI: sourcing, interviewing, references, comp negotiation, and closing. Read it here: michaelrbock.com/hiring/
4
113
Anthropic just released its 4th model in 2 months: Opus 5. It's half the cost of Fable 5. And surprisingly, for knowledge work tasks like tax filing: Opus performs even better than its more expensive counterpart. We tested Opus 5 on TaxCalcBench. Here are the results vs. Fable 5: - Opus 5 w/ web search: 40% tax returns computed correctly (strict) - Fable 5 w/ web search: 35% But only when we turned thinking all the way up. We ran both models with web search on the same 50 realistic federal and state tax returns at five thinking levels: - Lowest thinking: Fable won, 18% to 10% - Low: tie, 16% to 16% - Medium: tie, 24% to 24% - High: Fable won, 32% to 28% - Ultrathink: Opus won, 40% to 34% At ultrathink, the gap was even wider under lenient scoring: 56% for Opus versus 44% for Fable. So the headline isn’t simply “Opus is better than Fable.” It’s that Opus earns its lead when both models are pushed to their maximum reasoning setting. At every lower setting, Fable matched or beat Opus on strict accuracy. All of the results and model outputs are public in TaxCalcBench: github.com/column-tax/tax-ca…
3
12
797