ceo @matrix_build and @flowith to shape the interface for AGI. tech optimist, AI engineer, designer. previously founded X ACADEMY and Realm.

San Francisco, CA
Pinned Tweet
We've been quiet for 8 months. Because we've been busy building the infrastructure for a 100% agent-led companies. Still in the beta phase, but I can't hold back this preview. Introducing Matrix, where anyone can launch a 0-Person Company that actually earns. And yes, Matrix beta already achieved SOTA on the frontier harness, matching Fable's performance.
165
130
795
407,735
mid-autumn, a full moon, and a stanford comeback win. pretty good night. @StanfordFball
4
549
introducing flowith canvas powered by opus 5.5. we used to spend $10,000 on studio work for this kind of piece. we did it twice. opus 5.5 just made it for $10 in tokens in about an hour. one shot, zero extra apis. hard to overstate how fast things are moving.
10
6
59
8,728
wild that the thing with the potential to kill seedance isn't another video model, but an llm. opus 5.5 is unreal.😭
10
1,298
cloud llms got us here. 2026 is when local models take over.
4
14
817
spent the whole day playing with @typesafeai jev. built some pretty wild demos that would easily farm engagement if i posted them right now. instead i tried actually shipping it to production. swapped out gemini 3.8 flash/lite for jev on non-visual DOM execution in matrix browser use and modified our harness. real websites break it almost instantly, it's general reasoning just isn’t ready for browser use yet. the @browser_use jev demo is super cool, but that page was basically an ai-tailored sandbox, not the real messy web. the timeline is increasingly optimized for cherry-picked shock value over real utility, which ends up misleading a lot of builders trying to ship actual products. jev is a great concept and an impressive model. genuinely looking forward to the open source weights and multimodal release.
9
5
79
7,049
gpt 6 astra on codex feels noticeably degraded and usage limits took a dive. looks like it's time to move back to claude code/grok. @thsottiaux did the last fix not land or is this the intended behavior?
11
112
7,014
Derek Nee retweeted
introducing ═══════════ 𝚌𝚏𝚘.𝚊𝚒 ═════════════ tl;dr 1/ runway is now cfo.ai 2/ you can hire @arithecfo and he can start TODAY 3/ you can give @arithecfo a test drive by replying to this thread with any ridiculous thing you want him to model for you, and he'll come back to you with a beautiful, detailed, fully wired financial model ... in 30 minutes or less now, story time: 🧵 1/n
158
62
668
279,793
Derek Nee retweeted
GPT-6 Astra VS GPT-5.6 on animation coming soon to flowith.io
This is GPT-6 Astra. Anything you can do on a computer, Astra can do for you. Fast.
255
293
6,177
12,661,163
burning man is probably one of the last places left on earth completely airgapped from ai. dust was so brutal i barely took the leica q3 out, but managed to grab this.
2
1
18
2,639
the era of launching ai startups through x influencer campaigns is over. just watched a team pull 2m impressions on a launch and convert 300 followers. influencer feeds are so flooded with paid promos that users have total ad blindness now. roi is basically zero. founders need to stop burning budget here and find real distribution again
11
2
74
24,128
Derek Nee retweeted
flowith 开启新一轮的招募了,欢迎大家投递到 join@flo.ing ,并且带上感兴趣的岗位和办公地点。本次上海和 SF 均有比较多岗位开放,也欢迎大家可以找我私聊、推荐,推荐入职成功的朋友们会送出未来一年的 Agent token 额度。 简单补充下我们:Flowith 是一家面向全球的 AI 与智能体科技公司,团队位于上海与旧金山,旗下拥有 flowith 、Matrix 等核心产品,持续探索人与 AI 的下一代交互方式。目前,Flowith 已累计服务全球近五百万用户,并获得红杉资本、Vertex 、江远等一线基金千万美元投资。 除了创新的 C 端产品外,我们已与多家行业头部企业及顶级战略伙伴达成合作,将领先的 Agent 工程能力应用于制造、电商、消费品、娱乐、体育赛事等场景,支持智能体参与高复杂度、长链路的真实业务流程。并获得 微软、NVIDIA 、字节等全球科技公司的合作与支持,持续推动 AI 从前沿技术走向规模化生产应用。 下方评论区放了非常多详细介绍、招聘岗位,欢迎大家转发推荐,也在转发评论的朋友们里抽几位朋友送些会员。 我们在上海的办公室也欢迎大家参观,应该是在上海最独特的 AI 公司。
159
10
120
53,861
Derek Nee retweeted
we gave Matrix an iPhone and one job: find how indie developers get users. it built a TikTok research department, controlled the phone, searched the app, and returned the findings as a CSV. your next business researcher might already be in your pocket.
3
3
15
1,294
Derek Nee retweeted
OpenAI granted us a serious pile of free credits. so we're passing it on: every GPT model in Matrix is 90% off for a limited time. GPT 5.5 through GPT 5.6 Sol, all supported. that's GPT 5.6 Sol at $0.50 per 1M input tokens. receipts attached. come use it. heavily.
2
1
11
934
only 5 million people know ai doesn't have to be linear. flowith. branch out.
4
4
27
4,203
The best screenplays aren't written anymore. They're filed in federal court. A day ago, Apple sued a former engineer who allegedly grabbed the files, texted "LOL" about it, and joined OpenAI's hardware team the same week. A Matrix agent read the complaint and directed the film it deserved. Every beat from the filing.
Made with AI
6
2
46
10,337
GPT-5.6 Sol is already a very capable model. but its highest-leverage role may be as the CEO of models built by its competitors. put Sol in the CEO Office. give Fable 5 a permanent Strategy + Review department. let Grok 4.5, Luna, Codex, Claude Code, and browser workers execute inside the departments that own the work. Sol alone is a model. Sol + Fable + low-cost workers + Proof + Memory is a company. done right, the company can feel 10x smarter than Sol alone - and cost less per verified deliverable. the 10x is not hiding in another benchmark chart. it is in the architecture. the useful metric is not intelligence per API call. it is verified work per human operator. the 10x does not come from spending 10x more tokens. it comes from division of cognitive labor: > Sol chooses direction and allocates ownership > Fable attacks the plan's blind spots > durable departments preserve domain context > worker fleets execute scoped work in parallel > independent models catch correlated mistakes > Proof rejects output that only looks finished > Memory makes the next loop start ahead one model no longer has to plan, execute, criticize, remember, verify, and scale at the same time. that is where the intelligence multiplier comes from. and the cost advantage is just as important. at current API list prices per million input / output tokens: > GPT-5.6 Sol: $5 / $30 > Claude Fable 5: $10 / $50 > Grok 4.5: $2 / $6 > GPT-5.6 Luna: $1 / $6 do not pay a $30/M-output CEO to do extraction. do not pay a $50/M-output strategist to sit in every hot path. and do not ask a $6/M-output worker to make the few decisions where one mistake changes the company. the right split is: frontier intelligence where judgment changes the outcome. low-cost intelligence where the work is clear and parallelizable. cross-vendor review where correlated mistakes are expensive. proof everywhere. this is the company architecture underneath it: Workspace -> CEO Office / GPT-5.6 Sol -> durable department hierarchy -> Strategy + Review / Fable 5 -> Product / Engineering / Growth / Research departments -> same-owner worker seats / Grok, Luna, Codex, Claude Code, browser -> Criteria / Proof / Check-in -> department outcome reply -> CEO synthesis -> state + memory update -> next wake The CEO Office is the primary department and the user's entry point. Sol does not become a god-agent with every file, tool, permission, and task. It becomes the executive layer. It resolves ownership, routes work, creates an owner when none exists, arbitrates conflicts, follows up, and synthesizes the company-level result. Fable 5 is not a stateless side call. It becomes a durable Strategy + Review department with its own memory, skills, Key Results, task history, and accumulated taste. Product, Engineering, Growth, Research, and child departments can each choose the model that best fits their work. inside those departments, workers are temporary execution seats: > Grok 4.5 for efficient scoped execution > Luna for high-volume fan-out > Codex for repo-native GPT coding work > Claude Code for Claude-native coding and review > browser / computer workers for workflows that need visible proof the key distinction: a department message moves ownership. a worker parallelizes the current owner. if CEO Office creates Engineering and then secretly performs Engineering's work with CEO workers, that is not delegation. it is org-chart theater. Engineering should own the Key Result, decompose it into Tasks, dispatch its own workers, judge the returned artifacts, attach Proof, and send the outcome reply. workers return artifacts and traces. departments return outcomes. CEO Office returns one coherent company answer. that is the loop: company direction -> Workspace Objective -> department ownership -> Key Result + proof-bearing Tasks -> multi-model worker execution -> Criteria + Proof + Check-in -> outcome reply -> CEO synthesis -> memory update -> next wake the best model should not do all the work. it should make sure the right work is owned by the right department, executed by the right workers, and rejected until the proof passes. a one-model app gets smarter when its provider ships. an agent company gets smarter whenever any provider ships. a better OpenAI model can upgrade the CEO seat. a better Anthropic model can upgrade strategy and review. a better xAI or open model can upgrade the worker fleet and lower the blended cost. the model mix changes. the company compounds. that is loop engineering for an agent company.
7
2
54
4,979
Derek Nee retweeted
gpt image 2 x seedance 2.0 world cup quarter-finals, but make it a prompt template. we built a fan ootd system that turns each team into: 1. a pixar-style character poster 2. a clean outfit breakdown 3. a stable image-to-video transition pick a team. generate the look. rotate into the next one.
6
3
35
4,151
Derek Nee retweeted
a guy used to ship one app a month that nobody used. now he burns $1,500 of tokens a month and ships 37 apps that nobody uses. AI didn’t fix his problem. it scaled it. the bottleneck was never code — it was skipping the two ugly parts: finding out who actually needs the thing, and telling them it exists. hand-written code at least proved you could code. AI-written code proves you can afford tokens. building is the new comfort zone. distribution is the real work now. and the honest part: marketing is a trainable skill, same as coding. you learned to code by shipping bad code. you learn distribution by shipping awkward launches. stop optimizing the part AI already does.
1
1
10
706
new rule: no claude code, no codex on weekends. the models will still be there monday. probably smarter too.
4
1
23
2,405