AI dev • co-founder @polynternet • Building in progress

dm open
Mr. Buzzoni retweeted
this is pure f*cking gold for anyone running coding agents Jev founder Diogo Amogo wrote a PDF on building a Jev harness the promise: > 200x faster > 400x cheaper the model hasn't been the slow part for a while the speed and the cost sit in the harness wrapped around it • how to use it > drop this PDF and the article below into Claude Code or Codex > let it rebuild its own setup one evening of setup and your agent starts next week on a harness most teams still don't have 👇
42
73
610
91,618
Claude Opus 5.5 is scary good at motion design I asked for a showreel on Jev engineering and got studio-level work from one prompt stock Claude Code, Opus 5.5 on max reasoning, memory off: "Make a dynamic [24]-second motion graphics video that shows what an incredible motion designer you are, like it's your showreel for a resume. Go all out. Only about [TOPIC]. Style: [BRAND COLORS], sound synced to the cuts."
18
5
60
3,018
Mr. Buzzoni retweeted
THIS PLAYBOOK TURNS OPUS 5.5 FROM A CHAT WINDOW INTO A COWORKER THAT FINISHES YOUR WORK WHILE YOU'RE AWAY Addy Osmani wrote it for the Claude developer blog three of its habits changed: > it runs longer on its own > it says plainly what it did > it thinks before every reply your old prompts are fighting all three • how to ask > one message: the whole task, the finish line, when to stop > "every endpoint on the new client, old client deleted, tests pass" > then leave it alone. runs go longer than on Opus 5 > forgot something mid-run? type it and hit Enter > restarting costs more now • stop telling it to think hard > it already decides how much to think > deleting "think step by step" made replies start sooner > no clear drop in quality > for a short answer write "answer directly" • steering a long run > write the stops you want into CLAUDE.md > split big audits across subagents, one service each > check each one's evidence before accepting it > keep the checklist in TASKS.md > older turns get summarized, the file survives • checking the result > end every run with: Blocked on me, Changed, Found > ask it to review the diff > at its lowest effort it caught more bugs than Opus 5 at high > ask it to mark what it couldn't confirm • in the Claude apps > attach the chart itself, don't retype the numbers > hand it a long deck and ask for contradictions > in testing it caught a date on the wrong weekday rewrite your CLAUDE.md this week and the same subscription everyone uses for chat starts doing hours of work without you claude.dev/blog/getting-the-…
12
8
66
7,377
Mr. Buzzoni retweeted
this is unreal f*cking gold for Jev builders 20 repos people are building on Jev right now. browser agents, context tools, trading bots, even a drone 1. JEV-Ultrafast - a fast browser agent ↳ github.com/browser-use/jev-u… 2. Fast-JEV-Compaction - context compression ↳ github.com/tamaratran/fast-j… 3. JSON-Render - generative UI ↳ github.com/vercel-labs/json-… 4. Typesafe-MCP - use Jev with any client ↳ github.com/itsmostafa/typesa… 5. JEV-MCP - a judgment toolkit ↳ github.com/burnigtm/jev-mcp 6. Semdecide - a classifier that lives in your CLI ↳ github.com/sharziki/semdecid… 7. JEV-Codex-Router - routes each task to the right model ↳ github.com/0xNatoshi/jev-cod… 8. Winnow - garbage collection for your context ↳ github.com/GhalebDweikat/win… 9. JEV-Review - code review triage ↳ github.com/devagrawal09/jev-… 10. Blink - a repo navigator ↳ github.com/ellipsis-dev/blin… 11. Agent-Desktop - desktop automation ↳ github.com/lahfir/agent-desk… 12. Typesafe-Mario - an agent that plays Super Mario ↳ github.com/fhshaik/typesafe-… 13. JEV-Drone - drone control ↳ github.com/RomanSlack/jev-dr… 14. OneVOneJev - a browser FPS ↳ github.com/emrickgarrett/One… 15. JEV-Trader - HFT market making ↳ github.com/buberlo/jev-trade… 16. Prism - liquidity signal detection ↳ github.com/irfndi/prism-liqu… 17. Neo4Jev - knowledge graph traversal ↳ github.com/jexp/neo4jev 18. JEV-Curate - training data screening ↳ github.com/AkashPriyadarshii… 19. Canny - checks whether a task was actually completed ↳ github.com/qkal/Canny 20. KillMyIdea - scores startup ideas before you build them ↳ github.com/monteduro/killmyi… pick by what you do: > coding -> JEV-Review, Blink, Canny, JEV-Codex-Router > context -> Fast-JEV-Compaction, Winnow > automation -> JEV-Ultrafast, Agent-Desktop > clients and tools -> Typesafe-MCP, JEV-MCP, Semdecide > UI -> JSON-Render > trading -> JEV-Trader, Prism > data -> Neo4Jev, JEV-Curate > founders -> KillMyIdea > just for fun -> Typesafe-Mario, OneVOneJev, JEV-Drone grab the one closest to your job and ship something on top of it this week
21
41
224
27,599
Mr. Buzzoni retweeted
A FLASH-LEVEL MODEL WITH A 1M CONTEXT WINDOW, AND RIGHT NOW IT COSTS $0 I've been running Space Bunny Alpha, the stealth model on OpenRouter nobody knows whose lab it is first thing I noticed: I stopped batching my questions answers start inside 1.14s and finish in about 6 there's nothing to wait for, so I just kept asking then I gave it a real task a screenshot of a broken UI, the video of the bug, the repo behind it all in one window it came back with the fix in one pass • what I checked after that > 81 tok/s at P50, 183 at P99 > 99.53% uptime over three days > 95.15% cache hits, 3.71% tool call errors > reasoning I dial from low to max > 1M in, 524,288 out • who else is on it > Cline, Claude Code, Command Code, DeepSeek Harness, Kilo Code > all five are coding agents 711B prompt tokens in one day against 9.64B out people are feeding it entire repositories free through the preview, and the provider doesn't train on prompts try it before the preview closes 👇 ↳ openrouter.ai/stealth/space-…
10
4
50
7,060
#DreaminaCaughtTheVibe creating content usually means jumping between different steps so having everything feel more connected is a nice change @Dreamina_ai updated the Dreamina AI web experience and canvas together with curated AI video workflows and skills for: > film > brand ads > viral social media scenes > and more • get 90% OFF your first month of Dreamina Basic • limited offer available till October 9th I'd be curious to see how much smoother a real project feels with everything connected in one place
Why does a basic broom suddenly have more cinematic aura than most TV commercials? The newly updated Dreamina AI Web experience just dropped, and its also upgraded Canvas. They also added workflows and skills of many scenes, like film,brand ad, social media hit content and so on. I picked the most unglamorous, everyday object I could find—a humble broom—and built an artisanal, high-concept commercial around it. The AI orchestrated the entire shot sequence like a Hollywood director: sweeping wide shots, volumetric studio lighting, rich depth of field, and a dramatic slow push-in for the hero shot. The fact that a literal household broom now looks like an $8,000 architectural design piece is equal parts hilarious and breathtaking!
1
38
2,166
Mr. Buzzoni retweeted
I just claimed $250 in free Claude Code credits. if you're on Max or Pro, yours is sitting there too $250 on Max, $100 on Pro. for cloud sessions, which just came out of beta the part that matters: the credit sits outside your usage limits so when you hit your limit locally, you move to the cloud and keep working until the credit runs out claim it, connect GitHub, start a session. it applies on its own • one credit per account, cloud sessions only • or run /claim-credit in the CLI claim by Oct 7 -> claude.ai/code/claim-credit/…
Cloud sessions are officially available and out of research preview! They let you keep Claude Code working, even when your laptop is closed. Existing subscribers get a one-time credit to try them: $100 on Pro, $250 on Max.
8
8
79
31,905
the winner of an Anthropic hackathon open sourced his full Claude Code setup - and this is absolute f*cking gold 68 subagents, 286 skills, 94 commands, MIT license ECC turns one lone Claude Code assistant into an entire engineering department a blueprint lands before any build, a failing test lands before any fix, and every change gets read again by a context that never watched it get written • who does what > planning - give it one sentence, get back a plan you approve before a line of code exists > review - a fresh-context pass over your diff, with a separate reviewer per language > build repair - one dedicated fixer per toolchain, PyTorch and CUDA included > security - an OWASP sweep plus a scanner hunting injection holes in your agent config > architecture - catches design mistakes while they're still cheap to undo > domain work - database queries, ML pipelines, e2e tests, docs the security duo is the part almost nobody sets up an external OWASP audit runs four figures and a week of waiting this one finishes on your branch before lunch • what the skills cover > testing - tdd-workflow takes you red to green, eval-harness sits on top of it > language packs - Python, Go, Rust, C++, Django, Laravel, Spring Boot, Next.js > context - search-first reads the docs before writing, iterative-retrieval keeps your repo from flooding the window > shipping - Docker, CI/CD, health checks, rollbacks, migrations > beyond code - writing in your voice, market research, pitch decks fork it, cut it down, have your own version running by tomorrow start with one plan and one rules pack switching on all 286 skills at once is the fastest way to make everything worse whoever wires this in over a weekend spends the next quarter reviewing work instead of typing it
36
25
326
41,490
this is free f*cking gold for AI engineers 10 repos worth keeping, from python fundamentals to LLMs, agents and production AI. no more random tutorials 1. Python-100-Days - a 100 day run through fundamentals, data analysis, web dev ↳ github.com/jackfrued/Python-… 2. Generative AI for Beginners - LLM basics, prompting, RAG, agents, fine-tuning ↳ github.com/microsoft/generat… 3. LLMs from Scratch - build one yourself: tokenization, attention, transformers, training ↳ github.com/rasbt/LLMs-from-s… 4. ML for Beginners - 26 lessons of classical ML over 12 weeks ↳ github.com/microsoft/ML-For-… 5. OpenAI Cookbook - working examples that take you from reading to shipping ↳ github.com/openai/openai-coo… 6. Stable Diffusion - the original implementation and research code ↳ github.com/CompVis/stable-di… 7. AI Agents for Beginners - tool use, RAG, agent frameworks, multi-agent systems ↳ github.com/microsoft/ai-agen… 8. AI for Beginners - 24 lessons over 12 weeks: neural nets, vision, NLP ↳ github.com/microsoft/AI-For-… 9. LLM App - RAG pipelines, enterprise search, live data, vector search ↳ github.com/pathwaycom/llm-ap… 10. Segment Anything - Meta's foundation model for image segmentation ↳ github.com/facebookresearch/… start from where you actually are: > python -> Python-100-Days > ML -> ML-For-Beginners > LLMs -> LLMs-from-scratch > agents -> AI-Agents-for-Beginners > building -> OpenAI Cookbook, LLM App > vision -> Segment Anything take one, build something with it, then move on
7
27
251
19,704
A 680,000-LINE MIGRATION THAT WOULD HAVE TAKEN A TEAM WEEKS CLAUDE OPUS 5.5 FINISHED IT IN UNDER A DAY Anthropic shipped Opus 5.5 today frontier results, 40% cheaper than Opus 5 it leads on agentic coding, computer use and knowledge work: > Terminal-Bench 4.0 - 66.4% vs 52.3% for Opus 5 > GDPval-AA v2.1 - 1846 vs 1735 for Fable 5.1 > cache reads $0.20 per million, 60% cheaper > output comes out 30% faster but the story I keep coming back to is the audit one tester fixed a 200,000-line codebase in under three hours Opus 5 needed over 20 hours and 2.5x the tokens on the same job seventeen hours handed back to someone, on one task Deloitte ran it at its lowest effort setting and it caught 72% of known bugs Opus 5 at high effort caught 56% the cheap setting now beats the expensive one Opus 5.5 is live everywhere today. $4 in, $20 out I've stopped being surprised by the benchmarks. the price drop is what I'm still sitting with
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
13
3
64
5,104
I HAVEN'T OPENED CLAUDE AT MIDNIGHT SINCE I BUILT THIS FOLDER I used to wake up and find out what broke overnight -> now the receipts are already there. dated, graded, waiting on my review what sits inside the folder that took that shift: • the contract > CONTRACT.md - the shift rules, committed > contract.local.md - my overrides, gitignored • the harness (.claude/loops/) > settings.json - spend caps and timeouts. set once > schedule.yml - 05:00, pr-hunter. cron fires it, not me > rubrics/ - code.md reviews merges, writing.md checks voice, safety.md says what not to touch > pr-hunter/ - plan.md runs wake, read, act, verify > act.sh does the work. verify.sh is the gate to merge > tools.allow - what it may call, and nothing else • the state > receipts/ - one folder per shift. 17 shifts today > 5,382 receipts kept. none of them edited > trace.log - what happened, line by line > checkpoint.json - resumes exactly where it stopped > budget.json - $6.10 spent of a $40 day • the edges > kill.sh - the panic file. never used > .mcp.json - the tools it's allowed near > alerts.yml - who gets paged when a loop stops 2 loops. 3 rubrics. 8 months in service. 0 incidents nothing here is clever. every file exists because a run went sideways once I wrote the rule down instead of trying to remember it $6 a day for a shift I used to work at 1am the model doesn't wake. the folder does
31
100
747
69,933
send this to GPT-6 Astra before you type anything else today the bottleneck is your prompt, not the model so stop writing prompts. make it rewrite yours first it restates what you literally asked, then what you actually meant the gap between those two is where every bad output comes from then it stops and waits for your approval before doing any work full prompt in the image:
12
10
195
25,994
$140M ARR IN 90 DAYS IS AN ABSURD NUMBER UNTIL YOU SEE WHAT THEY AUTOMATED 500 hours a week of manual work at a $17B company, and an AI learned all of it by watching Deel's payments lead recorded himself reconciling one payment: bank portal -> billing -> sender history -> client statement -> ticket five systems, none of them connected Akai watched his screen, listened to him explain it out loud, and wired every system together on its own. even a bank portal with no API then it picked up the rules nobody ever wrote down: > reading broken invoice refs > never paying into a flagged org > closing bank interest without asking anyone every new edge case turns into a new branch. his 50 person team extends it just by asking the lesson for anyone stuck in a repetitive job: record yourself doing it once, out loud whoever does that first stops doing the work and starts running it watch it 👇
EXCITED TO LAUNCH: Akai (akai.run) Deel added >$140M ARR in 90 days without increasing headcount by automating~600 Full Time Employees' equivalent in work with Akai. Akai was an internal tool to automate our painfully repetitive operations in Finance, HR, Accounts Payable, and Compliance, etc. We never intended to make this a product. But we watched revenue per employee grow from $130K to $215K We built >8k agents that do the work of ~600 employees It had such a dramatic impact on our business that today we are launching it for everyone. How it works: Say you're automating payment reconciliation: 1. Record your screen while manually matching a messy transaction and Akai will capture your screen, voice, server requests 2. Akai will see that you pulled unformatted wire transfer info from an archaic bank portal, put it in some excel sheet, checked NetSuite invoices, payment history, and put a ticket on Zendesk 3. Akai reads between the lines and build a workflow + steps + conditional guardrails. It learns tacit edge cases, like resolving malformed invoice references without you writing a single regex 4. Simply connect NetSuite, your ledger, Zendesk, PSPs, and even legacy bank portals with zero API access 5. Run the workflow and tell it what to adjust in plain English: "strip slashes on wire memos and auto-apply partial payments." It adapts instantly 6. Once it works for you, add 100s of colleagues. Your entire payment ops team forks and extends the workflow for new PSPs, secondary ledgers, or regional settlement rules 7. We automated 85% of our payment reconciliation end to end, eliminating 500+ hours of soul-crushing manual grunt work every single week. Claude Code/Codex can't do this in multiplayer mode. Every person rebuilds the same skill from scratch in their own way. Deel built Akai to: 1. Understand backend operations edge cases (it had to work for our 7000 person team first) 2. Collaborative across 1000s of employees 3. Self-Learning from millions of runs 4. Optimises cost and gets cheaper every run We're so confident that we're announcing an Automation Guarantee: If our engineers can't automate a thousand of hours of work in your first 30 days, you get a full refund. Book a demo: akai.run if you're an exec at a company with hundreds of employees
7
56
6,624