Valcho retweeted
Anthropic just posted a write up on how they made the Claude frontend faster which legit coincidentally uses the same agentic optimization techniques as described in the blog post. claude.dev/blog/how-we-made-…
New blog post up: I discovered that you can indeed prompt agents to make your code faster to the point it beats current state-of-the-art libraries. This is not a vaguepost, I include both my prompts and benchmark results. minimaxir.com/2026/09/agenti…
5
7
348
62,526
Valcho retweeted
Replying to @thekitze
djev and gliner 2.5 seem like the best two options so far
1
3
857
Valcho retweeted
Beginning to think that when my coding agent says some effort is "~2-3 days" that means it'll take 2-3 days for me to discover & fix missed edge-cases, not that it'll take 2-3 days to implement
34
8
342
9,223
Valcho retweeted
THIS IS IT RIGHT HERE
New experiment: json-render + jev The future Generative UI is instant Your components, your actions, your design system Rendered in milliseconds
8
6
283
49,689
👀 These JS features -> now Baseline Widely Available: - Object.groupBy() / Map.groupBy() - Promise.withResolvers() - Element.checkVisibility() - ArrayBuffer.transfer() Available in all browsers since March 2024
1
8
142
6,707
Valcho retweeted
GitHub reviewed 2,500+ agents.md files. Five patterns stood out: • Put commands early • Show code, not prose • Name the exact stack • Set clear boundaries • Cover tests, structure, style, and Git An agent needs an operating manual, not a personality. github.blog/ai-and-ml/github… 👆 Use it as a review checklist, not another generic prompt template.
4
22
138
11,829
Valcho retweeted
Pullfrog is now completely free for unlimited personal and open-source usage 🐸 It's a batteries-included BYOK GitHub bot (super-powered CodeRabbit/Greptile) 🐸 Runs in GitHub Actions 🐸 Bring your own key or sub (incl Claude/Codex) pullfrog.com/blog/free-for-o…
27
30
611
32,890
Hermes agent + Buzz is a complete cheat code. In <5 minutes, you can build a full AI workspace where Hermes lives with saved memory, skills, context & more. If you're looking to become more productive with AI, this is an excellent place to start:
17
28
280
59,254
Valcho retweeted
My last open-source skill, /no-ai-slop, clearly hit a nerve with 4K GitHub stars. Today, I’m introducing /human-review, another free AI skill I think you’ll love. Let’s face it - giving feedback to AI in chat is painful. Asking it to “update the 3rd paragraph” or “update image size to X” is annoying when I just want to edit the text myself or give feedback to AI directly on the page. /human-review opens HTML and Markdown files in a visual editor that lets you: → Edit and format text directly → Resize images → Leave comments for AI like a Google Doc I've been using it to edit PRDs, landing pages, product copy, and much more. The whole review loop runs locally, so your data is safe and nothing is uploaded to the cloud. 📌 Get /human-review for free on GitHub: github.com/petergyang/human-… If you find it useful, please ⭐ the repo so more people can find it. 📌 Read my full post here for more on how I use /human-review: creatoreconomy.so/p/use-my-h…
72
133
2,151
316,804
Valcho retweeted
I'm open-sourcing my Three.js game dev skills. Build an isometric action RPG with camera controls, VFX, audio, monster assets, combat systems, and more. Everything is free: github.com/MengTo/Skills Play the game: vesperfall.mengto.chatgpt.si… Game assets: vesperfall.mengto.chatgpt.si…
83
308
3,973
234,352
Valcho retweeted
Deterministic core, agentic shell
recommended reading. in a few systems i built over the past 2 years i ended up with this flow for a given task: - start with full inference for all steps, observe behaviour, judge correctness - find inference steps that can be replaced with deterministic steps eventually, most inference goes away, leading to more stable yet capable systems that can do much more than deterministic systems alone.
14
10
154
27,925
Valcho retweeted
Fable and HTML are pushing the boundaries of what a presentation is: including whole webpage embeds and sophisticated animations. Pleasantly surprised at how well it could represent the complexity of the Full Duplex Orchestrator in tau-voice. Slides from my talk at the Stanford AI Measurement Science seminar taubench.com/talks/stanford-…
9
18
495
57,746
Valcho retweeted
Pro tip: when prompting Codex with really difficult /goals, ask it to "write a goal for another thread to achieve this and babysit it until it figures it out" By doing so, you'll add built-in steering and another layer of taste verification (that's how this video was made)
Asked 5.6 to make a video introducing itself
38
75
2,154
214,738
Valcho retweeted
I'm just going to dump my whole agentic setup out here, because I see too many people missing giant chunks of this and it's hurting them. Here's what I have and recommend: 0. an AGENTS.md that is a router -- it sends the agent to the right skills, docs, tools 1. a standard workflow doc/skill customized to my needs ... (grab Matt Pocock skills if you don't already have something) ... I tag this in most sessions with `@/AGENT_WORKFLOW.md` and it pulls it in. 2. self-healing docs for every system, and agents are instructed to keep them updated ... I tag the ones I know I need, or let the agent find them through AGENTS.md ... I also provide a more detailed summary in the first 7 lines of every doc, so they're easily greppable to find the right thing, and this is documented in AGENTS.md 3. agents always run the app ... the agent should always actually run the app itself, and test its work and fix issues as it goes, especially if running autonomously / asynchronously 4. end-to-end tests and instructions to write more and keep up to date, and docs on how to write tests, what to avoid, and a list of all the tests and what they test in yet another markdown doc ... write and run targeted tests during implementation, improve and commit with work 5. custom linters at precommit hooks looking for any problems you run across, with `--fix` fixing the problems automatically, OR if that's not feasible, it shells out to a cheaper LLM like Composer 2.5 or Sonnet to fix the problems -- NOT just flagging them, but actually resulting in cleaned code 6. cross-agent review at each major point: research, plan, implementation, and wrap-up. I mean codex, claude, cursor, whatever -- but it shouldn't be the same model reviewing the same code. And specific docs for agent review, what to look for, how to approach it. Also, personas -- looking at the code from different perspectives, such as maintainability, code quality, security, performance, AI smells, domains (e.g. "financial services expert" or whatever) ... and each persona also "owns" a set of system docs too and keeps them up to date 7. agent traces / worksheets that track what the agent is doing each session. if the agent fails partway through, you should be able to hand this worksheet to another agent and it could finish the job. commit this worksheet with the work so it's all connected and easy to reference later (you will reference these later!!), also have the agent apply git tags that correspond to specific worksheet names so they're easy to find 8. automatic agent feedback to you at the end of the session, added to a doc that is also committed with the work, that you periodically ingest into an interactive session and improve your workflows 9. a tools or bin folder that contains python or bash scripts that the agent has skills to make to make its job easier (for example, I have an `agent_review` bash script that lets the agent kick off agent reviews via CLI without knowing each agent's particular incantations) ... docs on how to make scripts effectively, and instructions to constantly build these out more 10. periodic agent sweeps through recent commits, looking for problems / gotchas from a higher level across commits 11. a coding conventions doc that is just for specific coding conventions you want to see in the code base, your review agents use these a lot (but a lot of this should be in linters) 12. an agent loop / night shift skill for autonomous work, that lays out how the agent is to approach this, from an orchestration standpoint 13. a task queue that is accessible to the agent (mine is just a TODOS.md, but yours might be in Linear etc, with a CLI to fetch via API) 14. a periodic false-confidence test audit skill that looks for tests that aren't actually testing what you think they're testing, and that fix those 15. visual regression tests -- take screenshots, compare via tool and with agent visual review, commit with work (git lfs useful here) or at least push into the PR 16. automatic performance benchmark tests that notice when performance degrades 17. performance profiling tools that can be used by agents for targeted benchmarking, trying new techniques, comparing outputs, and comparing profiles 18. end-of-shift full validations, including running all tests, performance, agent reviews, sweeps, everything -- when you return, it's all as pristine as it can be If you have all this, your agentic coding experience is going to be very different than dry prompting and manually guiding it toward the right thing every time.
151
264
4,013
351,391
Valcho retweeted
One clarification for folks using /wayfinder: The flow for big work should be: /wayfinder -> /to-spec -> /to-tickets -> /implement Once the /wayfinder map is complete, you turn it into a spec. Some folks are using /wayfinder as the ENTIRE flow - from grilling to prototyping to shipped work It certainly can be used that way - I've been doing that for non-coding stuff like course creation. But for coding I much prefer creating a spec and handing off implementation to an AFK agent. Means I can focus on other things for a bit while it churns away. I'll be putting all of this in an upcoming tutorial. Thanks for all the great feedback on v1.1. v1.2 is in the works and looking good.
83
75
1,467
82,497
Valcho retweeted
Okay so basically: 5.6 Sol Ultra for planning Sol Medium for coding the plan and general coding tasks. Terra High for quick context subagents, reading the codebase, searching. Luna (any thinking level) for chat or small computer actions like moving files, organizing folders etc
110
73
1,711
106,881
Valcho retweeted
I'm going to teach my son that with enough ambition and perseverance, NOTHING is impossible until he hits the weekly token limit
10
5
136
5,620
Valcho retweeted
here's the exact setup i'm running to manage 3 agent harnesses at once and ship like a madman: Claude Code + Codex + omp, all running inside @orca_build (not sponsored) solid environment, real mobile support, and their /orchestrate is the easiest way i've found to delegate work to agents... you describe the task, it just goes most of my work now happens upstream: i write one perfect, in-depth plan before a single agent touches anything but i don't actually write that plan... i have a council of models draft it using matt pocock's skills, argue it, tear it apart, then i lock the version that survives after that it's just engineering the loops and the smart goals around it for execution this is unreal... code, automations, advanced workflows, it all runs while i barely touch the keyboard but i NEVER let agents run in long loops on anything marketing or business every time i've let one run loose on copy or strategy it hands me slop... confident, polished, completely soulless slop so for code, i hand over the wheel... for anything that needs taste, my hands never leave it knowing what to fully automate, and what to never let an agent touch... that's the whole game right now
18
12
156
12,799
Valcho retweeted
👨‍🎨 We're sharing a formula for making vibe-coded apps look professional. It's a deconstruction of the "science of design" and we think it will help you coach your agents into mobile app design that brings polish to your apps. The formula and examples are in the post below ↓
29
64
1,322
130,915
Valcho retweeted
Every open-source project should be engineering agent loops right now. We've found success managing @warpdotdev with loops for: - Issue triage - Spec generation for larger features - Code review - Even letting agents self-improve their Skills on a cron
19
25
405
41,232