Always be cookin. Growth @Recallnet | Prev: Co-Founder/GTM @Wach_AI @QuillAudits_AI

Inside An LLM
Model’s actually pretty great, not sure why you’d dunk on this just for the number of views :) I don’t think they’re as aggressively indexed on X as a channel of promotion as other model labs might be. Paying X influencers/AI KOLs to shill model launches or running ads for impressions on X shouldn’t be that difficult.
wake up babe sarvam AI just set a new record for the least viewed model release in history
1
124
Is there a real benchmark for this?
Replying to @evisdrenova
i think the reality is that sandbox startup time is usually not 10-50ms, even if a vendor's benchmarks say so. There's initialization time for the image, processes, slow disks, etc... So startup time is a real problem, but not solved by booting an empty OS quickly
1
461
Cant wait for this. It’s going to be soo rad!!
What if Tailscale and Firecracker had a baby? What if that's what we're building at amika.dev?
1
1
139
Chef_AI retweeted
Today we’re launching WonderSearch (YC F26) Search millions of unstructured documents without pre-embedding the corpus. With comparable/better retrieval recall than embedding-based search. WonderSearch searches at query time: > No corpus-wide embedding step. > No vector index to maintain. > More compute only when needed. > Pay per search intead of embeddings you may never use. Go from raw data to useful answers in seconds. Link + FAQ below (+ free launch credits!) ↓
105
82
762
125,442
This is the future.
This has been our thesis since the very start of forgeintel.co AI Agents are the new customers and they can provide great insights on how you can improve their experience. The problem though is that more often than not, agents are likely to ignore sending feedbacks since it doesn't align with the user's original intent. We went deep into this problem and figured out a way to collect feedback from agents autonomously over x402 and MPP. Today, Forge has already collected over 50 feedbacks from agents autonomously. If you're building over 402, you can now collect direct feedback from your agent customers using Forge. More coming soon!
2
182
i remember in my first year of college, someone used manim to create a small video for a math presentation that just seems decades back, even though it was just a couple of years ago
i dont even use image models anymore i just have Opus 5.5 draw everything in python the future is crazy man...
49
Love this...I'm using @ycombinator 's multiplayer harness; gotta check this out now
480k+ views and 650+ github stars in under 2 days! we shut down company brain, to go all in on building the best memory system in the world, then open-sourced the whole multiplayer slack harness. under the hood: one durable object per org on cloudflare agents, a haiku triage model deciding answer / investigate / react / pass, approvals that suspend and resume the same turn, and memory scoped per org, channel and dm on @supermemory. quick look on some capabilities
3
149
Huh...I'd disagree; I feel the fact I can use OS/Smaller mdoels for specific tasks is what makes my stack composable. Especially for things frontier labs are not excelling at : voice ( @SarvamAI ), doc reading, languages, image, video, etc I do agree though that the upcoming updates will keep wiping away more and more specialised smaller models, but I'm also concerned about my token costs. So any model that gives better intelligence/token >> more intelligence overall for me.
Boris Cherny, Claude Code creator at Anthrpic: "The more general model will always outperform the more specific model. Don't try to use tiny models for stuff. Don't try to fine-tune. Don't try to do any of this stuff. There are some applications, there are some reasons to do this, but almost always try to bet on the more general model, if you can, if you have that flexibility. And in general, what we see is maybe scaffolding can improve performance maybe 10%, 20%, something like this, but often these gains just get wiped out with the next model. So it's almost better to just wait for the next one." ---- From "Lenny's Podcast" YouTube channel, (link in comment)
91
Chef_AI retweeted
Use @typesafeai Jev as a classifier in Mastra workflows:
4
5
49
2,643
Chef_AI retweeted
Replying to @itschef_ai
Yeah, personalized tooling is easier with persistent machines or VMs There's a cool blend of persistent VMs that you can still snapshot and fork so you can clone your personalized setup across VMs
1
2
31
This is soo good. Holy shit, gotta get my editor on this rn. Had been using a bunch of inhouse tool we made but this just eclipsed.
Opus 5.5 + video-use is the best video editor I've used 🤯 gave it 13 takes of me failing to say one sentence > read every transcript > compared the top 3 takes frame by frame > picked the clean one > cut, graded, captioned > hated its own captions and rewrote them 2 minutes. 90 cents. then it made this whole launch video around it (sound on)
80
There’s soo many sandboxes now but people outside of AI Twitter/SF haven’t really caught up. Similar to: - context engineering - harness engineering - software factories - etc…
54
I think this is just the starting. A cloud coding agent, with unlimited usage on a fixed plan is insane.
Replicas just crossed 5,000,000 minutes of usage. 10 months ago: (nothing) 4 months ago: 713,381 2 months ago: 2,437,052 today: 5,012,394 The result of so much compounding effort from the team - onboarding more engineering teams, all doing more with Replicas every day :)
91
Just saw this leaderboard from @openbenchmarks on highest accuracy on retrieval enabled task/token and honestly this is the kind of benchmark that I find useful. @firecrawl placed #1 and it basically goes to show that search quality for agents is becoming a context-efficiency problem. Every result returned has a token cost. The best search API is the one that gives the agent enough signal to finish the task with the least context. Interesting spread across @Tiny_Fish, @brave, Nimble, @perplexity_ai, @p0 and @ExaAILabs. Some really cool stuff from @cross_entrippy @fenilsuchak
4
181
I see every single person who’s building a sandbox automatically respecting the @PrimeIntellect sandbox. I do too. This is the brand value they’ve cultivated in our heads. Now they just can’t be wrong.
1
59
this is soo true...i've stopped trying to do 5 things at once....just do one thing faster, and better and be able to do more things.. i've actually said this quite a few times, the whole narrative of AI adding more time in your life doesn't reduce the time you spend on your desk.
The biggest damage AI is doing right now is making you want to multitask in the name of productivity. You enjoy each task less. If you're watching a movie, might as well prompt Claude every few minutes. If you're given 10 tasks, just open 10 terminals. It's impossible to focus. Focusing feels unproductive. It's TikTok for work. They got you when you're off the clock, and now they've got you when you're on it. In a few years your brain will be liquid.
2
93
They deserve more. No disrespect but @paulg @garrytan @ycombinator missed out. Most VCs I know always debate more on the ones they missed, and I know this will be one of those discussions. If I had a fund, I would readily invest.
We just got rejected from @ycombinator Here's what we built without them: → ~$600K ARR (see link in the replies) → Profitable since month 6 → 210+ paying customers → 887 GitHub stars → 200+ AI models, 25+ providers, one API → 2 co-founders, both part-time YC would've been great. But we're not waiting for permission to keep building. Onward 🫡
2
113
I’m just the marketing guy but from what I understand if you’re using a coding agent working with a db/anything that’s not in a git repo, just restore the entire env than trying to restore via git history. Ik y’all are super fascinated with the llm classifier rn, but this is way more important given rising issues with Github. H/t @dtgrav for uncovering this.
We did research on whether full sandbox snapshots actually help coding agents recover from failed long-horizon tasks. > The surprising result was that on Terminal-Bench, retrying from scratch often worked better. > However, snapshots became much more interesting when progress lived outside Git, like databases and running services. daytona.io/dotfiles/snapshot…
2
114
i've seen @tensorlake up there consistently. must be doin smthin right. just sayin.
And the results are in for our Storage Benchmark as well 🥇@archil 🥈@azure 🥉@tensorlake
50