Founding AI Engineer, Developer Relations @Wandb @Coreweave 🐝✨ Content on AI Engineering, MCP, Evals the stuff that’s actually moving right now

Lorenzo Porras retweeted
AI teams are moving from prototypes to production faster than ever, but rapid development can create new blind spots. Join us this Wednesday at 10:00 AM PT for the latest What's New Wednesdays session, and get valuable insights on vibe-coding, tracing, and deploying your agents. Reserve your spot here: utm.io/ur8ga
5
1
6
1,648
Congrats @CompleteSkeptic! I want to thank @typesafeai again for joining us at @CoreWeave Hacks, and giving users early access to Jev! I had a lot of fun using it and appreciate the swag!
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
1
6
146
Lorenzo Porras retweeted
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
4,033
8,251
76,172
39,620,279
Join me to discuss traces, and vibe coding! Really excited to discuss on What’s New Wednesday. Super excited to see y’all there!
Moving AI agents from prototype to production faster than ever? Don't let rapid iteration hide critical runtime behavior. Tune in to What's New Wednesdays: Vibe-code, trace, and deploy your agents to explore: • Production agent observability workflows in W&B Weave • Evaluating and self-improving agent loops at scale • Real-world takeaways from CoreWeave Hacks Register here: utm.io/ur8ga
73
Lorenzo Porras retweeted
Distillation isn't a research project anymore. It's a pipeline. Weave traces → dataset curation → Serverless SFT → eval → deployment. Join us for 40 minutes on turning real production usage into a smaller student model that meets your bar. utm.io/usxVH
1
3
38
3,066
Lorenzo Porras retweeted
Kickoff of day 1 at @CoreWeave Hacks. 200+ builders, one mission: build an agent loop that catches its own mistakes. @agihouse_org, @typesafeai and @marimo_io are here with $1k+ in credits and tons of compute. 24 hours to submit. One team is going home with a robot dog!
6
4
33
9,716
Thank you @typesafeai for helping sponsor @CoreWeave Hacks! And honestly one of the coolest prizes for best use of TypeSafe AI - 2 x @huggingface microducks!!
1
1
4
139
Next up Julia showcasing @CoreWeave @wandb ARIA! Winner for best use of ARIA gets $1,000 as well! Check out ARIA: wandb.ai/site/agent/
3
6
898
We are underway @CoreWeave Hackathon! @neutralino1 walking through @wandb Weave! With a prize of $1,000 for best Weave use this weekend. For those participating- Many more prizes like a robo dog and f1 tickets to come!
3
5
173
Hey y'all - Last week I got to speak to @MasterClass Executive Cohort 1! We walked through the @CoreWeave AI Loop, improving agents with @wandb , and covered how to build evaluations for AI systems in a business setting.
1
1
6
497
If this sounds like your kind of room, cohort 2 is on the way. Come see what it's about. mstr.cl/WeightsAndBiases
1
68
About to do some evals… brb
This is GPT-6 Astra. Anything you can do on a computer, Astra can do for you. Fast.
57
Lorenzo Porras retweeted
Muse Spark 1.3 is rolling out today with frontier performance almost too cheap to meter. This is the biggest jump we've made so far on coding and agentic work. Try it in Muse Code and our API. Next up 🍉 and Muse Spark open weights releases coming soon.
1,445
1,617
19,963
3,762,013
DeepSeek-V4-Pro-0813 is on CoreWeave Serverless Inference! Lots of interesting possibilities when you pair @CoreWeave’s sandbox with serverless inference. Just a thought 👀
DeepSeek-V4-Pro-0813 is live on CoreWeave Serverless Inference. 1.6T parameters. 1M context. Built for long-horizon agent work, and priced for it: cache reads run $0.044/M for all the context your agent re-sends every step. Start here: utm.io/usz4F
4
141
Come join me at Coreweave Hacks! Sept 12 - 13 in SF! I want to see what y'all build!
WeaveHacks is now @CoreWeave Hacks! Our latest agentic hackathon is September 12 and 13 in SF. Build agent loops that catch their own mistakes. $20k+ in prizes, a robot dog for Best Loop Design, and F1 tickets in play. Apply here: lu.ma/coreweavehacks
1
3
181
Finally
Resume terminal sessions in desktop app!
83
I started using @OpenAI's Codex more this week and realized I never played with pets. Decided to make my own @CoreWeave x @wandb Bee for my pet 🐝.
1
7
715
After working at Twilio I really grew the appreciation for Voice Agents. Really love seeing what @pipecat_ai is doing!
Introducing PhoneLLM, an open model for voice agents. GPT 5.6 Terra performance on typical voice agent tasks at 1/3 the latency and 1/18 the cost. For voice agents, we need models that are both very low latency and very good at tool calling and instruction following. There's a trade-off here, and we often have to compromise on either latency or capability when building voice agents. With PhoneLLM (and the training and data stack that made this model possible) we're fixing this problem. For the last couple of years, most of the effort in frontier model development has gone towards leveraging test-time compute. Which is awesome! Models of all shapes and sizes are available that perform really, really well ... if you have "thinking" turned on for your model. But if you need your agent to respond at voice conversation speed, you can't use thinking models. PhoneLLM is a full-weights fine-tune of NVIDIA Nemotron Nano 30B. We trained on a wide range of real-world telephone and customer support use cases. The training focused on taking the excellent Nano 30B base capabilities and teaching the model to do typical voice agent tasks with thinking disabled. The results are really good: accurate tool calling and concise, on-topic responses in long conversations. And fast: TTFAT measured server-side is <100ms if you run PhoneLLM on a lightly loaded B200. :-) But seriously, when we characterize model latency, we do it with full, end-to-end, batched request simulations using real Pipecat voice agent pipelines. You can serve more than 80 concurrent agents on a single B200 with P95 end-to-end TTFAT <600ms. Including network overhead. That's an LLM cost-per-minute around $0.0025. (1/4 of a cent.) At a latency lower than any third-party API offers today. More details about this model, including weights on @huggingface, how to spin it up with one click on Modal, and a starter project repo you can clone, are in the thread ...
1
90
Lorenzo Porras retweeted
1
1
9
1,732