Life expands to new territories. Painfully, perhaps even dangerously. But life finds a way.

Ciudad Autónoma de Buenos Aire
Farid retweeted
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks: Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better: Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better: Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better: Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work! In summary: - As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding. - Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
1,570
6,398
54,436
7,563,047
😢 ahí se va una de las mejores cuentas que X supo tener
Espero que cada vez que escuchen África de Toto en Aspen se acuerden de este bot. Hoy nos toca despedirnos: este antro apestoso pretende que pague para entretenerlos. "It's gonna take a lot to drag me away from you."
1
60
🙏
Don't be a meat proxy
42
Y yo que me sentía el peor padre porque ya lleve dos veces al pibe al jardin sin la mochila 🤣
🤣🤣🤣 JAJAJAJAJAJAJAJAJA. NO TE LA PUEDO... El marido se fue a llevar al nene para vacunar. Y en camino la esposa lo llama y le recuerda que se olvidó algo en la casa. Miralo por favor
2
41
Farid retweeted
Harnesses often get dismissed as just scaffolding, just prompt engineering, and not real research. But that couldn't be farther from the truth. The same model weights that score 30% on ARC-AGI score 95% with a better harness. So we gathered a group of researchers and founders working at the frontier to do a deep dive into the state of harnesses. We cover how we got to this point, the case for making your harness as expressive as possible, and what YC learned building an agent for every employee in the company. 00:00 - @FrancoisChauba1: Why harnesses matter 04:27 - Building an auto-researcher by accident 07:13 - A five minute history of harnesses 13:56 - Self-improving harnesses 18:35 - @sethkarten: Prime Agent, a self-improving RLM harness 21:50 - Context as an L1, L2, L3 cache 24:51 - From Turing machine to von Neumann computer 28:33 - Messaging between agents 30:04 - ARC-AGI results 33:09 - Emulator Bench and GPU kernels 37:30 - @JonSaadFalcon: OpenJarvis, personal AI on personal devices 38:26 - How far behind are local models 39:21 - The five primitives of a personal AI stack 42:47 - Letting cloud models optimize your local stack 43:53 - 800x cheaper than the cloud 45:58 - @josh__france and @jbellregan: QM, YC's agent harness for work 47:29 - A history of YC's internal agents 49:24 - OpenClaw and a fleet of 50 agents 51:04 - Pulling the brain out of the sandbox 54:43 - Letting the agent choose its own sandbox and model 57:16 - The grind tool: budgets on goals 58:50 - Agents don't understand social context
185
446
3,869
494,642
In the age of AI with less code more judgment and more flowers.
1
2
35
Farid retweeted
Se viene la invasión de calamares al Maracaná 🏟️⚽Platense va por la hazaña ante Fluminense 🔥 Fluminense vs. Platense | Mirá #LibertadoresEnTelefe Próximo martes a las 18.45hs con @giraltpablo y @jpvarsky 📺
35
176
769
185,903
Este chabon quiere ganar la libertadores con Platense es un genio Te quiero mucho Palermo
“Vamos a ir al Maracaná con la mayor de la confianzas para enfrentar a Fluminense. Vamos con la seguridad de hacer un buen partido para que después sea el definitorio en Vicente López”. Firma: Martín Palermo, en @TyCSports.
5
4
203
9,999
Farid retweeted
GLM-5.3 is now open-weight. Our most capable model for agentic coding and cyber defense is now available to download, run, and customize. Weights: huggingface.co/zai-org/GLM-5… Tech blog: z.ai/blog/glm-5.3
276
952
8,560
1,480,850
Vivo preguntándome cómo podemos hacer más, automatizar más e ir más rápido con IA. Pero cada tanto también vale frenar intelectualmente (no tecnológicamente, no es posible) y preguntarnos hacia dónde estamos yendo gatesnotes.com/a-turbulent-a…
19
😢
🟣 Murió el músico uruguayo Rubén Rada a los 83 años. corta.com
17
Lo veo más como una expresión de deseo del mercado que como una realidad. 9 de cada 10 devs/pms/ux/etc que conozco no están en condiciones de convertirse en product engineers, fdes, ai plats. Son roles importantes y que ganarán importancia pero escasos.
Roles are collapsing. If you're a "Software Engineer" today, you may be a "Product Engineer" soon - with many more responsibilities. So: - Learn design/UX/QA/PM skills. - Focus more on the user, and less on the syntax. Image via Gartner.
1
26
Platense es el equipo en el mundo 🌍 que menos tiempo (8 años) necesitó desde la tercera categoría (2018) para lograr todo esto: ✅ Ser CAMPEÓN en Primera División (2025) ✅ Jugar su PRIMER torneo internacional (2026) ✅ Estar entre los 8 mejores del continente 👉 Supera al Ipswich Town 🏴󠁧󠁢󠁥󠁮󠁧󠁿 que estuvo a un solo partido de conseguirlo (cayó ante el Milan 🇮🇹 en 8vos de la Copa de Europa en 1963, 6 años después de haber estar en la 3ra categoría (1957)
20
62
595
79,256
Farid retweeted
🧩 DeepSeek Harness v0.1 is now available in Developer Preview! 🔹 We’re opening it up to developers building agent harnesses worldwide and open-sourcing the codebase in MIT license. 🔹 Powered by the Cordis meta-framework, DeepSeek Harness is an agent harness built around one core idea: Everything is a plugin. Models, tools, skills, sessions, sandboxes, filesystems, loops, orchestration, and UI are ALL implemented as plugins, and can be mixed, matched, replaced, and extended. Try it now! github.com/deepseek-ai/deeps…
848
2,455
20,119
4,470,504
DeepSeek silently released V4-Pro 0813, up 15.8% on Terminal Bench from their April Preview model, with Fable 5 performance at ~57x cheaper cost. 1.6T param, 49B active, 1M context. This is the best price-to-perfomance model on the market right now. Available in ClinePass now!
63
El CoT estaba cifrado. Pero otros modelos podían leerlo. Muy buen paper sobre cómo convertir modelos en "decryption oracles" y extraer reasoning, secrets y PII de traces supuestamente opacos
We can finally talk about it: We found a way to extract hidden reasoning of frontier models using a vulnerability in the APIs of every frontier AI company. We verified that our reasoning token count matches billed API thinking tokens 1:1 for most of the prompts we queried.
1
40
We just launched Xirp, a vendor-neutral agentic development environment. One place to manage agent sessions across @ClaudeDevs, @GeminiApp CLI, and @OpenAI Codex. 1,300+ @Spotify engineers already use it. Now it's available for you to try. Learn more at xirp.spotify.com.
636
824
11,566
5,260,838
Fricción importante en el onboarding de la app nueva de @BancoNacion . En el campo de caracteristica nacional hay que colocar codigo de area (ej 11) o el codigo no llega, lo adiviné de pedo pero estuve una hora 🫣
1
2
144
ARC AGI Explained
8
6
187
10,547