Training AI Engineers on YouTube, Substack and our courses. Co-founder @towards_ai. Ex-Ph.D. student @Mila_Quebec. Author of Building LLMs for Production.

louis@towardsai.net
I am super excited to finally announce that we ( @towards_AI ) released our first independent industry-focussed course: From Beginner to Advanced LLM Developer. Put a dozen experts (frustrated ex-PhDs, graduates and industry) and a year of dedicated work, and you get the most practical and in-depth LLM Developer course out there (~90 lessons). It is a one-stop conversion for software developers, machine learning engineers, data scientists, aspiring founders or AI/Computer Science students. We think many millions of LLM Developers will be needed to build reliable customised products on top of foundation LLMs and achieve mass GenAI adoption at companies. We want to help you lead this new field! One of the reasons I quit my PhD in AI (one year + 8 days ago) was to build practical solutions that will help others in the real world and improve what exists. Along with our book “Building LLMs for Production,” this course is our attempt to achieve this goal. While I love the world of academia, when I first stepped into the startup world in 2019, it felt like I pretty much knew nothing all over again. I needed to understand the problems in the real-world application of AI and build solutions for it. Not just research but real models, real products, and real people using them. But here’s the thing: grasping these challenges is merely the first step. For the ‘how’ part of it, you need to get into the code, the architecture, the models, the APIs, the deployments, the trials and errors, and the complex and wide varieties of frameworks - you don’t have time to reinvent the wheel in a startup! This discrepancy between academia and the industry is even worse in the LLM era. So we’ve gathered everything we worked on building products and AI systems and put them into one super practical industry-focused course. Right now, this means working with Python, OpenAI, Llama 3, Gemini, Perplexity, LlamaIndex, Gradio, and many other amazing tools (we are unaffiliated and will introduce all the best LLM tool options). It also means learning many new non technical skills and habits unique to the world of LLMs. Even though the course is super practical (oriented towards building a real-world project), we believe the course teaches concepts that will stay relevant for a long time even as LLMs get better, such as reducing hallucinations, customising to specific companies and tasks, teaching how to work with them, some cool theory, practical tips and more. The only skill required for the course is some Python (or programming) knowledge. We cover the full stack of learning to build on top of foundation LLMs - from choosing a suitable LLM application to collecting data, iterating on many advanced techniques (RAG, fine-tuning, agents, audio, caching), integrating industry expertise and deploying. Our students will create a working product, which we certify, and we also provide instructor support in our dedicated Discord channel. This could become the seed of a startup, a new tool at your company or portfolio project for your LLM job applications. You can find all the lesson titles and more information on the course page (or DM me). I also want to thank some of our amazing team at Towards AI that helped achieve this, including @omar_solano1 , Fabio , Rucha , @_LouiePeters , Arjun, Jaiganesan and everyone in the team that contributed to this course. A big thanks to some of our reviewers and my friends for the support, @daansan_ml , @SerranoAcademy , @rohanpaul_ai , @tunguz , @antgrasso , @ceobillionaire , @Scobleizer , @swyx , @hasantoxr and many others... “From Beginner to Advanced LLM Developer” is now available on the Towards AI Academy: tinyurl.com/5n9y32y4 #llms #llm #llmops #llmdeveloper
29
30
186
67,335
Are open models catching up? The best open-weight model vs the best closed model on our internal writing benchmark, by release date. An open model, Kimi K2, briefly took #1 in September 2025. The gap shrank to just 101 Elo in February. Then Claude Opus 5.5 blew it back open to 447. Still, the best open writer today, GLM-5.3, costs 30x less per script.
3
391
Opus 5.5 managing ElevenLabs' API with a cloned voice is an extremely powerful and mind-blowing combination
8
617
Protesting is not the right move… AI will integrate our lives wether we want it or not. The thing we should control is how will it affect our lives. Something we can actually have more control over. We use calculators when we want to. I don’t use it for calculating tip as I learned how to do that mentally. But I use one for very important calculations (well, used to use one, now ChatGPT uses it for me). I believe AI is and will cause a big transformation. But so did the calculators, electricity and the internet. We cannot ask to pause the progress or to slow down. It’s simply unrealistic. But we can create the guardrails and shape laws to optimise AI’s integration into our society. Many companies are working in this, and hopefully governments will invest even more there. Not to push or restrict progress, but to evolve and iterate with it. We should and can act. It’s not too late, and is much more powerful than protesting (I think).
3
460
as per popular demand, here's every Opus 5.5 effort setting (on our internal writing benchmark, 167 model config in total:) max: #1, 2600 Elo, $3.43 and 17 min per script xhigh: #2, 2399 Elo, $0.86, 4.4 min high: #3, 2342 Elo, $0.34, 1.7 min adaptive (no effort flag): #6, 2264 Elo, $0.21, 1.0 min medium: #10, 2199 Elo, $0.21, 1.1 min low: #20, 2144 Elo, $0.17, 0.8 min The previous #1, Claude Fable 5.1 on max, sits at 2303 for $3.15 per script. Some high-level insights: - xhigh beats Fable by almost 100 Elo at about 1/4 of the price - high beats it too, at roughly 1/9 of the price - max is 200 Elo above everything else, but it costs 4x xhigh and takes 17 minutes per script - If you pass the no-effort flag with a Claude subagent, Opus 5.5 picks for itself and lands between medium and high For most writing work, high seems to be the sweet spot. Use Max when you want the best draft and can wait. I assume Max is great here because of our detailed, extensive prompt. 10 script tasks, 5 scripts each, scored blind by three AI judges from three different labs.
5.5 is IN-SANE. It almost broke our benchmark. I thought it was a bug. Claude Opus 5.5 is the new #1 on our writing benchmark, and it is not close at all. 2631 Elo. Second one, Fable, is at 2324. That is a 307 point gap, the largest single jump we have recorded since we started running this in June 2026. It is also the first model to clear 91 out of 100 on our rubrics. One thing to note: at max effort it takes 17 minutes and $3.43 to write one script. It is the slowest configuration on the entire board. By FAR. Not the most expensive one, but in the top 5 most expensive. More details on the effort and thinking levels of Opus 5.5 👇
12
9
220
39,024
Our internal writing benchmark uses 3 AI judges: Claude, GPT and DeepSeek. So do they favor their own family? The GPT judge scores OpenAI models 2.3 points higher (out of 100) than the other two judges do. The DeepSeek judge gives DeepSeek models +1.7. The Claude judge gives Anthropic only +0.4... but Google +1.3 😅 Hoping to add Jev as a fourth soon... waiting on @typesafeai ;)
5
634
Pick your LLM Writer 😂 (One-shot prompt to Opus 5.5 asking it to make a nice/fun and useful 8-bit animation from the benchmark; super unclear prompt, haha)
2
5
621
You can pay 429x more for writing that is 1.7% better. Early results from our internal creative writing benchmark: we have models write full YouTube scripts for our channel, 10 real tasks, 5 scripts each, scored out of 100 by three AI judges against our own edited references. GLM-5.3 Flash writes one script for $0.0074 and scores 88.2. Claude Fable 5.1 at max effort scores 89.7 and costs $3.15. That is the best score we measured under a cent, against the best score we measured at any price. Plenty of pricier setups score worse, so spending more does not buy that gap on its own. I would happily pay Fable money if it cut my editing time. That is the part worth measuring from your own experience: same prompts, model names hidden, then count the scripts you would actually publish and the minutes you spend fixing them.
3
1
10
1,164
Opus 5.5 is the worse release since ChatGPT I need to wake up at 5am tomorrow and can’t go to sleep It’s so fun to one shot tons of amazing stuff 🤯 It almost feels like we’ve just unlocked a next level of intelligence/understanding requiring much less detailed instructions
1
29
2,201
Opus 5.5 one-shotted this from our benchmark. This is a simple one, but so cool it's easy to do now haha It's all 95 models on our internal writing benchmark, each entering on its release date. The #1 spot changed hands 11 times in two years (well, I started doing this work in June, so all models are still available from June 2026 onwards). An open model, Kimi K2, held it for about three weeks in 2025, and 4 of today's top 10 are open weights.
4
1
6
834
prompt: can you generate an html animation of the model progress over time by model release date in an mp4 file frame by frame to sweetly show each model that got SOTA etc following our model orders we have on the benchmark with ELO taking the 161 models we currently have but showing them from their release date so that we see a nice animation of elo growing and models going up and down similar to this kind of animation if you can figure to take screenshots of this using gemini piped.video/riOhTYNJ2BY
1
450
5.5 is IN-SANE. It almost broke our benchmark. I thought it was a bug. Claude Opus 5.5 is the new #1 on our writing benchmark, and it is not close at all. 2631 Elo. Second one, Fable, is at 2324. That is a 307 point gap, the largest single jump we have recorded since we started running this in June 2026. It is also the first model to clear 91 out of 100 on our rubrics. One thing to note: at max effort it takes 17 minutes and $3.43 to write one script. It is the slowest configuration on the entire board. By FAR. Not the most expensive one, but in the top 5 most expensive. More details on the effort and thinking levels of Opus 5.5 👇
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
87
93
1,800
248,048
Opus 5.5 has three effort settings. All three beat the model that previously led the board. Default: 2283 Elo, $0.21, 1 minute. High: 2356 Elo, $0.34, 1.7 minutes. Max: 2631 Elo, $3.43, 17 minutes. The old #1 (Fable) scored 2324 at $3.15 and 10 minutes. So Opus 5.5 on its cheapest setting is already close to it at 15x less money and 10x less time, and the high setting passes it outright while still costing a tenth as much. Max effort buys a further 276 Elo. Big jump, but it also costs 10x more than high and takes 10x longer. The general lesson we keep seeing: start on the cheaper thinking setting, then make the expensive one earn the upgrade through your own prompts and needs.
7
2
81
28,330
Fable is making our work worse. Or have we just collectively accepted that better models need less review? Better models are getting very good at making 80% look like 100%. I've been noticing something uncomfortable in the work people send me for review. The technical people who use AI the most are producing great code, but the writing around it is getting blander. Repo descriptions, presentations, tutorials, course lessons and trainings. Complete, confident, and usually technically right. Still a pain to read. Meanwhile, some of the less technical people I work/interact with seem to be improving both quality and productivity. They use AI, but spend more time reviewing the output before sending it to someone else. I don't know how much of this is causation. Maybe I'm getting better at spotting the same generated patterns. But I think we're also getting too comfortable delegating the parts of the work that were supposed to come from us. Our upcoming book made this painfully concrete. We gave Fable a detailed plan, code, examples, and our own experiences. We iterated, refined the instructions, and regenerated. After multiple review rounds and 18 days on the first chapter, it was only starting to feel promising. That last 20% is the example you choose because you've actually lived it. The claim you disagree with. The paragraph you delete because it sounds nice but says nothing. It's basically a LinkedIn post. A reader gives you their time for those decisions. If a piece contains none of your experience or judgment, they could have asked a model for it themselves. Before you publish, read it closely enough to decide what deserves to stay, what needs your own example, and what you don't actually agree with.
3
11
2,659
I wrote the full newsletter about this, including why I think AI engineers may be particularly vulnerable: louisbouchard.substack.com/p…
1
309
GPT-6 Astra is OpenAI's flagship. GPT-6 Sol is the cheaper model positioned below it. On our writing benchmark, at each model's deepest effort setting: Sol scores 2182 Elo. Astra scores 2013. Sol costs $0.11 a script. Astra costs $0.85. Sol takes 171 seconds. Astra takes 526. The cheap model wins by 169 Elo while costing an eighth as much and running three times faster. Every effort setting of Sol we tested beats every effort setting of Astra. Astra's default setting sits at 1843, which is below where GPT-5.6 Sol was. this measures writing YouTube scripts in one editorial voice. Astra is built for agentic and engineering work, and nothing here says it is bad at those. It does say that paying the flagship price for prose is the wrong call. Check the cheap model against your own task before you assume the expensive one is better.
GPT-6 Sol is BETTER than Astra on our benchmark! Much better! We ran it over the same 10 YouTube scripts every other model on our board writes, five drafts each, scored blind by three AI judges. It is 6th out of 94 models, at $0.11 a script. Three things stand out… Its effort dial actually works. Default scores 2058, high 2086, max 2182. That is 124 Elo for roughly double the price, which is the widest effort spread we have measured on any model this month and almost the only one consistent getting better results the more it thinks. It beats GPT-5.6 Sol and Astra while costing a third as much than 5.6 Sol and much much less than Astra. And it beats GPT-6 Astra by far. More on that in the next post. if you were using GPT-5.6 Sol (or Astra) for drafting, this is a straight upgrade on both quality and price. Run it at max effort. Very nice upgrade by @OpenAIDevs !
6
3
63
23,566
GPT-6 Sol is BETTER than Astra on our benchmark! Much better! We ran it over the same 10 YouTube scripts every other model on our board writes, five drafts each, scored blind by three AI judges. It is 6th out of 94 models, at $0.11 a script. Three things stand out… Its effort dial actually works. Default scores 2058, high 2086, max 2182. That is 124 Elo for roughly double the price, which is the widest effort spread we have measured on any model this month and almost the only one consistent getting better results the more it thinks. It beats GPT-5.6 Sol and Astra while costing a third as much than 5.6 Sol and much much less than Astra. And it beats GPT-6 Astra by far. More on that in the next post. if you were using GPT-5.6 Sol (or Astra) for drafting, this is a straight upgrade on both quality and price. Run it at max effort. Very nice upgrade by @OpenAIDevs !
Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.
12
1
72
45,389
BIG NEWS!! @XiaomiMiMo V2.6 just surpassed DeepSeek V4.1-Flash AND GPT-6! It is BETTER and CHEAPER than both! #2 open model on our benchmark right now! MiMo-V2.6-Pro ranks 5th out of 93 models globally. It scores 2201 Elo for $0.011 per script. The model ranked above it costs ten times that. Fable costs 278 times more, for 129 more Elo. It also lands above Grok 4.7, which shipped at the same time and costs 27 times more. Its cheap sibling, MiMo-V2.6-Flash, lands 13th at under four tenths of a cent. It is now the cheapest thing in the top five by a factor of ten, and the top of this board is no longer a two-lab race. Thanks, China! If you are picking a writer on price, this is the new floor to beat.
Introducing Xiaomi MiMo-V2.6 — Pro & Flash. Frontier intelligence, all the modalities, built in public. 🔹 Two omnimodal models, advancing through scaled reinforcement learning 🔹 Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks 🔹 Pro scores 46 on the Artificial Analysis Intelligence Index — the highest among open-source models 🔹 Stronger coding, computer use, 3D reasoning and creative capabilities 🔹 Open model weights, technical report, RL environments and training code Blog:mimo.xiaomi.com/mimo-v2-6
9
1
34
6,039
@typesafeai could I get in please, testing Jev as a 4th "LLM" judge would be incredible ;p
296
An average score hides what a model is actually good at. So here are those same six lab leaders on eight axes: six writing metrics, plus what they cost and how fast they are. Nobody wins everywhere. The shapes are all very different. Grok 4.7 has the cleanest prose of the six and the best tone of voice. It is also last of the six on writing quality, and you can see that trade in the notch in its shape. GPT is the best of the group on screen cues and on factual substance, and the weakest on tone. Claude Fable 5.1 has the most complete shape and costs nearly 9x the next dearest model here. DeepSeek V4.1 Flash is extremely fast and cheap. The overall score separates these six by 2.2%. The individual axes separate them by up to 5.5%. If you are choosing a writer, choose on the axis you actually need, not on the leaderboard position.
4
11
2,687
We took the best writer each of the top labs has and ran them over the same 10 YouTube script generations (creative writing). Three numbers decide which one we'd actually use at scale: what it scores, what it costs, how long you wait. The whole frontier field is 2.2% points wide. The price range across it is 248x. DeepSeek V4.1 Flash scores 88.0 for 1.3 cents in 76 seconds. Claude Fable 5.1 scores 89.7, the best writing we have measured anywhere, for $3.15 and ten minutes. That is 1.7 points for 248 times the money and 8 times the wait. Grok 4.7 shipped this week, which puts xAI in the middle of this group on price and near the bottom on speed. The expensive end of this chart buys you very little average quality
1
1
7
1,437