Developer, interested in many different tech fields. Working at Ubisoft Paris.

Manu_TechAndGames retweeted
Replying to @SebDevLab
Oui tiens le lien github : github.com/bmad-code-org/BMA… Quand tu l’installes tu peux sélectionner « bmad game dev studio », c’est un lot d’agent, de skills, de workflow pour faire au propre ton projet avec un suivi via des fichiers .md etc.. tu as juste à te laisser guider ensuite
1
1
2
24
Manu_TechAndGames retweeted
Netflix replaced their 15 years old recommendation algorithm with an LLM. It’s called "GenRec" and it completely changes how recommendation algorithms are built. For over a decade, the Netflix recommendation engine was a masterclass in feature engineering. Data scientists built thousands of complex, handcrafted features to figure out what you wanted to watch next. It required bespoke architectures. Massive infrastructure. Constant manual tuning. Netflix threw all of it away. They built GenRec, an LLM-backed ranker. Instead of translating your behavior into complex math, they just turn your watch history and metadata into a natural language sentence. They feed that raw text into a foundation LLM. The AI simply reads your behavior like a story, understands your evolving tastes, and scores the entire catalog in a single forward pass. No manual feature engineering. No complex bespoke architectures. Here is the part that should terrify traditional data scientists. This text-based LLM didn't just match the highly tuned production system Netflix spent years perfecting. It beat it. And it achieved those statistically significant gains using roughly 40x fewer labeled training examples. We are watching a massive paradigm shift in real time. The most complex predictive algorithms in the world are being replaced by models that just know how to read. If Netflix can replace their core product engine with an LLM, what complex system in your business is about to become obsolete?
109
521
3,727
311,246
Manu_TechAndGames retweeted
"People won't play games made with a.i." - every anti-ai hater as they cope The data:
AI Game Database is Live! gamedb.ai Made for builders to check trending AI games on Steam. If you're developing a game, verify it to get featured!
59
37
658
141,149
The Russian AI campaigns to influence people are horrible.
Apparently the French, Dutch, German and Danish mothers all had the exact same thought, in the exact same order, on the exact same day. Remarkable.
5
Manu_TechAndGames retweeted
I tested Sonnet 5.5 and couldn't believe my eyes. 2 real repos, 105 bugs, find and fix what you can. It took it seriously. The results: - Sonnet 5.5 (max): 55.5 - GPT-6 Astra (max): 45 - GPT-5.6 Sol (max): 43.5 - Fable 5.1 (max): 43 - Opus 5.5 (max): 41.7 - Sonnet 5.5 (xhigh): 39 The secret? Sonnet 5.5 (max) might be cheap and fast for most tasks, but it's the least lazy model I tested. It leads in bug hunting by spending turns. It took about 6× Astra's turns and about 3× Opus 5.5's for 10–14 more bugs. xhigh cut its turns by more than half and dropped to 39, below Opus 5.5 max, which used fewer turns on average: - Sonnet 5.5 (max): 1,330 turns (n=2) - GPT-6 Astra (max): 222 turns (n=3) - GPT-5.6 Sol (max): 1,083 turns (n=2) - Fable 5.1 (max): 459 turns (n=1) - Opus 5.5 (max): 476 turns (n=3) - Sonnet 5.5 (xhigh): 588 turns (n=1) More effort levels for Sonnet 5.5 with n=3 drop in this thread today 🧵
114
101
1,242
116,708
Manu_TechAndGames retweeted
Destroy any website with a stickman. Have fun: destroy.spritefusion.com
953
6,492
62,653
4,394,400
Manu_TechAndGames retweeted
Anthropic has launched Claude Sonnet 5.5: it scores 56 on the Artificial Analysis Intelligence Index, just 2 points behind Opus 5.5 (max), but at the highest Output Tokens per Task we’ve seen With max effort, Sonnet 5.5 gains 18 points over Sonnet 5 and to #2 on the Intelligence Index behind only Opus 5.5 (max). Anthropic has priced Sonnet 5.5 identically to Sonnet 5 at $0.2/$2/$10 per 1M cache input/input/output tokens, however it outputs a higher number of Output Tokens per Task and costs $7.60 per task (~50% higher than Sonnet 5’s Cost per Task) Key takeaways: ➤ Meets leading models on agentic terminal use and knowledge work: in Terminal-Bench 4.0, Claude Sonnet 5.5 reaches 64% against 60% for Opus 5.5 and GPT-6 Astra. On AA-Briefcase (1811 vs 1822 Elo), GDPval-AA (1844 vs 1846 Elo), and AutomationBench-AA (71% vs 70% headline score), Sonnet 5.5 reaches parity with Opus 5.5, albeit with significantly higher token usage to achieve it ➤ Heaviest token use we have measured: at max effort, where it reaches performance nearing that of Opus 5.5, Claude Sonnet 5.5 used ~193k Output Tokens per Intelligence Index Task. This is the highest token use we have measured on around 60% higher than Opus 5.5 (max) or Sonnet 5 (max) and ~7x GPT-6 Astra (max) ➤ Pricing remains at $2/$10 per million tokens of input/output, matching GPT-6 Sol. At this pricing Claude Sonnet 5.5 sits off the Intelligence vs. Cost per Task Pareto Frontier. At high effort levels it sits behind Opus 5.5, while lower efforts have GPT-6 Astra or Sol configurations delivering equivalent performance for lower cost. The high effort setting is the most competitive on this basis, sitting very narrowly behind GPT-6 Sol on Intelligence at effectively the same Cost per Task ➤ Behind Opus 5.5 on factual knowledge and scientific reasoning: as a smaller class model, Sonnet 5.5 still lags on factual knowledge in AA-Omniscience compared to Opus 5.5. It scores 54% against 66% for factual accuracy, though with a lower hallucination rate (47% against 59%). It also sits ~6 points lower on Humanity's Last Exam and SciCode compared to Opus These evaluations were conducted on a pre-release deployment of Claude Sonnet 5.5, which Anthropic found to have a bug that can degrade responses to requests that use structured outputs. This is fixed for the public release and Anthropic expects minimal change or slightly understated performance, but we will be re-running relevant evaluations soon. Other model details: ➤ Context window: 1 million tokens with image and text input, unchanged from Sonnet 5 ➤ Pricing: unchanged from Sonnet 5’s latest $2/$10 per 1M input/output tokens; cache writes at $2.5, cache reads $0.2 ➤ Effort settings: five (low, medium, high, xhigh, max). Intelligence Index evaluations were run at all five with Anthropic's default fallback enabled. We see Sonnet 5.5 fall back in ~0.1% of tasks across the Intelligence Index, primarily in TerminalBench 4.0, falling back to Sonnet 5 in all cases.
164
292
3,245
371,641
Manu_TechAndGames retweeted
Claudes lie and betray more than any other model family in Diplomacy For all agents, lying correlates to being better at Diplomacy - except for GPTs GPT-6 Astra is in a league of its own wrt skill, and doing so without needing to betray/lie
49
87
1,225
75,239
Interesting
How exl3 works, maybe Instead of each weight owning 3 bits, we have 16 bit windows with a sliding window This gives us more possible values for weights, at the cost of computation spent: 1. quantising the model: convert bf16 exl3 2. reading weights during inference: unpacking
11
Manu_TechAndGames retweeted
ALERT🚨: Just 4 TPUv7s were able to get close to 1 giant @cerebras wafer in terms of interactivity due to the cracked @inferact 200IQ software using megakernels. Hopefully, the Cerebras DevX team won’t be cope-tweeting after seeing how great TPUv7 is.
58
84
1,458
208,274
Manu_TechAndGames retweeted
A Qwen3.8 Flash Next specialized for coding! huggingface.co/ISTA-DASLab/Q… 58 GB. They used RCO for expert pruning (not REAP). I'll evaluate it next week for long horizon agentic coding. It could be the best model for this type of task under 60 GB.
15
30
394
21,221
Nice framework to create videos from code, now that we have AI.
Not many remember but 5 years ago I started working on my FAST video rendering framework. I have never released it but when I've seen that people wait 12 hours I realized that I have to do it now Introducing fframes. This video was vibed in 48 minutes and rendered in 36 seconds
1
1
25
Strata is an inference engine made specially for ok device inference. It brings dedicated optimizations, and should be much faster than other engines.
Just doubled the prompt processing speeds from 539 token/s to 1290 tokens/s. Qwen3.8-Flash-Next on consumer GPU is now blazing fast. Next on the list is further optimizations on kernels and swapping which should boost results by 3X-4X. github.com/Niko1221/Strata
1
66
I'm not sure using a real human brings anything here , but probably more visibility. I still think this is the worst thing to do.
A robot whips a rope to knock an apple off someone's head, using a policy trained only in simulation. Introducing DeformX (Oral @ IROS 2026): Cosserat rod physics co-simulated with NVIDIA Isaac Sim, so cables bend, twist and collide like the real thing. deformx.github.io
64
Manu_TechAndGames retweeted
Ember-1 is a specialized model from Fireworks Research designed to make every token go further. Built on Kimi K3, it produces shorter reasoning traces, using roughly 40% fewer tokens while maintaining top-tier quality. fireworks.ai/models/firework…
21
34
536
158,044
Manu_TechAndGames retweeted
Introducing open-slide 2.0 🎉 The slide framework built for agents, now with: › A new visual editor › Editable pptx export › A redesigned UI Let your agent build the deck. Make the final touches yourself. Here's what's new ↓
48
130
1,832
123,553
Manu_TechAndGames retweeted
microsoft quietly open-sourced one of the best ai data analysis tools out there. it's called data-formulator. connect any source, csv, postgres, bigquery, even a live url, then build charts with a mix of drag-and-drop and plain english. the ai agent writes the sql and transforms underneath and hands back a chart you can actually edit, not a code dump. > drag fields onto x/y/color and it builds the chart > anchor a cleaned result so follow-ups don't drift back > branch any chart to explore a variation > connect live data with auto-refresh bring your own model key and run it locally.
11
106
574
31,607