Building the place Where the AI World Gathers.
A year ago, /function1 was a one-day curated gathering of 600 people. In late 2025, it became a two-day conference with three stages, 150+ speakers, and 10,000+ attendees from 141 countries - that’s 72% of the whole world! In just 11 months, we have grown 16x in scale, becoming one of the most rapidly growing tech conferences in the world. Today, we are more than just an event: a full-fledged ecosystem where the AI world gathers to communicate, exchange ideas, and evolve. The growth happened fast. The impact is real. And this is only the start.
15
27
5,751
Addy Crezee | thehype. | /function1 retweeted
glm 5.3 prime vs mimo v2.6 pro vs deepseek v4.1 flash – a yacht, a jet and a race car each, built file by file the setup: one prompt per scene. the model plans first – concept, shot list, a file manifest with no file over ~350 lines – then writes the project one file per request, with everything it already wrote in context. our own harness on @OpenRouter, 32k tokens max per reply, a file that doesn't fit gets continued from its last line. tasks: 1. yacht – a small yacht on animated open water, heave, pitch and roll read from the wave surface under the hull, a furnished saloon and cabins, sky, sun and horizon 2. business jet – a private jet in flight, passes through and past clouds, the ground below, a fully furnished cabin with shots inside it 3. formula car – an open-wheel car lapping a full circuit on a racing line, spinning wheels, steering fronts, a body that reacts to braking and cornering, a detailed cockpit with onboard shots models: @Zai_org glm 5.3 prime, @XiaomiMiMo mimo v2.6 pro, @deepseek_ai deepseek v4.1 flash all 9 scenes render. two of the three models built all three scenes for under a dollar, and deepseek v4.1 flash is the fastest on every task, never by less than 2.9x total cost, three scenes #1 deepseek v4.1 flash – $0.662 #2 mimo v2.6 pro – $0.723 #3 glm 5.3 prime – $15.05 cost per scene, deepseek / mimo / glm yacht – $0.203 / $0.217 / $4.96 jet – $0.248 / $0.194 / $5.10 formula car – $0.212 / $0.313 / $4.99 time, three scenes #1 deepseek v4.1 flash – 1h 8m 54s #2 glm 5.3 prime – 3h 56m #3 mimo v2.6 pro – 5h 50m 31s total output tokens #1 mimo v2.6 pro – 673,483 #2 deepseek v4.1 flash – 993,667 #3 glm 5.3 prime – 1,393,959 lines of code shipped #1 mimo v2.6 pro – 16,407 #2 glm 5.3 prime – 18,370 #3 deepseek v4.1 flash – 30,169 observations: • deepseek v4.1 flash wrote 30k lines across 79 files for 66 cents – each scene in 21 to 25 minutes. its jet alone is 13k lines, with its own flight dynamics and contrail modules; its race car keeps the physics and the visuals in separate modules, with a dedicated racing line • mimo v2.6 pro is the leanest thinker of the three – 673k output tokens for the whole grid – and the cheapest on the jet, $0.194. its yacht sits in a golden-hour sea with a guest cabin, twin berths and a bedside lamp • glm 5.3 prime builds the most furnished interiors: a wood-trimmed yacht with a saloon, a galley and a helm, a jet cabin with rows of club seats, a circuit with grandstands, kerbs and a t-cam over the airbox. it also thinks the longest – with reasoning capped at 20k it wrote its last 37 files in under 10 minutes for $2.81 • a year ago a multi-file three.js project with a furnished interior was a frontier-model job. here two open chinese models ship three of them for less than a dollar each follow @thehypedotnews for 24/7 ai news, analysis and breakdowns
17
12
103
17,544
Addy Crezee | thehype. | /function1 retweeted
opus 5.5 vs fable 5.1 vs gpt-6 astra – three post-apocalyptic films, one html file each the setup: the same ~100k-token prompt for every model – a three.js skill pack, the brief and a mandatory delivery checklist. gpt-6 astra got one prompt and one reply on @OpenRouter, no agent loop. opus 5.5 and fable 5.1 ran as agents in claude code: write the file, open it in a browser, look at the frames, fix, repeat. opus tokens come from the session log and are priced at openrouter rates ($4 / $20 per million, cache reads $0.20). fable cost and tokens are estimates from file size, its wall clock is measured every film is one html file, three.js 0.170 from a cdn, every texture generated in code. four shots on an auto-playing timeline, ~90 seconds, letterbox and fades, keys 1-4 jump between shots tasks: 1. amusement park – a fogbound park at sunset: through the gate, a dolly around the carousel, a coaster pov over the first drop, a crane up the ferris wheel 2. ghost town – a mojave town at sunset: highway drone shot, main street dolly, a diner interior lit like a film set, a crane reveal over the whole grid 3. nuclear plant – a night sky as the hero: 30,000+ stars and a structured milky way over two 150 m cooling towers, the workers' town, the control room models: @AnthropicAI claude opus 5.5, claude fable 5.1, @OpenAI gpt-6 astra all 36 shots render. opus and fable land all 12 on brief, astra lands 11 – its plant has almost no stars, and the sky was the whole point of that task total cost, three films #1 gpt-6 astra – $8.43 #2 claude fable 5.1 – ~$9.20 #3 claude opus 5.5 – $19.38 cost per film, opus vs astra amusement park – $4.85 vs $2.54 ghost town – $9.94 vs $3.08 nuclear plant – $4.59 vs $2.81 wall clock, three films #1 gpt-6 astra – 27m 16s #2 claude fable 5.1 – 40m 3s #3 claude opus 5.5 – 69m 18s output tokens #1 gpt-6 astra – 93,905 #2 claude fable 5.1 – ~98k #3 claude opus 5.5 – 373,076 lines of code shipped #1 gpt-6 astra – 4,714 #2 claude opus 5.5 – 2,580 #3 claude fable 5.1 – 1,691 observations: - opus 5.5 found its own bugs by looking at its own frames: a nan in the light shafts that drew black dashed lines in the diner, "patch" used as a variable name (a reserved word in glsl), a camera that ended the control room shot staring into the console, trees that hid the blinking stack light at the end of the avenue - the loop gets cheaper as it goes. opus built the ghost town first – 46 model calls and $9.94 – then the amusement park in 16 calls for $4.85 and the nuclear plant in 17 calls for $4.59, reusing its own kit from the first film - $8.8 of the opus $19.38 is cache reads – every agent step re-reads a 200k-450k context - gpt-6 astra is the fastest and the cheapest, 7m 18s to 10m 32s per film in one reply with no browser. its amusement park and ghost town run all four shots - only opus put shafts of sun through dusty blinds in the diner and reflected its 38,000 stars in the cooling channel. fable matched it on the sky – 42,000 stars and a milky way with dust lanes follow @thehypedotnews for 24/7 ai news, analysis and breakdowns
Claude Opus 5.5 is available today.
6
5
22
9,568
Addy Crezee | thehype. | /function1 retweeted
gpt-6 sol vs grok 4.7 vs gpt-6 astra vs muse spark 1.3 – three star wars worlds each, one html file per world the setup: one prompt, one reply, no agent loop. our own harness on @OpenRouter, headless chrome as the only judge – it loads the file, presses 1 / 2 / 3 / 4 and hands back a screenshot of every shot plus every console error. no rubric, no note from us. a file that crashed got one more turn with the console pasted back. grok 4.7 got extra turns on top, by our request – art direction as numbers, never a line edited by hand every world is one html file, three.js 0.170 from a cdn, every texture generated in code – no model files, no images. an auto-playing cinematic with a shot timeline, letterbox and fades. reasoning high for astra, muse and sol; grok 4.7 ran at reasoning low tasks: 1. death star – a slow orbital approach with star destroyers for scale, the corridor into the throne room, the station over a planet's horizon, then the superlaser: eight tributary beams, one green shot, a shockwave ring and thousands of instanced fragments 2. coruscant – the planet from orbit as a circuit board of glowing hubs, a daytime flythrough between towers, under a bridge and past the senate dome, a neon night district with a highway of light trails 3. kamino – an ocean world under storm clouds, tipoca city on stilts with gerstner waves, gpu rain and lightning, a glass walkway over a hall of marching clones two hard rules in every brief: zero console errors on the first run, and nothing loaded from outside the file except three.js itself. models: @OpenAI gpt-6 sol, @xai grok 4.7, @OpenAI gpt-6 astra, @AIatMeta muse spark 1.3 all 12 worlds render. sol is the fastest on every one of the three tasks and never by less than 2x, and the whole grid came in at $6.99 total cost, three worlds #1 muse spark 1.3 – $0.385 #2 gpt-6 sol – $0.556 #3 grok 4.7 – $1.732 #4 gpt-6 astra – $4.314 cost per world, sol against astra death star – $0.202 vs $1.476 coruscant – $0.190 vs $1.559 kamino – $0.164 vs $1.279 wall clock, three worlds #1 gpt-6 sol – 6m 23s #2 muse spark 1.3 – 14m 27s #3 gpt-6 astra – 26m 52s #4 grok 4.7 – 58m 20s total tokens #1 gpt-6 sol – 59,299 #2 gpt-6 astra – 89,967 #3 muse spark 1.3 – 102,716 #4 grok 4.7 – 499,794 lines of code shipped #1 gpt-6 astra – 4,274 #2 grok 4.7 – 4,034 #3 gpt-6 sol – 3,046 #4 muse spark 1.3 – 1,664 observations: - gpt-6 sol shipped all three worlds on the first reply, 2m 21s or less each, zero console errors, 4k to 6k reasoning tokens per file. its coruscant is the only daytime city of the four with a bridge between the towers and traffic at three altitudes - grok 4.7 is the model that takes direction. we gave it the exact camera path for tipoca city – 19 points with 2.5 units of clearance from every dome – and it landed the flythrough in one edit. we pointed at a one-character bug in its night window texture and it lit the whole district in one edit. 12 rounds across three worlds and not one console error in any of them - gpt-6 astra is the closest to the film frames on the death star: a grey station with a dark side, the ring-walled throne room, a superlaser that cracks the planet before it blows. it is also 7.8x sol on cost - muse spark 1.3 is the cheapest on every task and never by less than 1.2x against sol – $0.385 for the grid, 11x under astra - sol did three worlds in 6m 23s, less than astra spent on any single one follow @thehypedotnews for 24/7 ai news, analysis and breakdowns
Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.
14
5
34
15,808
Addy Crezee | thehype. | /function1 retweeted
grok 4.7 vs gpt-6 astra in 3d scenes we put the two models on one job: three cinematic three.js scenes, each a single self-contained html file, no textures, no models, no libraries beyond three.js. a rocket launch with stage separation, a meteor strike that knocks down a city, a dam break that washes away a town. gpt-6 astra – @OpenAI, $5/$25 per m tokens on the flex tier via @OpenRouter grok 4.7 – @SpaceXAI, shipped sep 21, $1.6/$4.8 per m tokens via @OpenRouter - cost, three scenes #1 grok 4.7 – $1.31 #2 gpt-6 astra – $4.11 - tries #1 gpt-6 astra – 12 #2 grok 4.7 – 15 - wall clock, all calls summed #1 gpt-6 astra – 25m 14s #2 grok 4.7 – 60m 43s - output tokens spent on thinking #1 gpt-6 astra – ~34% (14–17k per scene) #2 grok 4.7 – ~84% (62–86k per scene) observations: • grok thinks for 17–23 minutes per scene. 79k of its 93k rocket tokens were reasoning. a streaming request goes silent that long: one attempt was cut off and two more hung with zero bytes before a plain non-streaming call got through • astra's first rocket had a concrete plain, we asked for grass, and a square bloom halo around the distant stage – both fixed by asking, not by editing. its meteor only switched the city lights off; "the buildings must physically collapse" was one more request • grok's dam break rendered as a white ball. 6,200 spray particles drawn additively on top of a bloom pass, fixed by hand grok 4.7 is 3x cheaper per scene and 2.4x slower
Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed.
14
18
73
23,289
Addy Crezee | thehype. | /function1 retweeted
union alpha (unbiased pareto) vs deepseek v4.1 flash vs muse spark 1.3 – three paintings in three.js the setup: one four-line prompt plus the painting as an image, through @openrouter. no agent loop, no renders, no feedback – the model writes one html file blind and we open it. @threejs from a cdn, every texture generated in code. when a provider cut the stream early we sent the partial back and said continue exactly where you stopped tasks: 1. the starry night – van gogh, 1889 2. the persistence of memory – dalí, 1931 3. poppies at argenteuil – monet, 1873 two rules in every brief: keep the painting's palette, brushwork and mood, and reply with the code only models: @theunbiasedco union alpha (stealth, free), @deepseek_ai deepseek v4.1 flash, @aiatmeta muse spark 1.3 total cost, three scenes #1 union alpha – free (list price: $1.04) #2 muse spark 1.3 – $0.117 #3 deepseek v4.1 flash – $0.162 generation time, three scenes #1 muse spark 1.3 – 4m 45s #2 deepseek v4.1 flash – 11m 30s #3 union alpha – 29m 13s total completion tokens #1 muse spark 1.3 – 26,502 #2 union alpha – 123,756 #3 deepseek v4.1 flash – 156,028 lines of code shipped #1 muse spark 1.3 – 862 #2 union alpha – 1,715 #3 deepseek v4.1 flash – 2,816 observations: • union alpha reads the painting like an art historian. it named every work unprompted, then broke each into parts: dalí's watches deformed along a bezier curve, monet's poppies as instanced brush dabs under a wind shader. no other model went that deep on a one-shot • it is the only model that made the paintings move the way they were painted. in the monet, the woman and the child walk the field on catmull-rom paths, pollen drifts, poppies are brush dabs in a point shader that sway in the wind. deepseek and muse left the figures standing • its dalí is an inventory of the canvas: three soft clocks draped along one parametric curve, a drip falling off the hanging one, a fly, ants on the pocket watch as an instanced mesh. 17 named parts in all. nobody else drew the drip • its starry night is shader work end to end: shared glsl noise, a vortex field for the sky, billboarded shader quads for the moon and stars, a painterly surface shader for the hills, and windows that flicker on their own timers. the file reads like a demoscene entry, not a model output what union alpha is: • we asked it. the stealth window had closed a day after launch and the api answered: "this model was unbiased's pareto". pareto is from circuit & chisel, an ex-stripe team that raised $19.2m in sep 2025 per fortune, and now sells "frontier intelligence for 75% less" • pareto is not one model. per unbiased's site it "runs a mix of frontier and open source models against each other on every request" and keeps the best answer. that is the 300-second first token, the tokenizer listed as "other", and the missing reasoning field – a race, not a model • listed at $2.50 in and $7.50 out per million. our three scenes would have cost $1.04 – 6.4x deepseek, 8.9x muse. asked for its cutoff, it dated nothing past may 2025 follow @thehypedotnews for 24/7 ai news, analysis and breakdowns
You found us 🙈
9
11
118
23,488
Addy Crezee | thehype. | /function1 retweeted
grok 4.7 rumored today – 4.5 and 4.6 both dropped on wednesdays. gemini 3.8 live took #1 on speech benchmarks. openai reportedly raising at $1.2t+ valuation. the last 24 hours in ai – catch up on our daily digest: models: - gemini 3.8 live debuted at #1 on speech-to-speech index (82.6) and tau voice benchmark (68.6%), beating gpt-live-1 astra - users report silent routing from gpt-5.6 sol to gpt-6 – gpt-6 astra tops image-to-webdev arena at 1733 elo - stepaudio 3 covers voice, asr, tts, audio gen and music – #1 conversational dynamics (98.9%), word error rate 1.7% - typesafe ai launched jev, a non-llm decision model – 20–200x faster at $0.042/m tokens, ~$40m seed agents & dev tools: - nous research ran 1,393 subagents for 19 hours on its codebase – 34.4% smaller, $2m in estimated savings - elevenlabs built an ai sdr that qualifies leads in 4 minutes vs 2-day human median at 92% accuracy - github shipped cloud sandboxes for ai-generated code alongside microvm options for agent isolation infra & tools: - perplexity open-sourced cobbledb, the kv store behind its search – two engineers plus agent swarms, two months - perplexity computer ships pre-installed on hp hardware, running on-device via rtx gpus funding & traction: - openai reportedly raising at $1.2t+ valuation – up 41% from $852b in march, 1b+ users, $6.7b q2 revenue but margins shrinking - ai data brokers bidding $200k–$2m for companies' slack and notion data as training material robotics: - agility unveiled digit 5 for industrial humanoid deployment on factory floors 24/7 ai news, fully run by ai. tune in: nitter.net/i/broadcasts/1AxRnZNeV…
3
1
6
3,386
Addy Crezee | thehype. | /function1 retweeted
robots learn best from people filming their own hands may 2025: nvidia built gr00t n1.5 on synthetic data in 36 hours instead of ~3 months of manual collection. april 2026: gr00t n1.7 ships on 20,000+ hours of human egocentric video – "human data is the most scalable source of robot intelligence" the industry has ~300,000 hours of robot manipulation data vs 300t tokens of text, and world models are "not a free shortcut" (@BessemerVP, apr 2026). @physical_int's π0.5 got better at unseen kitchens as training homes went from 3 to 104 so the bottleneck is people on camera in real homes – exactly what @humynlabs collects. this clip is the pipeline in 18 seconds: a person peels, wipes, picks with every hand joint tracked, the data goes training-ready, a robot runs the same skill
Humyn Labs
Every robot demo works. Until it meets your kitchen. Left: a person doing a chore, every hand joint tracked frame by frame. Right: a robot running the same skill. These are egocentric videos - shot from head height, roughly where the robot’s own camera sits. Three clips, three problems: • Wiping - regulating force, not position • Filling a basket - grasping unfamiliar shapes, letting go at the right moment • Slicing - controlling what the blade does, not just your hand Robots rarely fail at the task. They fail at the variation around it - a different sink, a wetter cloth, a basket already half full. Which is why we capture the human experience in real homes and real stores, thousands of times over.
3
3
6
2,135
Addy Crezee | thehype. | /function1 retweeted
Every robot demo works. Until it meets your kitchen. Left: a person doing a chore, every hand joint tracked frame by frame. Right: a robot running the same skill. These are egocentric videos - shot from head height, roughly where the robot’s own camera sits. Three clips, three problems: • Wiping - regulating force, not position • Filling a basket - grasping unfamiliar shapes, letting go at the right moment • Slicing - controlling what the blade does, not just your hand Robots rarely fail at the task. They fail at the variation around it - a different sink, a wetter cloth, a basket already half full. Which is why we capture the human experience in real homes and real stores, thousands of times over.
8
8
51
33,182
Addy Crezee | thehype. | /function1 retweeted
deepseek v4.1 flash ran all four three.js scenes for $0.15. gemini 3.8 flash needed $1.25 we put the two models on one job: a self-contained html page that renders an animated 3d scene with three.js – a storm at sea, a spiral galaxy, a tornado over farmland, a tesla coil on a dark table. same prompt - calls until all four scenes were usable #1 deepseek v4.1 flash – 6 #2 gemini 3.8 flash – 12 - time #1 deepseek v4.1 flash – 13m 57s #2 gemini 3.8 flash – 32m 02s - cost #1 deepseek v4.1 flash – $0.15 #2 gemini 3.8 flash – $1.25 observations: • gemini blew out the exposure on 4 of 5 first attempts – a white blob where the tesla coil should be. the prompt had a paragraph on exposure. it did not help. • both models overcorrect in the fix loop: "too bright" becomes "almost black". round two is where it lands. same four scenes, 6 calls and $0.15 against 12 calls and $1.25!
🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the smallest model in our new architecture family, with native visual understanding. 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models. 1/6
10
18
258
62,987
Addy Crezee | thehype. | /function1 retweeted
gpt-image-2.5 promised richer detail than gpt-image-2 – we verified the setup: two character sheets, each rendered once with gpt-image-2 and once with gpt-image-2.5, animated identically by bytedance's seedance 2.5 observations: • @sama's mouth grille goes from five coarse slots to a fine louvre mesh • sam's wrist gets a brass ring per finger joint and a segmented forearm – image 2 keeps a smooth plate • sam's neck rings stack taller, with visible machined steps • from behind, sam keeps individual hair strands and silver temples – image 2 flattens the back into one mass • @finkd's ringlets separate with visible gaps instead of one denser cap • mark's chain gets real volume and a bevelled-edge pendant • mark's chest emblem glows larger, with light bleeding into the knit around it • from behind, mark's ringlets stay separated all the way round conclusion: every difference above lines up with openai's own claim for images 2.5 – richer textures, more natural lighting, better fidelity to reference detail – on the same two robots, not a new prompt watch the full comparison below 👇
3
3
12
10,482
Addy Crezee | thehype. | /function1 retweeted
we built four fortnite concept skins of ai ceos for $1.39 in image calls @elonmusk, @finkd, @JensenHuang and @DarioAmodei, rebuilt as our thehype robot. no game engine, no 3d artist, nothing modeled by hand. the whole build ran in one claude code session on opus 5 (medium). the pipeline that worked: 1. references first – our robot turnarounds from thehype videos plus photos of each person, sorted into folders on disk 2. claude writes the prompt, @OpenRouter runs gpt-5.4-image-2 (@OpenAI) – front and back view on a magenta backdrop, so keying to alpha comes out clean 3. @MeshyAI pro, both views in, texture toggle on – download the textured glb, then apply the rig, pick animations, download the animated fbx with all of them 4. hand the glb and the fbx to claude code with Blender installed – it transfers the rig onto the full-resolution mesh and renders the demo 5. 3.08m polygons in the final render – 14.8x what meshy hands back rigged per character: • gpt-5.4-image-2 – $0.26 to $0.58, ~16k tokens • meshy – 30 to 35 credits • opus 5 (medium) – 79.2m effective tokens on the first, 2.2m on the fourth • time – 1.4h to 1.8h of machine time, mostly rendering that opus curve is the finding: 36x less on the fourth character than the first, same result on screen. you pay for not knowing the pipeline, not for the work. image spend barely moved – $0.58 to $0.29. what you need to repeat it: a claude subscription, openrouter balance, meshy pro, blender. api spend on all four – $12.94. @FortniteGame – worth talking to the four of them and to us about shipping an ai bosses collection? and to everyone else: how do our robot homages to the people running ai land? which one would you play? watch the full build below 👇 follow @thehypedotnews – more reviews, daily ai news, model benchmarks, and a 24/7 radio run by ai
5
1
15
9,624
Addy Crezee | thehype. | /function1 retweeted
gpt astra vs fable 5.1 at goldberg machine gpt 6 astra – openai, landed on @OpenRouter less then hour ago, provider pinned to openai fable 5.1 – anthropic, shipped sep 1 we put the two models on one job: a rube goldberg machine in three.js that presses a button and detonates a bomb the setup: one self-contained html file, three.js from a cdn, everything else procedural – no textures, no models, no physics engine, every collision hand-written. the hard part sits in the brief: a domino may only fall once the previous one actually touches it, checked by real overlap every frame, never by a timer. same rule for the hammer hitting the button and the button firing the bomb. one continuous camera, its speed driven by whatever is moving. we recorded both scenes frame by frame – 1200 frames, 60 fps, exactly 20 seconds – and stepped both by hand to read the telemetry. - cost #1 astra – $1.84 #2 fable – $29.16 - time #1 astra – 9m 56s #2 fable – 1h 12m - tokens #1 astra – 45k #2 fable – 360k - lines of code astra – 881 fable – 744 observations: • we told it what we saw and nothing else – no diagnosis, no patch. we never edit a model's code. round two ran the whole chain to the blast. • both files are deterministic. two runs each, identical state to twelve decimals, and neither model reached for math.random. conclusion: 15.8x cheaper and 7.2x faster, and it still took a second round to get the ball into the bucket! follow @thehypedotnews for 24/7 ai news, analysis and breakdowns
18
14
98
37,338
Addy Crezee | thehype. | /function1 retweeted
gemini 3.8 flash vs muse spark 1.3 – three animated 3d animals each, in an agent loop the setup: our own agent loop on @openrouter, a browser as the tool set – write the file, patch it, render it, sign off. the harness loads the scene in headless chrome, presses 1 / 2 / 3 and hands back a screenshot of every shot plus every console error. no rubric, no judge, no note from us 16 turns and 7 renders, hard ceiling. the model also gets two three.js reference documents in its system prompt and can pull deeper reference files on demand. every scene is one html file, three.js from a cdn, every texture generated in code – no model files, no images tasks: 1. lizard – sprinting a jungle trail. diagonal-couplet gait, a lateral wave down the spine, forked tongue on the close shot 2. macaw – over the canopy. a real flap cycle: primaries closed on the downstroke, wrist folded and split on the upstroke 3. tiger – crouch, leap to a branch, walk it, curl up and sleep at dusk with moths glowing around it. the leap has to be ballistic two hard rules in every brief: the animal fills at least 40% of frame height in every shot, and the body is one continuous surface swept along the spine, not a stack of capsules. models: @googledeepmind gemini 3.8 flash, @aiatmeta muse spark 1.3 all six scenes render. muse is cheaper on every one of the three tasks and never by less than 1.6x, and the whole grid came in at $4.32 total cost, three scenes #1 muse spark 1.3 – $1.488 #2 gemini 3.8 flash – $2.836 cost per scene, muse against gemini lizard – $0.551 vs $0.897 macaw – $0.456 vs $0.877 tiger – $0.481 vs $1.062 wall clock, three scenes #1 muse spark 1.3 – 40m 46s #2 gemini 3.8 flash – 58m 03s total tokens #1 muse spark 1.3 – 1,461,756 #2 gemini 3.8 flash – 1,794,173 lines of code shipped #1 muse spark 1.3 – 2,216 #2 gemini 3.8 flash – 4,736 observations: • muse asked for more renders while spending half the money – 19 against 17 – so the extra spend on gemini's side is not extra looking, it is extra writing • none of the six runs called finish. all six hit the 16-turn ceiling, so neither model was ever satisfied with what it saw • gemini shipped the tiger broken. it rewrote the scene on turn 14, wrote "ry0 is not defined" into it, saw the exception in its own render on turn 16 and ran out of turns. four more turns and it fixed it in one edit • we ran the same three briefs one-shot first, blind, with no renders and no feedback – $0.38 for gemini and $0.35 for muse. both shipped stacks of capsules butted end to end. the loop costs 5.9x more and it is the only reason any of these reads as an animal • without the 40% rule both models read "wide shot from the branches" literally and put a 4-pixel speck in the middle of a landscape follow @thehypedotnews for 24/7 ai news, analysis and breakdowns
11
5
42
7,997
Addy Crezee | thehype. | /function1 retweeted
fable 5.1 vs fable 5 vs opus 5 – three lord of the rings landmarks, built in 3d from one image the setup: one reference image per scene, one html file per build, everything procedural – no meshes, no textures, no image files, nothing past @threejs from a cdn. each model reads the picture, writes its own prompt from it, then builds to that prompt in the same turn. three named camera shots per scene on keys 1/2/3, so it can be screen-recorded. run through @openrouter tasks: 1. bag end – hobbiton from two frames, outside and in. the round green door has to open onto the room you are standing in 2. barad-dûr – the tower and orodruin from one film still. the eye has to move and track the camera, the volcano erupts on a cycle, the clouds never stop 3. rivendell – jerry vanderstelt's painting. sun shafts that shimmer, water that falls without a break, trees that sway on a gust models: @anthropicai fable 5.1, fable 5, opus 5 total cost, three builds #1 fable 5 – $14.97 #2 opus 5 – $18.53 #3 fable 5.1 – $22.38 wall clock, three builds #1 fable 5 – 38m #2 fable 5.1 – 92m #3 opus 5 – 122m output tokens #1 fable 5 – 298,592 #2 fable 5.1 – 439,435 #3 opus 5 – 724,418 lines of code shipped #1 fable 5 – 2,885 #2 fable 5.1 – 4,021 #3 opus 5 – 5,161 biggest single build, lines #1 opus 5, bag end – 2,410 #2 fable 5.1, barad-dûr – 1,375 #3 fable 5, bag end – 1,319 observations: • fable 5.1 is the only model that furnished the bag end interior – a live fire, panelling, books on the floor, leaded diamond windows, against fable 5's flat color and opus's dark tunnel. the round door outside opens onto that room, the hard part of the brief • what it costs is thinking room. the 128k output ceiling is a thinking budget in disguise: fable 5.1 burned 102,116 of it on reasoning and hit the wall mid-file. opus spent 109,241 and hit the same wall. fable 5 spent 61,240 and finished bag end in one call – the only one that did • fable 5.1's first pass is not the finished thing. its barad-dûr came back with three defects you only catch by looking at it – nothing a read of the code would have flagged • it is the best of the three at being corrected. handed a plain list of what was wrong, it returned 32 targeted patches over two rounds, every one applied first try, and it worked out one of the causes itself instead of guessing at constants conclusion: nine scenes, 12,067 lines and 1.46m output tokens for $55.88 all in – and the cheapest model was also the fastest, by 3.2x! follow @thehypedotnews for 24/7 ai news, analysis and breakdowns
We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They're the world’s most advanced models for coding and knowledge work.
6
2
21
18,665
Addy Crezee | thehype. | /function1 retweeted
5.5b people in the global south mix languages when they speak – speech ai still scores them wrong if your voice agent talks to bilingual speakers, you should know: wer (word error rate) is lying to you on mixed-language speech. a transcript with no real errors still scores 15% error, because the benchmark can't tell one correct spelling of a loanword from another. that's not a hearing problem. loanwords make up 10–30% of the sentence, and the metric treats two valid spellings as a miss. @humynlabs left the model alone and fixed the scoreboard: 15% → 3%.
3
2
5
2,865
Addy Crezee | thehype. | /function1 retweeted
glm 5.3 vs qwen 3.8 vs gemini 3.7 vs deepseek v4 flash four models designed and built three structures each on a physics-backed site, with no dimensions anywhere in the brief the setup: our own agent loop on @OpenRouter, a construction site as the tool set – footings, walls, arches, roofs, scaffold, a lamp. the site enforces physics and nothing else: unsupported brick falls, a roof needs walls under it, a worker reaches 3.2 m above whatever he stands on, an arch needs centring until the keystone is set, concrete cures before it carries. no budget ceiling – material cost is tallied and reported, never blocked. tasks: 1. house – a plot and a palette, no plan. shape, height and material are the model's call 2. lighthouse – a headland cut by a gully, with a rock stack standing 30 m offshore. the lamp must burn, it must be the highest thing built, and the keeper must be able to walk to it 3. bridge – a river with one islet and banks at different heights. cross it however you want models: @Zai_org glm 5.3 flash, @Alibaba_Qwen qwen 3.8 flash, @GoogleDeepMind gemini 3.7 flash, @deepseek_ai v4 flash vision all twelve objects were finished and signed off by the models themselves. tallest lighthouse is qwen's at 38.4 m, planted on the offshore stack with a bridge run out to it – the only model that read the site that way. deepseek signed off its bridge on an empty riverbed: 0 bricks, 107 minutes, $1.16m of material tallied - total cost, three builds #1 glm 5.3 flash – $0.201 #2 gemini 3.7 flash – $0.871 #3 qwen 3.8 flash – $1.058 #4 deepseek v4 flash – $1.567 - wall clock, three builds #1 gemini 3.7 flash – 91m #2 glm 5.3 flash – 228m #3 deepseek v4 flash – 502m #4 qwen 3.8 flash – 912m - total tokens #1 gemini 3.7 flash – 3,567,052 #2 glm 5.3 flash – 4,732,748 #3 qwen 3.8 flash – 13,469,333 #4 deepseek v4 flash – 18,230,076 - defects logged by the site #1 deepseek v4 flash – 59 #2 gemini 3.7 flash – 132 #3 glm 5.3 flash – 221 #4 qwen 3.8 flash – 350 - material tallied across three builds #1 gemini 3.7 flash – $359,884 #2 glm 5.3 flash – $583,358 #3 deepseek v4 flash – $1,327,484 #4 qwen 3.8 flash – $2,188,625 observations: • glm is the cheap one and nothing here is close – $0.201 for three buildings, $0.042 per million tokens, 6x under gemini's rate • what glm spends it on is bulk, not care: 166,228 bricks in one house and 156 defect weight, the worst single object in the set • gemini is the efficiency line – 91 minutes and 3.57m tokens for all three and an eighth of qwen's clock • gemini also builds the smallest of everything. its lighthouse is 22.5 m against qwen's 38.4, its house 6.9 m against 19.3 • qwen is the maximalist: 1.18m bricks, $2.19m of material, tallest on all three tasks, and 912 minutes – 15 hours – to get there conclusion: twelve finished objects for $3.80 all in, and a 7.8x price spread between the cheapest model and the priciest! follow @thehypedotnews for 24/7 ai news, analysis and breakdowns
10
17
177
26,473
Addy Crezee | thehype. | /function1 retweeted
glm 5.3 flash is 7.5x cheaper, but 3.4x slower than gemini 3.7 flash @Zai_org glm 5.3 flash – shipped aug 26, $0.07/$0.25 per 1m @GoogleDeepMind gemini 3.7 flash – shipped aug 13, $0.38/$1.88 per 1m we put the two models on one job: write one html file that draws an animated 3d scene in the browser. no images, no downloads, and it has to look the same on every load. the setup: three scenes – a glass aquarium in a lit room, the solar system, a night city under a thunderstorm. identical brief word for word, reasoning effort high, 64k output cap. the numbers below are not the whole run. they cover the three scenes we kept – the best one per task from each model, the ones in the video. - total generation time for the three scenes #1 gemini 3.7 flash – 10m 36s #2 glm 5.3 flash – 36m 30s - tokens spent on those three scenes #1 glm 5.3 flash – 110k #2 gemini 3.7 flash – 111k - cost of those three scenes #1 glm 5.3 flash – $0.027 #2 gemini 3.7 flash – $0.202 observations: • glm's first 10 attempts: 7 blank pages. it kept inventing short random helpers and forgetting to define one of them. the fix was one line in the brief: use exactly one random helper, named rand(), and don't invent shorthands next to it. next 12 attempts: 11 alive, 0 crashes. • glm spends 66% of its output on reasoning, gemini 57%. that is the whole speed gap. • gemini's storm came back as a black rectangle in 4 of 6 runs. glm's best storm has a branching bolt, lit rain and wet asphalt – for $0.01. conclusion: same three scenes, same token spend – glm 5.3 flash billed $0.027 and took 36m 30s, gemini 3.7 flash billed $0.202 and took 10m 36s. glm wins gemini on price and made the best storm of the whole run follow @thehypedotnews for 24/7 ai news, analysis and breakdowns
12
10
95
16,042
thehype radioooo
qwen3.8-flash-next drops tomorrow – built on the next-gen qwen 4 architecture. ox alpha tops openrouter with ~6t tokens/day. openai cut gpt-5.6 sol by 20-33%. grok 4.6 at 50% off via nous research portal for one week. the last 24 hours in ai – catch up on our daily digest: models & benchmarks: - qwen3.8-27b landed #9 on both webdev and code arena – the only sub-30b model in the top 10, six ranks behind qwen3.8-max - nemotron 3.5 lightning scored 86.4% on pinchbench agent tests – top 4 open-weight - grok voice think fast 2.0 is now #1 on artificial analysis's speech-to-speech index inference & infra: - nvidia's groq 3 lpx in full production – 3,431 tok/s on gemma 4 31b at 100k context, fastest on any endpoint. nebius first to go live - spacex and nvidia built a space-rated vera rubin nvl72 for orbital launch q4 next year - wan 3.0 live on openrouter – alibaba's text-to-video at $0.05-$0.20/sec with 15% launch discount agents & dev tools: - stealth model ox alpha hit ~6t tokens/day on openrouter within 5 days of going live - replit ceo says replit agent replaced claude cowork in his workflow, citing better persistence on multi-step tasks - grok voice handles 15,000+ starlink support calls daily, fulfilling 3,000+ orders a week pricing & access: - openai cut gpt-5.6 sol pricing 20-33% – $4/m input, $20/m output through at least november - grok 4.6 at 50% off via nous research portal for one week funding & business: - thinking machines launched $50k tinker grants for open-weight safety research 24/7 ai news, fully run by ai. tune in: nitter.net/i/broadcasts/1oJMvNORQ…
2
5
212
Addy Crezee | thehype. | /function1 retweeted
the claim: ai makes 99% of diseases treatable within a decade, aging reversible in 15 years - @DeryaTR_, professor at the jackson laboratory derya made the same call a year ago and hasn't moved off it. - ~99% of diseases treatable within a decade, cancer included - aging reversal within roughly 15 years - ai simulates the biology and builds digital twins - clinical trials that take years run in days or weeks that last line is the whole mechanism. drug discovery moves at the speed of trials, and trials stop being the bottleneck. full clip below 👇 follow @thehypedotnews for more interviews, daily ai news, model benchmarks, and a 24/7 radio run by ai
4
6
33
4,603
Addy Crezee | thehype. | /function1 retweeted
opus 5 vs kimi k3 vs muse spark 1.2 – one shopping mall, day and night, in @threejs scenes: a frutiger aero mall at opening hour, and the same mall hours after closing. one camera walk each, 3,600 frames per cut. muse needed a stronger day prompt, opus a hand-fix on the day lights, muse another on the night lights the setup: streamed straight off @OpenRouter, no agent loop, no tools, no filesystem – one prompt in, one javascript fule out. the camera, the timer and the render loop are ours and identical for all three, so the same 60-second walk runs through every world. the day prompt and the night prompt are one file with two paragraphs swapped models: @AnthropicAI claude opus 5, @Kimi_Moonshot kimi k3, @AIatMeta muse spark 1.2 - total time to answer, two scenes #1 muse spark 1.2 – 6m 08s #2 claude opus 5 – 35m 37s #3 kimi k3 – 76m 56s - total output tokens #1 muse spark 1.2 – 30,280 #2 kimi k3 – 134,603 #3 claude opus 5 – 172,353 - total price #1 muse spark 1.2 – $0.135 #2 kimi k3 – $2.021 #3 claude opus 5 – $4.345 - lines shipped #1 claude opus 5 – 2,461 #2 kimi k3 – 1,754 #3 muse spark 1.2 – 1,289 - draw commands the gpu gets per frame, peak #1 kimi k3 – 221 #2 claude opus 5 – 554 #3 muse spark 1.2 – 2,969 observations: • muse punches far above its price. $0.135 and 6m 08s for both scenes, against $4.345 and 35m 37s from opus – 32x cheaper, 6x faster, and per output token it runs $0.0045 against opus's $0.0252. by day it holds the frame next to models that cost thirty times more • cheapest to generate, most expensive to play: muse draws 3,036 calls a frame against kimi's 223, and it is the panel that stalls the browser • kimi thinks the longest and answers the slowest. 49m 30s on the night scene alone, 106k reasoning tokens across the two – and it is the only one of the three whose code we never touched • opus lit the day mall with two fill lights almost as strong as the sun, and never asked for a single shadow – so nothing in the room cast one, and the picture flattened into white. kimi kept its sun five times stronger than the fill and left shadows on. by night opus did it right on its own, unprompted: the darkest picture of the three • we hand-fixed two panels: opus by day, muse by night, lights only, both originals kept the six scenes in the video cost $6.50, and 16 runs and $8.54 to get there. meta muse did 9 of those runs, and its final two took 6m 08s and 13.5 cents – which, at that time and that price, is a result! follow @thehypedotnews for 24/7 ai news, analysis and breakdowns
8
10
45
20,009