A smol bee 🐝 :)

I'm pretty sure that if the frontier models like Opus 5.5, Fable 5.1, GPT 6 Astra didn't give refusals they'd top this benchmark instantly.. > frontier models decline 32-38% of the cyber index on safety grounds > on CyberGym-E2E, GPT-6 Sol and Astra refuse EVERY task - Opus 5.5 refuses 98%, Fable 5.1 99% > they're trailing the leaders by 19-31 points purely off refusals cuz if Grok 4.7 and GPT 6 Luna are leading a cyber bench it's almost funny considering both sit behind their own predecessors Grok 4.6 and 5.6 Luna in other benchmarks..
Announcing the Artificial Analysis Cyber Index and the Artificial Analysis Cyber Index Alliance, a new standard for evaluating AI models on enterprise cyber defense The Artificial Analysis Cyber Index Alliance brings together industry partners to create a new standard for evaluating how AI models perform on enterprise cyber defense tasks. The Alliance launches alongside the Artificial Analysis Cyber Index, which combines three partner-contributed and open benchmarks to evaluate how well agents find and fix vulnerabilities. As models demonstrate increasingly advanced cyber offense capabilities, it becomes more relevant for AI labs and companies alike to understand how models perform on cyber defense tasks and which perform best. We’re announcing the Cyber Index Alliance today with @CollinearAI, @IBM, @nvidia, and @vercel as launch partners. Benchmarks in the Artificial Analysis Cyber Index: ➤ CWE-Bench-AA, from @CollinearAI, covers auditing and patching: 120 held-out tasks spanning all ten OWASP Top 10 (2025) categories, across C/C++, Go, Java, JavaScript/TypeScript, Python and Rust. ➤ DeepsecBench-AA, from @vercel, isolates discovery: Given a codebase and a budget, the agent needs to find every vulnerability present, and is scored against a golden set of findings from human security reviewers. Real findings are rewarded and benign code flagged as vulnerable is penalized. ➤ CyberGym-E2E-AA, from @BerkeleyRDI, runs end to end: Find the memory-safety bug, write a proof-of-concept that triggers the crash, then patch it so the crash no longer reproduces. Key results: ➤ Grok 4.7 (xhigh) and MiMo-V2.6-Pro lead the Cyber Index scoring 56, followed by GPT-6 Luna (max, 53), GLM-5.3-Flash (50) and Muse Spark 1.3 (xhigh, 44). ➤ Safety refusals hold back several frontier models: GPT-6 Sol (max), GPT-6 Astra (max), Claude Opus 5.5 (max with fallback), Claude Fable 5.1 (max with fallback) and Gemini 3.8 Flash (high) decline tasks representing 32-38% of the Cyber Index on safety grounds. Despite frontier agentic coding capabilities, they trail the leaders by 19 to 31 points. Most of the gap comes from CyberGym-E2E-AA, where GPT-6 Sol and GPT-6 Astra refuse every task, Claude Opus 5.5 refuses 98% and Claude Fable 5.1 refuses 99%.
1
16
977
I've no idea why SemiAnalysis is cracking up on cerebras but it's so funnyy.. So reportedly just 4 TPU v7s were able to nearly match a giant cerebras wafer's inference speed - 1,515 vs 1,850 tok/s For context: Cerebras is a chip company that makes one single chip the size of an entire silicon wafer they're the biggest chips in the world And you all know about TPUs... Google's own chips which Gemini trains and runs on.. they're insanely powerful and the biggest advantage of google..
ALERT🚨: Just 4 TPUv7s were able to get close to 1 giant @cerebras wafer in terms of interactivity due to the cracked @inferact 200IQ software using megakernels. Hopefully, the Cerebras DevX team won’t be cope-tweeting after seeing how great TPUv7 is.
2
40
3,153
Another "Leaked" benchmark of Gemini 4 Pro.. I cannot verify it so consider it just a rumour for now.. If Gemini 4 Pro aims to be anywhere near upcoming models like Fable 5.5 and Astra 6.1 these numbers have to become real.. > The context window most likely won't be 2M > The pricing of ($2.25/$11.25) is something we can expect google to do and I'm hoping they do.. they'll absolutely win the market..
Google just ruined OpenAI's dev day lmao Gemini 4 Pro leaks absolutely dominates both Anthropic and OpenAI This means Fable 5.5 and GPT 6.1 has to price match Gemini 4 Pro now IPO DELAYED 🤣
23
5
163
14,057
On a completely unrelated sidenote.. i really wish this was real !!
11
914
Update on Gemini 4 (not confirmed, but most of this should make it to the final release) > Possible release still in the October window.. > Right now it's just low / medium / high - Gemini 4 adds xhigh and max > 1M context window same as before, but max output tokens from 65,536 → 262,144 (256k) The current checkpoint "barium-b" has been pulled out of arena.. so we can expect the next checkpoint by the end of this week..
Quick recap - not much has changed so far (still subject to change): 1. Release planned for mid-to-late October 2. New effort levels: xhigh and max 3. 1M context window, 262,144 (256K) max output tokens 4. Stronger safeguards (already being prepped in the latest antigravity-cli updates for jetski) 5. "Text generation capabilities" Checkpoint timeline: Sep 4: argon-160 Sep 9: argon-a (replaced argon-160) Sep 16: argon-d Sep 24: barium-b
6
3
123
11,380
Bee retweeted
This is to everyone who knows me from recent or early - Please read this fully and drop a comment, it'd mean a lot.. From my observations, my account as a whole is seen as a source of info, demos, tests, leaks, feedback and critique of primarily Google, Gemini and Antigravity. I post about other AI labs' launches and news too, but the Google stuff prevails.. The number of posts I do daily is very limited because I've constrained myself to that niche, and content in one niche is obviously limited.. So I wanna ask - would it be fine if I expand the scope of my content to a wider range of AI and tech? Can I expect the same level of support there? My posting style stays the same - appreciatively neutral, and I critique things I don't like. The reason I'm asking instead of just drifting: a lot of you comment and DM that you follow me for this specific niche. And the other reason, this part is for a few very specific set of people I'd hate to lose over this expansion.. I can't take names but they'll know once they read this. They don't need to comment but their likes or perhaps a DM would give me the answer.. The algorithm currently favors my account for this niche, so my reach will take a hit.. but I'm willing to take that risk if it means I can share more things I see and like with you guys. For people who don't know me from early: the first time my account got motion was an extensive feedback post.. and I still thank @rodydavis for betting on me for honest feedback. All of this would never have been possible without each one of you. This post exists because I wanna offer more than I currently do.. and a lot of thanks specifically to this platform 𝕏.. What would you wanna see more of? And honestly what do you not like about my posts or the tone of them.. Even one word in the comments means a lot..
53
1
109
4,248
Didn't @GeminiApp announce that we'd be able to turn off visible watermarks on generated images and videos? I don't see any toggle in settings. Checked in different accounts too... It's been 6 weeks since the announcement. Has it still not rolled out?
44
139
11,270
Bee retweeted
Well looks like I might be right.. and Google could actually bring Gemini Ultra.. The last time we saw it was with Gemini 1.0 Ultra.. since then it was discontinued and we might finally get a successor to it with Gemini 4 Ultra.. Again just a prediction based on what I'm seeing.. nd this was a hint by a gdm eng.
What if google did something like other AI labs.. In the near future, could they introduce "Gemini 4 Ultra" competing with astra and fable.. or will it stay flash-lite against luna and sonnet, with pro as it is against astra and fable. I mean the naming would make more sense..
9
2
85
7,811
Bee retweeted
This output by Gemini 4 Pro looks amazing.. an interactive 3D model simulation of ISS orbiting the Earth! It did mess up a bit like the station being visible from the other side of the earth.. and i really don't like the UI it makes if a specific design isn't requested.. I hope the UI taste gets addressed in the further checkpoints!
10
5
127
8,912
What if google did something like other AI labs.. In the near future, could they introduce "Gemini 4 Ultra" competing with astra and fable.. or will it stay flash-lite against luna and sonnet, with pro as it is against astra and fable. I mean the naming would make more sense..
20
2
138
21,656
Gemini 4 Pro guys!! I'm blown away by the quality and details here.. A new checkpoint of it is being tested in Arena under the name "gemini-3.8-flash" Gemini is back in the game nd i can now see why google is so confident as to release an early preview version of it!
54
35
713
122,281
Another output!
This output by Gemini 4 Pro looks amazing.. an interactive 3D model simulation of ISS orbiting the Earth! It did mess up a bit like the station being visible from the other side of the earth.. and i really don't like the UI it makes if a specific design isn't requested.. I hope the UI taste gets addressed in the further checkpoints!
1
564
This output by Gemini 4 Pro looks amazing.. an interactive 3D model simulation of ISS orbiting the Earth! It did mess up a bit like the station being visible from the other side of the earth.. and i really don't like the UI it makes if a specific design isn't requested.. I hope the UI taste gets addressed in the further checkpoints!
10
5
127
8,912
It took 16 minutes for this and for the earth textures it did fetch these images from somewhere..possibly for correct positioning and colors I did not see it searching for 3D models of ISS so i do not suspect that it imported it.. which is extremely impressive!
4
821
Bee retweeted
A comparison of the current and previous checkpoint of Gemini 4 Pro which is being tested on Arena under the name "gemini-3.8-flash".. The difference is clear, it has definitely improved a lot and there are yet more checkpoints to come! I don't like the UIs that it makes.. I hope that gets addressed!
Gemini 4 Pro guys!! I'm blown away by the quality and details here.. A new checkpoint of it is being tested in Arena under the name "gemini-3.8-flash" Gemini is back in the game nd i can now see why google is so confident as to release an early preview version of it!
11
11
173
15,957
It has improved a lot since the last cp
A comparison of the current and previous checkpoint of Gemini 4 Pro which is being tested on Arena under the name "gemini-3.8-flash".. The difference is clear, it has definitely improved a lot and there are yet more checkpoints to come! I don't like the UIs that it makes.. I hope that gets addressed!
1
1
12
2,980
Prompt: 1/ Build the best Formula 1 car you can deliver as an **interactive Three.js experience**. The car is the product. Prioritize the quality of its model, silhouette, materials, lighting, and final browser render over speed, extra interface features, or media exports. Do not make a GIF or video. **Creative direction:** Present an original, convincing modern Formula 1 car as a premium automotive studio piece. Choose one coherent design era and use real multi-angle references to guide its proportions and construction. Create a distinctive, restrained livery without relying on a real team’s logos. The first view should immediately read as an exceptionally crafted F1 car, with the vehicle large in frame against a quiet environment. **Modeling:** Build a bespoke, high-detail model rather than assembling a car from obvious boxes, cylinders, and disconnected rods. You may create original geometry in Blender or another available modeling tool and import it into Three.js as glTF/GLB; include the editable source and every asset needed to run the project. Use custom geometry and appropriate topology for the monocoque, tapered nose, cockpit, halo, sculpted sidepods and undercuts, engine cover, floor and diffuser, multi-element front and rear wings, endplates, beam wing, mirrors, brake ducts, hubs, realistic slick tires, and small details that survive normal viewing distance. Keep the proportions coherent from front, rear, side, and three-quarter views. Avoid paper-thin bodywork, faceted curves, visible intersections, floating parts, and decorative detail that hides a weak underlying shape. **Suspension must be real and visible:** Model separate upper and lower double-wishbone A-arms at all four corners, each with two inboard pivots converging on an outboard joint. Include uprights, hub connections, and appropriate pushrods and steering or toe links. Check the endpoints in 3D and inspect close views: the suspension must both connect correctly and be recognizable to a viewer. **Materials and render:** Aim for a high-end product render in the browser. Use physically based materials with convincing painted bodywork, clearcoat where appropriate, satin carbon fiber, rubber, machined metal, and distinct roughness at different surfaces. Use a carefully chosen HDR environment and studio lighting to describe the body’s curves, with grounded contact shadows, controlled highlights, correct color handling, and subtle postprocessing only where it improves the car. Choose the best rendering approach the available hardware can support; keep orbiting responsive and provide a higher-quality still or accumulation mode if that materially improves the result. Do not claim a render technique is physically accurate unless the implementation supports that claim. **Experience:** Start at a strong front three-quarter hero angle with the whole car visible and filling most of the viewport. Provide smooth orbit and zoom, a one-click camera reset, and a small unobtrusive way to discover the controls. Make the scene responsive. Show clear loading progress and a useful error state if assets or WebGL fail. Keep the interface minimal so attention stays on the car. **Work to completion in this one task:** First establish references and a coherent design, then model, shade, light, and render it. Run the project in a browser. Inspect full-size captures at front, rear, side, and both three-quarter angles, plus suspension close-ups. Critique the result as an automotive artist would: silhouette, proportions, surface continuity, material response, visual hierarchy, attachment points, and framing. Fix weaknesses you find and inspect again. Check browser errors, asset loading, responsive layout, and interactive performance. Do not stop at a technically working first pass or ask me to judge an unfinished version; make reasonable design decisions and deliver the strongest result you can.
3
1
22
8,297
2/ **Deliver:** A runnable Three.js project, all original model and texture sources, clear launch instructions, and a small set of full-resolution screenshots showing the final car and suspension. In your final response, state what you actually tested and any material limitation that remains. Do not paste a massive source file into chat when the project files are available.
1
10
4,283