Dani Carmona retweeted
so great to spend the morning with Emilio from @SupersonikAI today ✨ probably THE coolest office in barcelona rn
20
6
214
7,234
you will get far! very cool to learn more about what you are doing
bumped into two guys at bos airport, they must be pretty smart and working on some cool things from the way they spoke
3
15
923
im at sfo airport to go to boston, flight is leaving in less than 2h, i see the flight in @perk_global app but apparently we dont exist in this flight. airport staff from @united cant do anything and im on the phone with perk but they also cant fix it? wtf is this?
4
7
1,187
we made it after a reschedule and an extra 2h delay. see you soon boston
60
20 min on the line with perk, not solved yet, potentially double booking where we got confirmation but they didnt? this is wild
2
127
all ubers in sf are recording you now. at this point maybe they should start selling data for training or something like that
2
194
my european mind cannot comprehend
1
11
336
best ramen in sf?
2
318
at some point we will saturate security through software and we will start to see problems in the hardware. those are much harder to fix
Greg Brockman says OpenAI pointed Astra at its own systems until it ran out of vulnerabilities to find: "We took 25% of our production engineers and said, 'Sorry, all your projects are on hold. You are now defending. You are now up-leveling our security architecture. You're going to use the models to find all the holes.' And we found a number of serious issues, and we fixed them." "We found some new problems, but eventually it saturated. We basically have found, to our knowledge, all of the P0s, all of the critical problems that Astra is smart enough to find. And of course, there will be a new model, there will be a new round." "You want to be in this tight loop of new cyber capability drops, you deploy it against your systems, you find the new holes, and ideally, you've managed to automate this, what we call defense factory. That's what we're building internally." "There are ideas, for example, formally verifying all of software, that are possible with AI." @gdb @bhorowitz
2
351
most people are past reviewing code and just review plans / tickets, but i was just wondering if maybe the best thing to review could be the conversations with the model that produced the plan? feel like its easier to read than specs and theres more insight?
2
7
433
I could listen to some Suno in this Rosalia
Welcome to the v6 era. Our best models yet, built in partnership with artists and producers. Here’s how the family works 🧵
1
1
500
Random thought: If we are seeing an explosion in AI research and models are going to become exponentially better, that probably means we will hit a wall much sooner than expected in hardware doing inference. Does that mean, assuming also that there will always be people solving problems that are worth more and willing to pay more for each token, that access to regular intelligence will not be that affordable or that well or evenly distributed (because all of the efforts and compute will be targeted and used by people with bigger, more expensive problems to solve)? I don't know, thinking about it.
3
4
442
this situation between openai and anthropic with the navier stokes proof starts to remind me of the deel and rippling drama
177
been there done that
if you are a programmer and you'd like to start a company it is helpful to internalize two big weaknesses you have 1. your instincts for anything outside programming are terrible. everything you believe is likely 100% backwards 2. you've always been treated like the smart person in the room so it will take you a very long time to understand #1 for me personally it took ~ 10 years before i started to be able to unwind this
2
332
my highest predictor of a high hrv is socializing with friends or family, even if i drink alcohol this was not on my list of healthy habits
6
344
apparently i’m a genius
one of the best signs of intelligence is definitive range. e.g. someone who can have a deeply serious intellectual convo & then immediately make a fart joke or say something completely regarded. you rarely ever encounter ppl who can effortlessly move between profound shit & the absurd. this is why the south park guys are geniuses.
2
1
8
1,182
with these jumps, maybe i start to agree that we are in the singularity and models are self improving rapidly
🚨 Gemini 4 Benchmark Leaks : Google might have a monster on its hands 🤯 A leaked benchmark sheet is making the rounds and the numbers are seriously aggressive. > Gemini 4 reportedly scores 72.8% on Terminal-Bench Science 0.1 > 48.7% on AutomationBench, >targeting complex Business workflows >99.1% on FrontierMath Tier 4 (v2) >62.3% on Terminal-Bench 4.0 >70.2% on HealthBench Professional >98.6% on BenchCAD And a ridiculous >99.9%+ on ARC-AGI-3 >The leaked sheet puts Gemini 4 ahead of GPT-6 Astra and Fable 5.1 on every listed benchmark If these numbers are even close to accurate... Gemini 4 isn't just another incremental upgrade. The biggest jumps appear to be in scientific reasoning, mathematics, terminal use and agentic workflows exactly the areas where frontier models are increasingly competing. And there's one huge caveat: The Gemini 4 column is explicitly marked “PREDICTED.” Google has confirmed Gemini 4 training is underway, calling it its most ambitious pre-training run yet, but these benchmark numbers have not been officially published by Google. So for now, treat the table as a leak/prediction, not verified benchmark results. But if Google actually ships something anywhere near 99%+ ARC-AGI-3 + 99% FrontierMath + 70%+ scientific workflows... the Gemini 4 launch could completely reset the frontier-model leaderboard 👀🔥
Community note
The benchmark image shows fabricated or AI-generated results for Gemini 4 with no match to any verified leaks, official data or current leaderboards; the model remains in early pre-training per Google with no such scores published. 9to5google.com/2026/07/26/goo… terminal-bench-science.ai/announcement snorkel.ai/leaderboard/te… benchlm.ai/benchmarks/arc…
3
365
robot duck fighting hackathon when?
Imagine two teams training their own policies and they meet in the arena 🤯 @RemiFabreRobot
206
i'm starting to see the matrix now everything makes sense its easier to play when you fully understand the rules of the game
2
15
1,002
attention is all you have small team = low attention seeking, you own your attention big team = high attention seeking, they own your attention (unless you are very careful) in a world where the best are multiplied and everything compounds, attention is your most valuable asset
6
214