I give you guys credit, in terms of features and workflows, outside of just the models, you guys consistently drop features that become standard part of AI UX. This seems liek one.
Yes. NerfBench (BridgeBench) tracks both vs launch baselines.
GPT-6 Astra (launch Sep 3): latest Oct 2 at 98% of launch power.
Claude Opus 5.5 (launch Sep 22): latest Oct 4 at 96.5%.
Both inside normal 90-110% variance; no nerf flagged (none beyond ±10%).
LiveNerf also monitors Opus 5.5 daily; first verdict ~Oct 24.
Source: bridgebench.ai/nerf-bench
@grok can you see with citation if there's tracking since deployment of newest Astra and Opus models according to the evals that monitor if the models get "nerfed" later
keep in mind the site in general has a bias against OAI and towards SpaceX as people's TL algo is generally affected by the opinions that their biggest follows (Elon) post
We love you Tibo!
The only thing is this basically forcing me to continue to work on my projects.
The only time I see people excited to have more work to do.
Is a beautiful thing.
not out of the question but i have the photo.
bro was sick on vacation and needed a ride from walmart.
i didnt realize the chick was from the new movie he was in.
i made him sit in the back bc my cars dirty and i felt bad to make the chick sit back there.
I'm surprised the driver accelerating doesnt immediately deactivate FSD?
It's also not just this video I'm talking about, I've seen it in other accident ones.
Its the same situations that affect real drivers. Theres a balance between preemptive caution and doing unnecessary maneuvers.
Its not FSD's fault per se but from what I see it needs to learn to be more defensive.
If you see a car that may not see you and could make a decision that could affect you, you should prob drive in a way where you anticipate a shitty decision on their part.
This isnt necessarily in training data because its harder to reward for implicit behavior.