Startup Founder. Interests: AI - Tech - Robotics - Space. Tech Support @ Hogwarts

Austin, TX
Asked Muse to book a dinner reservation. It couldn't get pasts Toasts Human 1 click. Asked Grok bot, which completed the task in 2 mins. Grok Bot - 1 Muse - 0
1
23
Blizzard kicks me out of game, to log back in to this seconds later... and the time is only going up
133
True story
it all makes so much sense
98
👀 Do we believe @GoogleDeepMind ? Gemini 4
Introducing Gemini 4 Argon – our new frontier model. It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense – rolling out today to a set of trusted testers through our Fairwind Program.
80
OpenAI and @thsottiaux, care to explain? I believe yesterday we bragged about 99% uptime?
61
Pure idiocy
JUST IN: Illinois moves forward with the nation’s first crypto tax, set to take effect January 1, 2027, even on trades that lose money.
104
You know what... if the agents are actually very good at coding and executing tasks and come as part of the sub with real usage limits... I'm here for it. Can't be worse than Grok @bot, which uses everyone else's model to do meaningful things.
147
I think BC and WTLK should be part of WoW Forever; just don't raise the cap. Also, remove any "retail" content that doesn't fit WoW Forever's Classic take. Think @Blizzard_Ent will incorporate these expansions?
106
I never use @cursor_ai but ULTRA comes with my grok sub. Cursors bugbot is connected to my github and runs on prs. Last week was pure documentation pr alignment before I started to code. Maybe 10-15 prs (all GPT-6 Astra and Fable) Combine those 10 - 15 prs and me having grok bot for the first time build something on the side that had nothing to do with the 10-15 prs. I get this email...after grok worked for 15 mins. Mind you this is Cursor ULTRA. 1. Usage is dogshit 2. 4.7 code is dogshit with the Cursor harness. Maybe better with Grok Build. 3. Cursor bugbot runs off of "other models" which is where most of the "other models" usage went. Aka grok bot doesn't even use their own model to pr bug. 4. When Grok Bot builds using Cursor they use cloud agents. Want to guess what models those cloud agents use??? "OTHER MODELS" which usage is dogshit. TL;DR turn off Bugbot and use CodeRabbit instead. These "1000" pr engineers aren't using the same usage quota and from the looks of it not even the grok models fully. Maybe im wrong but grok will be canceled before they swap me to $300 a money on Grokheavy, unless something changes.
3
3
294
Is it me, my algorithm? Am I seeing a lot of System One model drops after @typesafeai Jev?
122
I have in my PRD areas where I want the agent to do deep research before implementation of that slice. It uses Exa and Firecrawls Alexandria to get the most current data available. Ill have to make sure Agent Uktra is now a part of this workflow .
Introducing Agent Ultra - a step change in deep research Agent Ultra orchestrates swarms of agents to perform exhaustive research, build comprehensive lists, and answer questions requiring thousands of sources. In both evals and vibes, it's state of the art
106
Imo @thsottiaux silence on this matter is confirmation. OpenAI is planning for a hike in price. The next question is: will it be for 20x or a higher usage tier?
OpenAI is quietly cutting usage limits so they can launch a $500 plan. This is a pricing scam built on opaque subscription limits. Two months ago, burning through a $200 plan was hard. Last year, the $20 tier was more than enough for almost everyone. Now? A $200 tier fails after one day of real work. The $20 tier is just a free trial. We need real consumer protection for AI subscriptions.
90
Will you continue to pay for OpenAI max plan of usage and limitations stay the same but pricing goes up?
OPENAI 🔥: The upcoming ChatGPT Pro Max plan will cost $500. So far, this will be one of the most expensive AI subscriptions available on the market. I hope "it will be worth the wait" 👀
1
131
💥 The biggest trap in AI engineering? Building generic benchmarks instead of product-specific evals. Most teams waste weeks writing evaluators that miss actual edge cases. The team at ai-evals-course just open-sourced evals-skills, a framework to guide AI coding agents through building bulletproof evals. Here is a breakdown of the core skills: 🛠️ eval-audit & error-discovery Before writing a single eval line, you need to know where your system breaks. eval-audit inspects your active pipeline and flags flaws by severity. error-discovery eats your raw trace logs, spins up a zero-dependency HTML review app, and clusters data to isolate failure modes. 🤖 write-judge-prompt & validate-evaluator LLM-as-a-judge setups are notoriously fickle. write-judge-prompt designs custom evaluators tailored for subjective quality. validate-evaluator calibrates those judges against human labels using data splits, tracking TPR/TNR, and applying strict bias corrections. 🧪 generate-synthetic-data & evaluate-rag generate-synthetic-data goes beyond basic LLM prompting by using dimension-based tuple generation to build robust, diverse test inputs. evaluate-rag isolates and benchmarks retrieval precision vs. generation quality independently. 🚀 Build Your Own These pre-built skills protect against the most common footguns. But the real power comes from using them as a baseline to write domain-specific skills tailored directly to your engineering stack. Get started by injecting them directly into your agent workflows using npx skills. 🔗 Full repository: github.com/ai-evals-course/e…
4
1
196
Web agents just got an insane speed upgrade. ⚡️ Instead of burning time on slow screenshots or bloated site scripts, TypeSafe's Jev uses a brilliant architectural shortcut: one network round trip per decision cycle, predicting the operation and the target simultaneously. The results speak for themselves: -Zurich to London on Google Flights in just 7.1 seconds (at 1x speed, text generation and loading waits included). -25% reduction in median task time. -10x fewer browser protocol calls (slashed from 1,092 down to 101). It is taking off incredibly fast, already sitting at over 15.1k GitHub stars! 🚀 If you want to build the fastest, cheapest web agents possible, this architecture is a massive leap forward. Check out the repository below 👇github.com/browser-use/jev-u…
78
Do you think OpenAi hasn't addressed GPT-6 Astra because Sol 6.2 is dropping soon?
Rumors I’ve been hearing, not here on X. First, let’s start with OpenAI and I’ll go towards Anthropic. GPT-6 Sol is coming Tuesday. It’s both cheaper and more intelligent than 6 Astra, think of it like a 6.2 jump. The internal model “significantly more capable than Astra,” named Bel internally, helped with this release. Bel is considered “AGI” within OpenAI. They are very impressed with this model. OpenAI is growing very confident that their internal lead is so big that no other lab can catch up. Unbelievably confident. Anthropic is currently not in, let’s say, a “code red,” but is aware of OpenAI’s lead and doing everything in their power to catch up. Their new model Opus 5.5 is coming probably Monday rather than Tuesday due to OpenAI releasing on Tuesday.
198
Lies
No Dad can do all three. It's impossible.
105