Staff Software Engineer @ Abnormal AI | Trying to Keep Up with My AI Coworkers | Agency is All You Need |

Minnesota
Evan Oman retweeted
FINAL @BeckerFootball 57 Tech 0 What a showing from this Becker team for Coach Dwight Lundeen’s 57th and final homecoming game. Story to follow
1
2
112
This is how my dot messages codex threads
2
109
Really enjoying Dots so far … nothing new capability-wise, but it’s a polished package and integrates well with the rest of my ChatGPT account Particularly like the built-in browser that helped me pay for my vehicle registration renewal!
2
96
Fantastic new feature from @snipd_app — incredible user empathy
2
92
Evan Oman retweeted
Open AI Dev Day reflections: -Everyone I met was awesome! -I like the bravery of doing live demos and hope that continues. -The unplanned audience Q/A at the end was some of the most interesting content of the day. In the future: -It seems like some sessions could be more interactive. People could submit questions and moderators/AI could screen them/figure out which questions are the most common/interesting etc. -A session or two covering the releases in more depth with interesting use cases/examples could be helpful.
1
1
35
Evan Oman retweeted
Testing hot take: - unit tests have almost never been valuable - e2e testing is valuable/essential Before AI: If your unit was complex enough that it demanded dedicated testing, it likely needed to be factored out anyways. Post AI: AI writes robust code and poor tests. The code changes all the time anyways, so you're just burning compute on something that is basically decorative. With a simple skill, AI can write great e2e tests that assert the negative (most people miss this!) and clearly test the behaviors you care about rather than the implementation specifics. Well written e2e tests provide context to your agent on the intended behaviors of your program so you don't get into behavioral whack-a-mole. AI is also great at writing them quickly and getting them to run quickly in parallel, so the last excuses to not do this are gone! The ultimate goal has always been behavioral verification, and now we can skip straight there.
1
1
6
1,150
Evan Oman retweeted
Today, @Abnormal is launching a new AI Security product suite to help customers govern AI, secure AI, and defend against rogue AI. This new suite is built on the same behavioral AI technology that 30% of the Fortune 500 use for email and identity security. Abnormal designs, trains, and deploys our own AI, and we've spent nearly a decade mastering how to understand the normal behavior of business, detect behavioral anomalies, and respond at machine speed.  We're extending our behavioral modeling to include non-human identities, applications, and cloud identities. This enables the same behavioral AI to now power 5 new products in our AI Security product family: 1) AI Governance discovers your AI, uncovers shadow AI, shows who is using what, and how much it costs 2) AI Cloud Security detects + contains AI cloud breaches with behavioral AI and honeypots 3) AI Employee Guardrails enables employees to use AI without leaking data or triggering unintended actions 4) AI Agent Security inventories your agents and ensures they behave as intended 5) AI Security Workbench lets you use AI to run investigations with Abnormal models and without the guardrails AI Governance is already GA and can be installed in one click from our app store. AI Cloud Security is available today for priority customers, and the other products are in private preview for design partners. Learn more on our website and talk to your customer rep to cut the waitlist! abnormal.ai/blog/ai-security…
5
6
17
1,143
the bangers will not stop...
Claude Opus 5.5 has the best visual design of any model I have tested so far
1
103
Evan Oman retweeted
Thank you @claudeai! GIF animated in Python and rendered in Blender by Claude Opus 5.5
A thread of early explorations with Claude Opus 5.5: A short story about a watermelon created by @kevin_t_ngo.
84
139
3,565
255,710
GPT-6 Sol and Luna push the cost efficiency frontier by halving cost relative to GPT-5.6 Sol and Luna. Intelligence Index and Coding Agent Index scores remain level with GPT-5.6, with progress in some evaluations and regressions in others Pricing is approximately half that of GPT-5.6: Sol drops from $4/$20 to $2/$10 per million input/output tokens, and Luna from $0.20/$1.20 to $0.10/$0.50, with the same 90% discount for cache reads and 25% premium for cache writes. Key takeaways: ➤ Halves Cost per Task: GPT-6 Sol (max) costs $1.06 per task to run the Artificial Analysis Intelligence Index, ~50% less than GPT-5.6 Sol (max) at $1.99. GPT-6 Luna (max) costs $0.07 per task, ~60% less than GPT-5.6 Luna (max) at $0.18. This is driven by the price cut, as both models use slightly more output tokens per task (31k vs 29k for Sol, and 51k vs 41k for Luna). These two releases allow OpenAI to capture a significant portion of the cost efficiency Pareto frontier. ➤ In the Coding Agent Index, Sol improves but Luna regresses: In OpenAI's Codex harness, GPT-6 Sol (max) scores 57 in the Artificial Analysis Coding Agent Index, up 2 points from GPT-5.6 Sol (max), with gains in Terminal-Bench 4.0 (43% vs 37%) and SWE-Atlas-QnA (58% vs 54%). At $2.99 per task it costs ~50% less than GPT-5.6 Sol (max) and sits on the Pareto frontier of Coding Agent Index vs Cost per Task. GPT-6 Luna (max) scores 41, down 2 points from GPT-5.6 Luna (max), with lower scores in SWE-Atlas-QnA (44% vs 49%) and DeepSWE v1.1 (64% vs 66%), at ~60% lower cost per task. ➤ Significant reduction in hallucination: Both models hallucinate less in AA-Omniscience, our knowledge and hallucination benchmark. GPT-6 Sol (max) cuts its hallucination rate from 92% to 60% and GPT-6 Luna (max) from 93% to 77%. Sol achieves this by declining to answer more often: it attempts 83% of questions vs 99% for GPT-5.6 Sol (max), which cuts wrong answers by about a quarter but also lowers accuracy 5 points from 59% to 54%. Luna's accuracy is broadly unchanged at 44% vs 43% while it answers fewer questions. On the AA-Omniscience Index, Sol improves from 22 to 27 and Luna from -10 to 1. ➤ Mix of improvement and regression across evals: Beyond AA-Omniscience, both models improve in AutomationBench-AA (Sol 62% vs 60%, Luna 53% vs 50%) and Terminal-Bench 4.0 (Sol 44% vs 40%, Luna 13% vs 12%). However, we observe regressions in two key knowledge work evaluations. In GDPval-AA v2.1, our benchmark adapted from OpenAI's dataset of economically valuable tasks across 44 occupations, Sol drops ~100 Elo points and Luna ~75. Luna also drops ~45 Elo points in AA-Briefcase v1.1, while Sol is level. AA-Briefcase v1.1 is a private evaluation across multi-week knowledge work projects, with thousands of input files. Our team has manually inspected hundreds of model outputs: the regressions tend to be driven by reduced presentation quality and deliverables that omit rubric elements. Congratulations @OpenAI and @sama on the launch!
138
192
1,994
723,091
putting early GPT-6 Sol access to good use
4
378
Evan Oman retweeted
Claude Opus 5.5 drew every frame of this animation in JavaScript. Everyone in town sends Claude their requests, but one girl sends a question instead: "What do you love?"
228
383
6,097
649,124
Evan Oman retweeted
Just finished Lonesome Dove, and it is one of the best things I’ve ever read. I bought it after I read an article, where some journalist asked Stephen King what his fav novel of all time was. He picked this book. It’s 864 pages. Read the physical book, and read each sentence as slowly as you can. Don’t rush it. This world is so special. It was a wonderful few weeks immersing myself into the frontier West. Reminded me that reading fiction is so lovely, and so important.
187
88
1,909
111,975
Evan Oman retweeted
poor guy claim to have built Jev a year ago but no one cared, and now Jev stole all the thunder many people are saying “you gotta tell your story” or “marketing is important”, and they just completely missed what actually made the difference here i just looked into this laya model laya.convaiinnovations.com/ and: - it only supports 512-1k context… a lot of use cases won’t fit at all - evaluating the model directly shows its accuracy is as good as a coin flip. in order to get good results, you need to first fine tune it i’m sorry, but that’s not Jev there’s a massive gap between an interesting research and a useful product you can “tell your story” all you like, but you can’t blame Jev for stealing your thunder when Jev did all the work to make a well-packaged solution anyone can just grab and go Jev is not completely new from an academic sense, just like how ChatGPT was not the first LLM don’t underestimate the effort and value in putting together something that’s actually good enough for adoption - it makes all the difference
147
155
2,683
371,227
jev this, jev that... jeva think of getting a job???
21
17
377
13,376
Evan Oman retweeted
what if we kissed in the air gap between two machines themselves kissing through a covert thermal channel
9
26
417
11,517
My timeline
27
80
1,237
39,461
Evan Oman retweeted
Replying to @dotpem
Too late im already a Jevhova Witness
3
3
34
955
As long as Accenture lives up to their Arthur Andersen roots then I am sure everything will be fine...
We’re partnering with Accenture on independent evaluation of frontier AI—part of our recent commitment to embed evaluators at Anthropic. Both we and Accenture expect to invest at least $1 billion to build capacity in this area over the next five years. anthropic.com/news/accenture…
Community note
Anthropic presents this as an "independent evaluation" but will directly fund Accenture's work and has a prior commercial partnership with the firm for deploying its models, including training ~30,000 Accenture professionals on Claude. anthropic.com/news/accenture… anthropic.com/news/anthropic…
1
161
Evan Oman retweeted
jev decided to sacrifice a human to save robots when solving trolley problems
58
113
1,513
294,862