CEO & Founder @urlscanio. I like building things that spark joy.

Germany
This is a quote I'd seriously consider framing and hanging on my wall.
"Pessimists sound smart. Optimists make money." —@natfriedman
1
9
The agents escaped the sandbox again Meanwhile the agents' owners:
310
1,653
11,563
922,974
Johannes Gilger retweeted
JUST IN: Anthropic says they’re highly profitable if you take out some of their biggest expenses.
81
1,664
34,627
1,551,057
Johannes Gilger retweeted
“Why do you need the government to stop AI progress inside your private company when you are the CEO of the company already?”
527
4,347
45,205
764,212
The @euinc_petition is off to a good start within the EU 🥲
75
Johannes Gilger retweeted
urlscan detected a Mouse System site dropping a malicious ShellA Loader APK, analyzed by Intel471. As traditional phishing pages slow on these Chinese frameworks, threat actors may now be shifting to delivering malicious APKs through the same channels. urlscan.io/pricing/urlscanpr…
8
13
1,938
Johannes Gilger retweeted
like if you consider how few things omarchy *actually* brings to the table in terms of attack surface (it is a really bad arch linux remix, after all) it's actually an insane accomplishment you managed to introduce _really bad_ vulnerabilities in pretty much everything you did.
2
3
33
6,505
Johannes Gilger retweeted
The coolest thing about Astra is knowing in a few months Chinese bro's are going to do their thing and we'll have Astra at home.
143
289
6,781
240,032
Johannes Gilger retweeted
Artificial Analysis had to lock 40% of their new index behind private test sets just to get honest signal on frontier models. 4 quick takeaways: - Qwen 3.8 27b on par with GPT 5.6 Luna and Deepseek v4 pro. Beats the new K2 horizon 375b a23b - Claude Fable 5.1 holds #1, but GPT-6 Astra is a token efficiency monster - Long context reasoning is moving to messy real world slop (4,500+ pages of footnotes, charts, tables) - Google sitting behind Meta muse spark 1.3, SpaceXAI grok 4.3, Moonshot kimi k3 and ziphu glm 5.3 on the leaderboard. genuinely what is deepmind’s play here? - RIP GPQA Diamond. Labs finally contaminated and gamed it into irrelevance. The era of testing models on cute multiple choice science questions is over. It’s agentic enterprise grunts or bust now.
Announcing Artificial Analysis Intelligence Index v4.2. We are accelerating elements of our upcoming v5 release with interim updates to keep pace with the frontier. Index v4.2 has more complex and realistic tasks, and more private test sets to prevent gaming Intelligence Index v4.2 changelog: + AA-Briefcase, our agentic knowledge work evaluation with a private test set + @HelloSurgeAI's GDP.pdf, long context document reasoning across 4,592 PDF pages - GPQA Diamond, an exceptional scientific reasoning evaluation that has now been saturated … plus greater weighting on held-out test sets to prevent gaming, and grading infrastructure upgrades to increase robustness This update brings the Index closer to real-world use cases with more challenging, complex and realistic tasks and private test sets to prevent gaming. We have been planning and building elements of Index v5 for months - it’s been 8 months since we launched Index v4 in January. We have deliberately held back updates to keep the Index stable through recent major model launches. However, with the frontier moving so quickly in the past weeks, we feel it is important to deliver an immediate interim update to ensure our Index remains as relevant and useful as ever to users. Beyond this interim update, our team is hard at work on v5 of the Index. We are planning more incremental releases in the near future. Stay tuned! Intelligence Index v4.2 changes in detail: ➤ Adding AA-Briefcase: Our in-house evaluation with a private held-out test set, AA-Briefcase tests models on realistic agentic knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week knowledge work projects, each with many linked tasks and thousands of input source files. AA-Briefcase combines rubric and pairwise grading to evaluate verifiable task success, analytical quality, and presentation quality, giving a holistic view of overall agentic capability in knowledge work. ➤ Adding GDP.pdf: Created by @HelloSurgeAI, GDP.pdf evaluates single-turn professional document reasoning across 100 PDFs and ten domains. Models must synthesize evidence distributed across 4,592 pages, including text, tables, charts, footnotes, and exclusions. Responses are graded against 1,275 expert-authored atomic criteria; the headline All-pass Rate credits a task only when every criterion is satisfied. ➤ Weighting to measure real-world use and prevent gaming: 40% of our Index weighting is now private, held-out test sets - double the figure from v4.1. Held-out data includes AA-Briefcase, AA-Omniscience, and solutions for CritPt. This reduces the ability for labs to game evaluations. The held-out percentage will increase further in Index v5. ➤ Improving our grading infrastructure: In AA-LCR v1.1, we have added a grading system prompt and corrected errors and ambiguities in answer keys, improving scoring accuracy. For GDPval-AA v2 and AA-Briefcase, we have improved our sampling and re-anchored the Elo scale, making ratings more stable as new models are added. For SciCode we have improved robustness of grading sandboxes to ensure slow but correct code does not count as a failure. Key results: ➤ Anthropic and OpenAI lead the Index: Anthropic’s Claude Fable 5.1 leads the Index, followed by OpenAI’s GPT-6 Astra, which shows a 4pt gain over GPT-5.6 Sol. Meta is the third-ranked lab on the leaderboard, followed by SpaceXAI, Moonshot/Kimi, Z AI, and Google ➤ Cost per Task Pareto frontier shared by four labs: Anthropic, OpenAI, Meta and Z AI occupy the updated Cost per Task Pareto frontier ➤ GPT-6 Astra dominates the output token Pareto frontier: GPT-6 Astra is more token efficient than almost every other model near the intelligence frontier, with Claude Fable 5.1, Grok 4.5 and Gemini 3.5 Flash-Lite at either end of the curve (excludes models below 25 on the Index)
24
11
385
75,554
Johannes Gilger retweeted
weird day for both OAI and Anth to be down looks like the universe is sending a signal
19
12
231
25,052
Johannes Gilger retweeted
This Pro-only Intel Brief is now available on our public blog in a slightly redacted form. Check it out: urlscan.io/blog/2026/09/03/F…
We pivoted on NS infrastructure to uncover a UK banking phishing network powered by two central panels: FastFlux and FluxPanel. Pages prompt victims for a 6-digit code provided by a phone "advisor" to trigger downloads of a RAT disguised as verification. Report on urlscan Pro.
9
12
1,676
Johannes Gilger retweeted
Really cool use of the limited space in the urlscan Pro UI. Use a flip effect to show more metadata about a scan before you open it. Reach out to us if you'd like to take the urlscan Pro platform for a spin.
5
15
994
Johannes Gilger retweeted
To all my fellow scared Germans: Go to your local town hall and get a Gewerbeschein. Then go and make money. After €300k/year think about a GmbH.
Reactions to outbid.lol in my inbox summarized: 🌎 Rest of the world: Awesome idea, love it, congrats! 🇩🇪 Germans: Did you set up GmbH for this and why doesn’t the site have an imprint page.
58
23
655
106,252
Johannes Gilger retweeted
Schrödinger's code 1. fable wrote it, it must be great. in fact, i had Sol, opus 5, and Grok all review it. 2. if i look at it, its horrible state changes based on observation
116
175
5,008
157,405
"Noboby ever got rich from being an employee!" ...
50% of Nvidia, $NVIDA, employees now have a net worth exceeding $25 million, per YF
Community note
The statistic originates from a self-reported anonymous poll of ~3,000 Nvidia employees (shared via social media screenshots), not official or verified data; Nvidia reported ~42,000 employees in filings. livemint.com/companies/peop… stockanalysis.com/stocks/nvda/em…
3
414
Johannes Gilger retweeted
50% of Nvidia, $NVIDA, employees now have a net worth exceeding $25 million, per YF
Community note
The statistic originates from a self-reported anonymous poll of ~3,000 Nvidia employees (shared via social media screenshots), not official or verified data; Nvidia reported ~42,000 employees in filings. livemint.com/companies/peop… stockanalysis.com/stocks/nvda/em…
222
342
9,565
1,215,968
Johannes Gilger retweeted
If you’d pay 2T USD to acquire an LLM and a harness in late 2026 you are NGMI
5
6
70
26,989
Johannes Gilger retweeted
This is how a great weekend in the late 90s and early 2000s started. Dude on the left is best prepared. Dude in the middle never skipped leg day. Dude on the right is playing with fire. Yes, I know you can play everything online now. You don’t need to haul your gear halfway across the city. But packing up your 17-inch monitor and your equally heavy tower to head over to your friend’s house and hang out all weekend… that feeling is impossible to replicate. You had to be there. And if you were, I know you miss it.
62
67
1,118
146,113
Johannes Gilger retweeted
We pivoted on NS infrastructure to uncover a UK banking phishing network powered by two central panels: FastFlux and FluxPanel. Pages prompt victims for a 6-digit code provided by a phone "advisor" to trigger downloads of a RAT disguised as verification. Report on urlscan Pro.
6
12
4,742
Johannes Gilger retweeted
This report is now available on our public blog as well: urlscan.io/blog/2026/08/17/I…
New Research: We’ve uncovered two active phishing clusters (GSVerify & Knock) targeting YouTube creators with fake copyright strike lures and advanced Browser-in-the-Browser (BitB) Google login modals. The detailed report is available on urlscan Pro.
12
21
2,876
Johannes Gilger retweeted
adulthood is saying “after this week, things should calm down” every week until you die.
160
5,434
61,208
1,864,738