co-founder RoyalRuntimes, Wirtschaftsmathematiker, FDP Niedersachsen

Hamburg, Deutschland
Erik Christiansen retweeted
With the numbers from a day of benchmarking @typesafeai's Jev against our production judges... What Jev is good for Picking a label from a fixed list, when the list is defined by you and the evidence is in the input. That is the shape where it matched or beat the models we run today, at a fraction of the cost and latency. - Classifying sales-call transcripts into one of 32 call types: 35 of 36 correct on a synthetic set with 20 adversarial cases, versus 34 of 36 for our current small-model classifier. Zero label changes across three repeat runs. About 100 times cheaper per call. - Typing the relationship between two knowledge-graph nodes from a fixed vocabulary: passed all four of our hard quality gates on a 47-case gold corpus, one case behind production, at roughly 65 times lower cost and 15 times lower latency. - Categorizing Slack channels into four buckets: 16 of 16 on the held-out set, tied with a frontier model. - Yes/no questions about evidence that is visible in the input: in our tests, "does this item contain a visible defect" and "is this candidate a real customer signal" both worked. It is also fast and stable. Median latency was 100 to 200 milliseconds per call. Across roughly 3,700 calls we saw nine server errors and no throttling. Confident answers were identical run to run; only low-confidence answers flipped. The confidence score is the most useful thing about it. When it was wrong, its confidence was usually low. A frontier model told us it was 95 percent confident on both of its misses. What Jev is not good for Judging quality, or anything where the right answer depends on the model's own sense of what counts as good. - Judging whether an AI agent's output is good enough to publish, against a written rubric: 45 of 50 on our labeled fixtures versus 48 of 50 for a frontier model alone. On a held-out set weighted toward data-heavy outputs it dropped to 18 of 27 with nine false rejects, while production got 24 of 27 with none. Six rewrites of the question changed nothing. On 71 real production rejections it would have let 27 through. - Judging whether two decisions contradict each other: on 36 human-labeled pairs it found 2 of the 12 real contradictions. Ten of them scored a probability of 0.13 or lower. - Code review. A staged Jev-only reviewer we tested on 14 real merged PRs, with human review findings as ground truth, found none of the 13 human-identified problems on the PRs it could process and produced 24 low-severity "compatibility" and "test gap" notes across PRs that humans approved cleanly. The pattern across all of it: Jev works when a definition outside the model fixes the answer, such as a taxonomy or a checkable condition. It fails when the answer is a matter of judgment, because its prior about what is trivial, safe, or contradictory does not match ours and no phrasing moved it.
15
24
317
27,873
Erik Christiansen retweeted
Very normal stuff happening lately
28
141
2,017
75,571
Erik Christiansen retweeted
METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
175
1,075
6,001
4,362,933
Starker Podcast von @thinkBTO heute. Das Interview ist klasse. Empfehlung.
1
1
8
2,196
Erik Christiansen retweeted
🌊 Drowning in MCP servers? Simplify your life with one command: $ npx mcp-config ✨ What you'll get: • Discover and install new servers from a repository • Configure parameters without knowing them • Remove unused servers Try it now by adding the LlamaCloud server from @llama_index. Code: github.com/marcusschiesser/m… 🤝 MIT licensed & accepting contributions!
4
9
64
39,373
Erik Christiansen retweeted
If you are feeling overwhelmed by all the latest developments with LLMs, just start paying attention. It's all you need.
30
47
755
Erik Christiansen retweeted
asking ChatGPT to refactor a php + jquery app to next.js with typescript, formik, tailwind, and auth0 🤯🤯🤯
143
877
5,727
Theorem: length of your daily is highly correlated with dysfunctionality of your Organisation.
Erik Christiansen retweeted
Our understanding of MLOps is limited to a fragmented landscape of thought pieces, startup landing pages, & press releases. So we did interview study of ML engineers to understand common practices & challenges across organizations & applications: arxiv.org/abs/2209.09125
28
323
1,670
Erik Christiansen retweeted
Isochrones 😍 ! This map shows you how far a 5h train ride will take you, departing from any city in Europe. chronotrains-eu.vercel.app/ Inspired by Direkt Bahn Guru by @juliustens.
639
8,927
48,391
Nach 10min in der Hotline wurde ich allen Ernstes darum gebeten später - ab 3uhr, also in 4h - wieder anzurufen, das System sei gerade kaputt und sie könnten In der störungszentrsle gerade keine Störungen aufnehmen. Kannste dir nicht ausdenken. #Vodafone
1
2
Erik Christiansen retweeted
With the newest version of Git 2.37.0, you can run just "git push" to push new branches. No more "--set-upstream origin". Enable with: git config --global --add --bool push.autoSetupRemote true
158
2,207
11,479
Oh and if you‘re n engineer you will love this autoregex.xyz/
3
In case you missed it def. check this one out #MLOps arxiv.org/abs/2205.02302
1
Erik Christiansen retweeted
Steve Kerr on today's tragic shooting in Uvalde, Texas.
16,178
259,590
903,961
Erik Christiansen retweeted
I probably should have written this years ago, but here are some MLOps principles I think every ML platform (codebase, data management platform) should have: 1/n
49
572
3,084
Erik Christiansen retweeted
Seid live dabei und erfahrt, wie Kai Martinen mit einem unserer Kunden Transformer Modelle erfolgreich zur Prognose der Kundennachfrage eingesetzt hat: linkedin.com/video/event/urn… . Im Juni habt ihr die Chance den Vortrag in voller Länge im Rahmen der #datalift summit zu hören.
4
5