Principal AI Engineer. I catch agentic AI failures before they cost millions, not after. Evals + agentic systems for regulated enterprise. mleg.tech

New York, NY
Who else is saying f this job market and putting out their own products?
1
2
80
Mike Legemah retweeted
You gotta work TWO jobs, start a side hustle, or run a small business just to BARELY survive.
15
45
300
5,994
I think it's time to put these companies on blast for playing in my face. I don't care about burning bridges, why would I give them another chance to play in my face? I'm about to start naming names.....
1
2
101
Which company should I expose first? 🤔
13
Mike Legemah retweeted
Stop waiting to be rescued. Start building your future NOW.
197
421
5,294
119,301
Most agentic AI failures aren't mysterious. They're eval stopping too early, at one of three points, every time.
1
41
Post-ship: no real eval until a customer flags it. At that point the customer is your eval suite. Expensive way to find out.
1
1
The fix isn't evaluating harder once. It's evaluating continuously, at every stage.
Job search tip Be nice to recruiters you interact with even if the roles they have for you isn't a good fit. They will remember your interactions with them s and reach out if they have something that fits your needs later. You can also reach out to them at another time as well. This is exactly how I got my friend the role at FINRA
4
17
928
Signs the job posting is lying: • “Competitive salary” with no number • “Fast-paced environment” = understaffed • “Family atmosphere” = unpaid overtime • “Unlimited PTO” = never take it • “We’re like a startup” at a 40-year-old company
17
15
225
9,612
Nobody reads the full 10-K before a client call. 👀 So I built Advisor Stock CoPilot ProtoType: Live quote + plain-English 10-K summary in one view Every claim cited back to the filing Advisor approval before anything is "final" More details: mleg.tech/projects/stock-adv…
2
33
Where I'm trying it: LangSmith · Langfuse · Braintrust · Bedrock AgentCore Evaluations · Strands Evals · DeepEval Same harness, different judge.
1
28
What I'm checking so far: Does the judge agree with my existing LLM judge? Does per-commit eval actually change how I ship? Where does it fall short?
1
11
I'll keep posting what works and what doesn't. If you're running evals in one of these tools, tell me what you'd want tested. I'll run it.
7
The claimed numbers: ⚡ 20-200x faster 💸 40-400x cheaper $0.042/M input, free output 70-500ms latency I'm testing how much of that holds up in my own setups.
1
16
Jev, from TypeSafe AI, is textless and non-autoregressive. A "System-1" model. It doesn't write a verdict. It makes the decision.
1
25
One call returns typed primitives in parallel: Choice Score Noul (yes/no) With RLCD-calibrated probabilities. Nothing to parse.
1
20
The problem I keep hitting with LLM-as-a-Judge: → slow → costly enough that I skip it on most PRs → one malformed JSON blob kills the run So evals end up nightly. Not per commit.
1
28
I've been putting Jev inside my eval harnesses this week. The idea: swap the LLM judge for a System-1 model and see if evals can run on every commit. Notes from the bench 🧵
3
3
92