Principal Investigator | AIxCC Lead Architect (views my own)

Columbus, Ohio
In AI if you move the goal posts, own either team, or the refs, these evals aren’t wins or losses, they’re marketing theatre. Effective eval should naturally fall out of any system doing real work and subject to independent verification.
99
The ability of an LLM or agentic framework to refactor a code base into a new language, while impressive, does not equate to an effective software or product development lifecycle. The years of toil writing effective tests are key. It would be an interesting experiment for teams to attempt re-writing Bun or NodeJS with all tests removed first. The ability to replicate something that has decades of lessons learned does not mean a mechanism to autonomously create production grade innovative products. Effective tests and evaluations are probably the biggest pain I continue to see and not a void the LLMs seem to fill well.
1
2
109
Can I coin the term TADR - Too Agentic Didn’t Read?
2
2
121
Or TSNR… “Too Slop Not Reading”
13
Or maybe “Too AI Didn’t Read”
1
25
Forgot to share but here is the repo behind flaude.org github.com/obsecurus/flaude
1
162
Also a couple hidden Matt Damon movie references.
52
We need the @haveibeenpwned equivalent but instead as “haveibeencloned” and instead of leaked passwords it focused on leaked repo context/IP into frontier model providers. @troyhunt
1
172
Whoa 60,000x faster than Fable 5, essentially free, and open-source, check out QED-1 flaude.org
1
6
14,852
More about this here though: open.substack.com/pub/obsecu…
2
102
They should replace “pioneer day” with “data center day” for grade school. I had to go on a waste management facility field trip.
45
Fable 5 applying forced token maxing by converting to a Python strategy over RTK. I've personally found Python generation far less reliable and reliable than Golang, Rust, or TypeScript. I think this is primarily because of the strong typing, but perhaps even moreso the abundance of opinionated community content about the "right" way to implement something that likely went into training.
1
3
95
Forget "token maxing", I've derived this new agentic concept of "outcome maxing". This novel concept is where you focus on producing real world results, embrace stakeholder empathy, and maximize value by scaling desired outcomes for your customers.
62
The details of AIxCC challenges are now live at archive.aicyberchallenge.com… Huge thanks to all the challenge authors who contributed, and the teams who helped review content.
49
If you haven’t re-evaluated delivery services lately for Instacart, Uber, Door Dash, it’s probably a good time to do so. Seeing extreme price increases. +36% just for the item, then nearly 2-3x for delivery from 1/4 mile away. 4x if you take the max suggested tip.
1
1
149
generated company-profile.md
37
Does anyone know if there are any models specifically trained on data before GPT-3 was widely available? I know Sam has mentioned that GPT-4 could potentially be this but I think we need a site and a model like "before[.]ai" or similar that is specifically dedicated to continual training on ONLY datasets validated to be available before GPT-3 or some other clear line in the sand where we know models we're generating a significant volume of data on the internet. If we don't do this as some sort of public good, open-source, shared offering I think attempting to baseline understanding and separate raw human content from model generated content is going to be impossible. This is not about generating a large extremely capable model for free, but about human vs. AI/ML provenance. This seems like the sort of thing an academic institution or standards agency like NIST, etc. could help maintain with the support of large model vendors @sama @DarioAmodei @elonmusk
2
1
81
@huggingface might be an interesting endeavor for you also?
38