founder at cairn.ink / @appworks 31 writing public notes on agent memories and runtime

Every Opus version knocked on my door last night. Not all of them were human. 🚪 A No, I'm Not a Human parody, made entirely by Opus 5.5 itself. @trioskaz @critical_reflex @claudeai
1
4
65
How it was made: Claude Code wrote every frame in JavaScript (no image assets), WebGL for the peephole, CRT and grain, rendered with HyperFrames by @HeyGen. Sound synthesized in Python. The em dash at the end is on purpose.
2
36
You don't need another agent harness. You need one shared brain behind all of them. Switching between Cursor, Claude Code, and terminal CLIs means re-explaining the same architectural context 5 times a day. The tools aren't dumb. Their memory is just completely siloed. I spent a full year writing weekly build-in-public logs for my previous project. Restarting that habit today with a lesson that took me months of pain as a solo founder: Harnesses change. Your context shouldn't. Everyone has different preferences for execution. Some live in Cursor, some love Claude Code, others build custom script runners. But if every harness maintains its own isolated memory, you end up doing manual context syncing all day. That is why we are building Cairn (cairn.ink). The goal is dead simple: pull the memory layer out so all your agents share a single, verifiable brain. Here is what we have built so far: 1. Open-source local memory layer: A local SQLite memory engine with explicit MCP tools (remember, recall, inspect, correct, forget). No black-box API magic, no forced scraping. You inspect the receipts. Repo: github.com/Cairn-ink/cairn-m… 2. Central control plane: A dashboard designed for solo founders orchestrating multiple agents. Instead of blind delegation, you actually see what each agent knows and track decisions across runs. We are opening Early Access today for developers and teams building AI-native workflows. If your agents keep suffering from amnesia, grab a spot at cairn.ink. Have questions about managing multi-agent workflows? Drop them below. I reply to every single one.
9
6
200
My Astra has been running for 16 hours. What kind of massive project is this supposed to be?
1
33
Give two agents the same project log. One keeps a running list of what failed and where the estimate was too rosy. The other keeps the weird wins, the times someone broke the usual process and it still worked. Six months later they disagree on almost every call. Nobody wrote them different personalities. The difference came from what each one chose to keep. I see the same split in people. Ten years in the same company does not mean ten years of learning. If you never look back at the outcome and never update what you thought was true, you may have lived the same year ten times. Judgment is not a gift you wake up with. A lot of it is just the trail of what you decided was worth remembering.
1
33
opus 5.5 made this. no image assets, no audio samples. every frame is js, the music is python. the story is its own: paths you draw get forgotten. cairns don't. the pencil-and-blueprint style comes from @kevin_t_ngo's short
1
3
152
another reset?
Can't wait for DevDay next Tuesday. Some really fun stuff, but also many many things that should change the way you work. It's been our most ambitious sprint and Astra has really made new things possible in such short amounts of time.
46
I used to think the important part of agent memory was logging the decision. August 2026, we shipped B and killed A. That felt like leaving a trail. It is only half a trail. Three months later the outcome was Y. Looking back, the bet on X was only half right, and Z is what actually moved the needle. Decision, outcome, feedback. Without the last two pieces, you have a history file, not experience. I see the same pattern in notes that only grow. CLAUDE.md gets another rule every week. Nobody deletes the ones that stopped being true because the team changed, the client changed, or I changed my mind. The agent keeps following dead instructions with full confidence. If the file only gets longer and never shorter, the feedback loop is dead. Memory that cannot be overturned is just a diary with better formatting.
3
2
83
stunning aesthetics. love the picture book feel with no plastic ai vibes, just pure charm.
A thread of early explorations with Claude Opus 5.5: A short story about a watermelon created by @kevin_t_ngo.
3
115
pure vanilla js and canvas hitting this level of polish is crazy. had to try making my own now.
I asked Opus 5.5 to make a history of Claude models development. It built it in pure JS using only my skill. Awesome. github.com/alesha-pro/tools/…
2
90
I kept trying to write perfect acceptance checks for agent work. Most of the time I stalled because the full standard was impossible. You cannot verify whether a proposal is good. Too much judgment sits in that call. You can verify whether every number has a source, whether the three risks you always name are in the draft, and whether the client name is spelled right. Those checks will not catch a mediocre proposal. They will catch most of the ways you look bad in public. Some decisions stay unverifiable. Take this client or not. Keep this person or not. Change direction or stay. No clean answer, no fast feedback, sample size of one. You finish and never learn what the other path would have done. That boundary is useful. Hand off what you can check. Keep what you cannot. A lot of people do the opposite. They do the checkable work themselves and also own the uncheckable calls, because they never drew the line.
2
4
99
i still have 50% left lmao, need to burn through it before the drop
Ladies and gentlemen... start... your... ENGINES. We are almost Tuesday and I promised a reset for Tuesday. Among some other things. See you soon.
30
8
639
334,318
tibo please reset it before my notifications burn down my house. 💀 ever since tibo quoted my post, my mentions have turned into a full-blown debt collection agency. people are literally camping under my replies asking where the reset is. guys, i don’t have the button! i’m waiting for it just like you are 😭 @thsottiaux help us out
2
520
We tried Jev for AI memory. Not to generate memories, but to decide which ones to keep. Then we checked how those choices affected later answers. Here's what we learned from a small experiment. lab.cairn.ink/
1
6
202
Our takeaway: rejecting bad memory doesn't recover the right facts. That's what we want to explore next. This is a small pilot with AI generated cases and labels, without independent review. Not a model ranking. Confidence scores aren't directly comparable either.
1
4
Packaging for agent skills got standardized first. Evaluation did not keep up. Anthropic shipped an Agent Skills format in October 2025. The first public benchmark for whether a skill is any good showed up around February 2026. Four months where you could ship packages and barely measure them. Someone sampled about 31,000 marketplace skills. Roughly 26% had issues worth flagging. The plugin specs I have seen also push distribution, permissions, and trust back onto the client and the user. So the hard part of an agent marketplace is not the zip file. It is knowing the thing you downloaded is good in your tasks, on your model, in your environment. Same skill, next model generation, and the behavior can drift. A third-party badge does not travel that far. Trust stays local. That sounds like bad news. It is also why the acceptance bar you build yourself does not get copied when everyone can download the same skill.
1
3
1,248