engineerball retweeted
Anthropic 刚发了 Opus 5.5 的官方 Prompt 指南,非常值得看! 比如以前常写的「仔细思考」「一步步分析」,官方现在甚至建议考虑删掉。Opus 5.5 会自己决定思考多少,真正该调的是 effort,而且默认的 medium 在 Anthropic 测试里已经能达到甚至超过 Opus 5 的 high。 更有意思的是后半部分。 长任务要维护 checklist,跨 App 先主动探索 Context,多 Agent 甚至可以直接给它时间预算,让它自己调整并行策略。 感觉 Prompt Engineering 这件事也在变。 以前更像是在研究「这句话怎么写模型才听话」,现在更像是在设计模型工作的环境、反馈和约束。 Prompt 还重要,但 Harness 开始接管更多工作了。 官方指南: platform.claude.com/docs/en/…⁠
22
158
993
122,498
engineerball retweeted
JEV + Opus 5.5 is insane for live design... I built a live site redesigner with JEV + Opus 5.5 Paste any link → press Start → scroll, and Jev + Opus 5.5 rebuild every section of the site in front of you Full production ship in 20 seconds: 1. IntersectionObserver fires when a section is 30%+ in the viewport 2. Jev returns one typed decision in ~0.1s: { layout, copy, drop, type, palette, p } 3. Opus 5.5 writes the component (TSX) + a CSS patch for the chosen style 4. The new section wipes in with clip-path, the old one blurs out 5. Next section enters the queue, one at a time, no race conditions Output: 8 sections of a 2015 hosting site rebuilt in ~20s, streamed line by line in the terminal 3 styles, one renderer: orthographic globe + lambert shading → ASCII / 2-color halftone / ink stipple Scroll yourself and it redesigns whatever you land on Jev decides fast, Opus designs it
29
79
691
104,902
engineerball retweeted
准备面试 AI 岗位,想知道 Anthropic、OpenAI 这些大公司到底爱问什么,只能一个个帖子翻。 有人把 35 家公司公开流出的 AI 工程面试题,按公司逐家整理成了一个仓库,能找到解答的题目下面直接附了链接。 先看跨公司的高频题,注意力机制、KV 缓存、检索增强、Agent、系统设计这些分成 10 类,每题都标了哪几家问过。 GitHub:github.com/pallavi-shekhar/a… 每家公司还写了面试流程。像 Anthropic 是 30 分钟初筛、70 到 90 分钟编程测评,再加五轮左右,前后 3 周到 2 个月。 我觉得最有意思的是题目本身,比如「5 万份文档要跑一遍模型调用,接口只让开 100 个并发还会报错,代码怎么写」。 还有一题问,像 Claude Code 这样的编码 Agent,模型和外面那层框架哪个更重要,现在的面试确实很贴近实际工作了。 打算往大模型应用、Agent 方向投简历的,先把前面那 10 类高频题过一遍。
61
367
1,805
117,147
engineerball retweeted
𝗔𝗜 𝗔𝗴𝗲𝗻𝘁’𝘀 𝗠𝗲𝗺𝗼𝗿𝘆 is the most important piece of 𝗖𝗼𝗻𝘁𝗲𝘅𝘁 𝗘𝗻𝗴𝗶𝗻𝗲𝗲𝗿𝗶𝗻𝗴, this is how we define it 👇 In general, the memory for an agent is something that we provide via context in the prompt passed to LLM that helps the agent to better plan and react given past interactions or data not immediately available. It is useful to group the memory into four types: 𝟭. 𝗘𝗽𝗶𝘀𝗼𝗱𝗶𝗰 - This type of memory contains past interactions and actions performed by the agent. After an action is taken, the application controlling the agent would store the action in some kind of persistent storage so that it can be retrieved later if needed. A good example would be using a vector Database to store semantic meaning of the interactions. 𝟮. 𝗦𝗲𝗺𝗮𝗻𝘁𝗶𝗰 - Any external information that is available to the agent and any knowledge the agent should have about itself. You can think of this as a context similar to one used in RAG applications. It can be internal knowledge only available to the agent or a grounding context to isolate part of the internet scale data for more accurate answers. 𝟯. 𝗣𝗿𝗼𝗰𝗲𝗱𝘂𝗿𝗮𝗹 - This is systemic information like the structure of the System Prompt, available tools, guardrails etc. It will usually be stored in Git, Prompt and Tool Registries. 𝟰. Occasionally, the agent application would pull information from long-term memory and store it locally if it is needed for the task at hand. 𝟱. All of the information pulled together from the long-term or stored in local memory is called short-term or working memory. Compiling all of it into a prompt will produce the prompt to be passed to the LLM and it will provide further actions to be taken by the system. We usually label 1. - 3. as Long-Term memory and 5. as Short-Term memory. And that is it! The rest is all about how you architect the topology of your Agentic Systems. Any war stories you have while managing Agent’s memory? Let me know in the comments 👇
9
10
28
1,743
engineerball retweeted
La certificación oficial de Claude está disponible de forma gratuita hasta el 31 de agosto‼️ Será la certificación con mayor valor en el futuro. Destinatarios👇                                                                                                                                   El nombre oficial es «Claude Certified Architect Foundations» No es solo un curso sobre cómo usar Claude, sino una certificación técnica oficial de Anthropic Los destinatarios son Quienes implementan Claude en sus operaciones empresariales ・Desarrollo ・Construcción de agentes de IA Personas que los utilizan El alcance del aprendizaje es bastante orientado a la práctica, ・Comprensión básica de Claude ・Diseño de prompts ・Aprovechamiento de la API de Claude ・Operación de Claude Code ・MCP ・Integración de herramientas ・RAG ・Diseño de multiagentes ・Incorporación en sistemas empresariales ・Optimización de costos ・Diseño de implementación segura de IA Estos son los temas principales. En otras palabras, no es solo «puedo usar Claude», Sino una certificación que prueba que Puedes diseñar sistemas empresariales o Agentes de IA usando Claude Creo que está cerca de eso. Especialmente en el futuro, Consultoría de IA Apoyo a la implementación de IA Desarrollo de Claude Code Automatización empresarial Capacitación interna en IA Construcción de agentes de IA Para quienes quieran hacer de esto su trabajo, Creo que es fácil de usar como insignia oficial. En este momento, incluso en el sitio oficial de Anthropic, Se promociona como una certificación para La red de socios de Claude Las condiciones de examen o cupos gratuitos pueden variar según la persona, Así que es mejor verificar siempre en la página oficial. Sin embargo, como hoja de ruta de aprendizaje, Es bastante excelente No se trata solo de «probar Claude y ya», Sino de aprender de manera sistemática «Cómo diseñar y operar Claude en el trabajo».
16
228
2,130
188,167
engineerball retweeted
Vectorless RAG is quietly rewriting how we build retrieval systems 🔁 And most people haven't caught on yet. For 3 years the script never changed → RAG = embeddings + vector database. Chunk, embed, store, similarity search. It works. But the baggage is real 👇 ⚠️ Embedding drift when models change ⚠️ Chunk sizing that never feels right ⚠️ Semantic search missing exact keywords ⚠️ Vector store infra costs ⚠️ Re-indexing nightmares Vectorless RAG steps around all of it 🚫📦 🔑 BM25 / TF-IDF — the "old" stuff still hits 🔑 Knowledge graphs for entity relationships 🔑 SQL retrieval from structured data 🔑 Long-context LLMs, no indexing 🔑 ColBERT-style late interaction The real unlock in 2026? Hybrid systems ⚡ One query. Multiple retrievers. Route to the right method at runtime. Best result wins. Stop defaulting to vector search because the tutorial said so. Start asking what your data actually needs 💡 Save this 🔖 and tag a person still embedding everything. Credit: codewithbrij
3
11
37
1,305
engineerball retweeted
graph engineering explained (marketing edition) graph engineering is about designing the map your agents run inside, you draw the steps and the routes between them ahead of time, then they travel the path you put down it is the layer past looping, where one agent just circles a single task until the work meets the standard you set every agent graph is built from 4 pieces: > nodes: a stage the work passes through, research, draft, score, publish, a few run once, others are their own loop the agent circles until that step clears > routes: the paths you draw between the nodes ahead of time, every direction the work is allowed to travel > checkpoints: the check on each route that reads the result and sends the work forward when it clears or back to an earlier node when it misses > gates: a checkpoint the work cannot skip, nothing publishes until the draft clears the rubric if you have built a workflow in n8n you have already drawn one, nodes you connected, branches that fire on a condition, a step that loops until it clears. an agent graph is that same shape, each node holds an agent doing the work n8n would hand to a single api call the content graphs I run at my agency all take this shape, here is one you can build for SEO > 1 research: pull the keyword, the search intent, the competitors ranking for it, and the questions people keep asking > 2 brief: turn that research into a brief, the angle, the entities to cover, the queries the piece has to answer > 3 draft: an agent writes the article from the brief and nothing else in its context > 4 score: a critic grades the draft against your rubric, depth, intent match, originality. this node is a loop, it sends the weak drafts back to 3 and only releases one that clears > 5 publish: once the rubric clears and the brand rules pass, the agent adds internal links and the piece goes live each arrow between those is a route, every grade is a checkpoint that picks which route the work takes next, and the draft and score nodes form a loop inside the bigger map while the rest run once this is where the word graph starts to mislead. that draft and score loop can pass itself, the critic likes the draft, its rubric clears, it publishes, and still never ranks. the loop was grading the writing against another agent's opinion while the only thing that counts is whether it ranked so you add a checkpoint the agents cannot argue with, one that reads live search and AEO signals from outside the graph: > did google index it > is it climbing on the target query > are AI answers citing it > do people stay once they land if those move, the map keeps its shape and you feed it the next keyword. a stall sends the work back to research instead, because a miss this late usually traces to the angle or the intent you chose at the start, which a rewrite cannot fix anchor the map to results the agents cannot fake, and freeze the few rules they never rewrite, your brand voice and the claims you cannot make with the anchor in place, a failed piece shows you the exact node it broke on
what's the difference between a loop and a graph? (marketing edition) both are ways to run an agent, the difference is who decides the path, the agent or you. a loop still starts with you. you set the goal, the brief, and the bar it has to clear. what the agent owns is the path. take writing an SEO article: hand it the brief and it drafts, reads the draft back against that brief, rewrites the weak parts, checks again, and keeps circling until it clears the bar. the one thing you did not write is the step-by-step it took to get there. a graph is you drawing the steps and the routes between them ahead of time. same article, but now you set the map: research the keyword and the competitors ranking for it. draft from what you find. score that draft against your rubric. if it clears, add the internal links and publish. if it misses, back to the draft. the agent still decides how to handle each step, it just travels the routes you laid down. the shape of this has a name, a state machine. every node is a state the work can be in, and a check at each one decides where it goes next, forward when it clears or back to an earlier node when it misses. if you have built a workflow in n8n, you have already drawn one. nodes wired together, branches that fire on a condition, a step that loops until it clears, that picture is a graph. an agent graph is the same shape, the nodes hold agents doing the work instead of single api calls. the way I think about it, a graph is a map of loops and checkpoints. some nodes run once, others are their own loop where the agent works something out, and the checkpoints between them read the result and route the work. you keep laying down nodes and checkpoints until the map reliably gives you the output you want. the vault accelerator I run at my agency is one of these maps, 3 sessions that hand off in a fixed order: > research session: reads our company brain and past campaign results, pulls in competitor and market context, and builds the cohort we go after > landing page session: takes that research and builds the page from it > content session: uses the research and the page to write the copy, illustrations, and slides for the live sessions we run inside the content session runs a loop, a critic scores each draft against a rubric and sends it back until it clears the bar. that is one node on the map, the checkpoints between the sessions carry the work from one to the next a graph earns its extra setup on anything you run every week: > validation gates the work cannot skip > a fixed set of routes the job can take > a clear failure point, you see the exact step something broke on a loop on its own is enough for the work you only do once, where you don't know the path yet, let the agent find it. graphs earn their place on the jobs you repeat, the content pipeline, the SEO and AEO funnel step by step, the vault accelerator once the map works you reuse it, feed it the next cohort and the whole pipeline runs again past the loop, the next thing you design is the map it runs inside.
31
147
1,200
118,730
engineerball retweeted
ANTHROPIC ENGINEER: "DON'T PROMPT CLAUDE. BUILD A SYSTEM THAT PROMPTS ITSELF." In this free workshop, he explains why most people use Claude the wrong way. And how to turn one AI into an entire team of AI agents. He covers: • Proper CLAUDE.md setup • Plugins almost nobody uses • Advanced prompt caching (95% cache hit rate) • Why starting every chat from scratch is a mistake Worth more than most paid Claude courses. Watch this and bookmark it.
4
41
244
59,120
engineerball retweeted
Here's a decision tree for when you get to the end of a piece of work, and you're not sure how to continue: - Continue in the current session - /clear - /handoff - Use a subagent - /compact Posting for feedback. Which parts are confusing? What questions does it raise?
87
150
1,887
184,730
engineerball retweeted
How to turn your best prompt into a skill (in 10 min): 1. Don't write the Claude skill manually. 2. There's a faster way, and it takes 10 seconds. 3. Download a ready-made one: how-to-ai.guide. 4. Upload it: Claude settings → Capabilities → Skills. 5. Now the part nobody tells you 👇 6. Claude reads the 2-line description when it fires. 7. Not the prompt inside. The description. 8. So write it like a task, not a topic. Bad: "Helps with writing." Good: "Use when I ask for a LinkedIn hook." 9. Give each skill ONE job. 10. A skill that does 3 things fires for none of them. 11. Hooks & emails & posts? It reaches for nothing. 12. Boring and single-purpose wins. 13. Test it the right way: don't say "use my skill." 14. Just ask for the task like normal. 15. Fires on its own = you built it right. 16. Doesn't fire = Description is too vague. 17. Stuck anywhere? Type ELI5 ("explain like I'm 5"). 18. Claude drops the jargon instantly. I use this more 19. Run it on Cowork, Opus 4.8, effort on High. Same skill, better output. Here's why this beats any single prompt: A great prompt helps you once. A skill turns it into a reflex Claude has forever. It's the difference between driving a car & building a road you'll drive every day for the rest of your life. I built 27 of these - the ones I actually use daily and put them on how-to-ai.guide so you never start from a blank file. Download one, and if it changes how you work, send it to the person on your team who still thinks "Claude" is a guy's name. That's how this spreads. Grab all 27 free: how-to-ai.guide
4
26
161
14,825
engineerball retweeted
How OpenAI Built Its Data Agent We spoke with Emma Tang, OpenAI's Head of Data Platform Engineering, to get a firsthand look at how it works. In this video, we'll explain: - How it's built - How OpenAI uses Codex - 5 Key Learnings for Every Engineer
3
26
104
13,458
engineerball retweeted
Anthropic pays $750,000+ a year for engineers who can build LLM architectures from scratch. Stanford taught the entire thing in 1 hour lecture & released it for free. Bookmark & watch this today before someone takes it down and read this article below
3
30
104
27,958
engineerball retweeted
Covers 21 chapters of agentic design patterns including prompt chaining, routing, and reflection with hands-on code notebooks. github.com/evoiz/Agentic-Des…
3
81
494
22,570
engineerball retweeted
A few weeks ago everyone was talking about loops. Now it's graphs. Both live or die on one thing: your company brain. This week, Shann Holmberg shared how he runs his marketing on graphs. Here's the difference (resolving a support ticket): 𝗟𝗼𝗼𝗽𝘀 → you set the frame, the agent owns the path. Hand it a goal and a bar. It drafts, checks, fixes, and loops until it clears. You don't pick the steps. The agent does. 𝗚𝗿𝗮𝗽𝗵𝘀 → you draw the path first, the agent fills each node. read → triage → draft → QA → send, with a checkpoint at each step. The agent solves each box, but the map is yours. Every one of those nodes pulls from the same place. Triage needs past tickets. Draft needs the policy and product docs. That shared context (your company brain) is what each box reaches into as the work moves through. The brain has that map, the graph is that process, inferred explicitly. 𝗦𝗼 𝘄𝗵𝗲𝗻 𝗱𝗼 𝘆𝗼𝘂 𝘂𝘀𝗲 𝘄𝗵𝗶𝗰𝗵? → Reach for a 𝗹𝗼𝗼𝗽 when it's one-off work, no clear path yet. Let the agent figure it out. → Reach for a 𝗴𝗿𝗮𝗽𝗵 when it's repeatable work you already know the steps for. Lock them in. Loops are for exploring. Graphs are for scaling. Which one are you running right now? 👀
28
214
1,186
50,798
engineerball retweeted
ANTHROPIC ACABA DE PUBLICAR UN CONTENIDO PARA MONTAR UNA EMPRESA SIN EMPLEADOS - CEO: 1 persona. - empleados: agentes de claude dura 30 minutos y es gratis Guarda esto en favoritos para que no lo pierdas.
35
295
1,402
142,057
engineerball retweeted
Dario Amodei (Anthropic CEO): "You're only doing 5% of the task, the AI does the other 95%, and so you become 20 times more productive." What used to block that level of leverage was the code itself. In a recent talk, Dario described Cowork as "Claude Code for non-coders" - the terminal is no longer the barrier to entry. Most people are still doing the entire task by hand, one output at a time. The real skill now isn't prompting or building agents. It’s taste - knowing which 5% is worth keeping for yourself. No agent will hand that judgment back to you. The four-piece breakdown and two working builds that actually deliver this kind of leverage are in the article below.
2
6
27
5,459
engineerball retweeted
ANTHROPIC QUIETLY SHIPPED A FREE PROMPT LIBRARY FOR CLAUDE CODE it is copy paste, sorted by task and by role, and most people have never clicked it once so they keep writing every prompt from zero when the answer is already sitting on the page the whole thing runs across the full lifecycle: discover > design > build > ship > operate a few of the ready made ones, in my words: - what breaks if I delete this helper - plan this refactor, list the files, leave the code alone for now - write the tests, run them, fix whatever fails - this test is red, track down the cause and patch it it reads less like a cheat sheet and more like a map of everything Claude Code can already do on its own this is the same reason I keep tightening my CLAUDE.md file, the people who set these rules up early are pulling years ahead of everyone still typing prompts by hand I broke down the exact ones I use in 20 CLAUDE.md Rules for Getting Ahead of Your Competitors by 5 Years bookmark the library before you forget it exists
10
32
187
54,273
engineerball retweeted
Anthropic just dropped 5 workshops, revealing the latest capabilities of Fable 5: • 00:00 - deep look into Fable 5 • 11:22 - Fable 5 and the capability curve • 30:54 - building managed agents with Fable 5 • 44:29 - real use cases of Fable 5 by teams • 57:43 - how to deploy agents with Fable 5 These 1-hour of sessions will replace 100 articles on how to actually use Fable 5. Watch them today, then read the best practices from the sessions in the article below.
61
494
2,851
554,839