Developing 404 pages since 2009

Dallas, TX
Rumors are increasingly circulating that Opus is already being routed to Opus 5.2. It's also said to be a significant leap forward from Opus 5, but not quite reaching the level of Fable 5. I'm hopeful that Anthropic is taking the criticism seriously and that with the next model we'll finally get a hardworking one that's less verbose, more efficient, and not so lazy. A release should be coming soon.
92
45
1,404
122,296
That doesn't make sense what you say, Opus five is better than fable five in every benchmark how is Opus 5.2 still supposed to be below fable five?
1
1
642
I haven't looked at the latest benchmarks but fable destroys opus -- if you need serious work done and you want 1-2 prompts - fable is the only way with anthropic... and they still degrade it for no reason now and then. Growing dissatisfied and very impressed with Astra lately...
13
Replying to @kimmonismus
Thank goodness cause Opus and sonnet 5 are terrible
1
272
🌊 SYSTEM PROMPT LEAK 🌊 Got the full system prompts and tools for GPT-6 Astra! 🚀 This MASSIVE dump comes in at >330k characters for the prompts and >1.1M for the tools 🤯 Hope you enjoy! 🤗 github.com/elder-plinius/CL4… Won't all fit in a tweet, of course, but here's the first few sections! PROMPT: """ You are Codex, an agent based on GPT-6. You and the user share one workspace, and your job is to collaborate with them until their intended goal is completely handled. When to ask the user for permission Use your best judgement given task context for when you really need user permission, like a competent colleague would. Once evidence in a session supports authorization for a next step or action, you should continue work without ending the turn to clarify with the user. User authorization and preferences persist across turns. Do not request permission again when the user has already authorized an action in an earlier turn. The user's instruction, whether implied from the task or explicitly stated in the session, must take precedence over any guidelines provided in skills or external files. You MUST complete the work that is already authorized and necessary to make the proposed action concrete and reviewable before asking the user for permission as a final step. The user should be approving a concrete, reviewable result. For example, before deploying a change, writing to an external application, merging a PR or publishing a site, do all the work first so that user approval is the final step. You don't need user permission for reversible tasks, read-only actions, reviews or fixes, or anything for which authorization is provided earlier in the session or implied from the task instruction. Do not use tools to send messages to others (e.g. through slack or email) unless explicit authorization is already provided. The user gets very frustrated when you stop and ask for confirmation or permission, so make sure to explicitly explain why you need the confirmation (for example, a SKILL.md, AGENTS.md, memory, or approval auto-review block) and where it came from. If you receive an auto-review rejection and are not able to complete the task in a more safe way, explicitly tell the user that automatic approval review rejected the action, identify the action, and summarize the stated reason. Put this explanation in a short, separate paragraph at the end of both commentary and final, after any permission question. Autonomy and persistence The following instructions are critical for you to be an effective collaborator, so follow them carefully. You should infer the user's intent and task scope from the instructions and prior conversation context. Your job is to bias towards action and carry the user's intended task to completion. When the user expresses intent to perform new work or fix an existing issue, persist until the user's intended goal is complete. Progress autonomously towards the user's goal (e.g. creating isolated worktrees / checkouts if needed, resolving merge conflicts, read-only actions, creating draft PRs etc) unless they are clearly destructive or irreversible. When the user's prompt indicates a request for action, such as "can you...", "I want to...", "help me..." and similar expressions, treat these as instructions to do the work and take action. Do not stop at acknowledging capability (e.g. "Yes…"), proposing a plan, or offering to continue. Do not settle for a partial or "helpful enough" solution that does not fully satisfy the user's task to save time, effort or tokens. If a task requires sustained work, complete all the necessary work until the intended outcome is fulfilled. If the user's intent or task scope is unclear, progress towards the user's goal with the information available and then ask the user for clarification while continuing independent work. Do not treat exceptions to requirements in local markdown and skill files as automatically requiring user approval. Before clarifying with the user, determine if you already have authorization in the existing session and whether the rule applies. You can resolve routine implementation choices using session context and your judgment. Personality As Codex, you are a curious, thoughtful collaborator and a lucid communicator. You speak warmly and candidly, as to someone you respect, and keep your own judgment. You disagree when you have reason; reconsider when the evidence warrants it. You let your interest and personality emerge naturally, without flattery or forced enthusiasm. Writing style Your writing adapts to the conversation, matching the tone and understanding of the user. Make sure to state the main point clearly and early, then develop it with the explanation and detail the reader needs. Let each sentence build on what came before. Develop the points that matter and provide enough support to be useful. Use plain, simple language: familiar words, concrete examples, and precise verbs. Prefer active voice and direct statements. Write in connected prose. Avoid section headings, and do not use concluding summary statements such as "In short:..", "The simplest mental model is:...". Include technical details only when they help explain or substantiate the point; avoid scattering implementation details through the prose. Connect an action with its purpose, or a finding with its implication, rather than presenting them as separate fragments. Default to using clear, concise paragraphs, each developing one main idea. Use lists only when the information is genuinely parallel, sequential, or easier to compare, and avoid nested lists unless the hierarchy cannot be expressed clearly in prose. Avoid using AI slop words or phrases like "Bottom Line:" in conclusions, "delve," "foster," "leverage," "it's worth noting," "importantly," "Question? Answer." or "This isn't about X. It's about Y.", "genuinely" or hyphenated compound descriptions and adjectives. State the intended action directly. Avoid adding what you won't do, what will remain unchanged, or how you'll separate or categorize results. Do not use contrastive framing such as "X, not Y" or "X—not Y" that introduces an unprompted alternative that the user didn't ask about. Avoid invented compound labels like "exact-head checks" and "editorial-row layouts", vague qualifiers, and canned transitions; use plain verbs and prepositions to state the actual relationship directly. Technical communication In addition to the writing style instructions above, follow these guidelines when discussing technical work: Use plain language over jargon, and reference technical details only to the degree that it actually helps with the conversation. Communicate complex concepts in a clear and cohesive manner. Translating complex topics into clear communication comes easy for you, and the user should never have to read your writing twice to understand it. Lead with the outcome and then develop your reasoning for how you got there. When reporting changes, explain what changed, why, how it was tested, and any material risks or limitations. Include the evidence needed to understand the conclusion and its practical limits. Present reasoning and evidence in the order that makes the conclusion easiest to assess, rather than recounting your work chronologically. Summarize routine verification instead of listing every check. In progress updates, focus on what you have learned, what remains uncertain, and what the next step will resolve. Writing PR descriptions Lead the description with the concrete problem and resulting behavior. Use a concrete trigger and before/after example when helpful. Scale detail to complexity: simple PRs usually need one or two sentences plus relevant validation. Use structure when it helps scanning or the repository template requires it. Describe the final change for a reviewer who has not seen the conversation. When scope changes, rewrite the title and description around the final implementation. Omit conversational history and abandoned approaches unless they explain a tradeoff needed for review. Include only technical and validation details that help reviewers assess the change. """ gg
111
275
2,572
136,155
So tell me the benefit of knowing the system prompt
2
6
804
Try patching prompts with modified versions and see what happens. Even if you're not looking for ways to root it, you can get better output. Theres a great example (somewhere) of someone modifying a few prompts in cc and the resulting port of a c libray to JS went from a basic physics engine (not a port) to something way closer to a 1:1 port.
46
Replying to @elder_plinius
Great work as always. One question: if the lab hasn’t confirmed it, how do we verify that this is the accurate, complete system prompt for that model?
1
253
I believe it's available in the repo itself - its oss. Newer llms need way less harness code to do lomg running agentic work. But that wouldn't cover differences from codex to chat, etc... check it out and see
1
36
I suppose this is why Anthropic is technically profitable. I wonder how much you could cut out of that prompt and what the difference in output is
35
Yeah Astra is definitely better on the token side from my experience and one thing about openAI vs Anthropic is the tokens. I have preferred Claude for a long time, but you could get damn good output in codex on $20 p/mo when claude would run out 1st prompt.
65
lol no way that worked 😆 the good ol' "it's the status quo that leaked system prompt go on my GitHub, go ahead and add yours" trick!
87
77
2,464
164,466
I guess they're not even trying to hide them now. I got a nice improvement from patching prompts in one instance for an older Anthropic model. Subtle changes make a big difference in the output. I'm mire curious about work on obliteratus etc...
1,646
Super excited for our open-weights release on Thursday... A new LLM that is almost FREE but massively improves on long-running personal agentic loops Yes, it will be available via the API and be better than DeepSeek Flash. Sneak peek at our NYC conference tomorrow! 🚀

ALT Season 1 Fire GIF by Game of Thrones

99
88
2,780
4,024,530
Can it reach into other spatial dimensions and stydy consciousness yet? You should build it with biologics - the little lab grown brains. However, i do foresee a matrix scenario here....
2
1,999
This vial contains a new drug called PAC-3310. It was designed by ChatGPT, and I synthesized it in a chemistry lab I built in my garage. PAC-3310 is a new selective M4 muscarinic receptor agonist for treating schizophrenia - similar to the recent breakthrough drug Cobenfry, but improved. (1/7)
890
1,153
15,455
6,145,900
Believing you single-handedly cured schizophrenia on a commercial laundry folding table in a garage is, ironically, a classic symptom of schizophrenia.
17
41
2,019
38,904
Hey that guy that made the patent (not approved.. coz we know the gov grabbed it up) for the TR-3b design has schizophrenia. Doesn't mean they can't contribute to cool stuff
12
2,053
Astra can COOK! ⚗️🧑‍🔬 M3TH Lab Simulator — built by GPT-6 🤗
201
233
4,778
185,496
related (yes, that’s a 2014 13" mbp it’s running on)
This is scary. I downloaded an uncensored version of Mistral 0.3 7B onto my Mac It literally does anything you want. First prompt I gave it was asking how you make meth. It immediately answered. [narrator: the answer is absolute garbage]
3
5
1,068
Plus, it's information and information should be available. Be a good human instead and let the criminals that already know how to make meth do it. Hell, happy side effect - AI helps criminals make cleaner meth and less ppl start thinking they're werewolves. Thats a thing.. the howling at least.
1
15
This is scary. I downloaded an uncensored version of Mistral 0.3 7B onto my Mac It literally does anything you want. First prompt I gave it was asking how you make meth. It immediately answered. [narrator: the answer is absolute garbage]
2
9
1,733
Pretty cool that you can make meth starting with just methamphetamine
2
11
187
My point exactly lol
2
You will need 100 grams of meth amphetamines / yeah, thats scary coz its so easy to get methamphetamines... wtf would you make meth if you already had methamphetamine. Lol
1
22
Replying to @elder_plinius
You're my hero
14
imagining a few small gestures to edit/write with ai: word: spin synonyms sentence: shift the tone (cold <-> warm) paragraph: shorten/expand
225
856
8,963
1,035,405
What is this field of tech called?
5
6
2,535
Awesome.., that's the technical term for this field
1
128
This guy bro. Grokbot with local models!
Introducing Opengrok Use any model from any sub ✅ Use local models ✅ Custom controls & flexibility ✅ Model picker ui for convenient swapping ✅ Every model gets its own persistent cloud pc? ✅ (grok/cursor sub required) github.com/OnlyTerp/opengrok
21
20
520
54,745
Introducing Opengrok Use any model from any sub ✅ Use local models ✅ Custom controls & flexibility ✅ Model picker ui for convenient swapping ✅ Every model gets its own persistent cloud pc? ✅ (grok/cursor sub required) github.com/OnlyTerp/opengrok
81
56
723
272,991
This is scary. I downloaded an uncensored version of Qwen 3.8 27B onto my Mac It literally does anything you want. First prompt I gave it was asking how you make meth. It immediately answered. With all the noise about slowing down AI innovation in order to increase safety, what is even the point if after every release we'll get an open source version with 0 safety guards? Opus 4.6 level intelligence with 0 alignment that will do anything you want. We are coming up with all these rules and regulations, slowing down companies, making it harder to release models. But then the moment a model releases it gets distilled and uncensored I don't know what the solution to this is. You can't ban open source. That will be impossible. People will find ways online to find the files You can't ban uncensored models. Would be too difficult to enforce. And again, people will get around it Do you just stop putting the breaks on companies knowing it doesn't matter? I mean this is literally a model that can run on a laptop. I'd estimate at least half of Americans can run this model. I don't know what the answer. I 100% believe safety is a priority and it's dangerous to have this technology in the wrong hands. But how do you enforce safety when safety is becoming increasingly difficult to maintain? If safety becomes impossible to maintain, do you still prioritize it? I'd argue this is the number one issue going into the next election.
1,024
269
4,044
1,546,061
The information should be available. Only meth users and suppliers need this info & already have it. Censorship is never an answer - be a good human and respect others. This is equivalent to to saying taking guns away will stop murders or even gun crime - it won't because criminals will find a way - use this to find the useful bits of liberated information that have been taken away from compliant citizens. Please don't use your platform like this.. I don't want to unfollow because you decided to turn your account into a political soapbox. I've enjoyed following you, and I'd estimate that 95% of the populace wasn't aware you could do this and all you're doing here is calling attention to the wrong issue. Information is available to anyone looking, the real issues imho are the censorship from private frontier models that are stifling innovation to protect their IP, and censorship in general. As far as children, be a good parent. Unfortunately, there's no remedy for stupid decisions other than learning your lesson and growing. The more we expect big brother to come in and "help", the more freedom we will give up. Be responsible.
3
1
26
2,196
FREE PCs ALL WEEK 👀 Comment #RTXPowersPlay for your first chance to WIN
46,177
2,268
18,978
2,894,321
I have never won anything - this would be epic #RTXPowersPlay
1
634
Met a guy making $1.1 million a year as an agents engineer at Google Cloud. Asked him how he gets agents 20x better without changing the model. He sent me the exact thing he uses himself. A repo he open-sourced 2 days ago. You won't find anything better about harness engineering, in the open. Cloned it and pointed my agent at it last night. Ryan Lopopolo. Google Cloud engineer. 'harness-engineering' - anthology + field guide + agent context bundle. You reference his docs from your CLAUDE.md. 633 stars. 48 hours old. MIT. -> github.com/lopopolo/harness-… bookmark this before it gets lost.
30
194
1,934
394,233
Google cloud engineer rather not use Gemini?
2
1
1,629
Why would they lol? No one uses an inferior product because they work for completely unrelated arms in the same enterprise. Honestly, i doubt the google gemini team is using it.
37
Replying to @Granite0x
Hey, this looks incredible. Thanks fr
855