Meet the coding agent built for open models. Command Code. We do tool repairs, /design, #1 on leaderboards, and the world's best $10 GOAT plan w/ $70 usage.

San Francisco, CA
New in Command Code: /loop command Run a task, then repeat it on a fixed or self-paced schedule until the job is done. Try now: $ npm i -g command-code
11
160
9,673
You can now use Shift+Tab to bypass permissions. · Switch mid-session · No need to restart with cmd --yolo
21
5
158
9,092
Tested Mimo-V2.6-Flash, Grok-4.7, and DeepSeek-V4.1-Flash with 3D SubwayBench using the same /design prompt. 🔹 DSV4.1-Flash: 9/10 · $0.024 · one-shot · strong 3D UI/UX + gameplay 🔹 Mimo-V2.6-Flash: 8/10 · $0.018 · playable 2D result in ~4 iterations 🔹 Grok 4.7: 3/10 · $0.35 · slower UI + gameplay, needs more iteration
19
15
322
30,691
Command Code supports Git worktrees. Use /worktree to create or switch and run parallel sessions, each on its own branch.
7
1
87
5,704
Ran FlappyBench on Qwen3.8-Omni-Flash, DeepSeek-V4.1-Flash, and Gemini 3.8 Flash with the same design prompt. 🔹 Qwen 3.8 Omni Flash: 9/10 · $0.013 · smooth gameplay and cheaper 🔹 DS V4.1 Flash: 9/10 · $0.0089 one-shot with the lowest cost 🔹 Gemini 3.8 Flash: 5/10 · $0.205 · 4-5 iterations but still hard-coded gameplay and costs 23x more
37
4
196
26,299
Replying to @jayleaton
GOAT and above plans have API. Use where ever you like. We are so confident in the harness we've built and the inference engine, we hope to win you over. check all the recent updates nitter.net/MrAhmadAwais/highlight…
11
2,005
Replying to @huh_go
it's much better actually!
4
1
29
2,662
Stop copy-pasting errors. IDE diagnostics are quite strong. Open Command Code. Select your code. Type fix this
4
3
112
12,245
Introducing Command Code Desktop App! ⌘ Available today in beta for Mac, Linux, & Windows. Building the most powerful agentic app, starting with code. Multiple agents, files, Git, terminal, browser previews, plans, and instant edits with /design. This is just the beginning.
81
48
710
215,320
Mods are more powerful than plugins. Command Code can build mods and modify any behavior from programmatic hooks to a completely different inference provider. We bet you haven’t experienced something as powerful before. Build a mod now and make Command Code do anything you want.
14
8
128
13,944
commandcode.ai/models we got free models for you :) but honestly, our GOAT plan has the best of all on the market
Introducing the GOAT plan 🐐 Command Code GOAT plan is the best low-cost coding plan on the market today. - $10/mo to get $70 credits. 7x what you pay - 30+ open weight and closed models $1 Go got you started. GOAT gets you shipping. Subscribe — be a GOAT.
1
132
GOAT plan is GOATed! 🐐
1
9
460
Ran FlappyBench on DeepSeek V4.1-Flash, GLM 5.3, and Kimi K3 with the same /design prompt. 🔹 DSV4.1-Flash: 9/10 · $0.0089 one-shot with the lowest cost 🔹 GLM-5.3: 9.5/10 · $0.0184 · nailed the UI but cost 2x more 🔹 Kimi K3: 8/10 · $0.0740 · hard-coded gameplay and 8x more
17
13
225
21,892
DeepSeek V4.1 Flash is GOAT 🐐
24
10
312
15,160
The /fork command creates a new session from the current one with the same context.
7
1
48
9,660
Connect to MCP servers to query data, use tools and automate workflows.
6
2
54
9,685
everyone using open models today

ALT Football Im Him GIF

13
4
186
12,666
Replying to @lickthedip
what's the issue? working fine.
2
168
Ran FlappyBench on Muse Spark 1.3, Gemini 3.8 Flash, and Grok-4.6 with the same /design prompt. 🔹 Muse Spark 1.3: 9/10 · $0.072. Nailed UI in 2 iterations with lowest cost 🔹 Grok-4.6: 9/10 · $0.095 · Matched Muse’s output, but cost 2x more 🔹 Gemini 3.8 Flash: 7/10 · $0.205 · Most expensive and not fully finished
30
15
252
38,903
Use Ctrl+G to open an external editor for your multi-line prompts.
7
3
83
8,714
Use /plan-review mode to comment on the plan, revise, and approve it when it’s ready.
13
1
64
8,770
Tested FlappyBench with /design on Hy4 Preview, Kimi K3, and GLM 5.3. 🔹 Hy4 Preview ($0.0480): Most distinct UI and gameplay 🔹 Kimi K3 ($0.0740): Hardest gameplay and the most expensive 🔹 GLM 5.3 ($0.0184): Smooth gameplay and the cheapest of the three
9
13
199
19,930
Hy4 preview looks like a super strong open model.
Tencent Hy4 Preview is live in Command Code 🐐 770B mode, 49B active, 1M context. `cmd update` available in GOAT and above plans. Our early internal benchmarks show strong perf yet cheaper per task compared to GLM 5.3 Flash. Y'all share how it goes. 🐐
21
17
355
24,938
Type /resume and get back to your previous conversation • $ cmd --resume opens a picker for earlier sessions • $ cmd --continue resumes your latest session
7
4
62
10,029
Paste images as your chat input. • Take a screenshot • Paste it into the chat • Describe what’s wrong • Watch Command Code fix the issue Let your screenshots be the context.
14
4
122
11,983
Binge code this weekend with free "Ox Alpha" 🐐
20
7
126
12,610
Replying to @MrAhmadAwais

ALT copycat GIF

4
399
Tested FlappyBench on Qwen 3.8 27B, DeepSeek V4 Pro 0813, and Gemini 3.7 Flash with /design command. All scored 9/10, it comes down to cost, quality, and iterations. 🔹 DeepSeek V4 Pro ($0.0036): Low cost, solid, but movement and spacing need work. 🔹 Qwen 3.8 27B ($0.0068): One-shot, playable, with minor UI finish. 🔹 Gemini 3.7 Flash ($0.0083): Needed 2–3 iterations to catch up. Qwen for the best one-shot. DSV4 for the best value.
12
3
102
12,632
make sure you're on the latest version. taste model default is super cheap LLM. /config to set it or change. This is how it works 1. Your selected LLM for taste, sends your entire session to the taste-1 model (LLM cost of reading context) 2. taste-1 learns things (healthy free limits, we don't charge) and it sends back the learnings to your LLM 3. Your selected LLM for taste receives the changes and adjusts the taste files locally (LLM cost of reading/writing taste files) /config has options to enable/disable taste and select different models for it.
1
78
Text-only models now have vision in Command Code. • A vision model reads the attached image • Your main model gets the context in the same thread • /config lets you choose the vision model Stop describing screenshots. Just attach them.
25
19
339
19,564
Launching `a11y` mode in `/design` Run `/design a11y` to catch: • Tap targets that are too small • Focus management problems • Broken keyboard navigation • ARIA and accessible name gaps • Missing or incorrect form labels and errors Run `cmd update`, then try `/design a11y`
5
6
126
8,984
Replying to @flowingfl0wer
🐐
1
3
2,428
Replying to @Nexus_526
This should help you draw one.
1
5
1,969
Writing tests is now the easiest part of your workflow. Pick a file, and Command Code will: - Analyze the code - Generate the right test cases - Run the test suite - Resolve failures - Re-run until all tests pass Try it on a file that hasn’t been tested yet.
15
5
115
20,580
Tested FlappyBench with GLM 5.3, Fable 5 and GPT-5.6 Sol 3 models. Same prompt with /design command. Scored on features, UX/UI, and cost. 🔹 Fable 5 → 9.5/10 · $0.420 🔹 GLM 5.3 → 9/10 · $0.018 🔹 GPT-5.6 Sol → 9/10 · $0.150 Results: → Fable 5 wins on quality, but costs 23x more than GLM 5.3 → GLM 5.3 gives better output than GPT 5.6 Sol at 8x cheaper → Optimizing for cost? Go for GLM 5.3. Otherwise, Fable 5 if quality matters
19
7
195
22,870
GLM-5.3 is the most capable open-weights model for coding.
9
8
264
17,805
Pull requests are easier now. Just ask to open a PR, It will: • Read your changes • Draft the description and test plan • Commit and push your branch • Open a well-documented PR Try it with our $10/mo GOAT plan.
4
1
64
5,171
Tested FlappyBench with Grok 4.6, GPT-5.6 Sol, and Opus 5. 3 frontier models. Same prompt with /design command. Scored on gameplay features, UX/UI, and cost. 🔹 Grok 4.6 → 9.5/10 · $0.095 🔹 GPT-5.6 Sol → 9/10 · $0.150 🔹 Opus 5 → 9/10 · $0.253 GPT-5.6 Sol and Opus 5 gave comparable output, but Opus costs 1.7x more. Grok 4.6 gave best output with lowest cost: smooth motion & clean UI.
12
4
92
11,352
Tested DeepSeek V4 Pro 0813, Kimi K3, and GLM 5.2 through FlappyBench. 3 models, same prompt with the /design command. Scored on gameplay features, UX/UI, and cost. 🔹 DeepSeek V4 Pro 0813 → 8/10 · $0.0005 🔹 Kimi K3 → 9.5/10 · $0.0740 🔹 GLM 5.2 → 9/10 · $0.0480 DeepSeek is 148x cheaper than Kimi and 96x cheaper than GLM, for 84% of the quality. But $0.0005 for a playable one-shot game is insane. Kimi K3 had the best output, with GLM 5.2 close behind.
29
34
561
43,607
Code reviews should help you ship faster, not block you. Run /review [PR number] and get your code reviews directly in the terminal • scans the diff • flags risky changes • finds missing tests • posts a ready-to-ship PR review Try now with our $10/mo GOAT plan.
5
4
83
6,278
GOAT plan won developers' 💜 switching from claude/openai.
Introducing the GOAT plan 🐐 Command Code GOAT plan is the best low-cost coding plan on the market today. - $10/mo to get $70 credits. 7x what you pay - 30+ open weight and closed models $1 Go got you started. GOAT gets you shipping. Subscribe — be a GOAT.
20
2
135
10,536
Use /plans to align your workflows - Browse, review, and annotate saved plans. - Select plans from current session vs all. - Search across all plans.
3
48
4,031
Tested Muse Spark 1.2, Kimi K3, and GLM 5.2 with our FlappyBench 3 models, same prompt with the /design command. Reviewed gameplay features, UX/UI, and cost. 🔹 Kimi K3 → 9.5/10 · $0.0740 🔹 GLM 5.2 → 9/10 · $0.0480 🔹 Muse Spark 1.2 → 8/10 · $0.0187 Muse Spark 1.2 delivered a completely different game style despite being super cheap. Kimi K3 and GLM 5.2 outputs are comparable and close in cost.
22
10
194
23,623
Plan mode removes the guesswork. When the change is too big for a quick fix: • Your code stays untouched • Plan it first, then review it • Apply exactly as planned No surprises, just clarity.
6
5
80
6,630
It's super subjective to what your design looks like. Theory behind it was discussed here nitter.net/MrAhmadAwais/status/20… and in the live stream here piped.video/watch?v=a1C2GHJC… We have several benchmarks open sourced github.com/CommandCodeAI/sla…
how did we fix the ai design slop problem in llms - DeepSeek/Kimi/Qwen or Claude/GPT?! i've been thinking about "why do all ai-generated designs look the same?" is it a model problem or a harness problem? context: we're fixing the llm design problem with `/design` for @CommandCodeAI - atm it has 16 modes, 24 reference documents, ~4,500+ lines of encoded design taste from some of the best designers in the world. it reads your codebase, identifies what's broken, and edits real files. no figma. no markdown mockups. the output stops looking like ai slop. i've been staring at ai-generated uis for a while now and noticed something that i think is underappreciated: llms can write css fluently but have essentially zero design taste. and the failure mode is not random, it's a very specific, very small distribution. let me explain. when you ask a model to build a landing page, it reaches into the mode of its training distribution. the mode of all landing pages on the internet is: centered hero, gradient text, glassmorphism card, three identical feature tiles, indigo accent, Inter font, bounce animation. this is the "average website." the llm is doing exactly what we trained it to do - predicting the most likely next token given "build a landing page." the most likely landing page is the average landing page. the average landing page is mediocre by definition. this is not a capability problem. the model knows oklch(). it knows prefers-reduced-motion. it knows golden ratio. it knows how to set a 65ch measure. it just doesn't know when to use these things, because "when" is taste, and taste is not well-represented as a statistical prior over internet css. so we thought what if we gave every llm a design taste with `/design`. here's what we found: 1/ the failure design dataset is surprisingly small. we talked to a bunch of designers with great design taste and asked them to label AI-generated UIs. what are the tells? turns out there are basically ~10 and they account for ~90% of the "this looks AI-generated" signal: - tech gradient (blue-violet glossy energy on everything) - generic tech hue (indigo because "software" not purple btw) - feature tile grid (icon + heading + sentence x N, all equal weight, nothing prioritized) - accent rail (colored stripe on card edge = decoration pretending to be organization) - unearned blur (glassmorphism without a depth system) - stat monument (oversized numbers filling space where a product story belongs) - icon topper (rounded-square icon above every heading as template filler) - bounce everywhere (elastic easing because the API has it, not because it's purposeful) - default type (whatever font the training distribution likes this year) - center stack (everything centered because no composition decision was made) this is super similar to what we see in other llm tool failures. tool calling errors? 4-16 types. fixing that made deepseek outperform opus 4.7, i wrote about that before! so i started researching maybe a dozen common patterns are design tells? 10. the failure distribution is narrow and we could repair ai design. this means it's a tractable and deterministic problem. `/design smell` hunts all these and scores severity on a /10 scale. 2/ the deeper problem is compositional, not cosmetic. the more interesting thing i found was that most of these tells are symptoms, not causes. the actual bug is that the model chooses layout before it chooses purpose. a dashboard and a landing page have completely different jobs. a dashboard is a Monitor surface - status, alerts, metrics, live data. a landing page is a Decide surface - proof, risk reduction, one clear action. these need fundamentally different spatial compositions. but the LLM reaches for the same centered-hero-plus-cards layout for both, because that's the mode of the training distribution. so we built work-pattern-first composition. before the agent touches any visual property, it must identify which of 7 patterns the surface serves: - Monitor: status boards, alerts, metrics, live priority - Operate: command bars, canvases, inspectors, direct manipulation - Compare: tables, matrices, split views, ranked lists - Configure: grouped settings, forms, previews, commit areas - Learn: article flow, walkthrough rhythm, progressive sections - Decide: focused pitch, proof, risk reduction, one dominant action - Explore: search, filters, maps, galleries, reversible discovery this is essentially chain-of-thought for design - force the model to reason about the *purpose* of the layout before generating the layout. i think there's a general lesson here. when an LLM is generating something compositional (code, UI, writing), forcing it to commit to a structural frame *before* generating tokens within that frame helps a lot. it's the same reason chain-of-thought helps with math. you're reducing the entropy of the generation by conditioning on a high-level plan. this single constraint eliminated more generic-looking UIs than any aesthetic rule we wrote. many phenomenal skills exist in the space, i bet they had the taste for great design but didn't know they were fixing the chain-of-thought problem instead of the style problem. i think that's why their skills are super loopy instead of being reliably good. 3/ validate-then-repair, again. my first version tried to audit and fix design simultaneously. this what many design skills do and fail. it's the "preprocess" approach and it fails for the same reason it failed in tool calling: you're encoding a prior about what's broken, and you get false positives that silently corrupt things. it would recolor something that needed relayout, or polish typography on a composition that was fundamentally wrong. the thing that worked: separate diagnostic from treatment, but make them a mandatory pair. audit modes (`checkup`, `smell`, `review`) produce structured reports. treatment modes (`redesign`, `relayout`, `recolor`, `typeset`, `motion`, `interaction`, `responsive`) consume those reports before making changes. the audit localizes the problem. the treatment mode only spends "repair budget" where the audit actually disagreed. same shape as tool calling repair. let the design system complain first, then fix only what it complained about. the validator does the localization work for you. cheap-then-careful, fast-path-then-evidence. i keep seeing this pattern everywhere. treatment modes don't just do report cleanup. they run their own full pass after absorbing the report. the report is more context, it's not a todo list. 4/ why oklch() color fn matters for llms personally, i always struggled a bit with the oklch() css fn but llms understand it super well. this one is fun. llms default to hsl because that's what's in the training data. HSL lightness is perceptually nonlinear - hsl(60, 100%, 50%) (yellow) and hsl(240, 100%, 50%) (blue) have the same L value but look completely different to a human eye. so when the model tries to build a "consistent" palette by keeping L constant, the result looks wrong in ways the model can't diagnose from the css alone. oklch has perceptually uniform lightness. this means the model can reason about color mathematically and have the result match perceptually. equal steps in the number space produce equal steps in the visual space. it's the right abstraction for an llm to work in, because it makes the optimization landscape smooth small changes in the parameters produce small changes in the output. hsl has cliffs and plateaus everywhere. i think this generalizes: when you're designing an interface for an llm to work through (whether it's a color space, a schema, or an api), choose representations where the distance in parameter space correlates with the distance in output space. the model optimizes over parameters. if the mapping from parameters to outputs is nonlinear and full of discontinuities, the model will struggle even if it "knows" the right answer in principle. we go further: the agent picks emotion before hue. calm vs urgency vs trust vs momentum. then it builds the palette in oklch with constraints - clamp chroma at lightness extremes, tint neutrals toward brand hue, 60-30-10 distribution. the agent can't default to indigo. the system requires a reason before a hue. no more indigo slop. and it's indigo, not purple. 5/ state coverage is the most honest metric. the most quantitative signal we found: count the number of interaction states per component. a human designer ships 7-9 states (idle, hover, active, focus, loading, empty, error, disabled, overflow). an AI agent ships 1-2 (idle, maybe hover). this is a clean, measurable proxy for design quality that requires zero subjective judgment. we just... count. does this button have a focus state? does this form handle empty? does this list handle overflow? the median AI-generated component has 1.5 states. the median human-designed component has 6+. roughly an order of magnitude. the gap is enormous and trivially detectable. 6/ a meta-observation beats an infinite loop. the biggest failure mode of AI design tools i found is you detect problem → attempt fix → the fix creates a new problem → attempt fix → loops forever. the agent re-runs the same mode hoping for a different result. it never converges. we solved this by reward model written in plain English. after each mode completes, the system recommends 2-3 specific next modes: redesign → checkup, review (validate the change) smell → finish, refine (fix what was found) recolor → responsive, motion (test viewports, add transitions) finish → typeset, recolor (fine-tune the details) the flow is: build → audit → refine → style → frontend → ship. the agent knows what to do next instead of re-running what it just did. this is a trivial intervention - a lookup table, basically but it eliminated the looping problem almost entirely which is super common in most design skills out there. 7/ truthful completion is the hardest constraint. the most insidious AI design behavior: claiming work that isn't visible. "added hover states" when no hover CSS was written. "improved spacing" when margins didn't change. "enhanced motion" when no keyframes exist. every mode has a "bar" - the minimum visible change required for the mode to count as complete. `typeset` must change body text, heading scale, labels, button text, form text, metadata, and responsive behavior. changing only the hero headline is not enough. `motion` must add animation to at least 8 transition moments. changing one easing value is not enough. the agent can't claim "motion improved" because it changed a duration from 200ms to 250ms. the user must be able to see new or clearly better behavior. this is surprisingly hard to enforce and the single most important quality constraint in the system. 8/ finally here's my meta-observation about design taste in general what we built is basically a reward model for design, implemented as structured english instead of a neural network. it defines what good looks like across 24 reference documents, gives the llm a rubric, and lets it self-evaluate. the 10 smells are negative rewards. the 9 states are a completeness check. the 7 work patterns are a structural prior. i'm sure this will grow. this is taste engineering in the limit. you're not writing instructions. you're writing a curriculum. the model already has the capability (it can write any CSS). what it lacks is the policy on when to use which capability, and what "good" looks like. i find it interesting that the policy is so compact. ~4,500 lines to encode "design taste" well enough that the output passes designer review. that suggests taste (at least for UI design) is lower-dimensional than it feels. it's not an infinite space of subjective preferences. it's a finite set of principles, applied consistently, with a small catalog of common violations. the model didn't change. we told it what good taste looks like. same lesson as tool calling: "capability gap" is usually "contract gap." the model knows how to write css. it just hasn't been told what good css looks like for *this specific surface*. i now believe that different llms have different baseline design capabilities, but it's your coding agent, the harness, that makes the difference in the end. the model didn't get better at design. the harness taught it what designers actually look for. i'm sharing my learnings so every harness out there can benefit not just our agent. try it yourself with what we built in Command Code. `npm i -g command-code && cmd` then `/design smell` on any project. read the md or html report. i care about design more than most engineers do, and seeing this work feels super good. a lot of what looks like a model capability gap is actually a contract gap. fix your harness. design slop is your "coding agent skill issue," not the model's.
1
1
51
We ran a test between the new Qwen3.8-Max, Opus 5 and GPT-5.6 Sol. 3 models. same prompt. one-shot with the /design command. Reviewed gameplay features, UX/UI and cost. 🔹 Qwen3.8-Max → 9/10 · $0.0248 🔹 GPT-5.6 Sol → 9/10 · $0.150 🔹 Opus 5 → 8.5/10 · $0.253 Qwen3.8-Max is approximately 4.2× cheaper than GPT-5.6, Sol, Opus 5 and has the same level of UI, UX, and gameplay.
31
39
591
180,179