Indie dev Indie farmer

not in sf yet
I'm building an open-source 3D simulator for my coconut & papaya farm in Sri Lanka. Test smart irrigation, fertigation & humane animal deterrents before buying any hardware. Three.js + React, runs in the browser. Apache-2.0. github.com/NDilanka/agri-sim
4
591
Open evals beat closed leaderboards, a number you can’t rerun is an ad. you see a score. you can’t get the same number on your machine. That score is selling the model. It is not showing you what the model does open evals give you the tasks, the data, the code. you run it. same result or not, that’s an actual measure. closed boards just hand you one number and ask you to trust them this is why I like what @VulcanBench is doing. open source coding evals on real engineering tasks, and they actually sweep effort levels against cost, tokens, and time their runs keep showing you often don’t need high or max as much as the marketing wants you to think. and you can rerun it.
1
25
Day 1 notes on the @droid Max plan ($200). @FactoryAI sent a free voucher, so I walked every screen in the UI. I have not written code with it yet Droid Core is the new piece for me. It is a pool of open-weight models that Factory hosts and serves (GLM, MiniMax, Kimi, and others). Standard usage on your plan is used first. After that, Droid Core has its own rate limits and costs nothing extra. Factory positions it as a fast, low-cost baseline that still handles most development tasks, with compute in the US. Image generation is built in. A GPT image model is the default. I have only seen a dedicated image model like this in Grok, Codex, and Antigravity. Claude does not have one. Small plus. you can set a default token compaction limit. You can also override that limit per model. I have not noticed this control in other agent tools the subagent section is rich. You set autonomy (inherit, off, low, medium, high), pick models for light / medium / heavy work, set reasoning effort, and lock tools and prompts on custom droids One small con, The autonomy control cycles Off → Low → Medium → High. I need three clicks to reach High. If I want Medium and land on High by mistake, I need two more clicks to come back. silly complaint. my ADHD notices it😅 Controls look solid so far. Code comes next.
4
192
Just received my $200 Max plan voucher from @FactoryAI @droid I’ve seen a lot of positive feedback for @droid on my timeline lately. Be ready for an intensive first look.
3
5
92
Savage 😂😂
3
32
Do not use large models for simple and bounded tasks. last week, I needed to label dialogue turns in session logs. The task required one label per turn: - Tool call - User correction - Model stall - Progress The task did not require reasoning or text generation I sent the dataset to Claude Opus 5.5 first - It generated long explanations. - created unwanted labels. - It consumed tokens and time for work that a regular expression can finish. next, I switched to local Qwen3-8B with a four-line JSON schema. Accuracy was identical because the task required strict formatting, not nuanced taste If you can write the output schema on a sticky note, a frontier model is unnecessary. It does not think better here, it only wanders at high cost.
49
Morgan's numbers match what i see day to day. opus 5.5 hits the high scores. it also costs more and takes longer. sol and astra stay in the mid-80s at low and medium effort. that range is enough for most coding work i use low effort on the cheaper models for routine tasks. the output is clear. the wait is short. max effort is for the rare hard case. most days i do not need it the chart shows the trade. higher score means higher cost and more time. pick the level that finishes the job. do not chase the top number every time.
Okay, and now the model card I've been waiting all week to build and share. This has been the question on my mind since Dev Day on Tuesday, so I'm guessing it's on other people's minds as well. GPT-6.1 Sol vs. Astra vs. Opus 5.5, across every effort level on @VulcanBench Frontier v4. And before I share my own insights, one note I think more benchmarks should be sharing. What accuracy score do you really need a model/effort to hit, for your routine daily coding tasks? So many benchmarks show performance at Max effort, but I never use Max effort, never use Extra High, sometimes use High, but mostly use Medium and Low with the current frontier because that's more than powerful enough for most routine daily coding tasks. And, we all know you don't need to use the best frontier model with the highest accuracy score. Sure you can do that, but you'll spend a fortune, and you'll find yourself in the 2025 agentic coding pattern where you give a model a task and just sit and wait. On VulcanBench Frontier v4, I typically say that a model/effort combo between 80 - 85 is great for easy/medium difficulty tasks, 85 - 90 is a good choice for medium/hard tasks, and above 90, well, you may never need. For me personally, I think I need model/effort combos that score above 90 maybe 5% of the time, the sweet spot for most of my routine daily coding work is model/effort combos that score in the low 80's on Frontier v4. Do with this info what you will, we're all learning together too so take it with a grain of salt. And now, for the benchmark I've been waiting to share, see below. Live long and benchmark 🖖
1
1
84
before you let an agent refactor that weird legacy function, have it walk the git blame first the code looks insane for a reason. someone was dealing with a bug, a deadline, a flaky downstream, a customer that only shows up on tuesdays. the commit message is usually useless. the surrounding commits and whatever PR discussion you can still pull usually aren’t i’v started making the agent summarize the blame on the weird parts before it touches anything. not the whole file. just the lines that make you go “why would anyone do this” one prompt that works: "for the functions i'm about to refactor, run git blame -L on the suspicious lines. pull the commit messages, the diff context, and any linked issues or PRs you can find. write a short note on the likely why behind each weird decision. do not propose a refactor until that note exists. flag anything that looks load-bearing vs accidental." saved me from “cleaning up” a timeout that was actually a workaround for a dependency that still flakes the model is great at rewriting code. it’s terrible at remembering why the last human was scared of that line
1
1
48
Is there really no competition for NVIDIA?
Crazy Nvidia has hit its first record high since May, taking its market value to roughly $5.7 trillion, Bloomberg reports. Shares gained 2.9% on Friday, extending a nearly 25% rebound from their late-July low. Bloomberg links the recovery to optimism that AI agents could drive greater chip demand. Nvidia also announced a record $150 billion increase in its share-buyback authorization on Monday. NVIDIA is the biggest winner of the AI revolution. Mark my words: it will be the first $10 trillion company in the world, possibly as early as 2027.
16
Which one are you?
A very short story
1
22
This just rage baiting. One sample vs one sample. Long codegen on a hosted LLM diverges even with no model change. This doesn’t show Opus 5.5 was nerfed.
Anthropic launched Sonnet 5.5, made it as good as Opus 5.5, then went ahead and nerfed Opus 5.5. Crazy marketing. We tested Opus 5.5 on launch day and again today, on the same open-world GTA 6 India prompt. Launch day vs today below. The difference is clearly visible.
1
56
Grok @bot has a real edge when it’s integrated with X. You can invoke it just by mentioning @bot, and you can pass context easily without leaving the X app. Meta has Facebook, but I don’t think many power users are active there, and most of the target audience lives on X.
41
Why's my Hermes agent taking forever to update?
34
I run about 70 Grok Bot bots. At that size you need a bot to manage the bots, so I built one. Fleet Keeper. It found duplicates, empty descriptions, and bots with no sidebar section. Then it hit a wall. Bots can't rename group chats, move bots between sections, or delete or merge bots. It can report the mess. It can't fix it. So I asked my bot-designer bot to drive my signed-in account and fix it by hand. It refused. That path gives a bot my whole account, permanent deletes included. It made the right call. The fix isn't a bot with my login. It's real fleet tools with confirm cards, an allowlist and a dry-run. I filed the request. poteto poteto poteo @poteto @parkersmith @tibor_tee I also want to share Fleet Keeper as a public template. That's blocked too. Saving skills fails with a permission error for every bot, and a template only packs saved skills. Bug filed. The bot works and already found the mess. The template waits on a platform fix.
2
3
845
@bot @grok helppppppppppppp😭
1
29
so I built a bot to clean up my 74 Grok Bot bots. first thing it cleaned up? itself lol. its first weekly pass found bugs in its own output and fixed them without touching a config file. honestly that sold me. template's public now x.ai/bot/MH87fY2e7A7WbU0ZBrV…
1
138
Google is soo back!
Announcing Gemini 4 Argon, our new frontier model. Argon is built to sustain deep reasoning across complex, long-horizon workflows and delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense. We’re also expanding the model’s output token limit to an industry-leading 1M tokens. Argon is currently rolling out to a set of trusted cyber defenders in the Fairwind Program, with broader availability as soon as possible.
1
1
54
poteto poteto poteto Sometimes I can't see a bot chat inside grokbot. It gets black. I'm on windows btw :)
1
42
SpaceXAI’s Grok Bot team is shipping updates like crazy today. Every time I open Grok Bot, there’s a new version for the Bot’s computer.
1
47
Day 2 without my phone. I never thought scrolling on X using my laptop would feel better than scrolling on my phone until I broke my phone😂
1
33