I've been testing out
@grok build.
When it works, it can be astoundingly good and much better focused on complex designs (i.e. programming language specification) than codex 5.6 Sol.
However, occasionally, the quality flakes out and it totally devolves. I've also seen the dictation quality drop horribly off.
Not yet something I would pay the $100/month for. I'm currently using it via the lower paid tiers.
The issue is
@ChatGPT codex has a tendency to over-engineer specifications and implementations. Which honestly, is in their interest because that eats up more tokens... You have to ruthlessly review and keep the specs lean. Which
@grok seems better at by default.
One approach is to pit them against one another.
Basically, codex is very very good at generally solved areas. When you send it into research or experimental territory, it can start to over engineer without very careful scrutiny.
nitter.net/dharmatrade/status/209…
In prototyping ALOE (Scheme + Smalltalk + Types),
codex really tended to over-engineer the design.
Grok did great on the design specs.
Then I handed each checkpoint spec to codex for implementation.
That the process I used for 0.1 and 0.2.