Software Engineer Mac, iOS, AI. Usually building something open source

what is wrong with Astra? Dumb as a rock, the task is not that hard..
1
1
40
My Claude subscription normally resets shortly and i couldn't use all of it. The amount of usage is really incredible, ands for the past 2 horus i was using Max hoping to use up more.
1
2
92
Can anyone please test the following, GPT-6 Luna appears not to do any reasoning while on low or medium on Codex, but it does on the API. Can anyone measure the reasoning tokens on Codex with this low juice?
3
114
Opus 5.5 attention to detail is top notch, this recommendation was from something we had worked hours ago, it had split my mind completely.
1
4
90
I just made a plug in for jev-use for Devin and this is great , now Devin has the only thing it was missing a really fast local computer use tool. This can be even faster than codex with astra, although it can be model dependent of course. Great work!
1/ Fast Computer Use is now solved with @typesafeai Jev + Cua Driver. Available in development preview for macOS, Windows, and Linux. We call it jev-use. Draft #3943: github.com/trycua/cua
2
288
My house is mostly automated, stuff like lights, cameras, curtais and so on. I wanted to install a presence sensor on every room but these mmWave sensors for Homekit are very expensive. Around 100 each. So I went with the DYI way. Bought a 5$ ESP32 dev board for each room, and wrote a firmware to use my WiFi signal to detect the presence of people in each room. It works so well, no delays after a tunning. Leave the suggestion for someone who wants to save a lot of money. Right now.. training a local model to autonomously pilot my small Parrot drone activated by the security cameras to go check out any occurrence…. a watch drone . This drone thing is just a wacky fun project, but leave here the suggestion for the cheap presence sensors.
2
1
6
377
My wife is abandoning the Mac after many years. This against my advice...but she wants a touch screen laptop. So she got it and immediately we ran into a problem. We used to send files and copy and paste stuff from one mac to another no problem with airdrop. Was looking for a tool to share clipboard contents from Mac to windows but did not find, only from Android. So i quickly made a web one. Don't know if anyone is in need of such a thing or if there are already more mature solutions. It only works on the local network for, and no certificate for now. For now it works well, if this proves to be useful i will go deeper. github.com/Joaov41/mac2windo…
1
1
158
Funny story that combines the “This is AGI” , with fear mongers of “AI will kills us all” and the opposite side that think AI systems are the answer to everything. Funny store at my company today.. we were wired to replace a system for a parking lot, you know, those system that count how many spaces are available and update the screen, with the added difficulty of the system having to monitor cars that entered the garage while there were free space but meanwhile those got occupied and close gates to divert drivers to another level. Long story short, some other company installed a system with ultrasonic sensors, radar sensors for the entrance and exit, fine up to this point, but the decisions instead of being done by a algorithm in code, usually it is used PLC’s with ladder logic and a dashboard written in C, python or Typescript , they decided to give control of the system no an AI model instead of being done directly by low level code. The result was disastrous in the first 24 hours. The model made wrong decisions from corrected data, hallucinated states and actions and the costs was of course huge with thousands of API calls. It had to be shutdown because related caos in that parking lot. I don’t know yet what model was used but we were told that was a frontier model ( i will know once i get the logs) The moral of the story is, this is real world task that humans have solved using ordinary algorithms and code. There shouldn’t be a rush to replace every task with AI models just because it is trendy for some, and at the same time shows that none of these models are AGI yet, another example of a failed real world engineering task ( it is not the first example of a fail i tell you about here). Let’s not even talk about traditional code for these systems, a single human, like a keeper, could have done this manually just fine, fed by the information from the sensors. Let’s keep ourselves grounded on reality and away from the hype but also away from the AI will kill us all.
1
3
106
What the hell my codex usage just went from 51% to 0, when no task was even running? Anyone else? Barely used it today as i was using Devin How can I appeal this?
2
3
447
There may be a Astra orchestration bug behind the huge token burn some people are seeing. The parent agent appears to wake itself roughly every 30 seconds just to check whether sub-agents have finished. Those empty polling cycles can re-enter the full parent context even when there’s no new worker state. So you end up burning millions of input tokens on repeated are we done yet checks. Can anyone to whom OAI listen to ask them if they looked at this?
59
Lets put aside Blender demos for a while. Here is a real nasty engineering, algorithmic coding test, putting Astra against Fable 5.1 Problem: A network of 100,000 cities connected by possible roads, each with a different construction cost. The task is to always keep the cheapest possible set of roads that still connects every city that can be connected. Then we make the network dynamic Roads can appear. Roads can disappear. Old road IDs can be reused with completely different costs. Up to 1 million changes. After every single change, the program has to immediately return: the minimum total cost of the network how many disconnected groups remain And it can’t just recompute the whole thing from scratch every time. It is forbidden. That’s a fully dynamic minimum spanning forest problem. I won't bother you with the math but I will tell you that Astra solved it in 4 hours including test iterations. Fable 1 got surprisingly close, but failed on a nasty state-management edge case. It cached an old version of edge #3 and only checked whether that ID was active again. Same ID. Different edge lifetime. Old cost resurrected from stale state. What is interesting is that this wasn’t a syntax error or a crash. Most of the difficult dynamic graph machinery was working. One incorrect assumption buried deep in the replacement-edge logic was enough to corrupt the final answer. Astra handled the same adversarial case correctly. This was surprising because I made sure to keep the prompt as simple as possible (in this math context of course) and Fable is much better at gathering the intent then Astra so I though the results would be the other way around. But it is not a fault if intelligence, Fable was victim of its caching, which demonstrates Astra compaction is more advanced. So bottom line, for a real hard math/coding problem, both models are very similar in capability but Astra seams to retain more of the small details of the context when the task is huge. The dashboard U created by both was of course better on Fable 5.1 As for costs, Astra used 790M tokens and Fable 5.1 only around 540M. On last comment, Fable was more pleasant to work with, it communicated better than Astra, almost like having a human partner solving the problem and also asked fewer questions. Astra just took and run with with with very few, apart from the essential questions To give you and idea this is usually a week's worth of work for a human. Can't share any screenshots or more details as this is part of a professional project I am involved in, but his is the kind of coding experiment I find far more revealing than just testing UI, 3D, computer use and so on. If anyone sees this 😀 hope this is a slight different insight at testing models than usually seen here.
3
127
Astra Ultra just gave up on me.... yes the task is a hard real engineering task.. at least it is honest.
59
The GPU are on fire..
3
117
What kind of sorcery is this, and I am not saying this because I got it, because I will be posting my honest opinion, but right out of the box, while everyone is complaining about the limits on Fable 5.1 using Anthropic tools directly, I have been using it on Devin for 3 hours now with excellent usage consumption. I did create a custom profile to choose the subagent model I want and not the default, I also created several reviewers agents, while Fable 5.1 is only planning and orchestrating, but then again, all this should be adding up very fast with so many agents. So first impressions are excellent, I was already familiar with the UI as I used to be a windsurf subscriber 2 years ago. More details on the review to come.
thank you so much, I'll share my results.
2
1
5
871
oh and since Google it self is comparing to Opus 5, it is much much better, tried it with 5 tasks Opus 5 had me pulling my hair out.
Replying to @Da7_Tech
Very good at front end development. Extremely fast and intelligent. As I suspected from the agentic bench, it is its weak spot, although not as bad as the agentic bench score might suggest, but agent coordination it is not its strong suit. But for one shot task it is an excellent, cheap and so fast model.
1
281
Fable 5.1 should be a specialist upgrade, not a default switch. Move only difficult, long-running agent jobs first, then measure accepted output per dollar.
120
Fable 5.1 one day, Gemini 3.8 Flash the next, Astra reportedly “soon.” There is no time to test every release properly. What is the minimum evidence a model needs before it gets added to your eval suite?
156
Anthropic says Claude Code users should see about 60% fewer cyber interventions with Fable 5.1. Good. The next number should be false-positive rate by workflow. Safety that blocks benign work is still a broken product path.
1
1
154
Johnny retweeted
best local AI subreddit?
50% r/LocalLLM
50% r/LocalLlama
146 votes • Final results
5
2
9
11,382
There’s a subtle but important difference between building for yourself and building for an audience. When you build for yourself, every decision starts with friction you actually felt. You’re not guessing what the user wants,you are the user. When you build for an audience, it’s easy to start copying what successful products look like instead of solving a problem you deeply understand. That’s how apps become polished but hollow: the feature list grows while the original reason the product needed to exist slowly disappears. The best products often come from someone with a very specific vision who builds the thing the way they believe it should work, and then finds people who want to follow that vision. Not everyone has that. And you can’t manufacture it by studying competitors.
113