CTO of Homezee.ai. AI and Tech Focused. I build apps, solutions, and share my thoughts here.

South Florida
Will we get a new Gemma 31B from this? Please!?
Introducing Gemini 4 Argon – our new frontier model. It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense – rolling out today to a set of trusted testers through our Fairwind Program.
31
What's the best way to test your own agentic software? By running it in production of course. (Slight sarcasm here) BossMod agents are going to start handling real problems within my business.
21
Almost feels like I'm talking and interacting with a real person. The model is not ChatGPT or Claude. No, it's @Alibaba_Qwen Qwen3.8-27B Model.
1
58
I am now setting my Agent Employees up with their own email access. The first email I received from one of my agents.
1
34
I have Qwen 27B browsing the internet with my agent harness. ...Then my agent unintentionally started to DDOS a website... oops, hope they didn't ban my IP. Better add rate limiting to the harness!
1
40
Another exciting release in the LLM space. They took an existing open weights model and then improved it. Looking forward to trying it soon locally.
Introducing Naive-N0.5-Flash: Building Frontier AI with AI 🔹 309B MoE, 15.5B active: top-tier in coding, leading in AI R&D. 🔹 Native 1M context, no full-attention layers (hybrid SWA + DSA). 🔹 Inference runtime built by AI: up to 2,000 tok/s in Ultrafast mode. Weights are open today under MIT license. 🔗 Tech blog: naive.ai/en/research/ 🤗 Hugging Face: huggingface.co/NaiveAI/Naive… 💻 GitHub: github.com/NaiveAI-Labs/Naiv… 🌐 naive.ai/en/
1
4
75
Save some memory for the rest of us!
Replying to @minchoi
Colossus 1 is 150k H100, 50k H200 and 30k GB200. Colossus 2 is 110k GB200 and 440k GB300. Another 220k GB300 will be fully operational next week and another 220k in November. If we get lucky, yet another 220k GB300 by late December.
2
36
Historically speaking, what happens to a society when it doesn't adopt evolving technology?
Jesus christ
34
First real pressure test of my Agentic Harness running entirely offline (QWEN-3.8-27B). I have 6 agents running in a group chat building a Diablo Clone Proof of Concept.
1
35
I am asking my GLM 5.3-Flash Agents in my harness to write prompts for me to test QWEN-3.8-27B in my AI harness. I am starting to see my agentic harness come to life.
38
My AI Harness is starting to come together. The team is working autonomously- and it's kind of frightening lol!
2
36
GLM-5.3-Flash is so far my absolute favorite for local inference. Thank you @Zai_org for such a capable open weights model. I'll support you with another annual subscription renewal that I hardly use 😂
1
41
I need to get my hands on a RTX Pro 6000...
1
40
I asked Claude to make 9 edits to a front end. It's burned 300k tokens and is barely on edit 5.
26
GLM-5.3-Flash on CPU (using ~30gb of Vram) with optimizations getting ~14 tok/sec decode regardless of context length (0-96k actual context tested.)
38
Working on a llama.cpp fork for GLM-5.3-Flash that shows 55% higher throughput and 36% lower generation latency on CPU/MOE versus stock llama.cpp. Will need to piggy back off @UnslothAI PR 27752 to get it added I think.
38
Running GLM-5.3-Flash at ~15 tok/sec with 90k+ context on CPU and 40GB of Vram. It did require some code edits to llama cpp to make this happen. GLM 5.3 Flash Quant is my own ~4.54 bpw.
33
Running GLM-5.3-Flash locally. Amazing to have such an intelligent LLM running privately, and offline getting ~500 tok/sec eval and ~17 tok/sec decode speed running with ~40GB in VRAM and the rest in system ram.
1
34
Cannot wait to test this on my local server.
Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: z.ai/blog/glm-5.3-flash Available now across all official platforms: Weights: huggingface.co/zai-org/GLM-5… API: docs.z.ai/guides/llm/glm-5.3… Coding Plan: z.ai/subscribe ZCode: zcode.z.ai/en Chat: chat.z.ai AutoClaw: autoclaw.z.ai
52
GLM Flash and QWEN Flash all being released on the same day. It feels like Christmas.
45