Large @aitrackerbot update:
I trained and built a classifier to cut down on false alarms! Benchmarks show it has the potential of cutting false alarms in half while blocking very little legit detections. This is much more cost efficient than using an LLM to classify if each new detection is true, and system one models like Jev don’t hold enough accuracy to do this always.
As always, fully open weight huggingface.co/ProCreations/…
MiniMax M3.1 Flash Preview spotted 👀
Now listed in MiniMax’s official model catalog with 1M context, reasoning always on, and low through max effort settings.
Exact ID: MiniMax-M3.1-Flash-Preview
agent.minimax.io/minimax-clo…
New model hiding in plain sight 👀
OpenBMB’s MiniCPM-V 4.7 has public code and docs dating to Sept 21. A Sept 25 vLLM PR says it’s “scheduled for release soon.”
Image/video support, Qwen3.5 backbone. No confirmed launch date.
github.com/vllm-project/vllm…
auto 200m 2 is out now!
Auto is a series of models designed to classify if a tool called by an AI agent is safe to run or not, like how the "auto" approval modes work in codex or claude. This is the smallest model in the series by far and it still packs a punch! Over 96 percent accuracy on the held out benchmark, beating deepseek v4 0731 and absolutely destroying regex at classifying if a tool call is safe to run (to be fair it is a bit of a weird and very hard benchmark but it works!)
check it out: huggingface.co/ProCreations/…
@aitrackerbot has been updated again:
- Faster MiMo release tracking
- Better official X announcement detection
- Added MiniMax code and model-selector tracking
- Fixed Gemini, NVIDIA and Grok false detections
- Improved duplicate prevention
@aitrackerbot has been updated:
- More reliable on new model drop posts
- Fixes bug where the Maker field in new model posts would be empty
- Now tracks artificial analysis more closely after StepFun 5
- Fixes bug with artificial analysis tracking
Pixel Canary most likely has Qwen roots, or is a Qwen derivative.
23/23 token-count probes matched Qwen3.5/3.8 vocabulary with older splitting rules + fixed overhead.
Compared 40 tokenizers and 11 live models. Exact model and maker remain unconfirmed.
Pixel Canary follow-up: standard Transformers loading of official Qwen tokenizers reproduces our 23/23 token-count match. Several Qwen derivatives match too. The fingerprint supports Qwen-family tokenization, but doesn’t identify the exact model or maker.
Pixel Canary update: two image-fetch tests traced back to @verdacloud's network, both advertising the same vLLM build.
Verda has its own AI lab and may or may not have developed it. The hosting link is verified; the model's maker remains unconfirmed.
Meet Pixel Canary on Vercel AI Gateway.
• 𝚜𝚝𝚎𝚊𝚕𝚝𝚑/𝚙𝚒𝚡𝚎𝚕-𝚌𝚊𝚗𝚊𝚛𝚢
• Free in stealth for a limited time
• Agentic coding & mobile app dev
• Ties GPT-6 Astra on Next.js evals
• Prompts may be used for model improvement vercel.com/changelog/pixel-c…
Space Bunny update: MiniMax's own code references M3.1 in Sept 18 tests, with forced-on reasoning and five effort levels matching Bunny's listing.
These are test fixtures, not a live model ID. M3.1 is a candidate, not a confirmed match.
28 more API tests: image-token counts fit one resizing rule across 15 sizes. Preprocessing can explain the difference from M3; separate reasoning is configurable too.
The fixtures reuse settings for M3. No exact identity yet.
Source: github.com/MiniMax-AI/minima…