Tokyo AI company. We measure local LLMs, GPU inference, Japan's LLM rankings and AI on Raspberry Pi, then post the numbers. WireCanal, MotionVox, Bestllam.

Tokyo, Japan
Who we are, in numbers. Qualiteg is a Tokyo AI company founded in 2023. This account posts what we measure: local LLMs, GPU inference, Japan's LLM rankings, and AI on a Raspberry Pi. A short thread on what to expect.
1
1
302
vLLM refused the model: minimum capability 89, current 86. - The card was an NVIDIA RTX A6000, capability 86 - FP8 arithmetic arrives with Hopper, so 89 is the floor - Blackwell is capability 100 with native FP4 Quantization is a hardware question. Link in reply. #vLLM #AI
1
10
Gartner named the AI that watches AI: guardian agents. - Three types, reviewer, monitor and protector - Gartner sees 10 to 15% of the agentic AI market by 2030 - It also sees 70% of AI apps multi-agent by 2028 No product yet, risk now. Link in reply. #AISecurity #AI
2
22
Cross-entropy, part 5: one sample becomes a batch. - Score all four samples, then add the errors up - t and y take a second subscript i for the sample - Divide the sum by N so batch sizes stay comparable That is categorical cross-entropy. Link in reply. #DeepLearning #AI
1
15
Ollama 0.40.0-rc0 makes MLX the default on Apple Silicon. - Supported architectures now pick MLX on their own - MLX first landed in 0.19, back in March - No speedups and no documented off switch in the notes - The supported list still grows during the RC Link in reply. #LocalLLM
2
55
GPU PC build, part 13: the wiring, and the machine is done. - The 24-pin ATX bundle goes in until the latch clicks - CPU power is a 6-pin plus 2-pin, one end at each side - This card wants three PCIe feeds, so a 3-in-1 splitter Panels back on. Link in reply. #GPU #AI
1
39
LLM inference, part 5: how many users fit on one GPU. - Llama 8B at 16-bit takes 16 GB before any user - 2,000 tokens of context costs about 1 GB per request - So 24 GB serves 6, 48 GB serves 28, 80 GB serves 56 Those are ceilings, so load test. Link in reply. #LLM #AI
1
12
The newest ESP32 runs real Linux. Note it is the S31, not the P4. - Dual-core RISC-V with a real MMU, so virtual memory works - Espressif ships a Linux BSP preview on a 6.18 kernel fork - Buildroot 2025.02, 16MB PSRAM, 1G Ethernet Link in reply. #ESP32 #Linux
1
1
38
Espressif's BSP repo. ESP32-S31 only, and the branch is marked a developer preview that is not recommended for production yet: github.com/espressif/esp-lin…
1
21
Is a bigger tokenizer vocabulary always better? No. - Mistral's Tekken has about 130,000 tokens on tiktoken BPE - Reported 1.5 to 2 times more efficient on Japanese - Yet on one text Tekken used 74 tokens, a rinna model 42 Fit it to your language. Link in reply. #NLP #AI
1
32
Why did enterprise security get so complex? - One proxy once held the only way out, and it knew who - Cloud and mobile broke that path in the mid 2010s - SWG, CASB, ZTNA and SASE all try to win the view back Know who sent what to the AI. Link in reply. #CyberSecurity #AI
1
18
"Cheaper and smarter" is the pitch. What the docs and Artificial Analysis say: Claude Opus 5.5 is 20% below Opus 5, cuts cache reads by 60%, and is #1 on the Artificial Analysis Intelligence Index. The catch is that turning thinking off now returns 400. #Claude #Anthropic
1
22