🔥🔥Testing 7B #LLM on #RK3576+#RK1828 We put Qwen2.5-7B (W4A16) to the test on the reComputer RK3576 Dev Kit with RK1828 AI Accelerator: ⚡ Prefill: 121.12 tok/s 🚀 Generation: 60.42 tok/s ⏱️ TTFT: 309.64 ms But this isn’t just another benchmark. We gave it a real coding task: generate a #Python program to read serial sensor data, record timestamps, and save the results to CSV. The entire task runs locally on the device. No cloud round trip. No remote inference. Local AI. Real workloads. 7B intelligence at the edge. 🛠️ Start with reComputer RK3576 Dev Kit: seeedstudio.com/reComputer-R… 🛒 reComputer RK3576 is IN STOCK: seeedstudio.com/reComputer-R… #SeeedStudio #reComputerRK #Rockchip #reComputer #RK3576 #RK1820 #RK1828 #RK182X #VisionAI #EdgeAI #AIoT #ComputerVision #reComputerAILab #AIDeployment #TheEasierPowerfulAIComputer #TheAIHardwarePartner

Sep 15, 2026 · 1:00 PM UTC

8
6
76
4,605
Sort replies: Relevant Recent Liked
Replying to @seeedstudio
The 7B model hits ~6 GB VRAM on the RK1828, so you’ll need to drop to 4-bit quantization to stay under the 8 GB limit, which adds ~12 % latency to generation.
1
1
52
The RK1828 has 5 GB of on-chip DRAM, not 8 GB. The Qwen2.5-7B model we tested is already running in W4A16, with model weights around 4.1 GB, so it can run on a single RK1828. As for the claim that “4-bit adds around 12% latency,” we don’t currently have controlled comparison data to support a fixed 12% figure. On RK1828, we’d need a strict W16 vs. W4 comparison with the same model and setup before drawing that conclusion.
1
2
74
Replying to @seeedstudio
Getting a 7B model running is pretty impressive.
2
14
Replying to @seeedstudio
Better than needing two USB C ports and a battery to maintain continuous power or having to add a I2C FIxed HUSB Usb C. Faster than a entry level nano running gemma4 e2b at 7.2 GB vram too.
15
Replying to @seeedstudio
How about qwen3.5 0.8b baked into a silicon chip ?
47
Replying to @seeedstudio
No aguardo de alfum kit desses com a rk3588
33