qwen3.5 0.8b Q8 #GDS file This is how an LLM brain looks like if it had its own silicon chip. Comment if you have the means to produce a silicon chip (yeah FABs are NOT cheap 馃ゲ) #llm #qwen
9
Designing an LLM silicon chip with Verilog made me understand the level of engineering required to build the modern GPUs! - You are limited by physics! unless you go 3D (Stacking ROM vertically) making this MUCH more complex!
3
5
Based on my tests, #Taalas had to do some heavy quantization on the llama model (1~2 bit) to b able to fit the ROM and the logic into ~900mm^2 chip , that also explains the low ~50b transistor count! the ROM alone would take 20x that number at 16bit
3
- You can get 1000x speed by using the model weights AS the logic, but that is 5x~6x the number of transistors - the larger the lithography feature size, the slower it becomes (damn you physics)
2
- Even small models (~1B params) will end up needing between 6~9 cm square of silicon which is reaching the edge of what is possible with one mask (So either go wafer-scale or multi-layer-masking) - Running multiplications in parallel still require careful placement for speed
4
Been testing GLM5.3-flash and Qwen3.8-flash-next locally on 2 DGX Sparks on multiple SoA projects, each project has at least 20 services, 20 DBs, and combination of frontends (VueJS/Android/iOS). GLM5.3 did far worse than qwen3.8-flash-next. never had to correct qwen!
1
14
Partial market scan: - Above bollinger 15m & 1H & H1EMA200 $PAXG - Crossing EMA200 on H1 $SHIB - Approaching Standard OrderBlock on H1 $SHIB
1
38