🇺🇸🇮🇱 Icy fairy goofing around with embedded systems for fun!

Israel
Ok so i didn’t expect to get this much attention, But damn those are a lot of eyes. 👀 I’m out for holidays right now so sorry for the lack of updates. Thank you Bee fumo! 🐝
2
29
I have some pretty telling evidence that Union Alpha is some sort of GLM model. @Zai_org Can we please "pace the frontier"? /s And while you're at it could you kindly fix the glaring security hole? I wanted to work not dive down this rabbit hole, k? thx.
1
97
System64 retweeted
4
84
779
10,304
I don't know what others think but i find it quite amusing that i made a 12 CPU (CD8180) go toe to toe with a 20 core iGPU. (M4 Pro mac mini) Wish i had proof for this but the proper comparison with GLM 5.2 gave me: 0.38PP and 0.32TG Numbers from colibri:
1
136
Also please don't get me wrong, I'm not throwing shade at colibri or llama.cpp. They're great projects, I just can't benchmark properly since i don't have disk space nor memory for either. Sorry if that came off wrong. 😅
66
Humm maybe claiming to have a performance oriented inference engine followed by showing sub token performance on a pretty banger aarch64 system doesn't quite showcase the improvements.. Fine here are, Some direct 1:1 comparisons with llama.cpp on rk3588. 🧵
1
1
119
I also have an NPU backend that does indeed run here but unfortunately i can't publicly ship that code since my license is M.I.T and the NPU code is based on GPL.
1
30
Note: Just like my last post i had to re-create this post too, Again sorry!
25
Before i conclude i want to be crystal clear about some things. 1. I am not an AI/ML engineer, I am just a dude who likes to optimize things. 2. The codebase is primarily generated using @Zai_org's amazing models and i couldn't have done it without them.
1
1
126
PS: Had to re-upload the images and re-create the post because twitter's drafts made them super low resolution, Sorry about that! 😭
1
111
@orangepixunlong Orange pi 5 plus 16GB (DDR4) running gemma 4 E2B in kappai vs llama.cpp:
1
30
Or better yet a comparison with llama 3.2 3B IQ3_S:
15
Kappai is a performance/efficiency oriented LLM inference engine. Most people neglect the CPU, I didn't, Results? 1.3x to 8x perf difference in comparison with llama.cpp.
26