Ok so i didn’t expect to get this much attention, But damn those are a lot of eyes. 👀
I’m out for holidays right now so sorry for the lack of updates.
Thank you Bee fumo! 🐝
I have some pretty telling evidence that Union Alpha is some sort of GLM model.
@Zai_org Can we please "pace the frontier"? /s
And while you're at it could you kindly fix the glaring security hole?
I wanted to work not dive down this rabbit hole, k? thx.
I don't know what others think but i find it quite amusing that i made a 12 CPU (CD8180) go toe to toe with a 20 core iGPU. (M4 Pro mac mini)
Wish i had proof for this but the proper comparison with GLM 5.2 gave me: 0.38PP and 0.32TG
Numbers from colibri:
Also please don't get me wrong, I'm not throwing shade at colibri or llama.cpp.
They're great projects, I just can't benchmark properly since i don't have disk space nor memory for either.
Sorry if that came off wrong. 😅
Humm maybe claiming to have a performance oriented inference engine followed by showing sub token performance on a pretty banger aarch64 system doesn't quite showcase the improvements..
Fine here are, Some direct 1:1 comparisons with llama.cpp on rk3588. 🧵
I also have an NPU backend that does indeed run here but unfortunately i can't publicly ship that code since my license is M.I.T and the NPU code is based on GPL.
Before i conclude i want to be crystal clear about some things.
1. I am not an AI/ML engineer, I am just a dude who likes to optimize things.
2. The codebase is primarily generated using @Zai_org's amazing models and i couldn't have done it without them.
Kappai is a performance/efficiency oriented LLM inference engine.
Most people neglect the CPU, I didn't, Results?
1.3x to 8x perf difference in comparison with llama.cpp.