Humm maybe claiming to have a performance oriented inference engine followed by showing sub token performance on a pretty banger aarch64 system doesn't quite showcase the improvements.. Fine here are, Some direct 1:1 comparisons with llama.cpp on rk3588. 馃У
1
1
124
@orangepixunlong Orange pi 5 plus 16GB (DDR4) running gemma 4 E2B in kappai vs llama.cpp:

Aug 30, 2026 路 6:24 AM UTC

1
97
Or better yet a comparison with llama 3.2 3B IQ3_S:
1
26
I also have an NPU backend that does indeed run here but unfortunately i can't publicly ship that code since my license is M.I.T and the NPU code is based on GPL.
1
31
Note: Just like my last post i had to re-create this post too, Again sorry!
26
Sort replies: Relevant Recent Liked