Humm maybe claiming to have a performance oriented inference engine followed by showing sub token performance on a pretty banger aarch64 system doesn't quite showcase the improvements..
Fine here are, Some direct 1:1 comparisons with llama.cpp on rk3588. 馃У
I also have an NPU backend that does indeed run here but unfortunately i can't publicly ship that code since my license is M.I.T and the NPU code is based on GPL.