if you're getting into local ai, the first mistake is taking the easy route. ollama and lm studio are good apps, but they hide exactly the parts you're supposed to learn, so you end up running models for months without knowing why they're fast, slow or out of memory.
take the llama.cpp route instead. build it yourself, download one gguf, start llama-server and read the log. every line in it tells you something about your own card.
the flags that taught me everything:
> 1. -c, the context window you ask for, and what every extra token costs in vram
> 2. -ngl, how many layers live on the gpu, and why speed falls off a cliff when they don't all fit
> 3. -ctk / -ctv, kv cache quantization, the reason my bonsai 2 build fits a 262k window in 12gb
> 4. -fa, flash attention, which a quantized kv cache needs
> 5. -np, parallel slots, one slot for one user saves vram, the default split cost me 454 mib
> 6. --spec-type draft-mtp, speculative decoding, the flag that took a 3060 from 40 to 50 tok/s
> 7. --jinja, the chat template, the difference between tool calls that work and tool calls that don't
> 8. --reasoning-effort, leave a thinking model on xhigh and it can spend your whole token budget and hand you an empty answer
change one flag at a time and measure tok/s every time. once you understand how real serving works, every wrapper and agent harness on top of it makes sense, because you know what's underneath.
and when a wrapper does offer these flags, you'll finally know what they do, and when it doesn't, you'll know exactly what it's hiding from you.
start here:
github.com/ggml-org/llama.cp…