Local AI is very practical.
One prompt: download this arXiv paper PDF and save it into my AI Knowledge. Then summarize the core idea.
That is a real workflow. A tool call, a fetch, a file written into my knowledge base, then a read and a summary. Not a toy prompt.
Running deepseek-v4-flash locally in Open WebUI, with tools and web access on. Start to finish, from my first message to the finished summary, two minutes. I recorded it unedited so nobody has to take my word on the speed.
That is the part people still get wrong about running models on your own hardware. The assumption is that you trade away everything for privacy. Slow tokens, weak reasoning, no tools. That was true a year ago. It is not true now.
No API bill. No rate limits. The paper never touched a third party server, and neither did anything I asked about it.
Two minutes, on my own box (dgx spark).