Jeffrey Morgan retweeted
Added this little dude to my personal site
1
6
400
Jeffrey Morgan retweeted
You can now add usage credits for paid cloud models without an Ollama subscription. Add usage credits and pay as you go. ollama.com/settings
22
7
135
18,049
Jeffrey Morgan retweeted
Build agentic workflows completely offline. The Antigravity SDK now supports local execution with Gemma 4 and LiteRT. Run agents entirely on your local machine with: 💵 Zero token costs 🔒 Total data privacy 🔌 Offline reliability Bonus feature: Support for OpenAI-compatible endpoints. Use Ollama, llama.cpp, vLLM and more to serve Gemma 💪 Get started: pip install google-antigravity litert-lm Read the details: developers.googleblog.com/in…
46
195
1,548
100,247
Jeffrey Morgan retweeted
AI models are improving faster than we can understand them. Dario Amodei wants to narrow that gap by pacing development, but the consequences extend well beyond the models.
8
2
35
4,995
Many concerns this weekend about slowing down AI, some targeting open models. We need to be responsible. But we can't let this slow down, or worse, prohibit open models. It's the wrong risk. Open models have tremendous power to democratize AI and make it more personal. The larger risk I see with open models is right in front of us: in the last week I've read about how many popular platforms quietly send data to foreign jurisdictions, or worse, sell or train on it to gain an advantage. And this is becoming more and more mainstream. Now more than ever open model vendors must act in the user's best interest, not their own: zero data retention or training, and hosting in the user's region vs sending data overseas. This has been our belief and commitment with @ollama. Nobody needs to slow down open models. We need to distribute and run them in a way users can trust.
7
11
78
8,423
Small models are truly capable of the majority of conversational use cases and even a majority of difficult reasoning ones. Amazing work @Avanika15 @JonSaadFalcon + team and great feature in @FT !
dreams do come true 🥹. excited to see our work (w/@JonSaadFalcon, @HazyResearch, john hennessy and @Azaliamirh) feat. in a major way in @FT. the world is becoming increasingly less dependent on centralized cloud ai. we are just getting started 🚀🌖
6
4
34
10,504
Jeffrey Morgan retweeted
Amp users can now use Ollama's cloud models with Amp's new BYOK model routing. No limits or fees for BYOK. Build remote agents, controllable from everywhere!
Amp is now free to use when you bring your own compute and model subscriptions/keys. No more limits or fees for BYOK. ampcode.com/news/free-agent
35
13
248
47,062
Jeffrey Morgan retweeted
DeepSeek-V4.1-Flash is now fully rolled out and available on Ollama's cloud: - Hosted in US & Europe - Zero data retention: prompts and responses are never logged or trained on - Per-token pricing matches the DeepSeek API, including off-peak pricing - Get started with Ollama's Pro, Max, and Team plans, or pay as you go with a free account with no service fees This new model by DeepSeek is more capable, faster, and more cost effective than all prior DeepSeek models including DeepSeek-V4-Pro 🚀.
🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the smallest model in our new architecture family, with native visual understanding. 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models. 1/6
66
47
759
66,490
Jeffrey Morgan retweeted
DeepSeek-V4.1-Flash is now rolling out for Pro plan subscribers.
DeepSeek-V4.1-Flash is now being rolled out on Ollama's cloud, starting with Max and Team accounts. We are quickly adding more capacity to roll it out to all subscribers.
37
20
503
44,880
Jeffrey Morgan retweeted
ChatGPT Desktop (the Codex app) can now be configured to use Ollama models. Download or update to Ollama 0.34 to get started.
71
105
1,430
125,459
Jeffrey Morgan retweeted
DeepSeek-V4.1-Flash is now being rolled out on Ollama's cloud, starting with Max and Team accounts. We are quickly adding more capacity to roll it out to all subscribers.
🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the smallest model in our new architecture family, with native visual understanding. 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models. 1/6
30
39
502
87,684
Jeffrey Morgan retweeted
we've built a local-cloud harness that is 800x cheaper & 4x faster than cloud-only harnesses. @JonSaadFalcon gives the tl;dr at the @ycombinator paper club. ff to 37:30 for more deets 🙂. link to full paper in comments 👇
Harnesses often get dismissed as just scaffolding, just prompt engineering, and not real research. But that couldn't be farther from the truth. The same model weights that score 30% on ARC-AGI score 95% with a better harness. So we gathered a group of researchers and founders working at the frontier to do a deep dive into the state of harnesses. We cover how we got to this point, the case for making your harness as expressive as possible, and what YC learned building an agent for every employee in the company. 00:00 - @FrancoisChauba1: Why harnesses matter 04:27 - Building an auto-researcher by accident 07:13 - A five minute history of harnesses 13:56 - Self-improving harnesses 18:35 - @sethkarten: Prime Agent, a self-improving RLM harness 21:50 - Context as an L1, L2, L3 cache 24:51 - From Turing machine to von Neumann computer 28:33 - Messaging between agents 30:04 - ARC-AGI results 33:09 - Emulator Bench and GPU kernels 37:30 - @JonSaadFalcon: OpenJarvis, personal AI on personal devices 38:26 - How far behind are local models 39:21 - The five primitives of a personal AI stack 42:47 - Letting cloud models optimize your local stack 43:53 - 800x cheaper than the cloud 45:58 - @josh__france and @jbellregan: QM, YC's agent harness for work 47:29 - A history of YC's internal agents 49:24 - OpenClaw and a fleet of 50 agents 51:04 - Pulling the brain out of the sandbox 54:43 - Letting the agent choose its own sandbox and model 57:16 - The grind tool: budgets on goals 58:50 - Agents don't understand social context
3
12
75
10,134
Jeffrey Morgan retweeted
Introducing off-peak hour token rates. DeepSeek-V4-Flash and Pro are now half price outside of 12:00 to 18:00 UTC on weekdays (5am-11am pacific), and all day on weekends! Off-peak pricing will be available soon for more models. DeepSeek models on Ollama's cloud are hosted in the US & Europe with ZDR and fast performance.
100
50
969
104,017
Jeffrey Morgan retweeted
OSS models are quite viable we’ve seen not just large F100 adopt it but also many of the s26 startups use it for coding too thx for sharing inaights @jmorgan
🦙 @ollama is used by 9 million developers and 85% of the Fortune 500, giving co-founder and CEO Jeffrey Morgan (@jmorgan) a unique view into which AI models people are actually using and how that’s changing. Right now, the biggest shift he sees is toward open models, driven by coding agents, falling costs, and capabilities that are rapidly catching up to the frontier labs. On Ollama Cloud, that shift has driven a 150x increase in token usage since the start of the year. In this episode of @LightconePod, Jeff joins @garrytan, @snowmaker, @sdianahu, and @harjtaggar to talk about the future of open models and the story behind Ollama, from two years of searching for the right idea to building one of the most widely used AI developer tools in the world. 00:43 — The Shift to Open Models 03:03 — How AI Agents Are Driving Token Usage 05:31 — Are Open Models Catching Up? 08:26 — What Happens When a New Model Launches 11:31 — Ollama as an Operating System for AI 14:05 — The New Opportunities Above the Model Layer 18:19 — Why 80–90% of Enterprise Tokens Could Be Open 20:57 — The Future Is Local and Cloud 26:40 — Why AI Is Coming Back to Your Computer 28:56 — The Coming Era of Unlimited Tokens 32:30 — Do We Still Need a “God Model”? 33:41 — Open Models and Geopolitics 36:14 — The Origins of Ollama 40:36 — Two Years Lost in the Wilderness 42:39 — The Pivot That Changed Everything 47:02 — How Ollama Found a Business Model 49:43 — Why Second-Time Founders Did YC
7
5
61
18,844
Jeffrey Morgan retweeted
Congrats to the team! Bullish on open models and all the partners we make along the way! ❤️ Let’s go
Exciting day for NVIDIA and @huggingface. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. They allow every developer, startup, university, industry and country to build with, customize and benefit from AI. Thank you @ClementDelangue for coming to me. NVIDIA is going to be a great home for Hugging Face, its community and the future of open models. 🤗 blogs.nvidia.com/blog/nvidia…
19
11
406
39,450
Jeffrey Morgan retweeted
Ollama’s Pro, Max, and Team plans now use transparent per-token pricing. Based on your feedback, every plan includes a monthly pool of usage credits. If you’re on an existing Pro, Max, or Team plan, your plan continues to work as-is. You can upgrade to the new pricing anytime in your Ollama account settings. Every plan includes: - High-performance access to the latest open models, at published per-token rates - Works with popular coding agents, including Claude Code and Codex, plus an API for your own tools - Monthly usage credits included with every plan - Zero data retention, hosted in the US and Europe, plus Singapore for a limited set of Qwen models - No service fees or hidden limits Pro: $20/month, includes $60 of monthly usage Max: $100/month, includes $300 of monthly usage Team: $500/month, includes $1,000 of shared monthly usage for unlimited users The free plan now includes a small amount of monthly usage for a set of starter models. Learn more directly from Ollama's pricing page: ollama.com/pricing
242
68
839
320,931
Jeffrey Morgan retweeted
Ollama already support the Muse Code harness out of the box? You can try it with models via Ollama (local and cloud): ollama launch muse
Muse Code is out of beta and now built to handle bigger, more complex engineering tasks. Developers can get started with one command today: curl -fsSL dev.meta.ai/install.sh | bash
30
12
220
34,931
Jeffrey Morgan retweeted
welcome to the open source + model agnostic era
35
34
375
66,481
Jeffrey Morgan retweeted
Try GLM 5.3!
Terminal-Bench 4.0 just dropped. Benchmark iteration is catching up with model development.
29
12
403
33,582