Inference systems designed for agents

Cambridge, MA
Subconscious Systems retweeted
The club LOVES long-horizon agents
2
4
157
Subconscious Systems retweeted
Sharing the good word of Subconscious at @getpostman last night. Great questions and energy from the crowd 🤝
2
5
149
Subconscious Systems retweeted
Pumped to be a part of this event with @baseten and @CapitalG. Baseten has been a killer partner for us, and I'm always thankful for a platform to spread the word of Subconscious 📣📣
3
2
12
754
Subconscious Systems retweeted
GLM 5.2 live on the Subconscious API - Frontier coding ability - Powered by our inference system optimized for agents 95%+ cache hit rate, fast token throughput, longer context reasoning with automated nearly-lossless context compression. Check it out.
2
2
7
830
Subconscious Systems retweeted
We designed an inference runtime specifically tuned to run agent workloads with a prefix *and* suffix cache, aka when you prune stuff in them middle you still get a 100% cache hit on the tokens around it. We see a 95% cache hit rate for our team's coding agent instance w/ GLM.
How to keep AI spend flat while token usage grows exponentially: Not with friction and spend alerts. With better defaults, routing, and caching. Better Defaults (not Usage Caps) – Engineers can choose any model they want, but defaults matter. We’re experimenting with defaulting to open weight models like GLM 5.2 and Kimi 2.7 through our LLM gateway, while still encouraging engineers to choose the right model for the task. 91% of our employees were never hitting their usage caps, so instead of lowering caps and driving up alerts, we're moving to cheaper defaults. Note that code reviews use a diversity of models, so they can check each other's work. Better Routing – In our custom harnesses, we preprocess prompts and route to the best model for the job, considering cache hits and model pricing. For instance, you may want a frontier model for planning, but not for execution where they can be overkill. Ultimately, humans shouldn't be choosing models - AI can automate this task. Better Caching – Cache misses are the easiest way to drive your cost up. All of our requests are cache aware, so we’re reusing a warm cache wherever possible. For example, our cache hit rate went from 5% → 60% in LibreChat once properly implemented. Keep Context Lean – Start fresh sessions when switching tasks. Scope file context narrowly. Disconnect unused tools. Don't just compact. The goal isn't fewer tokens used, it's fewer tokens wasted. Better Visibility – Our engineers can use as many tokens as they want, from whatever model they want, but we’ve made usage visible – and the more you spend on AI, the more impact we expect. The goal isn't to suppress usage. It's to build the infrastructure that makes exponential growth sustainable. Putting this into practice has cut our AI spend nearly in half, while our token usage continues to grow.
1
1
356
Subconscious Systems retweeted
GLM-5.2, the latest frontier coding model, can now control and compress its own context in inference time, powered by Subconscious Cache. Model-driven context engineering will be the future of AI Agents. subconscious.dev/blog/subcon…
2
13
78
7,042
Subconscious Systems retweeted
Won the Beat The Clock Agent Hack at Wayfair HQ. 2 hours to build. Live demo. 200+ hackers. Thanks @Wayfair, @subsysdev, @baseten, and @Cloudflare for hosting 🚀 #BosTechWeek
1
8
853
Subconscious Systems retweeted
Last week at NY Tech Week, founders and builders filled the room for our BYOA (Bring Your Own Agents) workshop. Everyone brought new ideas and their own agents to showcase and try new tools from us and our partners. Big thank you to our co-host @Mercury, and to @paysponge and @subsysdev for joining us!
5
13
1,647
Subconscious Systems retweeted
This company puts GPUs on top of water heaters, and uses the GPU to heat the water. It's an incredible company, and we're excited to power agents on this distributed inference network. More info below!
Made with AI
1
2
3
190
Subconscious Systems retweeted
Launch of our TIM-Qwen3.6-27B Model today. We took a great, small Qwen model and substantially improved it's performance: - 1M+ context window (10x extension) - Run 3x as many concurrent runs - 75% less memory required - 1.5x token throughput - OpenAI / Anthropic SDK compatible
Made with AI
1
1
3
202
Subconscious Systems retweeted
If you haven't tried the Subconscious playground, here's a look. I gave the model search tools, asked it to tell me about our research paper, and boom, we have a personalized deep research agent. Took 15 seconds to make.
2
6
350
Subconscious Systems retweeted
Beyond Context Limits Subconscious Threads for Long-Horizon Reasoning
2
13
51
25,249
Subconscious Systems retweeted
Today we're launching Subconscious: a new platform for building agents with long-horizon reasoning and tool use, backed by MIT research. One API call. Tool use. Context beyond existing limits. If you're building agents, let's talk.
119
22
256
12,115
Subconscious Systems retweeted
Coming soon...
2
5
244