Founder @tensorlake, building sandbox infra for agents. Past - Built Nomad @hashicorp, AI Infra @meta, Container Scheduler tech lead @linkedin and @netlfix

San Francisco, CA
We are dabbling in CI. Some of our users were already building DIY CI runners on the platform. We want to offer infinite compute for your CI, with pricing that’s transparent and based on the compute and storage you actually consume instead of opaque runner shapes, committed usage, or buying X compute minutes per month. Our biggest learning so far: caching is the biggest unlock for CI workloads. Reliable build caching plus OS/NPM/pip package caching is just as important as having fast CPU cores and lots of RAM.
As cost of coding has plummeted, long builds (CI) has become a bottleneck for many software cos. Three of the biggest startups solving this: 1. Blacksmith (@useblacksmith) 2. Namespace (@namespacelabs) 3. Depot (@depotdev)
7
2
46
5,456
An interesting aspect of this autoscaler is that it uses Jev at the core to decide which cloud and hardware type to bring up (metal 23xl/48xl on AWS; Z3/C4 metal instance types on GCP), based on demand, cost, and where the sandbox snapshots are present. Traditional autoscalers like Karpenter use greedy, heuristic-based algorithms. We used a new open source Rust library called Reflex from @ArunP76475 and @nerdsane's team at @datadoghq. It uses the current cluster state, along with blocked work state in the sandbox scheduler to determine an autoscaling action. Reflex commits the decision only if it stays within the parameters we define, such as max fleet size.
We built a new cluster autoscaler specialized for stateful sandboxes to keep up with growth over the past few weeks. As demand has gone through the roof, we kept getting paged constantly because we were running at 90–95% utilization. Selling out 80% of RAM on machines is bad for p95 sandbox resume latency from memory snapshots. Autoscaling on on-demand capacity on hyperscalers has helped alleviate some capacity constraints while we continuously source longer-term compute contracts for steady-state demand. The autoscaler automatically cordons nodes as demand stabilizes and we add more reserved capacity, migrates sandboxes in some cases, and then scales the on-demand clusters back in. Here’s an example of a cluster scaling up on AWS in reaction to a spike in sandbox creation requests to maintain enough headroom.
7
5
68
12,001
We built a new cluster autoscaler specialized for stateful sandboxes to keep up with growth over the past few weeks. As demand has gone through the roof, we kept getting paged constantly because we were running at 90–95% utilization. Selling out 80% of RAM on machines is bad for p95 sandbox resume latency from memory snapshots. Autoscaling on on-demand capacity on hyperscalers has helped alleviate some capacity constraints while we continuously source longer-term compute contracts for steady-state demand. The autoscaler automatically cordons nodes as demand stabilizes and we add more reserved capacity, migrates sandboxes in some cases, and then scales the on-demand clusters back in. Here’s an example of a cluster scaling up on AWS in reaction to a spike in sandbox creation requests to maintain enough headroom.
8
39
11,345
Every incumbent developer platform will have sandboxes as an offering. It’s the new primitive for building applications, don’t expect anyone to sit out of this. Companies specializing in only sandboxes will have to offer more to continue being in the game. File systems, serverless functions and durable execution are the other primitives I expect most platforms to offer in the near future.
11
3
55
4,998
Diptanu Choudhury retweeted
The results are in for the Storage Lifecycle 10mb 🥇@azure 🥈@awscloud 🥉@tensorlake
1
2
3
614
Diptanu Choudhury retweeted
Replying to @LoopholeLabs
@LoopholeLabs is officially joining @livekit! In 2020 I decided to chase true live migration, a problem that folks had been trying to cracked for 15+ years. Now our breakthroughs will build the next generation of stateful infra for agents at LiveKit's scale.
LiveKit has acquired @LoopholeLabs, an infrastructure company whose technology will enhance our platform for building voice, video, and physical AI agents. With their team, we move closer to offering customers essentially unlimited concurrent agents. Welcome to LiveKit! fortune.com/press-releases/l…
18
6
61
2,704
Custom kernels coming to @tensorlake sandboxes in a day! Users will be able to build their own kernels and use them in sandboxes. This has been one of the most requested features from companies using and testing RL environments on our platforms.
1
13
852
We are ahead of the market here with TLFS. In the long term state of agents in a versioned file system which can be attached to compute on demand feels inevitable. TLFS has a lot of traction and getting love from users who are trying it out and figuring out what they could do with a file system that can be distributed between sandboxes, local machines and also versioned.
Replying to @diptanu @tensorlake
Also I think you’re ahead of most of Twitter. They’ll figure out they need this in a few months 😅
1
1
12
1,427
An interesting use case of @tensorlake's versioned file system is that users can decouple state of an agent or CI build cache from the sandbox where the process ran. TL FS supports any platforms where sandboxes support FUSE mounts such as E2B, Namespace, Vercel or any VM based sandbox environments. A note from a user last night about how they are using TLFS!
4
22
1,187
Rocket ship!
We are hiring a growth engineer. Our funnel is driving millions of new revenue a week. We’re looking for someone who can run 100s of experiments across onboarding, pricing, upsells and more. If you think you can crush this, please DM me!
6
1,332
Spent all morning playing with excel sheet and unit economics of buying compute from different vendors, considering reservations and benchmarking hardware. Selling compute for RL is as much of a financial engineering problem as building and designing for high throughput scheduling of VMs.
1
1
23
1,578
We got our hands on this new library @nerdsane’s team is building, we have a cool use case for this in relation to managing fleets of sandboxes! We will share the results when the library is open to everyone.
I feel delightfully called out. Gonna share a cool library with these use cases pretty soon.
1
14
2,847
We are heavily investing into @harborframework!
pr merged. congrats to @tensorlake on adding streaming support in harbor. streaming now supports 4 providers: @daytonaio, @modal, @tensorlake and @Docker. If you want to merge yours, comment or DM me.
1
12
1,900
We are moving all our x86 sandbox workloads to AMD to get some flexibility when we have to spill over from hyperscalers to bare metal providers and vice versa. Lack of intel bare metal providers makes it hard to use them even on hyperscalers.
1
13
874
Other than the sandboxes, every @tensorlake service now runs on ARM servers on GCP and AWS. We really want to launch ARM sandboxes too, waiting on someone to ask us for them!
6
1
14
1,995
Some of the best to build software for developers were either good at marketing or had partners who were A+ at telling the story.
Replying to @mholt6
yep, sucks! gotta tell more people. Life's not fair! Developers say they hate marketing and then don't market and then get frustrated when a flashy tool gets attention
1
6
1,474
Incredible! Amazed at how many cool use cases people are discovering for Jev.
1
7
2,678
FoundationDB is now powering all the major services in @tensorlake. 1. Our new sandbox orchestrator uses FoundationDB for storing the WAL for the cluster scheduler, route tables of sandboxes, watches for replicas which replays the WAL to in-memory indexes for fast responses to API calls. 2. Durable functions (which hasn't been released) uses FDB for storing the current execution graph of a request. 3. Versioned File System and SCM infra uses it for storing metadata of file systems and git repositories. We use the K8s operator, and it's been working pretty well.
10
11
148
9,541
This is why Tensorlake sandboxes are nothing but VMs, containers are not even an option you can chose on our platform.
We escaped Docker's hypervisor with three lines of bash. CVE-2026-77179: A container gets complete read and write access to the host filesystem. When you mount a folder into a container, Docker's VMM uses virtio-fs, and the file server runs on the host. Because of a TOCTOU bug, if a container opens a file, deletes it while holding its file handle open, and replaces the parent folder with a symlink, the kernel will follow the symlink to anywhere on the host. Full technical breakdown: accomplish.ai/blog/escaping-…
Community note
This affects Docker Sandboxes on macOS and Docker Desktop only if the new Docker VMM (beta, not default) is enabled; it is a virtio-fs shared workspace symlink/TOCTOU issue, not a general container or hypervisor escape. docs.docker.com/security/secur… cve.org/CVERecord?id=C… docs.docker.com/ai/sandboxes/r…
4
2
19
2,491