An interesting aspect of this autoscaler is that it uses Jev at the core to decide which cloud and hardware type to bring up (metal 23xl/48xl on AWS; Z3/C4 metal instance types on GCP), based on demand, cost, and where the sandbox snapshots are present.
Traditional autoscalers like Karpenter use greedy, heuristic-based algorithms.
We used a new open source Rust library called Reflex from @ArunP76475 and
@nerdsane's team at
@datadoghq. It uses the current cluster state, along with blocked work state in the sandbox scheduler to determine an autoscaling action.
Reflex commits the decision only if it stays within the parameters we define, such as max fleet size.
We built a new cluster autoscaler specialized for stateful sandboxes to keep up with growth over the past few weeks.
As demand has gone through the roof, we kept getting paged constantly because we were running at 90–95% utilization. Selling out 80% of RAM on machines is bad for p95 sandbox resume latency from memory snapshots.
Autoscaling on on-demand capacity on hyperscalers has helped alleviate some capacity constraints while we continuously source longer-term compute contracts for steady-state demand.
The autoscaler automatically cordons nodes as demand stabilizes and we add more reserved capacity, migrates sandboxes in some cases, and then scales the on-demand clusters back in.
Here’s an example of a cluster scaling up on AWS in reaction to a spike in sandbox creation requests to maintain enough headroom.