I've talked to a lot of platform teams that still use calendar reminders for certificate rotations, Slack threads for failed runs and spreadsheets for stale stacks. Those manual Day 2 ops won't cut it for much longer.
Think about what day-to-day infrastructure operations entails:
- When a run fails, someone has to gather context, find the right owner, file a ticket and follow up.
- When drift is detected, someone has to classify it, make a remediation decision and notify the right people.
- When a developer needs a new environment, someone has to review the request, get approval and provision it.
- When a certificate is expiring, someone has to remember, rotate it and confirm it still works.
On most teams, each of these workflows is handled by a different person in a different way. That was barely sustainable before AI. Now, with agents shipping more code faster and putting pressure on every layer of the pipeline?
It's becoming untenable.
Day 1 has defined workflows, clear ownership and audit trails through GitOps, IaC and CI/CD. Day 2 hasn't gotten the same treatment yet.
But that day is coming.
And the teams that aren't ready will see their ops collapse under AI's weight.