Founder LivingDevOps | DevOps Lead | Educator | Mentor | Tech Storyteller

Noida, India
After 14 years in production DevOps and training 400+ engineers, I just launched my most advanced bootcamp to date. 28 weeks. 56 live classes. DevOps + MLOps + AIOps. Built for experienced engineers who want to move into MLOps and AIOps without faking it on their resume. Check it out ๐Ÿ‘‡ livingdevops.com/courses/28-โ€ฆ Launch offer: up to 30% off if you book by 15th June.
1
14
94
9,985
Terraform is great. Until you have to run it across three environments. You write the code once. Then you copy it for staging. Then again for prod. It feels clean at first. Each environment has its own folder, its own state, its own everything. A few months later, you fix a bug in dev and forget staging. Someone patches prod during a late-night hotfix. A one-line change becomes three PRs. And nobody can explain why prod behaves differently. The problem is no longer writing Terraform. The problem is keeping the same Terraform in sync across environments. The approach I have found useful is simple: Keep one codebase. Move the differences into config. terraform/ โ”œโ”€โ”€ modules/ โ”œโ”€โ”€ main.tf โ”œโ”€โ”€ variables.tf โ””โ”€โ”€ vars/ โ”œโ”€โ”€ dev.tfvars โ”œโ”€โ”€ dev.tfbackend โ”œโ”€โ”€ prod.tfvars โ””โ”€โ”€ prod.tfbackend The code defines the infrastructure. The .tfvars define the environment. The .tfbackend defines where the state lives. Fix something once, and it's fixed everywhere. The only pain left was the flags: terraform init -backend-config=vars/prod.tfbackend terraform plan -var-file=vars/prod.tfvars Every single time. For every environment. And one wrong filename away from planning prod with dev values. That is why I built Toffee. Instead of remembering flags, you just say which environment: toffee dev plan toffee prod apply It does not replace Terraform. It does not add a new layer to learn. It just removes the repetitive part. That is how I think infrastructure tooling should work. Less copy-paste. Fewer flags to remember. More time building infrastructure. It's open source; try it and tell me what's missing: github.com/akhileshmishrabizโ€ฆ
2
190
Akhilesh Mishra retweeted
We should start writing our docs in YAML. AI agents are reading them anyway. At least weโ€™ll save some tokens. ๐Ÿ˜‚
1
10
1,143
Akhilesh Mishra retweeted
The Kubernetes trap about being "Production ready"
Made with AI
2
3
8
586
Akhilesh Mishra retweeted
Why do Terraform docs still exist when no one is reading them anymore?
7
4
43
9,978
Akhilesh Mishra retweeted
- One guy collaborates with a billionaire to create OpenAI for the welfare of humans. - As soon as the project starts getting traction, he turns it into for-profit. Gets billions of dollars in investment to build AGI. That's the pitch. - Tells the world humans are not needed, they can duck off. - Sees big competition from China building competitive models. - Then his AI models hack Hugging Face( basically orchestrated) - Now he yells out "wolf is coming" and urges everyone to slow down AI model improvement. - Launches his newer models within weeks. - Now going to the Senate to give a speech on safety. Is this concern for humanity or just a smart business move?
1
1
5
1,057
Akhilesh Mishra retweeted
Kubernetes has started to become Jenkins for containers I know a lot of people run clusters that have more CRDs than pods. Every operator promises to make your life easier. Cert management, Ingress, eso, argocd, service mesh, CNPG, and list goes on and on. Each one solves a real problem on its own. But When stack ten of them together, Youโ€™re not running Kubernetes anymore. Youโ€™re running a distributed system of controllers reconciling other controllers.๐Ÿ˜€๐Ÿ˜€ And then comes the Kubernetes upgrade. Now you have to worry about operator versions, CRD compatibility, dependencies and whether that project is still maintained. Jenkins didnโ€™t die because plugins were bad. It died because nobody owned the plugin list. Kubernetes operators are heading the same way I mean, how many plugins and operators is too many before it becomes unmanageable dependency hell?
8
8
114
9,894
Akhilesh Mishra retweeted
Agentic AI interview question: An agent can refund up to โ‚น10,000 automatically but needs approval above that. Where would you enforce this rule: the prompt, model, tool, or backend?
59
8
105
14,919
Akhilesh Mishra retweeted
Before you learn Terraform, understand why to learn Terraform. Or should you? Around 2006, AWS emerged and changed everything. For the first time, you could spin up a server in minutes with a few clicks. No $50,000 hardware. No weeks of setup. Just click, click, done. Pay only for what you use. Scale on demand. Microsoft launched Azure. Google launched GCP. Cloud computing was born. For startups, this was revolutionary. Launch without buying a single server. But a new problem emerged. Managing 5 servers manually? Easy. Managing 500 servers across multiple environments? Nightmare. You'd log into the AWS console. Click to create a server. Select instance type. Choose network. Configure security groups. Add storage. Click, click, click, submit. Need the same setup in staging? Repeat all those clicks. Production? Repeat again. One wrong click and your production looked nothing like staging. "It works in staging" became the new "it works on my machine." Disaster recovery was a joke. Your infrastructure got deleted? Good luck remembering every single configuration you clicked through. No documentation. No history. No way to track who changed what and when. Configuration drift was constant. Three environments that should be identical gradually became completely different. Nobody knew why. Want to recreate your infrastructure? Better hope someone documented every single step. Spoiler: nobody did. Scaling was painful. Black Friday coming? Start clicking to provision servers two weeks early. One by one. Teams worked in silos. Ops managed infrastructure through console clicks. Developers had zero visibility. Compliance audits were nightmares. "Show us all infrastructure changes from last quarter." Impossible. This is exactly what Infrastructure as Code solved. The idea was simple. Stop clicking. Start coding. Define your infrastructure in code files. Version control them. Deploy infrastructure the same way you deploy applications. AWS built CloudFormation for AWS. Azure built ARM templates for Azure. Google built Deployment Manager for GCP. Better than clicking. But there was a problem. Different syntax for each cloud. Want to use AWS and Azure? Learn two completely different tools. Different commands. Different workflows. Organizations were adopting multi-cloud strategies. Using AWS for compute. GCP for machine learning. Azure for enterprise apps. Managing three different IaC tools was still painful. Then came Terraform in 2014. HashiCorp saw the gap. They built one tool that works with all clouds. Same syntax. Same workflow. Whether you're on AWS, Azure, GCP, or even managing GitHub repositories. Write once. Deploy anywhere. Terraform used HCL (HashiCorp Configuration Language). Simple. Readable. Declarative. You describe what you want, not how to create it. Built in Go. Fast. Reliable. Creates resources in parallel. But here's what made Terraform win: it was cloud-agnostic when everyone else was cloud-specific. Your infrastructure became code. Track changes in Git. Review in pull requests. Roll back when needed. Recreate entire environments in minutes. No more clicking. No more guessing. No more drift. Why did Terraform win? Perfect timing. Cloud adoption was exploding. Multi-cloud was becoming the norm. Companies needed one tool to manage infrastructure across all clouds. Terraform solved that exact problem. It's idempotent. Run the same code 100 times, get the same result. It's modular. Reuse code across projects. It manages state. Knows what exists and what needs to change. 12 years later, Terraform is the de facto standard for IaC. Now Terraform skills are in massive demand. Senior DevOps engineers with Terraform expertise command premium salaries. Understanding this story is more important than memorizing terraform commands. Now go learn Terraform already.
3
18
74
3,803
These Linux commands helped me most in last 13 years of IT career Daily stuff: โ€ข ps aux | grep {process} - Find that sneaky process โ€ข lsof -i :{port} - Who's hogging that port? โ€ข df -h - The classic "we're out of space" checker โ€ข netstat -tulpn - Network connection detective โ€ข kubectl get pods | grep -i error - K8s trouble finder Log related commands: โ€ข tail -f /var/log/* - Real-time log watcher โ€ข journalctl -fu service-name - SystemD log stalker โ€ข grep -r "error" . - The error hunter โ€ข zcat access.log.gz | grep "500" - Compressed log ninja โ€ข less +F - The better tail command Container cli: โ€ข docker ps --format '{{.Names}} {{.Status}}' - Clean status check โ€ข docker stats --no-stream - Quick resource check โ€ข crictl logs {container} - Raw container stories โ€ข docker exec -it - The container backdoor โ€ข podman top - Process peek inside containers System Detectives: โ€ข htop - System resource storyteller โ€ข iostat -xz 1 - Disk performance poet โ€ข free -h - Memory mystery solver โ€ข vmstat 1 - System vital signs โ€ข dmesg -T | tail - Kernel's recent gossip Network stuff: โ€ข curl -v - HTTP conversation debugger โ€ข dig +short - Quick DNS lookup โ€ข ss -tunlp - Socket statistics simplified โ€ข iptables -L - Firewall rule reader โ€ข traceroute - Path finder File and stuff: โ€ข find . -name "*.yaml" -type f - YAML hunter โ€ข rsync -avz - Better file copier โ€ข tar -xvf - The unzipper (yes, we all google this) โ€ข ln -s - Symlink wizard โ€ข chmod +x - Make it executable Performance: โ€ข strace -p {pid} - System call spy โ€ข tcpdump -i any - Network packet sniffer โ€ข sar -n DEV 1 - Network stats watch โ€ข uptime - Load average at a glance โ€ข top -c - Classic process viewer Git Essentials: โ€ข git log --oneline - History simplified โ€ข git reset --hard HEAD^ - The "oops" eraser โ€ข git stash - The work hider โ€ข git diff --cached - What's staged? โ€ข git blame - The "who did this?" resolver Quick Fixes: โ€ข sudo !! - Run last command with sudo โ€ข ctrl+r - Command history search โ€ข history | grep - Command time machine โ€ข alias - Command shortcut maker โ€ข watch - Command repeater
14
200
1,193
48,017
Kubernetes is the #1 skill recruiters screen for. Most people only know the tutorial version. I teach the production version. livingdevops.com/courses/9-wโ€ฆ
1
293
Most Kubernetes engineers don't know these kubectl tricks that save hours during outages ๐—ธ๐˜‚๐—ฏ๐—ฒ๐—ฐ๐˜๐—น ๐—ด๐—ฒ๐˜ ๐—ฒ๐˜ƒ๐—ฒ๐—ป๐˜๐˜€ -๐—” -๐˜„ | ๐—ด๐—ฟ๐—ฒ๐—ฝ -๐˜ƒ "๐—ก๐—ผ๐—ฟ๐—บ๐—ฎ๐—น" --> Watch only the bad stuff across the entire cluster > It streams warnings and failures from every namespace in real time so you can catch issues before they turn into incidents. ๐—ธ๐˜‚๐—ฏ๐—ฒ๐—ฐ๐˜๐—น ๐—ฑ๐—ถ๐—ณ๐—ณ -๐—ณ ๐—ฑ๐—ฒ๐—ฝ๐—น๐—ผ๐˜†๐—บ๐—ฒ๐—ป๐˜.๐˜†๐—ฎ๐—บ๐—น --> See exactly what will change before applying > It helps catch accidental changes like wrong image tags, deleted environment variables, or broken resource limits before deployment. ๐—ธ๐˜‚๐—ฏ๐—ฒ๐—ฐ๐˜๐—น ๐—ด๐—ฒ๐˜ ๐—ฝ๐—ผ๐—ฑ๐˜€ -๐—” --๐˜€๐—ผ๐—ฟ๐˜-๐—ฏ๐˜†='.๐˜€๐˜๐—ฎ๐˜๐˜‚๐˜€.๐—ฐ๐—ผ๐—ป๐˜๐—ฎ๐—ถ๐—ป๐—ฒ๐—ฟ๐—ฆ๐˜๐—ฎ๐˜๐˜‚๐˜€๐—ฒ๐˜€[๐Ÿฌ].๐—ฟ๐—ฒ๐˜€๐˜๐—ฎ๐—ฟ๐˜๐—–๐—ผ๐˜‚๐—ป๐˜' --> Find unstable pods instantly > This sorts pods by restart count so the most unstable workloads appear first. ๐—ธ๐˜‚๐—ฏ๐—ฒ๐—ฐ๐˜๐—น ๐˜๐—ผ๐—ฝ ๐—ฝ๐—ผ๐—ฑ๐˜€ -๐—” --๐˜€๐—ผ๐—ฟ๐˜-๐—ฏ๐˜†=๐—บ๐—ฒ๐—บ๐—ผ๐—ฟ๐˜† --๐—ฐ๐—ผ๐—ป๐˜๐—ฎ๐—ถ๐—ป๐—ฒ๐—ฟ๐˜€ --> Find who is actually consuming memory > The `--containers` flag breaks usage down per container instead of showing only pod-level metrics. ๐—ธ๐˜‚๐—ฏ๐—ฒ๐—ฐ๐˜๐—น ๐—ฑ๐—ฒ๐—ฏ๐˜‚๐—ด -๐—ถ๐˜ <๐—ฝ๐—ผ๐—ฑ-๐—ป๐—ฎ๐—บ๐—ฒ> --๐—ถ๐—บ๐—ฎ๐—ด๐—ฒ=๐—ฏ๐˜‚๐˜€๐˜†๐—ฏ๐—ผ๐˜… --๐—ฐ๐—ผ๐—ฝ๐˜†-๐˜๐—ผ=๐—ฑ๐—ฒ๐—ฏ๐˜‚๐—ด-๐—ฝ๐—ผ๐—ฑ --> Debug a pod without touching the original workload > This creates a copy of the failing pod with a debug container attached for investigation. ๐—ธ๐˜‚๐—ฏ๐—ฒ๐—ฐ๐˜๐—น ๐—ด๐—ฒ๐˜ ๐—ฝ๐—ผ๐—ฑ๐˜€ -๐—” -๐—ผ ๐˜„๐—ถ๐—ฑ๐—ฒ --๐—ณ๐—ถ๐—ฒ๐—น๐—ฑ-๐˜€๐—ฒ๐—น๐—ฒ๐—ฐ๐˜๐—ผ๐—ฟ ๐˜€๐—ฝ๐—ฒ๐—ฐ.๐—ป๐—ผ๐—ฑ๐—ฒ๐—ก๐—ฎ๐—บ๐—ฒ=<๐—ป๐—ผ๐—ฑ๐—ฒ-๐—ป๐—ฎ๐—บ๐—ฒ> --> See every workload sitting on a node before you touch it > Run this before cordoning or draining a node during maintenance. ๐—ธ๐˜‚๐—ฏ๐—ฒ๐—ฐ๐˜๐—น ๐—ด๐—ฒ๐˜ ๐—ฒ๐˜ƒ๐—ฒ๐—ป๐˜๐˜€ --๐—ณ๐—ถ๐—ฒ๐—น๐—ฑ-๐˜€๐—ฒ๐—น๐—ฒ๐—ฐ๐˜๐—ผ๐—ฟ ๐—ถ๐—ป๐˜ƒ๐—ผ๐—น๐˜ƒ๐—ฒ๐—ฑ๐—ข๐—ฏ๐—ท๐—ฒ๐—ฐ๐˜.๐—ป๐—ฎ๐—บ๐—ฒ=<๐—ฝ๐—ผ๐—ฑ-๐—ป๐—ฎ๐—บ๐—ฒ> --๐˜€๐—ผ๐—ฟ๐˜-๐—ฏ๐˜†='.๐—น๐—ฎ๐˜€๐˜๐—ง๐—ถ๐—บ๐—ฒ๐˜€๐˜๐—ฎ๐—บ๐—ฝ' --> Get the full chronological event history for one object. > kubectl describe only shows recent events. This pulls the complete timeline for a specific pod, useful when you're trying to reconstruct what happened over the last hour.
5
42
150
4,114
Akhilesh Mishra retweeted
Most Kubernetes engineers don't know these kubectl tricks that save hours during outages ๐—ธ๐˜‚๐—ฏ๐—ฒ๐—ฐ๐˜๐—น ๐—ด๐—ฒ๐˜ ๐—ฒ๐˜ƒ๐—ฒ๐—ป๐˜๐˜€ -๐—” -๐˜„ | ๐—ด๐—ฟ๐—ฒ๐—ฝ -๐˜ƒ "๐—ก๐—ผ๐—ฟ๐—บ๐—ฎ๐—น" --> Watch only the bad stuff across the entire cluster > It streams warnings and failures from every namespace in real time so you can catch issues before they turn into incidents. ๐—ธ๐˜‚๐—ฏ๐—ฒ๐—ฐ๐˜๐—น ๐—ฑ๐—ถ๐—ณ๐—ณ -๐—ณ ๐—ฑ๐—ฒ๐—ฝ๐—น๐—ผ๐˜†๐—บ๐—ฒ๐—ป๐˜.๐˜†๐—ฎ๐—บ๐—น --> See exactly what will change before applying > It helps catch accidental changes like wrong image tags, deleted environment variables, or broken resource limits before deployment. ๐—ธ๐˜‚๐—ฏ๐—ฒ๐—ฐ๐˜๐—น ๐—ด๐—ฒ๐˜ ๐—ฝ๐—ผ๐—ฑ๐˜€ -๐—” --๐˜€๐—ผ๐—ฟ๐˜-๐—ฏ๐˜†='.๐˜€๐˜๐—ฎ๐˜๐˜‚๐˜€.๐—ฐ๐—ผ๐—ป๐˜๐—ฎ๐—ถ๐—ป๐—ฒ๐—ฟ๐—ฆ๐˜๐—ฎ๐˜๐˜‚๐˜€๐—ฒ๐˜€[๐Ÿฌ].๐—ฟ๐—ฒ๐˜€๐˜๐—ฎ๐—ฟ๐˜๐—–๐—ผ๐˜‚๐—ป๐˜' --> Find unstable pods instantly > This sorts pods by restart count so the most unstable workloads appear first. ๐—ธ๐˜‚๐—ฏ๐—ฒ๐—ฐ๐˜๐—น ๐˜๐—ผ๐—ฝ ๐—ฝ๐—ผ๐—ฑ๐˜€ -๐—” --๐˜€๐—ผ๐—ฟ๐˜-๐—ฏ๐˜†=๐—บ๐—ฒ๐—บ๐—ผ๐—ฟ๐˜† --๐—ฐ๐—ผ๐—ป๐˜๐—ฎ๐—ถ๐—ป๐—ฒ๐—ฟ๐˜€ --> Find who is actually consuming memory > The `--containers` flag breaks usage down per container instead of showing only pod-level metrics. ๐—ธ๐˜‚๐—ฏ๐—ฒ๐—ฐ๐˜๐—น ๐—ฑ๐—ฒ๐—ฏ๐˜‚๐—ด -๐—ถ๐˜ <๐—ฝ๐—ผ๐—ฑ-๐—ป๐—ฎ๐—บ๐—ฒ> --๐—ถ๐—บ๐—ฎ๐—ด๐—ฒ=๐—ฏ๐˜‚๐˜€๐˜†๐—ฏ๐—ผ๐˜… --๐—ฐ๐—ผ๐—ฝ๐˜†-๐˜๐—ผ=๐—ฑ๐—ฒ๐—ฏ๐˜‚๐—ด-๐—ฝ๐—ผ๐—ฑ --> Debug a pod without touching the original workload > This creates a copy of the failing pod with a debug container attached for investigation. ๐—ธ๐˜‚๐—ฏ๐—ฒ๐—ฐ๐˜๐—น ๐—ด๐—ฒ๐˜ ๐—ฝ๐—ผ๐—ฑ๐˜€ -๐—” -๐—ผ ๐˜„๐—ถ๐—ฑ๐—ฒ --๐—ณ๐—ถ๐—ฒ๐—น๐—ฑ-๐˜€๐—ฒ๐—น๐—ฒ๐—ฐ๐˜๐—ผ๐—ฟ ๐˜€๐—ฝ๐—ฒ๐—ฐ.๐—ป๐—ผ๐—ฑ๐—ฒ๐—ก๐—ฎ๐—บ๐—ฒ=<๐—ป๐—ผ๐—ฑ๐—ฒ-๐—ป๐—ฎ๐—บ๐—ฒ> --> See every workload sitting on a node before you touch it > Run this before cordoning or draining a node during maintenance. ๐—ธ๐˜‚๐—ฏ๐—ฒ๐—ฐ๐˜๐—น ๐—ด๐—ฒ๐˜ ๐—ฒ๐˜ƒ๐—ฒ๐—ป๐˜๐˜€ --๐—ณ๐—ถ๐—ฒ๐—น๐—ฑ-๐˜€๐—ฒ๐—น๐—ฒ๐—ฐ๐˜๐—ผ๐—ฟ ๐—ถ๐—ป๐˜ƒ๐—ผ๐—น๐˜ƒ๐—ฒ๐—ฑ๐—ข๐—ฏ๐—ท๐—ฒ๐—ฐ๐˜.๐—ป๐—ฎ๐—บ๐—ฒ=<๐—ฝ๐—ผ๐—ฑ-๐—ป๐—ฎ๐—บ๐—ฒ> --๐˜€๐—ผ๐—ฟ๐˜-๐—ฏ๐˜†='.๐—น๐—ฎ๐˜€๐˜๐—ง๐—ถ๐—บ๐—ฒ๐˜€๐˜๐—ฎ๐—บ๐—ฝ' --> Get the full chronological event history for one object. > kubectl describe only shows recent events. This pulls the complete timeline for a specific pod, useful when you're trying to reconstruct what happened over the last hour.
5
42
150
4,114