07 Aug 2026
A few days ago I had a conversation I was not expecting to be difficult. It was framed around non-technical topics: how I mentor and influence others, how I learn, how I handle conflicting needs across teams.
I work in infra. My day-to-day is Terraform, node pools, and YAML. I figured this would be a relaxed conversation. It was not. I ended up asking for questions to be repeated more than once, not because they were unclear, but because I genuinely needed time to think.
I am writing this because the reflection itself was worth it, whatever comes next.
More …
04 Aug 2026
The original bot was simple. Slash commands, whitelist check, run some kubectl or MySQL query, return the result. That was the whole thing.
Then someone asked: can I just ask the bot a question? Not a slash command. Just a question, in natural language, inside a Google Chat thread, and get a useful answer back.
That broke what I had.
More …
13 May 2026
We shipped faster with AI. We also shipped a security risk.
That sentence took a while to sit right. For the first few months after the team adopted AI coding assistants into the workflow, everything felt like a win. Pull requests were moving faster. Features that used to take three days were done in one. Someone on the team joked that we had finally become a “10x team.” And statistically, the numbers backed it up. We went from pushing to production roughly once a day to doing it three times.
Nobody stopped to ask whether our security posture had kept up.
More …
09 May 2026
There is a particular kind of anxiety that comes with being a DevOps engineer. Not the kind from outages or failed deployments, though those are present too. The quieter kind. The background hum of knowing that something could break at any moment and that when it does, people will be waiting for you to fix it.
The standard assumption baked into that responsibility is that you will be at your desk. That you have a terminal open, or can get one open quickly. That your laptop is somewhere nearby, that you can SSH into things, run commands, check logs, and do the actual work.
Most of the time that is true. And then sometimes you are on a commute, or standing in a queue, or halfway through a trip with your laptop sitting at home, and a notification arrives telling you something is down.
More …
01 May 2026
The first sign something was wrong wasn’t an alert. It wasn’t a spike in the error rate or a pager going off. It was a message from finance at the end of the month asking why room revenue was lower than expected.
The engineering team pulled up the logs. Order service: reservation created. Room service: room assigned. Payment service: charge processed. Everything green. Every service reporting success. But somehow, users were checking into rooms they hadn’t actually paid for. The money wasn’t there.
Nobody could explain it. Not because the system wasn’t logging, it was. Not because there was no observability stack, there was. Grafana was deployed. Loki was ingesting logs from every service. Tempo was ready for traces. The team had spent two sprints setting it all up and had proudly declared themselves production-ready.
More …