Ship It Weekly Host Commentaries
Host commentary is the written layer behind each episode: judgment calls, context the audio did not have time for, and links worth bookmarking. This archive collects every episode that ships with commentary so you can skim by week without opening the full player.
Commentary is distinct from show notes (RSS descriptions) and transcripts. Show notes summarize the episode; commentary is the host's editorial read on what mattered and why.
What this page is for
What host commentary is
Editorial context from the host — not a recap of the audio. Expect opinions, follow-up links, and the operational framing that does not fit in a headline.
Read inline or on the episode page
This archive shows full commentary text for browsing and search. Open any episode for audio, chapters, transcripts, and show notes in one place.
Pair with transcripts
Prefer the spoken word? The transcript archive lets you search episode dialogue without scrubbing audio. Episode transcripts →
Read host commentaries
This page is for you if…
- You want the host's take without listening to the full episode
- You are sharing operational context with your team in writing
- You prefer editorial framing over RSS show-note summaries
- You bookmark links and references from weekly news roundups
More ways to read
Skim transcripts, host commentaries, and show notes across every episode
Cloudflare Saves 100TB of RAM, AI Drives Server Prices Up, AWS Adds a Fourth London AZ, Route 53 DNS Self-Service, AKS eBPF Routing, Go 1.27, and the Danger of Hidden Infrastructure Assumptions
Cloudflare Saves 100TB of RAM, AI Drives Server Prices Up, AWS Adds a Fourth London AZ, Route 53 DNS Self-Service, AKS eBPF Routing, Go 1.27, and the Danger of Hidden Infrastructure Assumptions
The thing I kept coming back to in this episode is how much infrastructure depends on assumptions that stop looking like assumptions after they have been true for long enough.
Cloudflare’s DNS cache is a good example from the performance side. Saving a few bytes in a data structure sounds almost pointless until you multiply it by more than 250 billion cache entries. Then suddenly those bytes turn into roughly one hundred terabytes of RAM. That is the kind of optimization work I really enjoy because it is not flashy. Nobody replaced the entire system or introduced another platform. They looked closely at how the existing thing behaved at scale and found places where small inefficiencies had become enormous.
It is also a useful reminder that “premature optimization” does not mean optimization is bad. It means you need to know where optimization actually matters. There is no reason to spend days shaving bytes from something instantiated a few thousand times. But if the same structure exists billions of times, memory layout, allocations, and cache locality become architecture decisions. At that scale, changing one field can matter more than adding another server.
The OVHcloud story gets at the same idea from a completely different direction. Most teams think about AI infrastructure costs in terms of GPUs, model APIs, or whatever AI workload they are actually running. But the demand is starting to affect the hardware market underneath everything else. If manufacturers shift capacity toward high-bandwidth memory because that is where AI demand is, then a company running normal databases and web servers can still end up paying more for RAM.
That is where FinOps gets more interesting than just finding idle instances. Right-sizing and commitment discounts are still useful, but they cannot completely solve a changing cost basis underneath the service. Sometimes the infrastructure itself becomes more expensive. A database optimization nobody could justify last year might suddenly make sense when the memory underneath it costs fifty percent more.
The AWS London story is probably the cleanest example of a hidden assumption finally getting tested. AWS added a fourth Availability Zone. That should be good news. More capacity and another failure domain are things we usually want.
But if your Terraform discovers every available AZ and another part of the module only has three subnet CIDRs, you have a problem. The code looked dynamic. It just was not dynamic all the way through.
I think that distinction matters a lot in infrastructure code. Dynamic discovery is not the same thing as supporting arbitrary change. Sometimes explicitly saying “this module supports exactly three zones” is safer than pretending everything is flexible. At least then the failure is obvious. The dangerous version is when the input changes and your automation quietly creates an architecture nobody intended.
That same tension shows up in the Route 53 story. Central networking teams usually own DNS for a reason. It is shared infrastructure, and letting every account independently control global resolution would be chaos. But forcing every private hosted-zone association through a networking ticket does not scale either.
The shared DNS view model is a pretty good example of platform engineering when it works well. Centralize the part that actually needs governance. Delegate the routine action the application team can safely perform itself. You are not giving up control. You are getting a person out of the middle of something that should not require scheduling.
That is usually what I want from an internal platform. Not a giant abstraction that hides everything. Not a portal for the sake of having a portal. Just guardrails around the decisions that matter and self-service around the ones that do not.
Even the lightning stories fit the same pattern. AKS moving more networking behavior into eBPF can improve performance, but it also changes assumptions around iptables. CloudFront Functions putting custom context directly into access logs removes one of those annoying observability gaps that seems small until you are debugging something under pressure. Go 1.27 adds language and runtime improvements that are easy to ignore until they intersect with something your codebase depends on. And the memory-isolation research is basically a reminder that every security boundary depends on what the layer underneath it is actually doing.
The kubectl closer is probably the most relatable version of all of this.
You have two terminals open. You switch context in one of them. Later you use the other one, assuming it is still pointed where it was before.
The command succeeds.
That is what makes the mistake dangerous.
The system did not fail. Your mental model did.
We spend a lot of time trying to remove friction from engineering workflows, and most of the time that is the right goal. But production is one place where a little friction can be healthy. Showing the current cluster in your prompt, explicitly passing a context for destructive commands, using separate kubeconfigs, or requiring stronger permissions in production may slow you down by a few seconds.
That is a pretty cheap trade compared with deleting something from the wrong cluster.
And the bigger lesson is not really about Kubernetes. It is about where we place safety.
If the only thing preventing a production mistake is an engineer remembering which terminal tab they used twenty minutes ago, the control is too weak. The same applies to the AWS AZ story. If the only thing keeping the module correct is AWS continuing to return exactly three zones forever, the control is too weak.
Good infrastructure engineering is not eliminating assumptions. You cannot do that.
It is deciding which assumptions are important enough to enforce, validate, or design around.
Sometimes that means validating the AZ count.
Sometimes it means putting the cluster name in the command.
Sometimes it means letting application teams manage their own DNS association inside a boundary the platform team owns.
And sometimes it means realizing that eight wasted bytes are not eight bytes anymore when you have 250 billion of them.
Most incidents are not caused by something completely unknowable.
A lot of them happen when reality changes and the system keeps behaving according to an assumption nobody realized was still there.
Scroll inside the box to read the full commentary, or expand for a larger view.
Ship It Conversations: Justin Garrison of Sidero Labs on Kubernetes, Platform Engineering, AI, Golden Paths, and Knowing What to Say No To
Ship It Conversations: Justin Garrison of Sidero Labs on Kubernetes, Platform Engineering, AI, Golden Paths, and Knowing What to Say No To
The thing that stuck with me most from this conversation with Justin Garrison is that mature platforms are defined just as much by what they refuse to support as what they enable.
That sounds backwards at first.
When teams build platforms, the natural instinct is to think about capability. How many workloads can we support? How many teams can use this? How many tools can we integrate? How much can we automate?
But every additional capability creates another responsibility.
Every exception adds complexity. Every supported workload type adds another path to maintain. Every abstraction creates another thing engineers have to understand.
At some point, the platform that was supposed to reduce cognitive load becomes another source of cognitive load.
That is the trap Justin described throughout this conversation.
Platform engineering is not about building the biggest possible platform. It is about creating the right boundaries.
One of the examples that stood out to me was Justin’s experience at Disney Plus, where the platform made a deliberate decision not to support stateful workloads.
At first glance, that sounds limiting.
A platform that says “no databases” or “no stateful applications” might feel like it is making developers’ lives harder. But that constraint created simplicity everywhere else. Cluster upgrades became easier. Ownership became clearer. Operational concerns were pushed to the teams and systems that were actually designed to handle them.
The platform became better because it knew what it was not.
That is a lesson that applies far beyond Kubernetes.
We have a tendency in technology to take something useful and expand its scope until it becomes the answer to every problem.
Kubernetes is a great example.
Kubernetes solved real problems. It gave teams a consistent API for deploying workloads, managing containers, and automating parts of infrastructure operations. It became the foundation for modern platform engineering for good reasons.
But somewhere along the way, Kubernetes also became the default answer for almost everything.
Need to run an application? Kubernetes.
Need internal tooling? Kubernetes.
Need batch processing? Kubernetes.
Need AI workloads? Kubernetes.
Need a simple service that could run on a VM for five years without anyone touching it? Probably Kubernetes.
Sometimes that is the right answer.
Sometimes it is just the tool we know.
Justin brought up a great comparison with Ansible. People still use Ansible for things it was never really designed to do because it became familiar. It was the tool they trusted, so every problem started looking like an Ansible problem.
The same thing happens with every successful technology.
We do not just adopt tools. We build identity around them.
That is why the conversation around Kubernetes, platform engineering, and even AI is so interesting right now. The challenge is not learning another tool. The challenge is knowing when the tool is actually helping.
The discussion around cloud versus on-prem infrastructure was another area where I thought Justin brought some useful nuance.
For years, the cloud conversation was often framed as if moving to cloud was automatically the mature decision. The assumption was that companies moved from servers to virtual machines, from virtual machines to cloud, and eventually everything became cloud native.
But reality is more complicated.
Cloud solved a lot of problems. It gave teams access to infrastructure faster, reduced the need for massive upfront investment, and created incredible flexibility.
But cloud also introduced new problems.
Cost management became a discipline of its own. FinOps teams exist because cloud spending can become complicated quickly. Autoscaling sounds simple until you realize scaling infrastructure often means scaling databases, moving data, managing networking, and handling operational complexity.
At some point, some companies look at their environments and decide that more control, predictability, or sovereignty matters more than unlimited flexibility.
That does not mean cloud failed.
It means architecture is about tradeoffs.
The right answer depends on the workload, the business, the regulatory environment, and the operational capabilities of the team.
That same idea applies to Kubernetes distributions.
Justin’s experience working on EKS Anywhere was interesting because it highlighted something people sometimes overlook. “Kubernetes anywhere” does not necessarily mean the same thing everywhere.
Managed Kubernetes offerings are optimized for their environments. Running Kubernetes on-prem is a different problem. Running Kubernetes at the edge is a different problem.
The API might look familiar, but the operational model underneath can be completely different.
That matters because portability is often oversold.
You can move Kubernetes workloads between environments, but the deeper assumptions around networking, storage, identity, observability, security, and operations do not magically disappear.
The YAML might move.
The architecture usually does not.
That leads directly into the platform engineering discussion.
Justin’s advice for platform teams was simple but probably one of the most important points of the conversation.
Figure out what you are willing to say no to.
That is the part many teams skip.
They start with what they want to build. A developer portal. A golden path. Self-service deployments. Internal tooling. Kubernetes clusters.
But before building any of that, the harder question is: what problem are we solving?
Are developers struggling because they do not have enough automation?
Or are they struggling because approvals take two weeks?
Is the deployment process slow because there is no platform?
Or because nobody owns the testing strategy?
Is Kubernetes complexity the problem?
Or did the organization choose a tool that created complexity it did not actually need?
The platform is not the outcome.
The outcome is enabling teams to deliver software safely and effectively.
The same thinking applies to centralization.
There are some things that make sense to centralize.
Security controls. Compliance visibility. Logging. Monitoring. Cost visibility. Supply chain security.
Those are areas where consistency and visibility provide real value.
But centralizing every developer workflow, every deployment pattern, and every engineering decision can create a bottleneck.
A platform should remove friction.
It should not become the new approval queue.
The AI conversation followed a similar pattern.
One of the interesting things about AI right now is that it is not introducing an entirely new human behavior. It is accelerating a very old one.
We have always become dependent on tools.
People built strong opinions around Vim versus Emacs. Developers became attached to languages and frameworks. Teams built entire processes around CI systems and deployment platforms.
AI is another version of that relationship, except the speed of change is much faster.
Six months ago, one tool was the answer. Today, another tool is the answer. Six months from now, something else will probably replace both.
That creates a challenge.
The risk is not using AI.
The risk is forgetting how to think without it.
Justin’s point about fundamentals really resonated here. The engineers who will benefit the most from AI are not the ones who know how to generate the most code.
They are the ones who know when the generated code is wrong.
A Kubernetes manifest can apply successfully and still be a terrible architecture.
A Terraform plan can complete successfully and still create a disaster waiting to happen.
A generated application can compile perfectly and still have the wrong security model, scalability assumptions, or operational behavior.
The syntax is becoming cheaper.
The judgment is not.
That changes what experienced engineers bring to the table.
The value moves higher.
Architecture. Tradeoffs. Failure modes. Security boundaries. Understanding systems deeply enough to recognize when something looks correct but is actually dangerous.
That is the part AI cannot replace easily because it requires context.
And that connects to Justin’s final advice about finding the parts of engineering you enjoy.
Technology changes constantly.
The tools that define our careers today might be completely different a decade from now. Kubernetes replaced other platforms. Cloud replaced other infrastructure models. AI is changing how we build software.
But curiosity lasts.
The engineers who stay successful are usually the ones who keep asking questions.
Why does this work?
What happens when it fails?
What assumption is this abstraction hiding?
What problem are we actually solving?
Those questions matter more than any specific tool.
Because every new technology eventually becomes someone else’s old technology.
The fundamentals are what carry forward.
Additional Links
Justin Garrison:
https://justingarrison.com
Justin Garrison on BlueSky:
https://bsky.app/profile/justingarrison.com
Sidero Labs:
https://www.siderolabs.com
Talos Linux:
https://www.talos.dev
Fork Around and Find Out podcast:
https://forkaroundandfindout.com
Kubernetes:
https://kubernetes.io
Amazon EKS:
https://aws.amazon.com/eks/
Amazon EKS Anywhere:
https://anywhere.eks.amazonaws.com
Kubernetes Cluster API:
https://cluster-api.sigs.k8s.io
Kubernetes The Hard Way:
https://github.com/kelseyhightower/kubernetes-the-hard-way
Google SRE Book:
https://sre.google/sre-book/table-of-contents/
FinOps Foundation:
https://www.finops.org
Scroll inside the box to read the full commentary, or expand for a larger view.
GitHub Outage, PleaseFix Agentic Browser Vulnerability, AWS Certificate Manager Drops Email Validation, Cloudflare TypeScript CI Workflows, AI Observability Consolidation, and the Hidden Cost of “Simple” Platform Changes
GitHub Outage, PleaseFix Agentic Browser Vulnerability, AWS Certificate Manager Drops Email Validation, Cloudflare TypeScript CI Workflows, AI Observability Consolidation, and the Hidden Cost of “Simple” Platform Changes
The theme I kept coming back to this week is how many systems quietly become critical long before we start treating them that way.
GitHub is probably the clearest example. For years it was easy to think of source control as something developers used to write and review code. That is not really what GitHub is anymore. It sits in deployment paths, approval workflows, automation, identity, incident response, release management, and increasingly AI tooling. When it goes down, the impact is not just “developers cannot push.” In some environments, you lose part of your ability to operate production.
The question that matters to me is not whether you can keep shipping during a GitHub outage. Stopping deployments may actually be exactly what you want. The harder question is whether you can recover. If rollback depends on checking out a repository, starting a GitHub Action, or getting an approval through the same service that is unavailable, then your emergency path has the same dependency as your normal path. That is the kind of thing that looks completely reasonable until the day you actually need it.
I do not think the answer is necessarily maintaining a second Git platform and duplicating everything. That can create more complexity than it solves. But keeping known-good artifacts somewhere independent, understanding exactly which operational procedures depend on GitHub, and testing what happens when those dependencies disappear is pretty reasonable. The important part is knowing where the dependency exists instead of discovering it during the incident.
The PleaseFix story is a different kind of dependency problem, but it gets at something I think we are still learning with agentic systems. We keep focusing on whether the model itself can recognize malicious instructions. That matters, but it is not the security boundary I would want to bet everything on.
If an agent can read an untrusted webpage and also has access to privileged tools, credentials, files, or external APIs, then the important question is what sits between those two capabilities. A webpage should be allowed to influence what the agent thinks about. It should not automatically be allowed to influence what the agent is authorized to do.
That sounds like a subtle distinction, but it is really just an old security principle showing up in a new place. Untrusted input should not directly control privileged execution. We already know how to think about that in shells, web applications, CI systems, and APIs. Agentic browsers just make the path less obvious because there is a model in the middle translating one into the other.
I also think “human in the loop” gets treated as more protection than it sometimes provides. If the same agent summarizes what it wants to do, provides the explanation, and then asks you to approve it, the human is not necessarily making an independent decision. The safer architecture is one where dangerous actions are structurally different. Reading a webpage and uploading a file should require different authority. Looking at an issue and merging code should require different authority. The model should not be the component defining where that line sits.
The AWS Certificate Manager change is much less exciting, but it may be the most operationally familiar story in the episode. Certificates are one of those things that feel solved until one expires. The expiration date was always known. The certificate was always discoverable. The outage still happens because ownership, renewal, monitoring, or automation was not as clear as everyone assumed.
Moving away from email validation is a good default because email introduces a human process into something that is much better handled as infrastructure. Mailboxes disappear. People change roles. Distribution lists get forgotten. DNS validation is not magically perfect, but it gives teams a much more durable automation path.
The bigger lesson is that certificate inventory needs to be tied to endpoints and ownership. Knowing that a certificate exists is not enough. You need to know where it is actually being served, who is responsible for renewing it, what system performs that renewal, and how you know when that process stops working. The certificate itself usually is not the surprising part. The surprising part is discovering the one forgotten endpoint that uses a completely different renewal path.
Cloudflare’s TypeScript CI work is interesting because it pushes in the opposite direction. Instead of taking something complicated and making it more constrained, it takes CI and makes it more programmable.
I can see the appeal immediately. Types, functions, libraries, tests, reuse, normal programming constructs. Anyone who has maintained a giant YAML pipeline has probably had the thought that this would be easier if it were just code.
But code is not automatically simpler. We have spent years proving that.
Once pipelines become arbitrary software, they inherit software problems. Dependency management, abstraction layers, shared libraries, version compatibility, testing, security review, and eventually some internal framework that only two people completely understand. The interesting question is not YAML versus TypeScript. It is whether we are finally willing to acknowledge that CI/CD has become application software and operate it accordingly.
That means ownership. Tests. Observability. Release discipline. Documentation. And probably a willingness to delete clever abstractions when they start making the system harder to understand than the problem they were supposed to solve.
Even the lightning stories fit this broader pattern. Dynatrace buying Arize shows AI observability getting absorbed into the normal observability stack. Dogwood is another attempt to put policy between an agent and the tools it can use. Pulumi is adding stronger credential protection around infrastructure configuration. And the AWS compromise detected through egress costs is a reminder that useful operational signals do not always come from the security product.
Sometimes the first indication that something is wrong is the bill.
That is why I like having FinOps, security, SRE, and platform engineering increasingly overlap. They are all looking at different symptoms of the same systems. A cost anomaly might be a deployment mistake. It might be a runaway workload. It might be credential abuse. The more those teams can share signals instead of treating them as separate domains, the faster somebody is likely to notice that the system is behaving differently than expected.
If I had to boil this episode down to one thing, it would be that boundaries and dependencies both need to be explicit.
Know which systems your recovery path depends on.
Know what authority an agent actually has.
Know who owns the certificate.
Know whether your CI pipeline is configuration or software.
And know which signals might tell you something is wrong before the obvious alarm fires.
The things that cause the biggest incidents are often not mysterious.
They are usually the dependencies everybody knew existed, but nobody realized had become critical.
Scroll inside the box to read the full commentary, or expand for a larger view.
Ship It Conversations: Ned Bellavance of Ned in the Cloud on DevOps Beyond the Buzzwords, Terraform, AI, the Future of Infrastructure as Code, and Why Fundamentals Still Matter
Ship It Conversations: Ned Bellavance of Ned in the Cloud on DevOps Beyond the Buzzwords, Terraform, AI, the Future of Infrastructure as Code, and Why Fundamentals Still Matter
The thing that stuck with me most from this conversation is how easy it is for a team to look mature without actually being mature.
You can have standups, sprints, CI/CD, Terraform everywhere, a platform team, an internal developer portal, automated security scans, dashboards, SLOs, and a giant pile of tooling. None of that automatically means you are doing DevOps well.
That was one of the first things Ned got into, and I think it is a useful distinction because DevOps has accumulated so much ceremony around it that sometimes the ceremony becomes the goal. Teams start asking whether they have the right meetings, the right tools, the right titles, or the right platform. The better question is whether any of it actually improved how the organization works.
Are developers and operations communicating better? Is feedback getting back to engineers faster? Can you ship a change safely without waiting a week for five different handoffs? Can people understand what happened after something reaches production? Are you fixing the problems that actually matter to the business?
That is much harder to measure than whether everybody attended standup.
It connects to something I see constantly in platform engineering too. It is really easy to start building the platform before deciding what problem the platform is supposed to solve. You can spend six months creating beautiful abstractions, golden paths, self-service workflows, templates, and automation, then discover the thing slowing teams down was an approval process nobody challenged, a flaky test suite, an API rate limit, bad IAM boundaries, or a deployment process that required somebody to manually click three buttons.
The platform is not the outcome. The pipeline is not the outcome. Terraform is not the outcome. They are tools we use to get somewhere else.
I liked Ned's framing here: start by being honest about where you actually are. Then figure out where you are trying to go. Then measure whether the work you are doing is moving you in that direction.
That applies to reliability too. I brought up the example of a company wanting 100 percent uptime because that sounds great as an executive goal. Of course we want 100 percent uptime. But what does uptime mean?
Does the homepage respond? Can the user authenticate? Can the application talk to the database? Can the customer complete the transaction that actually makes the business money?
Those are very different things. You can cache a webpage and proudly show a green uptime dashboard while the application behind it is completely unusable. The number means nothing until you define what the number represents.
Security works the same way. You can say security is the priority, but the actual risks depend on the system. Are you handling PII? Do you have regulatory requirements? Are you exposing public APIs? Are you running AI agents with access to developer credentials?
That is why I liked the part of the conversation where Ned talked about basic security hygiene. There is always a new scary vulnerability, package compromise, supply chain attack, prompt injection technique, or headline about some model doing something terrifying. Those things matter, but there is also a reason the same basic classes of mistakes keep showing up year after year.
People still over-permission credentials. People still install packages without really checking what they are. People still expose things that should not be public. People still reuse secrets. People still grant applications far more access than they actually need.
The shiny new attack gets attention. The boring old mistake gets exploited.
The OpenClaw discussion was a good example of that. Giving an autonomous system access to your email, GitHub, cloud accounts, filesystem, browser sessions, or developer credentials should immediately change how you think about permissions.
The interesting part of Ned's setup was not that he was running an AI agent. It was that he was treating the agent like an untrusted automation system. Separate VM. Separate credentials. Read-only access where possible. Gradually expanding permissions instead of handing it everything on day one.
That is just least privilege.
There is nothing particularly AI-specific about the principle. What AI changes is the speed and autonomy of the thing holding those permissions. If I accidentally give a script too much access, the script can do whatever its code explicitly tells it to do. If I give an agent too much access, I am giving a probabilistic system a collection of tools and asking it to figure out how to accomplish a goal.
That should probably make us more careful about permissions, not less.
Then the conversation moved into infrastructure as code, which I think was probably my favorite part.
Terraform has been incredibly successful for a reason. Declarative infrastructure was a huge improvement over giant procedural scripts that had to manually check whether every resource existed before deciding what to do next. You describe the desired state, Terraform builds a graph, compares what you declared with what exists, and figures out what needs to change.
That model has worked really well. But successful abstractions eventually run into the edges of the assumptions they were built around.
State is one of those edges.
Anyone who has worked with a sufficiently large Terraform estate has eventually had the conversation about how many resources belong in a state file. Too large and plans become painfully slow. Too small and you create dependency and orchestration problems between dozens or hundreds of states.
Then you start introducing wrappers, dependency graphs, CI orchestration, remote state lookups, generated configuration, and conventions about where everything lives. It works, but you can feel the complexity accumulating around the original abstraction.
The GitHub example we talked about is a good one. Terraform might only need to change one branch protection rule, but before it can confidently decide what needs to change, it may need to refresh a huge number of repositories and related resources. Those reads count against API rate limits.
Suddenly the infrastructure problem is not actually creating the resource. The infrastructure problem is figuring out what already exists without exhausting somebody else's API.
That suggests the next generation of infrastructure tooling may not simply be Terraform with nicer syntax. The state and reconciliation model itself may evolve.
Ned talked about Swamp from System Initiative, and I think the interesting part is not whether Swamp specifically becomes the thing everybody uses. Nobody knows that yet. The interesting part is the model.
Terraform providers largely operate around CRUD-style resource lifecycle operations. Create, read, update, delete, maybe list. That maps nicely to provisioning infrastructure.
But operations work is bigger than provisioning.
Restart this virtual machine. Back up this database. Rotate this credential. Update these tags. Run this synchronization. Perform this one operation against this one part of the resource without pretending the resource itself needs to be recreated.
Those are normal operational tasks, but they do not always fit neatly into the original declarative resource lifecycle. That is where I think the next few years of infrastructure tooling get interesting.
AI makes it much easier to create new interfaces around infrastructure because generating the glue code suddenly becomes cheap. And that led into what I think was the most important point Ned made in the entire conversation.
Writing code is getting cheaper.
Understanding what the code should do is not.
That distinction matters a lot.
Terraform syntax used to be a meaningful part of the skill. You had to understand HCL, modules, expressions, loops, dependencies, providers, data sources, state, and all the weird edges of the language. Those skills still matter today, but an LLM can write a pretty decent Terraform module in seconds.
The same thing is happening with Python, TypeScript, Bash, Kubernetes manifests, GitHub Actions, Helm charts, CloudFormation, and Ansible. The cost of producing syntactically plausible infrastructure code has collapsed.
But somebody still needs to know whether that infrastructure makes sense.
Should this workload be a Lambda? Should it be Kubernetes? Should it be a virtual machine? Does it need a load balancer? Should this database be public? What network paths should exist? Where should secrets live? What should happen if the region fails? What data can be lost? How much availability does the business actually need? What happens when traffic increases ten times?
AI can suggest answers, but the person reviewing those answers needs enough context to know whether they are reasonable.
That is why I do not think AI makes fundamentals less valuable. I think it makes them more valuable.
If AI handles more of the syntax, human value moves upward into architecture, constraints, tradeoffs, security boundaries, failure modes, troubleshooting, and understanding how systems interact.
Knowing when the generated answer is technically valid but operationally stupid is going to matter a lot.
Generated infrastructure can look incredibly convincing. Everything can validate. The Terraform plan can look fine. The Kubernetes manifest can apply successfully. The pipeline can turn green.
And the architecture can still be wrong.
This also changes how I think people should learn DevOps and cloud engineering. There is a temptation right now to skip directly to prompting. Why learn Terraform deeply if Claude can write Terraform? Why learn Kubernetes if an agent can create the manifest? Why learn networking if AI can tell you which security group rule you need?
Because eventually something breaks.
And the clean abstraction disappears.
The SRE job is not just creating the thing. It is understanding the thing when the assumptions stop being true.
That is why Ned's example about making a network cable actually made sense to me. You probably do not need to make your own Ethernet cables anymore. That does not mean there was no value in understanding what was inside the cable.
The same thing happens when you build Kubernetes the hard way. Nobody should manually build every production Kubernetes cluster from individual binaries and certificates. That would be absurd. But doing it once teaches you what kubelet is, what certificates are doing, where control plane communication happens, how networking fits together, and which components exist underneath the managed abstraction.
Then when something breaks at 3 a.m., you are not looking at Kubernetes as one giant magic box that stopped working. You have some idea of what could actually be broken inside the box.
That is why the learning discussion near the end of the episode resonated with me. Sometimes the best way to learn something is to do it wrong.
Not intentionally wreck production. Not create chaos just for the sake of chaos. But build something in a place where failure is cheap. Try it. Break it. Misconfigure it. Read the error. Figure out why it did not work. Fix it. Do it again.
That debugging process creates a completely different kind of understanding than following a perfectly scripted tutorial.
Tutorials are great for getting started. They are terrible at teaching you what happens when step seven does not work.
Production is mostly step seven not working.
Experienced engineers sometimes forget how much of their judgment came from those failures. You remember the certificate issue because you spent four hours figuring out why TLS was broken. You remember the networking problem because you accidentally configured a route that black-holed traffic. You remember the IAM permission because the application worked everywhere except production and you eventually realized one role was missing a single action.
Those experiences become intuition.
You start recognizing the shape of failures before you fully understand them. Something feels like DNS. Something feels like permissions. Something feels like stale state. Something feels like a dependency timing issue.
That intuition is difficult to teach and difficult to speed-run.
AI can help you investigate it. It can search logs faster, explain an error message, suggest likely causes, and generate commands. That is useful. But there is still value in the engineer understanding why one hypothesis is much more likely than another.
That is the judgment we should be trying to preserve and teach.
I also liked Ned's three personal principles: make yourself uncomfortable, be kind, and be prepared to fail.
Those are surprisingly good engineering principles.
Make yourself uncomfortable means keep learning things slightly outside what you already know. Be kind matters because incidents and technical disagreements involve humans, and everybody is trying to solve the same problem under some amount of pressure. Being prepared to fail matters because failure is unavoidable if you are actually experimenting and building things.
I joked that I would be suspicious of somebody who had been a senior engineer for ten years and claimed they had never taken down production.
I still mostly believe that.
Not because taking down production is some badge of honor. It absolutely is not. But if you have spent enough years making meaningful changes to complex systems, eventually you will get something wrong.
The more important question is what happens next.
Do you hide it? Do you blame somebody? Do you panic? Or do you work the incident, understand what happened, restore the system, and make the next failure less likely?
That is where the experience comes from.
And that connects perfectly to Ned's Day 2 DevOps framing.
Day one is easy to make look good. The demo works. The POC works. The deployment succeeds. Everybody celebrates.
Day two is when reality arrives.
Users show up. Traffic increases. Certificates expire. Dependencies change. Security requirements change. Costs grow. Someone upgrades a library. A region has problems. The API behaves differently than you expected. The person who built the original system leaves the company.
Now somebody has to operate it.
That is when you find out whether the architecture was actually good.
Honestly, that might be the thread connecting this entire conversation.
DevOps ceremony looks good on day one. Platform demos look good on day one. Terraform looks good when the state is small. AI-generated infrastructure looks good when the code validates. New tools look great in a POC.
The real test is what happens after the novelty wears off.
Can we operate it? Can we troubleshoot it? Can we change it? Can someone other than the original author understand it? Does it actually improve the outcome we cared about? What happens when it breaks?
That is the work.
So my takeaway from this episode is pretty simple. Do not confuse having DevOps tooling with doing DevOps well. Do not confuse infrastructure as code with good architecture. Do not confuse AI-generated code with understanding the system. And do not optimize so hard for avoiding failure that nobody gets the experience required to handle failure when it eventually happens.
Use the new tools. Experiment with AI. Try the new infrastructure models. Let AI write the boring code. That part is useful.
But keep learning networking. Keep learning Linux. Keep learning security. Keep learning databases. Keep understanding how cloud services actually fit together. Keep building things. Keep troubleshooting them. And every once in a while, build something the hard way just so you understand what the easy way is hiding.
Because the syntax is getting cheaper.
The judgment is not.
Additional Links
Ned in the Cloud: https://nedinthecloud.com
Day 2 DevOps: https://day2devops.com
Ned in the Cloud on YouTube: https://www.youtube.com/c/NedintheCloud/
Ned Bellavance on LinkedIn: https://www.linkedin.com/in/ned-bellavance/
Swamp: https://swamp.club
Terraform: https://developer.hashicorp.com/terraform
OpenTofu: https://opentofu.org
Terragrunt: https://terragrunt.gruntwork.io
OWASP Top 10: https://owasp.org/www-project-top-ten/
Kubernetes The Hard Way: https://github.com/kelseyhightower/kubernetes-the-hard-way
Scroll inside the box to read the full commentary, or expand for a larger view.
Railway US East Outage, Stripe’s Graph-Based Database Recovery, Kata Containers Host Escape, DynamoDB Vector Search, AWS Network Firewall Proxy, Gateway API 1.6, and containerd 2.4
Railway US East Outage, Stripe’s Graph-Based Database Recovery, Kata Containers Host Escape, DynamoDB Vector Search, AWS Network Firewall Proxy, Gateway API 1.6, and containerd 2.4
The thing that stuck with me most from this episode is that fixing the thing that failed is not always the same as recovering the system.
Railway is a really good example of that. The original problem was understandable enough: an upstream network issue, a routing change, and suddenly the site had lost its last default route. But even after the route came back, the system was not really back. Connections had already moved onto a management network that was never supposed to carry that traffic. Some of them stayed there. Other private-network connections ended up blackholed. The triggering failure was gone, but the state created by the failure was still hanging around.
I think we sometimes underestimate how much state exists outside the thing we are actively repairing. We fix a route and assume traffic will normalize. We restore a database and assume clients will reconnect cleanly. We recover a dependency and assume retries will settle down. But connections, caches, sessions, queues, circuit breakers, DNS, and failover paths may all have reacted to the outage. Recovery has to include those reactions too.
That is probably the broader lesson here. Incident response cannot stop at “the dashboard is green again.” You have to ask whether the system actually returned to its intended steady state. Are clients still pinned to a degraded path? Did a temporary fallback quietly become permanent? Is a retry storm still pushing load somewhere unexpected? Did anything make a decision during the incident that it will not automatically undo?
The Stripe story approached the same problem from another direction, and I really liked it because it was an automation story without immediately becoming an AI story. Stripe modeled database recovery as a graph. The current condition is a state, remediation steps move you between states, and software can calculate a valid path back toward health. That is a very different idea from simply giving an agent production credentials and telling it to fix things.
There is a lot of useful space between a manual runbook and fully autonomous remediation. State machines, policy engines, dependency graphs, health models, and constrained automation are not as exciting to talk about as an AI agent running your infrastructure, but they give you something incredibly valuable: boundaries. You can define which transitions are allowed, test them, understand why a decision was made, and keep the recovery process explainable. Stripe says that approach cut database pages by roughly 30 percent. That is automation doing exactly what I want automation to do: remove repetitive toil without making the system harder to understand.
The Kata Containers vulnerability is another reminder that boundaries only mean something if you understand what actually crosses them. Kata gives workloads a much stronger isolation model by putting them inside lightweight virtual machines. That is useful. But the host and guest still have to communicate somehow. Filesystems, virtio devices, runtime components, and other interfaces become part of the trusted surface. In this case, a bug in that boundary could let guest root reach host root.
We use words like sandbox, isolated, private, and secure very casually in infrastructure. Those words are really shorthand for an architecture. A sandbox is only as strong as the interfaces leading out of it. A private network is only private based on the routes and controls around it. A container is isolated according to a collection of kernel, runtime, filesystem, and device boundaries. The useful question is rarely “is this isolated?” It is “what still crosses the isolation boundary, and what happens if that component is compromised?”
DynamoDB adding vector search is almost the opposite kind of story, but I think it fits the episode surprisingly well. Sometimes reliability comes from adding stronger boundaries. Other times it comes from deleting unnecessary ones. If DynamoDB already stores the application data and can now handle the vector workload you need, maybe you do not need another database, another synchronization process, another backup policy, another set of credentials, and another thing for somebody to understand at three in the morning.
There is always a temptation in platform engineering to solve a new requirement with a new box on the architecture diagram. Sometimes that is absolutely the right answer. Specialized systems exist for a reason. But every additional component creates operational surface area. The best architecture is not necessarily the one with the most purpose-built services. Sometimes it is the one where you can safely remove three arrows and a database.
And the human closer connects to all of this more than it might seem. Tutorials teach you what a system looks like when every assumption is correct. Engineering starts when one of those assumptions is wrong. The version is different. The collector runs but sends nothing. The permissions look right but are not. The documentation describes an older release. That is when you stop following instructions and start reasoning about the system.
That ability to reason is what connects every story this week. Railway had to reason about the state left behind after the route was repaired. Stripe encoded reasoning about recovery paths into software. Kata reminds us to reason about where isolation really begins and ends. DynamoDB forces the architecture question of whether another service actually buys enough value to justify its operational cost.
Tools matter. Runbooks matter. Automation matters. But the thing that keeps showing up in production is judgment.
Getting the obvious failure to disappear is one step.
Understanding what the system became while it was failing is usually the harder one.
Scroll inside the box to read the full commentary, or expand for a larger view.