AWS Retires DevOps Guru: What the End of Support Means, Kubernetes Cross-Namespace CVE-2026-2270, Node.js Undici WebSocket DoS & Cloudflare’s New CLI for AI Agents
Try reloading the player or open this episode directly on YouTube.
This episode of Ship It Weekly discusses AWS's retirement of DevOps Guru, urging users to transition to CloudWatch and Amazon DevOps Agent. It also covers Kubernetes vulnerabilities, a Node.js issue with Undici, and Cloudflare's new CLI for AI agents.
Now Playing
AWS Retires DevOps Guru: What the End of Support Means, Kubernetes Cross-Namespace CVE-2026-2270, Node.js Undici WebSocket DoS & Cloudflare’s New CLI for AI Agents
Ship It Weekly
0:0016:41
Chapters
Jump to a section in this episode.
Speed & share
Transcript
AWS is retiring DevOps Guru and pointing customers
toward a new generation of AI-assisted operations.
A Kubernetes controller vulnerability can let
namespace-scoped permissions create a pod somewhere
they should not be able to. And Cloudflare says
AI agents now account for almost half of the
usage of its Wrangler CLI. I’m Brian Teller from
Teller's Tech, and this is Ship It Weekly. Welcome
back to Ship It Weekly, the show about the DevOps,
SRE, cloud, platform, and security stories that
matter when you are the person keeping the thing
running at three in the morning. For the weekly
story list and source links, check out
OnCallBrief.com. For past episodes and show notes, head
over to ShipItWeekly.fm. This is episode 69.
Nice. Anyway, quick programming note before we
get into it. I’ll be at AWS headquarters in Arlington
on Friday for AWS DMV Community Day. I’m speaking
and recording some interviews while I’m there,
so you’ll see some of those conversations coming
to the show soon. This week, AWS is officially
putting an end date on DevOps Guru and recommending
customers move toward CloudWatch and the newer
Amazon DevOps Agent. Kubernetes disclosed a controller
vulnerability that can cross a namespace boundary
under some specific permissions. We have a Node.js
and Undici WebSocket issue where a malicious
server can crash the client process. And Cloudflare
has launched a new CLI as AI agents become major
consumers of its developer tooling. Then we have
a quick lightning round and a human closer from
SRE Weekly about why some categories of production
failure may never go away. Let's get into it.
First up, AWS is retiring Amazon DevOps Guru.
The service will stop accepting new customers
on October 29, 2026, and reach end of support
on September 30, 2027. Existing customers can
continue using it until then, but there will
be no new features. After end of support, the
console and APIs go away and previously generated
insights will no longer be retrievable. DevOps
Guru was AWS’s earlier generation of AI for operations.
The Operator Is No Longer Just a Person at a Terminal
One thread runs through almost every story in this week’s episode: more and more of the work inside our systems is being performed by something other than the engineer who originally asked for it.
Sometimes that is a managed AWS service. Sometimes it is a Kubernetes controller. Sometimes it is a library buried underneath an application. Increasingly, it is an AI agent interacting with infrastructure tooling.
That changes how I think about operations because the action you request and the action that eventually happens are not always the same thing.
The Kubernetes vulnerability is probably the cleanest example.
CVE-2026-2270 is not interesting because somebody suddenly discovered that Kubernetes RBAC is useless. The attacker still needs meaningful permissions, and exploitation requires a fairly specific set of conditions.
What makes it interesting is the role of the controller.
A user has permission to manipulate one resource. The StatefulSet controller sees that resource and then performs another operation using the controller’s own authority. Under the vulnerable conditions, that can result in a pod being created in another namespace.
The user did not have permission to directly perform that operation.
The controller did.
That distinction matters.
When we review access, we naturally focus on the principal making the request. Can this service account create pods? Can this engineer modify deployments? Can this CI job access secrets?
But modern platforms contain a lot of intermediaries.
Controllers reconcile state. Operators provision resources. CI systems assume roles. GitHub Apps receive permissions. Cloud services react to events. Infrastructure automation creates other infrastructure.
You are often granting somebody permission to ask a more privileged system to do something on their behalf.
That is not inherently bad. It is basically how platforms work.
But it means authorization analysis cannot always stop at the API request.
Sometimes you also need to ask what is going to react to that request afterward.
The Undici vulnerability has a similar shape, although at a completely different layer.
The Node.js application is the client. Normally, that sounds like the safer side of the connection.
We spend enormous amounts of time thinking about protecting servers from hostile clients. Validate the request. Sanitize the payload. Limit the body size. Authenticate the caller.
But clients consume data too.
In this case, a malicious WebSocket server can send compressed data that triggers an error inside Undici’s decompression path. The failure happens below the application's normal WebSocket error handling, and the Node process can terminate.
Your code initiated the connection.
Your process still consumed untrusted input.
That is an easy assumption to miss because developers often think about trust according to direction.
Incoming traffic is dangerous.
Outgoing traffic is something we chose.
Except the fact that your application chose to connect somewhere does not guarantee that the thing on the other end will always behave correctly.
The endpoint could be compromised.
DNS could change.
A third-party service could return malformed data.
Or there could simply be a parser bug in your own dependency.
That does not mean we should treat every outbound connection like an active attacker. It means the trust boundary is not automatically located at the ingress controller.
Sometimes it is inside the HTTP client.
Sometimes it is inside decompression.
Sometimes it is inside a parser nobody on the application team even knew they were running.
The AWS DevOps Guru retirement adds another version of the same problem.
When we consume a managed service, we give up a lot of operational responsibility intentionally.
That is the product.
You do not run DevOps Guru. AWS does.
You integrate with the API, define some resources, consume the output, and build processes around it.
That can make the service feel like infrastructure in the same way S3 or CloudWatch feels like infrastructure. It exists. You use it. You stop thinking very much about the software underneath it.
Until the provider puts an end date on it.
AWS is giving customers plenty of notice here. This is not a shutdown happening next week.
But the migration still illustrates something worth remembering about managed services.
You did not eliminate the lifecycle.
You transferred ownership of part of it.
AWS decides when the service changes. AWS decides when the API stops receiving features. AWS decides when the resource types disappear.
Your side of that contract is dealing with the dependencies you created around it.
That might include CloudFormation resources.
It might include Terraform.
It might include automation consuming insights.
It might include operational processes where somebody expects DevOps Guru to produce a particular signal during an incident.
The migration is not simply:
“Turn off DevOps Guru and turn on DevOps Agent.”
AWS itself is pointing different pieces of the old capability toward different destinations.
Monitoring and anomaly detection may belong in CloudWatch.
AI-assisted investigation may belong in DevOps Agent.
If your team used the service heavily, you first have to understand what job it was actually doing for you.
That sounds obvious, but I have seen enough infrastructure migrations where the first conversation is about replacing a product instead of replacing a capability.
Those are not always the same project.
Then there is Cloudflare.
The number that makes that story worth paying attention to is 48 percent.
Cloudflare says AI agents now account for almost half of Wrangler usage.
That is a much more concrete signal than somebody predicting that agents are eventually going to change software development.
They already changed the usage pattern enough that Cloudflare is changing the tooling.
Wrangler exposes a few hundred operations. Cloudflare has thousands of API operations. Maintaining a manually designed CLI command for everything becomes a very different problem when software is consuming the interface at scale.
So Cloudflare built cf closer to the API itself and open-sourced Forge, the generation system behind it.
I think this is where the agent conversation becomes much more interesting for platform engineers.
The first wave was mostly about putting an agent in front of tools designed for humans.
Let it type shell commands.
Let it call Terraform.
Let it inspect Kubernetes.
Let it operate the same CLI somebody would have used manually.
The next wave is going to involve changing those interfaces because machines use them differently.
Humans care about discoverability in a command hierarchy.
Humans tolerate interactive questions.
Humans read formatted tables.
Humans remember weird command names because they have used the tool for years.
An agent benefits from consistency, complete coverage, predictable schemas, structured output, and fewer special cases.
If a meaningful percentage of the consumers of your developer platform become software agents, that absolutely can influence API and CLI design.
That does not necessarily mean we need a separate “AI version” of every tool.
In fact, the better outcome may be the opposite.
Design interfaces that are predictable enough that both humans and automation can use them without maintaining two completely different worlds.
There is also an authorization question hiding behind all of this.
If an agent is operating infrastructure on behalf of an engineer, whose authority is it using?
What can it do?
What can it discover?
What happens when it chains several individually reasonable operations together?
The Kubernetes story is a useful reminder here.
Security is not always about whether the original actor has direct permission to perform the final action.
Sometimes the important question is what another component will do after receiving the request.
That becomes even more important when the requester is software capable of generating hundreds of those requests very quickly.
The human closer from Lorin Hochstein fits into this better than I initially expected.
His argument is that some categories of availability risk are always going to be present.
Resources are finite.
Networks fail.
Security controls can affect availability.
Production systems have to change.
And the mechanisms we add to improve reliability create additional states of their own.
Controllers are one of those mechanisms.
Managed services are one.
Retries are one.
Failover is one.
Automation is one.
Agents are going to be one too.
We add these things because manually operating everything would be slower, less reliable, and often impossible at the scale we run today.
But every layer capable of making a decision also becomes another place where the system can behave differently than the person at the keyboard expected.
That is why I do not think reliability work eventually converges on a world where enough automation eliminates operational surprises.
The automation gets better.
The platforms get better.
The interfaces get better.
The failure modes change.
Then somebody still gets paged at three in the morning because the system entered a state nobody anticipated.
The job is not to avoid automation or abstraction.
We could not operate modern systems without them.
The job is to understand where authority lives, what can act on our behalf, and how much visibility we have when those actors make a decision.
The operator is no longer just the engineer sitting at the terminal.
It is the controller reconciling the object.
It is the managed service behind the API.
It is the library parsing the response.
It is the CI runner assuming a role.
And increasingly, it is an agent issuing commands for us.
We still own what happens next.
📝 Notes
Show Notes
This week on Ship It Weekly: AWS is retiring Amazon DevOps Guru and pointing customers toward CloudWatch and the newer Amazon DevOps Agent. Kubernetes disclosed a vulnerability where StatefulSet and ControllerRevision permissions can allow cross-namespace pod creation under specific conditions. A vulnerability in Undici can let a malicious WebSocket server crash a Node.js process through compressed data. And Cloudflare launched a new CLI as AI agents grow from 25 percent to 48 percent of Wrangler usage.
The bigger theme this week is how the systems around our infrastructure are changing. Managed cloud services still have lifecycles that eventually become migration work. Kubernetes authorization can depend on what controllers do with the resources users are allowed to manipulate. Applications acting as clients still process untrusted data. And infrastructure tooling is starting to treat AI agents as first-class users rather than humans who happen to automate commands.
In the lightning round: another Kubernetes vulnerability affecting Windows nodes can expose NetNTLMv2 credentials through NTLM coercion. GitHub now supports custom runners for Dependabot version and security updates. And external systems like a CMDB or internal developer portal can push repository properties into GitHub while remaining the source of truth.
And the human closer comes from Lorin Hochstein and SRE Weekly. Some availability risks are probably never going away. Resources are finite, networks fail, security controls can affect availability, and production systems have to change. Preventing individual failures still matters, but incident response is part of reliability engineering too. Sometimes improving reliability means getting better at handling the failures you cannot eliminate.
The Operator Is No Longer Just a Person at a Terminal
One thread runs through almost every story in this week’s episode: more and more of the work inside our systems is being performed by something other than the engineer who originally asked for it.
Sometimes that is a managed AWS service. Sometimes it is a Kubernetes controller. Sometimes it is a library buried underneath an application. Increasingly, it is an AI agent interacting with infrastructure tooling.
That changes how I think about operations because the action you request and the action that eventually happens are not always the same thing.
The Kubernetes vulnerability is probably the cleanest example.
CVE-2026-2270 is not interesting because somebody suddenly discovered that Kubernetes RBAC is useless. The attacker still needs meaningful permissions, and exploitation requires a fairly specific set of conditions.
What makes it interesting is the role of the controller.
A user has permission to manipulate one resource. The StatefulSet controller sees that resource and then performs another operation using the controller’s own authority. Under the vulnerable conditions, that can result in a pod being created in another namespace.
The user did not have permission to directly perform that operation.
The controller did.
That distinction matters.
When we review access, we naturally focus on the principal making the request. Can this service account create pods? Can this engineer modify deployments? Can this CI job access secrets?
But modern platforms contain a lot of intermediaries.
Controllers reconcile state. Operators provision resources. CI systems assume roles. GitHub Apps receive permissions. Cloud services react to events. Infrastructure automation creates other infrastructure.
You are often granting somebody permission to ask a more privileged system to do something on their behalf.
That is not inherently bad. It is basically how platforms work.
But it means authorization analysis cannot always stop at the API request.
Sometimes you also need to ask what is going to react to that request afterward.
The Undici vulnerability has a similar shape, although at a completely different layer.
The Node.js application is the client. Normally, that sounds like the safer side of the connection.
We spend enormous amounts of time thinking about protecting servers from hostile clients. Validate the request. Sanitize the payload. Limit the body size. Authenticate the caller.
But clients consume data too.
In this case, a malicious WebSocket server can send compressed data that triggers an error inside Undici’s decompression path. The failure happens below the application's normal WebSocket error handling, and the Node process can terminate.
Your code initiated the connection.
Your process still consumed untrusted input.
That is an easy assumption to miss because developers often think about trust according to direction.
Incoming traffic is dangerous.
Outgoing traffic is something we chose.
Except the fact that your application chose to connect somewhere does not guarantee that the thing on the other end will always behave correctly.
The endpoint could be compromised.
DNS could change.
A third-party service could return malformed data.
Or there could simply be a parser bug in your own dependency.
That does not mean we should treat every outbound connection like an active attacker. It means the trust boundary is not automatically located at the ingress controller.
Sometimes it is inside the HTTP client.
Sometimes it is inside decompression.
Sometimes it is inside a parser nobody on the application team even knew they were running.
The AWS DevOps Guru retirement adds another version of the same problem.
When we consume a managed service, we give up a lot of operational responsibility intentionally.
That is the product.
You do not run DevOps Guru. AWS does.
You integrate with the API, define some resources, consume the output, and build processes around it.
That can make the service feel like infrastructure in the same way S3 or CloudWatch feels like infrastructure. It exists. You use it. You stop thinking very much about the software underneath it.
Until the provider puts an end date on it.
AWS is giving customers plenty of notice here. This is not a shutdown happening next week.
But the migration still illustrates something worth remembering about managed services.
You did not eliminate the lifecycle.
You transferred ownership of part of it.
AWS decides when the service changes. AWS decides when the API stops receiving features. AWS decides when the resource types disappear.
Your side of that contract is dealing with the dependencies you created around it.
That might include CloudFormation resources.
It might include Terraform.
It might include automation consuming insights.
It might include operational processes where somebody expects DevOps Guru to produce a particular signal during an incident.
The migration is not simply:
“Turn off DevOps Guru and turn on DevOps Agent.”
AWS itself is pointing different pieces of the old capability toward different destinations.
Monitoring and anomaly detection may belong in CloudWatch.
AI-assisted investigation may belong in DevOps Agent.
If your team used the service heavily, you first have to understand what job it was actually doing for you.
That sounds obvious, but I have seen enough infrastructure migrations where the first conversation is about replacing a product instead of replacing a capability.
Those are not always the same project.
Then there is Cloudflare.
The number that makes that story worth paying attention to is 48 percent.
Cloudflare says AI agents now account for almost half of Wrangler usage.
That is a much more concrete signal than somebody predicting that agents are eventually going to change software development.
They already changed the usage pattern enough that Cloudflare is changing the tooling.
Wrangler exposes a few hundred operations. Cloudflare has thousands of API operations. Maintaining a manually designed CLI command for everything becomes a very different problem when software is consuming the interface at scale.
So Cloudflare built
cfcloser to the API itself and open-sourced Forge, the generation system behind it.I think this is where the agent conversation becomes much more interesting for platform engineers.
The first wave was mostly about putting an agent in front of tools designed for humans.
Let it type shell commands.
Let it call Terraform.
Let it inspect Kubernetes.
Let it operate the same CLI somebody would have used manually.
The next wave is going to involve changing those interfaces because machines use them differently.
Humans care about discoverability in a command hierarchy.
Humans tolerate interactive questions.
Humans read formatted tables.
Humans remember weird command names because they have used the tool for years.
An agent benefits from consistency, complete coverage, predictable schemas, structured output, and fewer special cases.
If a meaningful percentage of the consumers of your developer platform become software agents, that absolutely can influence API and CLI design.
That does not necessarily mean we need a separate “AI version” of every tool.
In fact, the better outcome may be the opposite.
Design interfaces that are predictable enough that both humans and automation can use them without maintaining two completely different worlds.
There is also an authorization question hiding behind all of this.
If an agent is operating infrastructure on behalf of an engineer, whose authority is it using?
What can it do?
What can it discover?
What happens when it chains several individually reasonable operations together?
The Kubernetes story is a useful reminder here.
Security is not always about whether the original actor has direct permission to perform the final action.
Sometimes the important question is what another component will do after receiving the request.
That becomes even more important when the requester is software capable of generating hundreds of those requests very quickly.
The human closer from Lorin Hochstein fits into this better than I initially expected.
His argument is that some categories of availability risk are always going to be present.
Resources are finite.
Networks fail.
Security controls can affect availability.
Production systems have to change.
And the mechanisms we add to improve reliability create additional states of their own.
Controllers are one of those mechanisms.
Managed services are one.
Retries are one.
Failover is one.
Automation is one.
Agents are going to be one too.
We add these things because manually operating everything would be slower, less reliable, and often impossible at the scale we run today.
But every layer capable of making a decision also becomes another place where the system can behave differently than the person at the keyboard expected.
That is why I do not think reliability work eventually converges on a world where enough automation eliminates operational surprises.
The automation gets better.
The platforms get better.
The interfaces get better.
The failure modes change.
Then somebody still gets paged at three in the morning because the system entered a state nobody anticipated.
The job is not to avoid automation or abstraction.
We could not operate modern systems without them.
The job is to understand where authority lives, what can act on our behalf, and how much visibility we have when those actors make a decision.
The operator is no longer just the engineer sitting at the terminal.
It is the controller reconciling the object.
It is the managed service behind the API.
It is the library parsing the response.
It is the CI runner assuming a role.
And increasingly, it is an agent issuing commands for us.
We still own what happens next.