AWS is retiring DevOps Guru and pointing customers toward a new generation of AI-assisted operations. A Kubernetes controller vulnerability can let namespace-scoped permissions create a pod somewhere they should not be able to. And Cloudflare says AI agents now account for almost half of the usage of its Wrangler CLI. I’m Brian Teller from Teller's Tech, and this is Ship It Weekly.
Welcome back to Ship It Weekly, the show about the DevOps, SRE, cloud, platform, and security stories that matter when you are the person keeping the thing running at three in the morning. For the weekly story list and source links, check out OnCallBrief.com. For past episodes and show notes, head over to ShipItWeekly.fm. This is episode 69. Nice. Anyway, quick programming note before we get into it.
I’ll be at AWS headquarters in Arlington on Friday for AWS DMV Community Day. I’m speaking and recording some interviews while I’m there, so you’ll see some of those conversations coming to the show soon. This week, AWS is officially putting an end date on DevOps Guru and recommending customers move toward CloudWatch and the newer Amazon DevOps Agent.
Kubernetes disclosed a controller vulnerability that can cross a namespace boundary under some specific permissions. We have a Node.js and Undici WebSocket issue where a malicious server can crash the client process. And Cloudflare has launched a new CLI as AI agents become major consumers of its developer tooling.
Then we have a quick lightning round and a human closer from SRE Weekly about why some categories of production failure may never go away. Let's get into it. First up, AWS is retiring Amazon DevOps Guru. The service will stop accepting new customers on October 29, 2026, and reach end of support on September 30, 2027. Existing customers can continue using it until then, but there will be no new features.
After end of support, the console and APIs go away and previously generated insights will no longer be retrievable. DevOps Guru was AWS’s earlier generation of AI for operations. It used machine learning to analyze operational data, identify abnormal behavior, correlate signals, and help engineers investigate problems. AWS is now telling customers to split those capabilities into two directions.
For monitoring, alerting, and anomaly detection, move toward CloudWatch. For AI-assisted investigation, AWS recommends looking at Amazon DevOps Agent. So if you use DevOps Guru today, this is more than swapping one AWS service for another. You need to figure out which capabilities you are actually using and where each one moves. There is also an infrastructure-as-code problem to deal with.
If you have DevOps Guru resources declared in CloudFormation, CDK, Terraform, or automation built around its API, you cannot leave those definitions sitting around after the service disappears. AWS specifically warns that CloudFormation operations involving withdrawn resource types can fail after end of support. You have a year, so this is not an emergency. But managed services are still dependencies.
You may not maintain the software underneath them, but you still build infrastructure, automation, and operational processes around their APIs. Eventually, some of those APIs disappear. Next, Kubernetes disclosed CVE-2026-2270. This is a vulnerability involving the StatefulSet controller and ControllerRevisions.
Under specific conditions, somebody with namespace-scoped write permissions for StatefulSets and ControllerRevisions can cause the controller to create a pod in another namespace. That sounds particularly bad because namespace-scoped permissions are supposed to stay inside the namespace. There are some important limitations.
The resulting pod will normally be deleted by Kubernetes garbage collection unless the attacker can construct a valid StatefulSet owner reference. To do that, they need the UID of an existing StatefulSet in the target namespace. The vulnerability is rated Medium with a CVSS score of 5.9 and requires substantial existing permissions. So this is not somebody outside the cluster suddenly getting arbitrary pod creation.
Kubernetes describes it as a confused deputy attack. The user does not directly have permission to create the pod in the target namespace. The controller does. The user manipulates an object they are allowed to control, and the controller performs the more privileged operation for them. That is worth considering when reviewing Kubernetes RBAC. Permissions on an object do not always end with that object.
Controllers react to resources and perform additional operations using their own authority. Sometimes, the important question is not just, “What can this identity create?” It is, “What can this identity convince a controller to create?” Third, we have a denial-of-service vulnerability in Undici.
Undici is the HTTP client used by Node.js, and its WebSocket implementation backs Node's built -in globalThis.WebSocket. CVE-2026-85024 involves WebSocket compression. A malicious or compromised WebSocket server can send a specially constructed compressed message that exceeds the decompressed payload limit and contains malformed DEFLATE data.
Inside Undici, cleanup for the size limit removes an error listener from an internal InflateRaw stream. That stream can still generate an error. When it does, Node sees an unhandled error event and terminates the process. Adding an error handler to your application's WebSocket does not fix this.
The failure occurs inside the decompression object, so the public WebSocket error and close handlers never get an opportunity to catch it. The vulnerability is rated Medium with a CVSS score of 5.9. The advisory says that roughly 130 kilobytes sent over the network can expand beyond the internal limit and terminate the process. If the application automatically reconnects, it could potentially end up in a crash loop.
Affected Undici branches have fixes in 6.28.1, 7.29.1, and 8.10.2. The operational point here is pretty simple. Your application, being the client, does not make the server trusted. Clients still parse, decompress, deserialize, and process data coming from the other side. If you use WebSockets from Node, especially against endpoints outside your control, check which version of Undici you are running.
Fourth, Cloudflare has launched a new CLI called `cf`. Normally, another CLI would not make the main story list. The numbers behind this one are why it did. Cloudflare says AI agents accounted for about 25 percent of Wrangler usage in March. Last week, that reached 48 percent. Agents also use almost twice as many distinct commands per day and are nearly four times as likely to use six or more commands.
Wrangler was primarily built for people. It exposes around 280 operations, while Cloudflare has thousands of API operations. Instead of manually adding commands for all of those, Cloudflare built `cf` around its API definitions. Cloudflare also open-sourced Forge, the generation system behind the tooling. We have spent a lot of time talking about agents using infrastructure tools.
Now we are starting to see infrastructure tools change because agents are using them. If almost half of your CLI usage is coming from software, things like consistent commands, structured output, discoverability, and complete API coverage start becoming more important. The CLI is no longer only an interface for the engineer sitting at the terminal.
It is also becoming an interface for software operating on the engineer's behalf. Quick lightning round. First, Kubernetes disclosed another vulnerability affecting Windows nodes. CVE-2026-76654 involves a pod volume mount using `subPath` with a symbolic link pointing to an attacker-controlled network share. The Windows kubelet can follow that link to a UNC path and attempt NTLM authentication to the remote share.
That can expose the NetNTLMv2 hash of the account running the kubelet, which could potentially be cracked or relayed if the node is domain joined. It is rated Medium with a CVSS score of 5.8. Second, GitHub now lets repository administrators configure custom runners for Dependabot version and security updates. You can select the runner type, a custom label, and an optional runner group.
That is useful when Dependabot needs access to private package registries or other resources only reachable from your own environment. And third, GitHub has added external custom properties for repositories. An external system of record like a CMDB or internal developer portal, can now push information like ownership, service tier, lifecycle stage, or compliance status into GitHub.
Those values remain read-only in GitHub while the external system stays the source of truth. Because custom properties can also be used with filtering and rulesets. Service catalog metadata can now feed directly into repository governance. The human closer this week comes from Lorin Hochstein and a piece featured in SRE Weekly called Omnipresent Availability Risks in Cloud Software.
The argument is that there are some categories of production incidents we probably are not going to engineer away. Take saturation. Every resource is finite. CPU, memory, disk, database connections, queue depth, network capacity. Eventually, something can hit a limit. Then there is networking. Cloud software is distributed software, and distributed software depends on networks.
Networking failures can also have enormous blast radius. Security creates another conflict. Availability says legitimate users should be able to access something. Security says illegitimate users should not. Sometimes the security control doing exactly what it was designed to do can cause the outage. An expired certificate is probably the simplest example. Then there are changes you cannot simply stop making.
Sometimes you have to migrate the database. Sometimes you have to replace infrastructure. And sometimes production is already broken and making another change is the only way to recover it. Even the systems we build to improve reliability add complexity. Retries, failover, replication, health checks, circuit breakers. They help us recover from failures, but they also create more states and interactions that can fail.
That changes how I think about one of the questions that comes up after almost every incident. How do we make sure this never happens again? For a specific preventable failure, absolutely. Fix it. Add the guardrails. Improve the alert. Remove the dangerous manual step. But we are not going to eliminate finite resources. We are not going to eliminate networks. We are not going to eliminate security boundaries.
And we are definitely not going to eliminate change. Incident response itself has to be part of reliability engineering. How quickly do we notice the problem? How quickly can we understand the blast radius? Can we get the right people involved? Can we safely make changes while the system is already behaving unpredictably? And can we learn from what happened afterward?
Sometimes improving reliability means preventing the next incident. Sometimes it means getting much better at handling the incidents you cannot prevent. That is it for this week's Ship It Weekly. We covered AWS retiring DevOps Guru and pointing customers toward CloudWatch and DevOps Agent. A Kubernetes vulnerability that can cross namespace boundaries. A WebSocket bug that can crash a Node.js process.
And Cloudflare redesigning its CLI as AI agents become a major part of its usage. Plus, another Kubernetes security issue on Windows nodes, custom Dependabot runners, and GitHub pulling service-catalog metadata into repository governance. Follow or subscribe wherever you are watching or listening.
You can find the weekly story list and source links at OnCallBrief .com and past episodes and show notes at ShipItWeekly .fm. I’m Brian Teller from Teller's Tech. Thanks for listening. And remember, you cannot engineer away every failure, but you can get a lot better at what happens next.
Scroll inside the box to read the full transcript, or expand for a larger view.