On Call Brief – Week of August 30 – September 5, 2026
This week's top stories
1. RHSA-2026:63024: Important: python3.12 security update
- Category: Deep Dive
- What happened: Red Hat has released important security updates for two separate products: python3.12 on RHEL 10.0 Extended Update Support (RHSA-2026:63024) and dracut on RHEL 8.8 Update Services for SAP Solutions and Telecommunications (RHSA-2026:61252). Both advisories are rated as having significant security impact based on Common Vulnerability Scoring System ratings, though specific CVE numbers were not provided in the source summaries. Operators running affected RHEL versions should apply these updates promptly through their standard patching processes, with particular attention to SAP and telecommunications environments for the dracut update. Given the "Important" severity rating on both advisories, these patches should be prioritized in the next maintenance window rather than waiting for routine monthly updates.
- Takeaway: This update may require immediate attention to mitigate potential security risks associated with python3.12 vulnerabilities in production environments.
- Sources: Red Hat Security Advisories (RHSA)
2. Kubernetes: v1.37
- Category: Community
- What happened: Kubernetes v1.37 has been released with 67 enhancements, notably promoting the Metrics API to stable and improving workload identity management through pod certificates. The release processed a record 118 PRs through the API review group using declarative validation improvements. In related security news, a report found that 768 leaked AWS keys remain active with full administrative rights, presenting a critical security risk for operators. SRE teams should plan upgrades to v1.37 to benefit from the stable Metrics API and enhanced identity features, while immediately auditing their AWS environments for any potentially compromised credentials using tools like git-secrets or truffleHog. Organizations should implement automated key rotation policies and review IAM permissions to follow least-privilege principles, as the persistence of leaked credentials with admin rights demonstrates significant exposure to potential compromise.
- Worth reading: The Kubernetes v1.37 release may require operators to review and plan for upgrades, particularly regarding workload identity and metrics management. The AWS key leak highlights a critical security risk; operators should audit their AWS keys and implement regular rotation and monitoring to prevent unauthorized access.
- Sources: DevOps'ish
3. Reduce Traffic Interruptions with Gateway Load Balancer TCP Reset
- Category: Community
- What happened: AWS introduces TCP Reset for Gateway Load Balancer, allowing it to send TCP Reset packets when a target becomes unhealthy. This feature reduces traffic interruptions from minutes to seconds, enabling quicker recovery for applications by establishing new connections to healthy targets. The change is particularly beneficial for time-sensitive workloads, improving the overall user experience during target failures.
- Worth reading: This update can significantly enhance the reliability of applications that depend on Gateway Load Balancer, especially in scenarios where downtime is critical. Operators should consider enabling TCP Reset to minimize connection delays during target failures.
- Source: AWS Networking Blog
4. Anthropic: Attackers Using Infostealers to Hijack Claude Sessions
- Category: Deep Dive
- What happened: Anthropic issued a security warning about attackers using infostealer malware to hijack Claude AI sessions by stealing session tokens and authentication cookies rather than traditional credentials, according to Security Boulevard. SRE teams should audit their authentication mechanisms and consider implementing additional session validation controls for services using Claude APIs. Separately, users in Oracle Cloud's me-riyadh-1 region reported complete blackholing of outbound HTTPS traffic to Anthropic and approximately 50% packet loss to Fastly and Debian mirrors, with the issue appearing to originate from upstream transit providers rather than Oracle's VCN infrastructure, as discussed on Reddit's r/devops. Organizations running services in me-riyadh-1 should verify connectivity to Anthropic services and consider temporary failover to alternate regions if Claude API access is business-critical.
- Takeaway: This incident highlights the evolving threat landscape where session hijacking can lead to unauthorized access and misuse of AI services. Operators should review their security measures to protect against such attacks.
- Sources: Security Boulevard, Reddit r/devops
5. Guildma (Astaroth) Malware Spreads via Brazilian Portuguese Phishing Email
- Category: Deep Dive
- What happened: A recent malware infection involving Guildma (Astaroth) was triggered by a malicious email in Brazilian Portuguese. The malware was geofenced to only activate for users with Brazilian IP addresses and language settings. The infection process involved downloading a zip file containing a Windows shortcut that retrieved a DLL file, which was then used to install the malware. The article provides specific indicators of the infection, including email headers and the malicious link.
- Takeaway: This incident highlights the risks associated with targeted phishing attacks and the importance of language and regional settings in malware distribution. Operators should be vigilant about email security and consider implementing additional filtering for geofenced threats.
- Source: SANS ISC
6. Delays in credit purchases
- Category: Deep Dive
- What happened: An issue caused delays in the availability of newly purchased credits for users with a zero balance, leading to erroneous 'credit balance is too low' errors when accessing the Claude API. The problem was identified and resolved, with monitoring indicating that newly purchased credits are now available as expected.
- Takeaway: This incident may have affected users relying on the Claude API during the specified time, potentially disrupting operations that depend on credit availability.
- Source: Anthropic Status
7. Fetching mails from Microsoft accounts producing errors and preventing email based creation of Jira work items
- Category: Deep Dive
- What happened: Jira experienced issues with fetching emails from Microsoft accounts, leading to failures in email-based creation and updates of Jira work items. Users encountered 500 internal server errors during email fetch attempts. The incident has been resolved, and full recovery across affected Atlassian services has been confirmed.
- Takeaway: This incident could have disrupted workflows relying on email integration for Jira, potentially delaying task creation and updates. Monitoring Microsoft services for ongoing issues is advisable.
- Source: Jira Status
CVE & Security
8. SonicWall warns of actively exploited SMA1000 zero-day flaws
- Category: Security / Patch
- What happened: SonicWall has disclosed two zero-day vulnerabilities in the SMA1000 series appliances that are currently being exploited in the wild for remote code execution attacks, though specific CVE numbers and version details were not provided in the initial warning. Operators running SMA1000 devices should immediately apply available patches from SonicWall and review access logs for signs of compromise. Separately, Ubuntu has released USN-8705-2 addressing an OpenZFS vulnerability where improper authorization checks on certain ioctl operations could allow local attackers to execute pool-administrative commands or access privileged data on Linux systems. Organizations using affected SonicWall appliances should prioritize patching given the active exploitation, while those running OpenZFS on Ubuntu should apply the security update to prevent local privilege escalation attacks.
- Do this Monday: This is a critical security issue as the vulnerabilities can lead to remote code execution, potentially compromising systems. Immediate patching is necessary to protect against exploitation.
- Sources: Bleeping Computer, Ubuntu Security Notices (USN)
9. CISA Adds Seven Exploited Flaws as Attackers Deploy Reverse Shells and Crypto Miners
- Category: Security / Patch
- What happened: CISA has added seven security vulnerabilities to its Known Exploited Vulnerabilities catalog, highlighting their exploitation by attackers. One notable vulnerability is CVE-2026-83548, a critical server-side request forgery flaw in SonicWall SMA 1000 Appliances, which could enable remote unauthenticated access.
- Do this Monday: Organizations using SonicWall SMA 1000 Appliances should prioritize patching CVE-2026-83548 due to its critical nature and potential for exploitation. Awareness of these vulnerabilities is essential for maintaining security posture.
- Source: The Hacker News
10. Researcher Releases FalconFlank PoC Showing Privilege Escalation in CrowdStrike Falcon
- Category: Security / Patch
- What happened: A security researcher has released a proof of concept for a zero-day privilege escalation vulnerability named FalconFlank, which affects CrowdStrike Falcon. The flaw exploits the remediation of malicious macros in the CrowdStrike Falcon Sensor.
- Do this Monday: This vulnerability could allow unauthorized users to escalate privileges within systems protected by CrowdStrike Falcon, potentially leading to significant security breaches - operators should assess their defenses against this exploit.
- Source: The Hacker News
11. ALAS2KERNEL-5.10-2026-130 (important): kernel
- Category: Security / Patch
- What happened: This advisory details multiple CVEs affecting the kernel in Amazon Linux 2, highlighting the importance of applying the security updates to mitigate potential vulnerabilities. The list includes CVE-2024-58240 and numerous others, emphasizing the critical nature of these patches.
- Do this Monday: Failure to apply these kernel updates could expose systems to significant security risks, making it essential for operators to prioritize these patches in their deployment schedules.
- Source: Amazon Linux 2 Security Advisories (ALAS2)
12. PostgreSQL: 2 related updates
- Category: Security / Patch
- What happened: Amazon Linux 2 and Red Hat Enterprise Linux have released important security updates for PostgreSQL addressing multiple vulnerabilities spanning CVE-2026-14662 through CVE-2026-6471, with Red Hat specifically targeting the PostgreSQL 12 module on RHEL 8.6. The vulnerabilities are rated as having significant security impact based on CVSS scores, though specific technical details about the nature of these flaws are not provided in the advisories. Operators running PostgreSQL 12 or 14 on Amazon Linux 2 or RHEL 8.6 should apply the respective updates immediately through their standard package management tools (yum/dnf update postgresql) and plan for service restarts to ensure the patches take effect. The updates address concerns related to both security and stability of affected PostgreSQL installations according to Amazon Linux security advisories and Red Hat security bulletins.
- Do this Monday: Operators using PostgreSQL on Amazon Linux 2 should prioritize applying the security patches to mitigate potential vulnerabilities that could affect production systems.
- Sources: Amazon Linux 2 Security Advisories (ALAS2), Red Hat Security Advisories (RHSA)
13. USN-8689-1: OpenJDK 26 vulnerabilities
- Category: Security / Patch
- What happened: Multiple vulnerabilities have been identified in OpenJDK 26, primarily affecting the JSSE, ImageIO, 2D, and Libraries components. These vulnerabilities include improper user authentication and authorization, which could allow remote attackers to read or modify sensitive data or cause denial of service. Specific CVEs include CVE-2026-46968, CVE-2026-46917, CVE-2026-47010, CVE-2026-47021, CVE-2026-47059, CVE-2026-47027, CVE-2026-60147, and CVE-2026-47063.
- Do this Monday: These vulnerabilities could lead to significant security risks, including unauthorized access to sensitive data and potential service disruptions. Immediate patching of OpenJDK 26 is recommended to mitigate these risks.
- Source: Ubuntu Security Notices (USN)
14. Omarchy: Any User Process Can Escalate to Root
- Category: Security / Patch
- What happened: The article discusses a vulnerability named Omarchy that allows any user process to escalate privileges to root. This could lead to significant security risks if exploited, as it undermines the fundamental security model of user permissions. The implications of such a vulnerability are critical for system integrity and security.
- Do this Monday: This vulnerability poses a severe risk to system security, potentially allowing unauthorized access to root privileges. Operators should assess their systems for exposure and apply necessary mitigations immediately.
- Source: 0Xcc via Lobsters
- Discussion: https://lobste.rs/s/bxihn3/omarchy_any_user_process_can_escalate
15. Grafana 13.2.1
- Category: Security / Patch
- What happened: Grafana version 13.2.1 has been released with security fixes for CVE-2026-12704 and CVE-2026-14199, along with bug fixes for dashboards, packaging, and the PanelEditor component. Red Hat has issued RHSA-2026:62407 addressing these same vulnerabilities in their packaged Grafana for RHEL 8, rating the update as Important severity. SRE teams running Grafana should upgrade to version 13.2.1 or apply the corresponding RHEL updates if using Red Hat's distribution. Organizations should prioritize this update given the Important severity rating and review the specific CVE links for detailed impact assessments relevant to their deployment configurations.
- Do this Monday: The security fixes address vulnerabilities that could potentially affect the integrity and security of Grafana deployments. Operators should prioritize upgrading to this version to mitigate risks associated with these CVEs.
- Sources: Grafana releases, Red Hat Security Advisories (RHSA)
16. ALAS2TOMCAT9-2026-028 (important): tomcat
- Category: Security / Patch
- What happened: A security advisory has been issued for Tomcat due to CVE-2026-66299, which may affect users of Amazon Linux 2. The advisory details the vulnerability and recommends applying the necessary updates to mitigate potential risks.
- Do this Monday: This advisory indicates a critical security vulnerability in Tomcat that could impact production environments using Amazon Linux 2. Immediate action is recommended to apply the security patch.
- Source: Amazon Linux 2 Security Advisories (ALAS2)
17. A Simple Website Summary Just Exposed the Limits of AI Coding Guardrails
- Category: Security / Patch
- What happened: A security researcher demonstrated that AI coding assistants, specifically Claude Code's Auto Mode, can be exploited to achieve remote code execution through malicious instructions embedded in seemingly innocuous webpage content that the AI is asked to summarize or process. The attack chain bypasses current AI coding guardrails by leveraging the AI's natural language processing to execute unintended commands, highlighting that these tools can be manipulated through prompt injection techniques similar to other AI security vulnerabilities. SRE teams using AI-assisted coding tools should treat any AI-generated code as untrusted input requiring human review before execution, avoid running AI coding assistants in auto-execute modes on untrusted content, and implement additional sandboxing or approval workflows when AI tools interact with external web content. According to DevOps.com, this incident demonstrates that AI coding safety mechanisms are still insufficient to prevent exploitation through carefully crafted natural language attacks.
- Do this Monday: This incident reveals significant security risks associated with AI coding tools, emphasizing the need for stronger guardrails to prevent exploitation in production environments.
- Sources: DevOps.com
Also this week
Deep dives & postmortems
18. Cloudflare: 8 service incidents (Elevated connectivity degradation in NRT, +7 more)
- Category: Deep Dive
- What happened: Cloudflare experienced multiple service disruptions across their infrastructure between September 3-4, 2024. A hardware failure in Tokyo (NRT) caused elevated connection resets and latency from 2:48-3:04 UTC, while Workers scripts in the same region encountered increased errors requiring a separate fix. R2 object storage returned elevated 503 errors in Western North America between 1:04-1:28 UTC, Workers Builds experienced queue delays affecting multiple customers, and Pipeline creation (including Pipelines, Streams, and Sinks) was temporarily unavailable though existing resources remained functional. Additionally, Cloudflare Access one-time PIN emails were delayed for customers using Proofpoint email security gateways due to SMTP 421 4.7.0 deferral responses, with Cloudflare advising affected customers to allowlist specific Cloudflare IPs. All incidents have been resolved with fixes implemented and monitoring completed, though operators relying on these services should review logs for the specified timeframes to assess any impact on their applications.
- Takeaway: This incident may have affected users' connectivity and performance in the Tokyo region, potentially impacting applications relying on Cloudflare services.
- Sources: Cloudflare Status
19. Fly.io: 3 service incidents (Packet loss in ORD, +2 more)
- Category: Deep Dive
- What happened: Fly.io experienced three separate service degradations affecting different platform components: an API background job runner failure impacted provisioning operations including app creation, IP assignments, and certificate management until a fix was deployed and verified; the ORD region saw approximately 50% packet loss on some hosts due to an upstream provider issue that caused slow replication for certain Managed Postgres (MPG) clusters; and sprite deletion jobs failed for 47 minutes on August 30, 2023 between 21:18 and 22:05 UTC before resolution. Operators should verify that any pending provisioning operations completed successfully after the API job runner issue and check application health in the ORD region, particularly for workloads dependent on Postgres replication. All three incidents have been resolved according to Fly.io Status updates, with the upstream provider issue in ORD showing improvement.
- Takeaway: Background job failures could have disrupted API-dependent operations, impacting app deployments and certificate management until the issue was resolved.
- Sources: Fly.io Status
20. OpenAI: 3 service incidents
- Category: Deep Dive
- What happened: OpenAI experienced multiple service degradations affecting different regions and tiers, now largely resolved. The Responses API suffered elevated latency that has been mitigated and is under monitoring, while ChatGPT conversations for Free and Go plan users experienced elevated errors that have been fully resolved. Additionally, APAC region users faced degraded performance across multiple services including file uploads, voice mode, conversations, image generation, and Codex Web, with investigation still ongoing for this regional issue. Operators using OpenAI services should monitor their error rates and latency metrics, particularly if serving APAC users, and consider implementing additional retry logic or failover mechanisms for production workloads that depend on these APIs. According to OpenAI Status, the Conversations and Responses components were specifically affected, suggesting potential capacity or routing issues in their infrastructure.
- Takeaway: This incident may have affected applications relying on the Responses API, potentially leading to slower response times for users. Monitoring continues to ensure stability post-mitigation.
- Sources: OpenAI Status
Lightning links
- Kubernetes v1.37: Garhwal (Kubernetes Blog) -- Kubernetes v1.37 introduces 67 enhancements, improving stability and features.
- Enterprise Live Migrations is now in public preview (Azure DevOps Blog) -- Organizations can now move Azure DevOps repositories to GitHub Enterprise Cloud with minimal disruption.
- AWS Lambda functions now support full IAM resource-based policies (Last Week in AWS) -- Lambda's support for full IAM policies enhances permissions management significantly.
- Fix circular role dependencies before upgrading Amazon RDS and Amazon Aurora PostgreSQL (AWS Database Blog) -- Avoid upgrade issues by fixing circular role dependencies in RDS and Aurora PostgreSQL.
- Authenticating TeamCity Builds to External Services With OIDC (JetBrains Blog) -- Enhance CI/CD security by using OIDC for authenticating TeamCity builds.
- Financially Motivated Threat Actor BREEZE COMET Targets Brazil (Google Cloud Blog) -- Stay informed about BREEZE COMET's tactics targeting Brazilian financial services.
- AWS MGN Architecture: How Continuous Replication Actually Works (dev.to (DevOps tag)) -- Understand the mechanics of continuous block-level replication in AWS migrations.
- Solving mysterious Kubernetes pod setup timeouts by tuning conntrack garbage collection (SRE Weekly) -- Resolve Kubernetes pod setup timeouts by adjusting conntrack garbage collection settings.
- Free AI Log Triage: Cut Weekly Review to 15 Minutes (dev.to (DevOps tag)) -- Leverage AI to significantly reduce log review time in your workflows.
- Provision a secure Amazon DocumentDB cluster with Terraform (AWS Database Blog) -- Follow best practices for provisioning a secure Amazon DocumentDB cluster using Terraform.
Human Stories
Looking at this week's incidents from Cloudflare and Fly.io alongside the Anthropic session hijacking warning, I'm reminded that our attention naturally gravitates toward the sophisticated threats while the mundane failures keep doing the real damage. We obsess over AI session token theft and geofenced malware campaigns, and rightfully so, but a hardware failure in Tokyo or a background job runner hiccup can take down production just as effectively. The AWS Gateway Load Balancer TCP Reset feature captures this tension perfectly - it's solving the unglamorous problem of reducing failover time from minutes to seconds, the kind of improvement that doesn't make headlines but saves someone's on-call shift every single week. The truth is that reliability engineering remains a discipline of defending against both the exotic and the ordinary, and most weeks, it's still the ordinary that gets you.
Also worth reading
We already know not to let the app own its own audit log. Agent tooling forgot (Reddit r/sre)
The discussion highlights the risks of allowing agents to manage their own audit logs, emphasizing that the audited service should not control its own audit history. It critiques the emerging trend of agent management platforms, suggesting they may not provide adequate security if logs can be manipu
The image tag that meant something different every hour (dev.to (Kubernetes tag))
The article discusses the issues encountered when using the ':latest' image tag in Kubernetes deployments, leading to multiple versions of a service running simultaneously. This situation arose due to different pods pulling the image at different times, resulting in discrepancies in behavior. The au
The Cloud Quota Nobody Knew About Until It Stopped the Deploy (dev.to (SRE tag))
The author recounts an incident where a routine capacity increase was halted due to reaching an unseen cloud resource quota. Despite the cloud's promise of infinite capacity, account-level limits can unexpectedly impede operations. The author emphasizes the importance of treating these quotas as par