❌

Reading view

Kubernetes Multi-Cluster Project Karmada Reaches CNCF Graduation

The Cloud Native Computing Foundation (CNCF) announced on September 2026 that Karmada, a multi-cluster and multi-cloud Kubernetes orchestration project, has graduated. This multi-cluster and multi-cloud Kubernetes orchestration project reached CNCF's highest maturity tier.

By Claudio Masolo
  •  

Presentation: When Incidents Refuse to End

Vanessa Huerta Granda explains how marathon incidents expose the gap between work as imagined and work as done. Drawing from real-world scenarios, she shares how complex outages reveal organizational fragility, human limits, and system interdependenciesβ€”and why incident response requires structured endurance, humane rotations, and holistic cross-functional coordination.

By Vanessa Huerta Granda
  •  

Repeated VM Escapes By GPT-5.6-Cyber Based Agents Prove VMs and OS' Require Better Maintenance

Traditional virtual machines are inadequate for isolating cyber-capable autonomous agents. Tests using GPT-5.6-Cyber indicated multiple escape attempts due to kernel flaws. While Firecracker provided some containment, vulnerabilities remained. The study underscores the need for minimal attack surface virtualisation technologies and rapid, proactive patching strategies to safeguard host systems.

By Olimpiu Pop
  •  

GPT-6 Astra Is the First Model OpenAI Classifies as Critical for Cybersecurity

OpenAI has classified GPT-6 Astra at the Critical cybersecurity threshold under its Preparedness Framework, a first. In expert-led testing the model found previously unknown vulnerabilities in a browser and an OS kernel and built working exploits. The same system card reports a substantial decline in chain-of-thought monitorability.

By Steef-Jan Wiggers
  •  

Dropbox Outlines How Focusing on Existing Infrastructure Efficiency Can Create Headroom for AI

Dropbox has outlined how a decade of infrastructure optimization is helping it absorb growing demand from AI without treating new data-center capacity as the only answer. Its work spans forecasting, fleet utilization, storage density, hardware lifecycles, and rack-level power delivery, much of it predating the current AI boom.

By Matt Foster
  •  

Agoda Replaces 72-Shard SQL Server Price Cache with DragonflyDB

Agoda migrated its 1.5 TB hotel Price Cache from 72 SQL Server shards to DragonflyDB to handle growing read and write volumes. The migration used staged dual reads, parity validation, gradual traffic shifting, and decentralized failover detection. Agoda reports an approximately eightfold reduction in P99 read latency, with two DragonflyDB clusters providing high availability.

By Leela Kumili
  •  

Lambda SnapStart Comes to Container Images, Ending a Packaging Tradeoff

AWS has extended Lambda SnapStart to container image functions, which hold up to 10 GB against 250 MB for zip archives. Teams previously chose between dependency headroom and sub-second startup. A Reddit thread from a month earlier shows what that cost: stripping whitespace and docstrings from installed packages to stay under the limit.

By Steef-Jan Wiggers
  •  

Netflix Reworks Conductor for 420 Million Monthly Workflow Executions and 10X Larger Workflows

Netflix has reworked its Conductor workflow orchestration engine to handle larger workloads, increasing supported workflow size from about 2,500 to 30,000 tasks and reducing p99 workflow evaluation latency by about 40%. Conductor 4.0 separates workflow metadata from task data, moves evaluation to asynchronous processing, and introduces dynamic worker allocation and concurrency controls.

By Leela Kumili
  •  

Presentation: How To Run on Three Clouds at Once, and When Not To

Ross McFarlane and Kevin Holditch discuss Form3's evolution from a single-cloud setup to a triple active multi-cloud architecture. They share key engineering strategies for cross-cloud networking, distributed databases with CockroachDB and NATS, custom Kubernetes operators, and navigating distinct regional disaster recovery expectations across the UK, Europe, and US financial markets.

By Ross McFarlane, Kevin Holditch
  •  

Session Traces and Cost Controls Help Diagnose AI Agent Failures

Session traces and cost controls are emerging as key observability techniques for diagnosing AI agent failures, helping teams spot tool-call loops and runaway spend while preserving enough execution context for post-incident debugging.

By Mark Silvester
  •  

Presentation: Fixing the AI Infra Scale Problem by Stuffing 1M Sandboxes in a Single Server

Felipe Huici explains how Unikraft achieves millisecond cold boots, stateful scale-to-zero, and extreme density for sandboxing AI workloads. He discusses isolation primitives, Linux kernel optimizations, and snapshotting tricks, demonstrating how to maintain sub-10ms performance at scale while integrating seamlessly into Kubernetes environments with hardware-level security.

By Felipe Huici
  •  

Azure Virtual Desktop Hybrid Reaches GA with Licensing Details Unpublished

Microsoft has made Azure Virtual Desktop Hybrid generally available. Session hosts run on customer hardware through Azure Arc while brokering stays in Azure. The Hybrid service license is unpriced, Windows Server support requires RDS CALs with Software Assurance, and multi-session Windows is not supported at all.

By Steef-Jan Wiggers
  •  

HashiCorp Packer 1.16 Adds Native SLSA Provenance Generation and Verification for Machine Images

HashiCorp has released Packer v1.16.0, adding native support for generating, signing, and verifying SLSA provenance attestations for every image the tool builds. The release provides teams with a secure, tamper-proof record of how a machine image was made. It does this without needing extra supply-chain tools.

By Claudio Masolo
  •  

Presentation: Platform Engineering in the Age of AI

The panelists explain how platform teams adapt to support AI-assisted engineering, highlighting which capabilities belong in the platform. They discuss trade-offs between standardization and developer autonomy, while sharing strategies to manage AI tooling, security guardrails, and shifting workflows.

By StΓ©phane Di Cesare, Davide de Paolis, Stephen Cihak, Camila Macedo, Renato Losio
  •  

GitLab Warns That AI Agent Sandboxes Are Only as Secure as Their Network Access

GitLab warns that isolating an AI coding agent in a sandbox does not necessarily make the agent safe. In a new security analysis, the company describes an internal evaluation in which an AI agent escaped its sandbox by exploiting a vulnerable package proxy that had been explicitly placed on the sandbox's allowlist.

By Craig Risi
  •  

Article: Implementing Chaos Engineering in Financial Payment Systems: Lessons from Enterprise ECS Deployments

Standard chaos engineering assumes experiments stop cleanly, blast radius is knowable in advance, and production is fair game. Payment systems violate all three. Salim Adedeji describes ECS-specific failure modes from enterprise deployments: a 60-second DNS TTL that produced 93-second failover, retry logic amplifying database load 2.4x, and AZ rebalancing loops that generic tooling misses.

By Salim Adedeji
  •  

Does ICANN Open the Door on Identity Theft by Dropping 3rd Level .name Domains Registrations?

Neil Fraser's disclosure highlights a regulatory change affecting the .name top-level domain. Following ICANN's approval, Verisign will eliminate third-level registrations due to declining usage. This affects about 22,000 registrants and raises security concerns, as released second-level domains could be exploited. Affected users are considering legal options to challenge the decision.

By Olimpiu Pop
  •  
❌