eBPF-based Networking, Observability, and Security
eBPF-based Networking, Observability, and Security
eCHO news 113
eBPF-based Networking, Observability, and Security
eCHO news 114
eBPF-based Networking, Observability, and Security
eCHO news 115
eBPF-based Networking, Observability, and Security
eCHO news 116
eBPF-based Networking, Observability, and Security
eCHO news 117
High-quality, ubiquitous, and portable telemetry to enable effective observability
Exploring the OpenTelemetry Instrumentation Ecosystem
OpenTelemetry has a lot of pieces: APIs, SDKs, a protocol, semantic conventions, instrumentation, and tools like the Collector. The APIs and protocol define how telemetry is created and exchanged, while instrumentation is what actually observes what happens inside an application and turns it into telemetry. Semantic conventions give the people writing that instrumentation a shared way to describe what they’re observing.
The beauty of this approach is that completely separate auth
Kubewarden is a Policy Engine powered by WebAssembly policies. Its policies can be written in CEL, Rego (OPA & Gatekeeper flavours), Rust, Go, YAML, and others....
SBOMscanner 0.13 Release: OpenTelemetry Traces and Metrics
We are happy to announce SBOMscanner
v0.13.0!
The highlight of this release is
OpenTelemetry
support. SBOMscanner now exports traces and metrics for every component,
from the admission webhook down to the SQL statements of the storage layer.
The release also brings plain-HTTP support for insecure
eBPF-based Networking, Observability, and Security
Cilium at KubeCon + CloudNativeCon and CiliumCon North America 2026
Cilium is heading back to Salt Lake City this November for KubeCon + CloudNativeCon and CiliumCon North America 2026. Since coming…
MySQL-compatible, horizontally scalable, cloud-native database solution.
Hardening EmergencyReparentShard in v25
EmergencyReparentShard operations are being hardened in upcoming release v25. In this blog, we cover how ERS works and the upcoming changes that make recovery safer, faster and less brittle
What is EmergencyReparentShard? # EmergencyReparentShard (ERS) is the Vitess failover process used when a shard's current primary is dead or unreachable. While PlannedReparentShard gets a clean handoff from a healthy primary, ERS has to pick a replacement using only surviving tablets. It compares their transa
wasmCloud 2.10: Same-host routing, outbound identity, and OpenTelemetry that follows the spec
wasmCloud 2.10 ships opt-in same-host routing for HTTP between co-located workloads, mTLS client identity for outbound HTTPS with live rotation, spec-conformant OpenTelemetry configuration, enforced plugin egress, and Wasmtime 48.
Confidential Containers is an open source community working to enable cloud native confidential computing by leveraging Trusted Execution Environments to protect containers and data.
PQC Support in KBS Protocol
Introduction
Those who are interested in Confidential Containers as a framework, already have a high bar for security. The delivery of resources from the trusted Trustee services into untrusted guest components requires resources to be secured in transit once the request has been attested to.
The change that this blog entry describes relates to the steps which have now been
Kubernetes is an open-source system for automating deployment, scaling, and management of containerized applications
Spotlight on SIG Apps
As Kubernetes adoption has grown, the conversation has shifted beyond running containers to managing increasingly complex application lifecycles. Modern platforms support stateless web services, stateful databases, batch processing, AI workloads, and platform services. At the same time, they must remain reliable during upgrades, scaling events, and infrastructure failures.
Every Kubernetes user relies on SIG Apps, whether they realize it or not. Deployments, StatefulSets, DaemonSets, J
High-quality, ubiquitous, and portable telemetry to enable effective observability
Prometheus and OpenTelemetry interoperability in 2026: Survey results
We ran a survey asking users of OpenTelemetry and Prometheus how they collect, process, and store metrics. The goal was to understand, with real usage data rather than assumptions, how far the ecosystem has moved and whether the interoperability still causes friction.
Key takeaways
- Interoperability has measurably improved since our
Connect is a family of libraries for building browser and gRPC-compatible HTTP APIs.
Protobuf, JSON Schema, and OpenAPI
Learn how to generate JSON Schema and OpenAPI specs using Protobuf.
Keycloak is an open-source identity and access management solution for modern applications and services, built on top of industry security standard protocols.
Speak at KeycloakCon Europe 2027 in Barcelona!
The Call for Papers for KeycloakCon Europe 2027 is now open!
KeycloakCon Europe 2027 is a half-day, single-track conference in Barcelona on March 15, 2027 — co-located with KubeCon + CloudNativeCon Europe. After successful events in Amsterdam and Yokohama, KeycloakCon is back with talks, discussions, and community connectio
Etcd is a distributed, reliable key-value store for the most critical data of a distributed system. By using etcd, developers can ensure that their applications have access to up-to-date configuration data, even as they scale up or down, and can maintain consistency, fault tolerance and coordination across multiple instances of the application.
Etcd Patch Releases: v3.7.2, v3.6.15, and v3.5.34
SIG-etcd has distributed patch updates for all three supported release branches. These releases update dependencies, fix a file-handle leak during file cleanup, correct etcdctl endpoint status output, and improve version detection in v3.7. Users on v3.5, v3.6, and v3.7 should update at the next scheduled maintenance window after the releases become available.
Obtain the updates here:
Apache Kafka running on Kubernetes
Beyond mTLS: Configurable Security for Internal Apache Kafka Cluster Communication
For a long time, one thing has been common to every Strimzi-based Apache Kafka cluster. It uses TLS encryption and mTLS authentication for all internal cluster communication. Data replication between brokers, KRaft controller communication, Strimzi operators talking with Kafka … all of this always uses TLS encryption and mTLS authentication. But with Strimzi 1.3.0, this is going to change!
This blog post previews an unreleased feature that will be part of the upcoming Strimzi 1.3.0 release.
When we created Strimzi, we wanted it to be secure out of the box. So TLS encryption and mTLS authentication were baked into it from the beginning. Being secure sounds like a good idea and it is what is desired in most cases. But there are always some situations where hardcoded TLS encryption is not the optimal choice.
TLS encryption does not come for free. It can cost a lot of CPU and impact the performance. It also prevents you from using zero-copy when reading data from disk and sending it to consumers. And in air-gapped environments, one can argue that performance might sometimes be more important than encryption. And even when TLS encryption is desired, it might be done at a different level. For example, when encryption is already provided by Istio or some other service mesh, you do not need it to happen again at the Strimzi level. These are some of the reasons why Strimzi users have been asking for a long time to be able to disable the TLS encryption.
Disabling encryption sounds simple. Unfortunately, it is a bit more complicated than that. Remember, Strimzi does not only rely on TLS encryption, but also on mTLS authentication. And it cannot use mTLS authentication when TLS encryption is disabled. And without authentication, we would also lose authorization and the Kafka cluster would be completely insecure. And that would usually not be acceptable. Even in air-gapped environments, you want the services using your Kafka cluster to be authenticated and properly authorized. You do not want anyone to be able to connect to the Kafka cluster and consume or produce messages without any restrictions. And since Istio does not really understand the Kafka protocol, it cannot give you any real security either. It can only encrypt the communication.
So the task was not just to allow disabling TLS encryption. But also to introduce a new authentication mechanism that is independent of the TLS encryption. And that was one of the reasons why it took us so long to ship this improvement.
Configurable Internal Cluster Security
The solution was introduced in Proposal 150 — Configurable Security for Internal Kafka Cluster Communication. The proposal introduces two new configuration options:
- Encryption configuration
- Authentication configuration
The encryption configuration allows you to enable or disable TLS. And the authentication configuration lets you choose between mTLS authentication (supported only when TLS encryption is enabled as well), no authentication, and Service Account-based authentication.
This configuration affects only the internal communication within the Kafka cluster. It has no impact on the listeners you configured in
.spec.kafka.listeners.
Service Account authentication is the new authentication type we are introducing, and it does not depend on the TLS encryption. Already before this proposal, every Strimzi component had its own Service Account. And now we use these Service Accounts for authentication at the Apache Kafka level as well. The Service Account identity is based on the Service Account name and its namespace. And the Service Accounts are represented by JWT tokens that can be validated using OIDC and JSON Web Key Sets (JWKS). To isolate the Service Accounts belonging to different applications, each Kafka cluster is using its own unique audience in the JWT tokens.
We use two different ways to get the Service Account JWT tokens.
In the operands — Kafka nodes, Topic and User Operators, Kafka Exporter, and Cruise Control — we use projected serviceAccountToken volumes.
The projected volumes mount the Service Account token into the container file system and automatically refresh it before it expires.
The tokens are short-lived and are typically valid for one hour.
But you can configure how long they are valid.
The Cluster Operator cannot use the projected volumes.
It needs to use a different Service Account for each Kafka cluster in order to make sure the different Kafka clusters are properly isolated.
And it cannot add a new projected volume to itself every time a new cluster with Service Account authentication is created.
So instead, it uses the TokenRequest API to get the tokens directly from the Kubernetes API.
At the Kafka level, we use the Strimzi OAuth library. On the client side, it loads the token from the file mounted through the projected volume and uses it for authentication with the Kafka server. And on the server side, we use the Kubernetes JWKS keys to validate the tokens. In the Cluster Operator, we use our own callback handler to be able to get the token from the Kubernetes API itself.
The Service Account tokens are also used for authentication in the Kafka Agent that is running inside the Kafka brokers. The Kafka Agent provides some additional information about the state of the Kafka cluster through an HTTP API. It is used by the Cluster Operator.
So, how do you configure it?
Configuring the Internal Cluster Security
In Strimzi 1.3.0, the cluster security configuration will be available as an early access feature.
We hope you will give it a try and let us know whether it works for you or not.
While in early access, we use an annotation to configure it.
And once we — with your help — manage to validate that the model works, we will move the configuration into the .spec section of the Kafka CR.
The annotation used for the configuration is strimzi.io/internal-cluster-security and it contains a JSON structure with the authentication and encryption configurations.
The annotation is placed on the Kafka CR.
The following example disables both the TLS encryption and the mTLS authentication:
apiVersion: kafka.strimzi.io/v1
kind: Kafka
metadata:
name: my-cluster
annotations:
strimzi.io/internal-cluster-security: |
{
"encryption": {
"type": "none"
},
"authentication": {
"type": "none"
}
}
spec:
# ...
For the authentication, you can choose one of three types: none, mtls and service-account.
For the encryption you can use the types none and tls.
The mTLS authentication can be used only with the TLS encryption.
But otherwise you can combine the options in any way you want.
The following example shows Service Account-based authentication with TLS encryption:
apiVersion: kafka.strimzi.io/v1
kind: Kafka
metadata:
name: my-cluster
annotations:
strimzi.io/internal-cluster-security: |
{
"encryption": {
"type": "tls"
},
"authentication": {
"type": "service-account"
}
}
spec:
# ...
For the full list of configuration options and additional examples, please follow the Strimzi 1.3.0 documentation once it is released.
Without the annotation, TLS encryption and mTLS authentication will be used as before. So if you are happy with how Strimzi worked until now, you can just ignore the annotation. But when you are deploying a new Kafka cluster and want to use the new options, just add the annotation.
But what about the existing Kafka clusters?
Migrating Existing Clusters
You can of course change the cluster security configuration for existing clusters as well. However, you cannot do it on the fly without any interruptions. You have to:
- Pause the reconciliation of the Kafka cluster
- Shut down the Kafka cluster by stopping all its pods
- Update the security configuration
- Start the Kafka cluster again by unpausing the reconciliation
For the detailed steps, please follow the full migration documentation once Strimzi 1.3.0 is released.
Stopping the whole Kafka cluster makes the migration significantly easier. For example, we do not need a complicated multi-step process that would add and remove internal Kafka listeners to change the authentication or encryption. So it saves us a lot of implementation and maintenance effort. And since this is not the kind of configuration that you would be changing every day, we do not expect this to be a major issue and do not have any plans to support the migration on the fly.
Early Access and What’s Next
As mentioned above, this feature will be released in Strimzi 1.3.0 as early access. Please give it a try and share your feedback with us. With your help, we should be able to validate this feature in one or two releases and mark it as generally available.
There are also still some limitations to be aware of. For example:
- In Strimzi 1.3.0, even if you disable TLS encryption, Strimzi will still maintain its Certificate Authorities.
- While disabling the TLS encryption makes it easier to run Strimzi within a service mesh such as Istio, it does not provide full integration.
These and other improvements might be added in future Strimzi releases.
If you find any issues with the implementation, or have suggestions related to the internal cluster security configuration, you can share your feedback with us on Slack, or by opening a discussion or an issue on GitHub.
Kubernetes is an open-source system for automating deployment, scaling, and management of containerized applications
Kubernetes v1.37: Tracking When a PersistentVolumeClaim Was Last Used (Beta)
Kubernetes v1.37 promotes the PersistentVolumeClaimUnusedSinceTime feature gate to Beta (enabled by
default). With this feature, the PersistentVolumeClaim (PVC) protection controller adds an Unused
condition to each PVC, telling you whether any running pod currently references it — no custom
tooling or cross-referencing required.
For the API definition of PVC conditions, see the
High-quality, ubiquitous, and portable telemetry to enable effective observability
Announcing the 2026 OpenTelemetry Governance Committee Election
The OpenTelemetry project is excited to announce the 2026 OpenTelemetry Governance Committee (GC) election. Nominations are due by 16 October 2026 23:59 AoE. The list of eligible candidates will be shared on 19 October 2026. Voting will take place between 26 October 2026 12:00 UTC and 28 October 2026 end of day AoE (29 October 2026 11:59 UTC), and the final election results will be announced 30 October 2026.
Vote!
The Prometheus monitoring system and time series database
Prometheus and OpenTelemetry interoperability in 2026 - Survey results
We ran a survey asking users of OpenTelemetry and Prometheus how they collect, process, and store metrics. The goal was to understand, with real usage data rather than assumptions, how far the ecosystem has moved and whether the interoperability still causes friction.
Key takeaways
- Interoperability has measurably improved since our 2024 survey: the average ease-of-use rating rose from 3.1 to 3.6, the equivalent of one in two respondents rating a whole category higher, and the share of respondents finding the two hard to use together fell from 29% to 10%.
- In infrastructure instrumentation, Prometheus exporters remain the most-used method (72%) with OTel receivers close behind (57%), and nearly half of respondents run both at once rather than migrating from one to the other.
- In application instrumentation, OTel SDKs are the most-used method at 65% with Prometheus SDKs at 52%, and 41% use only the OTel style of application instrumentation.
- Prometheus relabeling rules (54%) and the open source OTel Collector (53%) are the two most common processing steps, and 65% of respondents run a "vanilla stack" of one or both with no vendor transformation or custom Collector build anywhere in the pipeline.
Demographics
From 186 people who responded, 81 passed our screening for active OpenTelemetry-for-metrics users on a Prometheus-adjacent backend. We also filtered out observability vendor employees to focus on end users. In the analyzed sample:
- All respondents are active OpenTelemetry users.
- All respondents use some flavor of Prometheus – Prometheus itself (46%), an open source Prometheus-compatible backend such as Thanos, Cortex, or Grafana Mimir (42%), or a PromQL-compatible vendor product (12%).
- Respondents' observability maturity is high. 48% describe their organization as having "a well-established observability practice" (Expert), 41% are "setting up an observability practice" (Intermediate), while only 11% consider themselves beginners in observability.
- Organizations skew large. 42% have 1,000+ employees, 31% have 100–999, 15% have 50–99, and 12% report having under 50.
Ease of use change over time
How easy or difficult is it to use OpenTelemetry and Prometheus together?
This year, we asked the same question as in the similar 2024 survey to see whether end users saw progress in interoperability.
The average rating rose by 0.5 point, from 3.1 to 3.6 — as if every second respondent had moved up a full category. The clearest movement is at the difficult end of the scale: the share of respondents who found the two hard to use together dropped to roughly a third of its 2024 level. Also, nobody this year picked "Very difficult".
Two years of work on interoperability is paying off. At the same time, since the single largest group of responses sits at "Neither easy nor difficult", there is still a lot of work to be done in this area.
Note: The 2024 survey didn't ask respondents whether they worked for an observability vendor, so this is not an exact apples-to-apples population match. However, putting vendor employees back into the 2026 sample (n = 108) would barely change the result for the ease of use rating (0%, 10%, 40%, 33%, 17% → 0%, 10%, 41%, 33%, 16%). To keep this year's results consistent, we decided to stick with filtering vendor employees out.
Infrastructure metrics
How do you instrument infrastructure metrics collection?
Prometheus exporters are the most common single instrumentation method for infrastructure metrics but OTel receivers are close behind. Built-in /metrics endpoint, built-in OTLP push, and OpenTelemetry eBPF instrumentation (OBI) follow.
When looking at how these methods combine, the picture is clearly hybrid, not either/or. Nearly half of respondents are mixing Prometheus and OTel instrumentation styles at once for infrastructure metrics, rather than doing a full migration. Among respondents using a single instrumentation style, Prometheus-only style is twice as popular as OTel-only style.
Note: Instrumentation style describes whether a respondent uses methods native to one project only, or a mix of both. OTel-style includes using OTel receivers, Built-in OTLP push, or OpenTelemetry eBPF Instrumentation (OBI). Prometheus-style includes Prometheus exporters or Built-in /metrics endpoint (no exporter). The 4 "Other" responses are write-ins: Zabbix, Heorku Telemetry (likely "Heroku Telemetry"), textfile collector, Telegraf. All 4 respondents also selected a real Prometheus/OTel method alongside their write-in — but in the style chart above, a write-in places a respondent in "Other" regardless of what else they selected.
Work in progress: The Prometheus and OTel communities are working on making Prometheus exporters run as an OTel Collector distribution. The conversations are still ongoing. The discussion is open in this issue.
Application metrics
How do you instrument application metrics collection?
Preferences swap for application instrumentation. OTel SDKs come out on top with Prometheus SDKs following behind them. OBI holds roughly the same share as in infrastructure instrumentation.
Instrumentation styles shift as well. The largest share of participants (41%) use only OTel style instrumentation, nearly twice as common as only Prometheus style. Fewer than a third mix styles.
Note: In application instrumentation, OTel-style includes using OTel SDKs or OpenTelemetry eBPF Instrumentation (OBI). Prometheus-style includes Prometheus SDKs. Again, there are 4 write-ins that we categorized as "Other": already built exporters, Micrometer, textfile collector, jvm-exporter. 3 of the 4 also selected a real Prometheus/OTel method. One respondent's original write-ins, "Self instrumentation" and "manual instrumentation for OTEl," were recoded to plain OTel SDKs.
Transformation
What do you use to process or transform metrics before sending them to storage?
Prometheus relabeling rules and the open source OTel Collector are the two most common processing steps with neither of them leading clearly.
Most respondents run a vanilla stack: only Prometheus relabeling rules and/or the plain OTel Collector, with no vendor distribution and no custom-built Collector in the pipeline. The three vanilla patterns come out close to even.
Note: "Other" combines respondents who do no transformation at all (15%, n=12) with those using a vendor distribution or custom-built Collector (20%, n=16).
What practitioners want improved
What would you like us to improve to make OpenTelemetry and Prometheus work better together?
We received 19 open-ended responses with suggestions on what to improve. Three themes emerged from this data: unification of Prometheus and OTel's data models (attributes/labels), better handling of resource attributes and metadata, and naming and formatting friction. There were also a few individual asks. Prometheus maintainers György "Krajo" Krajcsovits and Arthur Sens went through the responses and addressed each point below:
- Unifying Prometheus and OTel's data models (attributes/labels)
- This is a valid ask that we recognize. We will raise it for a discussion at the Prometheus Dev summit in October.
- Resource attributes and metadata gaps
- This should be addressed by the native metadata design doc. One thing that we have to wait for is finishing the OTel Entities spec.
- Naming and formatting friction
- Several relevant things already exist — the OpenMetrics 2.0 exposition format lets OTel-style names be used directly in code, PromQL already supports UTF-8 metric names, and Prometheus's OTLP receiver has configurable translation strategies. The pieces exist; they're just not the default yet. We have to work on this.
- Using Prometheus native recording rules in the Collector
- There's an open Prometheus proposal and proof-of-concept PR for scrape-time recording rules, which wouldn't need a full TSDB the way recording rules do today. Since the OpenTelemetry Collector's Prometheus Receiver uses Prometheus code as a Go Library, this proposal would also benefit the Collector.
- Enable MCP or agentic AI workflows
- Prometheus just onboarded the Prometheus MCP project repository to its GitHub org. This should enable MCP workflows for Prometheus. The Prometheus community would love to see people start using it and get feedback. Also, the native metadata design doc explains how we plan to make agentic AI workflows even better in Prometheus.
Interesting observations
Mid-size organizations may be furthest into OTel-native tooling
In our data, organizations with 100-999 employees have the highest OTel SDK adoption for application metrics and OTel receiver adoption for infrastructure metrics. eBPF-based instrumentation (OBI) doesn't follow the same pattern — there, it's the 1,000+ organizations that stand apart from every smaller band.
Adoption by organization size:
| Organization size | OTel SDKs(application) | OTel receivers(infrastructure) | eBPF / OBI(infrastructure) |
|---|---|---|---|
| 1–49 (n = 10) | 40% | 20% | 20% |
| 50–99 (n = 12) | 58% | 58% | 17% |
| 100–999 (n = 25) | 84% | 76% | 20% |
| 1,000+ (n = 34) | 62% | 53% | 3% |
Our hypothesis is that mid-size organizations — big enough to have a dedicated platform effort, small enough to move without a multi-year migration plan — might be pushing furthest into newer OTel-native tooling.
Note: This is an interesting observation and a hypothesis, not a confirmed finding: with 10–34 respondents per band, none of these gaps is big enough for a survey this size to confirm.
Team type tracks backend choice
Platform Engineering and SRE teams lean heavily toward OSS Prometheus-compatible backends (Thanos, Cortex, Mimir), while Dev teams lean the other way, toward plain Prometheus.
Here, the dividing line looks like operational ownership rather than preference. Teams running metrics for a whole organization eventually outgrow a single Prometheus deployment, whereas teams instrumenting their own service generally don't.
Backend choice by team type — OSS Prometheus-compatible (n = 30), Prometheus (n = 35), PromQL-compatible vendor (n = 8):
| Team type | OSS Prometheus-compatible | Prometheus | PromQL-compatible vendor |
|---|---|---|---|
| Dev | 24% | 71% | 6% |
| DevOps | 23% | 62% | 15% |
| Observability | 29% | 41% | 29% |
| Platform Engineering | 69% | 31% | 0% |
| SRE | 69% | 31% | 0% |
Note: Sysadmin (n = 6) and Operations (n = 2) respondents are excluded from this table — both groups are too small to interpret — leaving n = 73 of the 81 respondents. As with the previous breakdown, the per-band numbers here (8 to 35) are too small to draw firm conclusions.
Get involved
Interoperability is measurably easier than it was two years ago, but the open-ended answers point to concrete gaps — data model differences, resource attributes and metadata gaps, and naming and formatting friction. There is still a lot of work to do on both the OpenTelemetry and the Prometheus side.
Everyone is welcome to contribute. The discussion happens in the #otel-prometheus channel in the CNCF Slack.
NOTE: This blog post was also published on opentelemetry.io/blog (canonical version).
Kubewarden is a Policy Engine powered by WebAssembly policies. Its policies can be written in CEL, Rego (OPA & Gatekeeper flavours), Rust, Go, YAML, and others....
Admission controller 1.38 Release
Welcome to the monthly release of Admission Controller. On the menu for this month we have a handful of security fixes and a scalability improvement.
Let’s go through each serving!
Hardening namespaced policies
The issue has been found by @Pinguladora, who filed this GitHub security advisory<
Falco, the cloud-native runtime security project, is the de facto Kubernetes threat detection engine
Blog: Introducing Falco 0.45.0
Dear Falco Community, we are happy to announce the release of Falco 0.45.0 today!
This release brings raw byte matching in rule conditions, new reload status and control endpoints, and fixes for package upgrades and container metadata collection. It also improves modern eBPF capture on preemptible kernels and adds sandbox rules for detecting suspicious GPU activity in containers.
We upgraded libs to 0.26.0 and drivers to 11.0.0+driver. We also ship
High-quality, ubiquitous, and portable telemetry to enable effective observability
Dual-exporting .NET metrics with OTLP and Prometheus
Many applications export their metrics directly to Prometheus. If you’re unfamiliar with Prometheus, in a nutshell it’s a time-series database for storing metrics, like counters and histograms. Applications that store their metrics in Prometheus typically use a popular Prometheus client as part of the integration.
Now that OpenTelemetry is
Real-Time Communication Fabric for Distributed Agents and Applications. NATS unifies messaging, streaming, and state into one real-time system — connecting services, devices, and AI agents from cloud to edge.
NATS Server 2.15 Release
Previous NATS Server releases, like 2.12 in September of last year and 2.14 in April of this year, focused primarily on the addition of new and exciting features: atomic & fast batch publishing, counters, schedules, etc. The reliability of the server has been gradually increasing throughout as well, with better error handling and the squashing of tons of bugs, but it otherwise was mostly work performed in the background.
This release largely forgoes new features and instead focuses entirely
Connect is a family of libraries for building browser and gRPC-compatible HTTP APIs.
A faster Protovalidate
Go, Java, and TypeScript now run Protovalidate's standard rules as native code, while protovalidate-py 2.0 replaces its pure-Python core with a much faster native extension.
Provision bare metal hardware via k8s-native APIs, including integration with the Cluster API.
Release-1.14.0 Highlights: What’s New in CAPM3, IPAM and BMO
The Metal3 community has published new releases across three core projects: Cluster API Provider Metal3 (CAPM3) v1.14.0, Bare Metal Operator (BMO) v0.14.0, and IP Address Manager (IPAM) v1.14.0.
This post highlights the key changes in each release, with a focus on major breaking changes and user-facing improvements. It is intended as a concise summary rather than a full changelog.
Full release notes are available here:
Full release notes remain the source of truth for the complete list of fixes, dependency updates, and maintenance changes.
CAPM3 v1.14.0 Highlights
CAPM3 Breaking Changes
-
Deprecate
Metal3DataTemplatespec.clusterName(#3520) — AMetal3DataTemplateis no longer tied to a single cluster, so theclusterNamefield is now optional and unused by the controllers. Setting it produces an admission warning, and the template is no longer garbage-collected with the cluster. To keep it moving duringclusterctl move, the template now carries theclusterctl.cluster.x-k8s.io/movelabel. Drop the field from your manifests:<div class="language-yaml highlighter-rouge"><div class="highlight"><pre class="syntax"><code><span class="na">apiVersion</span><span class="pi">:</span> <span class="s">infrastructure.cluster.x-k8s.io/v1beta2</span>kind: Metal3DataTemplate
metadata:
name: nodepool-1
spec:
# clusterName: my-cluster # deprecated: remove this line
metaData:
ipAddressesFromIPPool:
- key: provisioningIP
name: provisioning-pool
Add blockmove annotations to allow M3M owner references to be removed from BMH
(#3459) —
CAPM3 no longer sets aMetal3Machineowner reference on theBareMetalHost;
theConsumerRefis used instead. To keepclusterctl movesafe while a host
is being claimed, CAPM3 now sets theclusterctl.cluster.x-k8s.io/block-move
annotation on the BMH until its pause and status annotations are applied, then
removes it. Note that this annotation blocks the entire move operation while
present on any object, so a BMH stuck in this state can stall the whole move.<div class="language-yaml highlighter-rouge"><div class="highlight"><pre class="syntax"><code><span class="na">metadata</span><span class="pi">:</span>annotations:
# set automatically by CAPM3, removed once the BMH is paused
clusterctl.cluster.x-k8s.io/block-move: ""
Remove all code related to the legacy form of ProviderID
(#3373) —
Only the currentProviderIDformat (metal3://<namespace>/<bmh-name>/<m3m-name>)
is supported. Clusters that still carry nodes with the oldmetal3://<bmh-uid>style provider IDs must be reconciled onto the new format
before upgrading.
CAPM3 New Features
-
Fetch DNS servers from IPPool for the CAPI
IPAddressClaimpath (#3534) — A CAPIIPAddressdoes not carry DNS information, so when allocating through the Cluster APIIPAddressClaimflow, CAPM3 now reads DNS servers directly from the referenced Metal3IPPool. It resolves the allocated address to the matching pool entry and uses that entry’s per-subnetdnsServers, falling back to the pool-leveldnsServerswhen there is no override:<div class="language-yaml highlighter-rouge"><div class="highlight"><pre class="syntax"><code><span class="na">apiVersion</span><span class="pi">:</span> <span class="s">ipam.metal3.io/v1alpha1</span>kind: IPPool
metadata:
name: provisioning-pool
spec:
dnsServers: ["8.8.8.8"] # used as fallback
pools:
- subnet: "192.168.1.0/24"
dnsServers: ["9.9.9.9"] # used for addresses in this subnet
BMO v0.14.0 Highlights
BMO Breaking Changes
-
Deprecate the BMH
Taintsfield (#3465) — Thespec.taintsfield onBareMetalHostis now deprecated because it was never actually implemented. Apply node taints through your Cluster API machine templates instead.<div class="language-yaml highlighter-rouge"><div class="highlight"><pre class="syntax"><code><span class="na">apiVersion</span><span class="pi">:</span> <span class="s">metal3.io/v1alpha1</span>kind: BareMetalHost
metadata:
name: node-0
spec:
# taints: # deprecated
# - key: dedicated
# effect: NoSchedule
online: true
Provisioner: drop the separate
TryInitcall
(#3327) —
Removes the standaloneTryInitcall from the provisioner initialization flow,
simplifying how the Ironic provisioner is brought up. This is an internal
refactor with no user-facing configuration change.
BMO New Features
- Add fast Redfish-based inspection mode (#3482) — Introduces a faster hardware inspection mode using Redfish.
-
Add provisioning retry limit to prevent infinite loops (#3468) — A host that repeatedly fails to provision the same image no longer retries forever. A new
--max-provisioning-retriesflag (default5,0disables the limit) caps consecutive failures, tracked in a newstatus.provisioningFailCount. The counter resets on success or when the image spec (URL, checksum, format) changes:<div class="language-console highlighter-rouge"><div class="highlight"><pre class="syntax"><code><span class="gp">#</span><span class="w"> </span>disable the limit <span class="o">(</span>previous behavior of unlimited retries<span class="o">)</span>baremetal-operator --max-provisioning-retries=0
- Add webhooks to validate URLs
(#3333) —
New admission webhooks validate URLs before they are accepted. - Implementation of the
HostClaimcontroller
(#3406) —
Adds a controller for the newHostClaimresource. - Add BMH and HNA types and cross-resource validation
(#3338) —
Introduces new types and validation across related resources. - Add TLS curve preferences support
(#3423) —
Allows configuring preferred TLS curves. - Implement structured logging pattern from CAPM3
(#2943) —
Adopts the structured logging approach already used in CAPM3. - Cache Ironic client and its status across reconciliations
(#3343) —
Reuses the Ironic client and status to reduce overhead between reconciles. - Migrate to a go-plugin system for provisioners
(#3166) —
Provisioners are now loaded through a go-plugin based system. - Reduce and simplify requeue delays in the Ironic provisioner
(#3302) —
Streamlines requeue timing for faster provisioning progress. - Add
BareMetalSwitchCRD and controller for switch config generation
(#3046) —
Introduces a new CRD and controller to generate switch configuration. - Allow exiting the externally provisioned state
(#3181) —
Hosts can now leave the externally provisioned state. - Validation webhook for
HostClaims
(#3196) —
Adds a validating webhook forHostClaimresources.
IPAM v1.14.0 Highlights
IPAM New Features
-
Add random IP allocation strategy to
IPPool(#1359) —IPPoolnow supports aspec.allocationStrategyfield. The default,sequential, allocates the first available address; the newrandomstrategy picks a random free address within a pool. The strategy is immutable after creation, andrandomrequires bounded pools (start/end or subnet) so the pool size can be computed:<div class="language-yaml highlighter-rouge"><div class="highlight"><pre class="syntax"><code><span class="na">apiVersion</span><span class="pi">:</span> <span class="s">ipam.metal3.io/v1alpha1</span>kind: IPPool
metadata:
name: provisioning-pool
spec:
allocationStrategy: random # "sequential" (default) or "random"
pools:
- start: "192.168.0.10"
end: "192.168.0.100"
Notes
As always, review the full release notes linked above before upgrading, and pay particular attention to the breaking changes in CAPM3 and BMO. If you run into issues or have questions, reach out on the Metal3 community channels.
Thanks to all our contributors who made these releases possible!
Kubernetes is an open-source system for automating deployment, scaling, and management of containerized applications
Kubernetes v1.37: Hardening Container Storage with Bind Mount Options and EmptyDir Permissions
Kubernetes v1.37 brings important storage security features: emptyDir permission modes and bind mount options. They help application programmers and security professionals implement rigorous security policies, for example, prohibiting deletion of files across containers or execution of arbitrary binaries from writable volumes, directly in Kubernetes without any complicated circumvention.
Linux storage and permission fundament
Kubernetes-native tools to run workflows, manage clusters, and do GitOps right.
Argo CD v3.6 Less Memory, Faster UI, Better Rollout Visibility
Argo CD v3.6 Release PosterWe’re excited to announce the first release candidate for Argo CD 3.6!
This release candidate brings multiple enhancements to existing features, some brand new ones along with a ton of bug fixes. In this blog post, we’ll highlight the most significant updates you should be aware of.
Better Web UI User Experience
Thanks to @agaudreault (Intuit) Argo CD got the following three great UI enhancements:
The New Resources Explorer View
Now you can list, search, filter and view details on managed resources from all your Applications in the global Resources Explorer view:
Resource Explorer View ScreenshotThe view aggregates managed resources from all accessible Applications while respecting existing RBAC. It supports standard sorting and filtering options with filter state persisted in view preferences and URL query params. It also supports Summary view: pie charts for sync status, health, namespaces, clusters, and resource kinds.
From the Resources page, users can open the owning Application and highlight the resource in list or tree view. Clicking a resource opens a sliding panel with the same detail view used on the Application page. From the resource details panel on the Resources page, a Details button opens the same resource on its Application page.
Note, that only managed resources are listed, that is, the resources directly declared in Git and deployed by Argo CD. Dependent resources created by controllers (for example, the `ReplicaSet` and `Pod` objects generated from a `Deployment`) are not shown, because Argo CD does not deploy them directly. To inspect those child resources, open the owning application and use its resource tree.
Settings List Pages Overhaul
The Settings list pages (Repositories, Clusters, Projects, Accounts, GPG Keys, Certificates) now share a unified and consistent search/filter/sort/pagination experience:
Settings List Screenshot- Search and URL synchronization: every settings list now has a search bar whose state (search text and filters) is reflected in the URL, consistent across pages.
- Filter sidebars: added per-list filter panels with live result counts.
- Sortable columns: click-to-sort headers on all lists.
- Pagination supported on all settings lists.
The Advanced Settings View
The new Settings/Advanced page allows to view global instance configuration options in a convenient form. The JSON tab lets you quickly view the full configuration in textual form.

Web UI Navigation improvements
Thanks to @blakepettersson (Akuity) you now can right click on rows/tiles of Applications, Appsets, Resources as well as on sidebar links and open them in a new browser Tab or Window.
Regex search in Application(Set) Lists
In larger Argo CD installations plain substring search might be limiting. Now, thanks to @rickbrouwer, you can enable regular expressions in the Applications and the ApplicationSets list search bar.
Regexp Search UI ScreenshotFaster Filtering and Searching
Thanks to the great work of @jwinters01 (Intuit) Argo CD is now listing Applications faster: The UI is now compiled with React Compiler and lots of Rules-of-React were fixed. This allows caching of rendering work that would otherwise be repeated on every interaction. Toggling filtering and searching on large installations now do roughly 20–30% less rendering work.
UI Lazy loading
@jwinters01 also reduced the delay of page rendering by implementing lazy-loading of top-level UI routes, settings panels, and heavy applications-view features (tree view, terminal, logs, diff, panels), which previously were downloaded as a one big bundle.
Virtual Scrolling for Application Lists/Tiles when all is selected
Thanks to @aali309 (Red Hat) this release introduces virtual scrolling for the Applications list and tiles views, solving a long-standing problem for large-scale deployments. Previously, selecting “All” as the items-per-page option on an Applications view with thousands of apps could cause the UI to slow to a crawl or even crash outright, because every row or tile was rendered to the DOM at once. Now, when the app count exceeds the threshold of 50 applications Argo CD automatically switches to virtualized rendering: only the rows or tiles currently in the viewport are drawn.
Other New Web UI features
- Thanks to @gcadoret you can clear logs in the Pod Logs Viewer.
- To keep the robots away from publicly available instances Argo CD now sends HTTP header “X-Robots-Tag: noindex,nofollow” (@yugstar).
- The ui.loginButtonText field in argocd-cm now allows overriding SSO login button label (@NotKiwy)
- Event lists now report the name of the component that generated the event (@Suven-p).
Appset Controller Improvements
New OCI Generator
Thanks to @robinlieb ApplicationSets gain the OCI generator, bringing OCI artifacts up to parity with the existing Git generators.
Instead of pointing an ApplicationSet at a Git repository and scanning for directories or config files, you can now point it at an OCI artifact and have Argo CD discover directories and parse files inside it, using the same include/exclude pattern semantics you already know from the Git directory and file generators. So the teams who have moved their manifests, Helm charts, or config bundles into an OCI registry no longer have to keep a Git repository around purely to drive ApplicationSet generation.
Progressive Sync Enhancements
Thanks to @ranakan19’s (Red Hat) relentless work on the Progressive Sync feature it got three important improvements in correctness, diagnosability, and observability:
Progressive sync no longer acts on stale data
When the ApplicationSet controller detects a change in any of its owned Applications, it now adds a refresh annotation to all owned Applications and waits for them to finish reconciling before starting the rollout, so the controller is never making step decisions on stale state. A new--refresh-grace-period-seconds flag lets operators define a minimum grace period before a progressive sync starts refreshing outdated Applications, giving webhooks time to land first.
Progressive sync was previously hard to observe in aggregate. This release adds four new metrics:
argocd_appset_progressive_sync_app_status, a gauge of how many Applications sit in each step by status
argocd_appset_progressive_sync_syncs_triggered_total, a counter of sync operations triggered per step
argocd_appset_progressive_sync_detection_to_trigger_seconds, a histogram measuring the time between detecting that a sync is needed and actually triggering it.
argocd_appset_app_refresh_total, a count of application refreshes triggered by the ApplicationSet for progressive sync.
All the metrics are labeled by namespace and AppSet name (and step, where relevant), so you can now build a dashboard that answers “which step a rollout is stuck on and for how long,” and configure alerts on stuck rollouts.
Until now, progressive sync exposed only one condition, and it reported that the rollout had completed even in cases where progressive sync never started at all. That made a successful rollout and a rollout that never began look identical. This change introduces two new reasons, ApplicationSetRolloutError and InvalidRolloutConfig, plus a new ApplicationSetConditionInvalidRolloutConfig condition type for situations that were previously only reported as warnings in the log, and it marks the ApplicationSet health as degraded when a RollingSync strategy has no steps defined.
`tpl` Function Support in ApplicationSet go-templating
Thanks to @twobiers go-templates now allow to use the Helm-like `tpl` function, which renders a string as a Go template at runtime. It allows to put template expressions into values that are created by generators, for example from yaml files that are parsed by the Git generator. This closes a feature request that had been open since 2023.
ApplicationSet Workqueue Rate Limiter flags
Thanks to @om7057 the ApplicationSet controller now exposes the same workqueue rate limiter flags that the Application controller has long supported. Previously, argocd-applicationset-controller relied on the controller-runtime default exponential backoff, which capped retries at roughly 16 minutes and offered no way to change it.
CLI enhancements
Support for server-proxy-url flag in cluster add command
@ppapapetrou76 (Octopus Deploy) added a new--server-proxy-url flag to argocd cluster add, decoupling the proxy used by the CLI from the proxy stored for the Argo CD server.
Previously, a single proxy setting, taken from--proxy-url or kubeconfig, was used for both the local connection the CLI makes to the target cluster to provision RBAC and to the connection used by the server to reach the cluster, with no way to override one without affecting the other.
New option to Omit status when Exporting Resources
@nitishfy (Akuity) added a new--strip-status option for argocd admin export that allows to export Application/AppProject/etc. resources without their full status. This allows the use of import/export as Gitops-style backup, which can be cleanly applied by kubectl.
Improved CLI signal handling for graceful cancellation.
Thanks to @Eduard-Voiculescu and @blakepettersson Argo CD now handles termination signals properly. Previously, interrupting a command didn’t always stop it correctly: the command could hang more than a minute or leave behind orphaned running server tasks. @Eduard-Voiculescu’s PR fixed this by adding proper context handling into CLI commands. @blakepettersson’s PR extends this development to the “argocd admin dashboard” command and to the server components of ARgoCD.
Application Controller Improvements
New “rollback” RBAC action
Thanks to @ppapapetrou76 Argo CD got a dedicated “rollback” RBAC action for Applications.
The action allows rolling back applications to previously synced revisions with some restrictions: it works only for non-autosynced Applications, can only rollback to revisions in the Application’s revision history, only allows to override options ‘dryRun’ and ‘prune’ options and cannot perform partial sync on specific resources.
This feature is disabled by default for backwards compatibility. To enable it, set the server.rbac.rollback.enforce.enable flag in argocd-cm configmap.
Label-based resource filtering
Thanks to @alexmt (Akuity) Argo CD now supports the resource.selectors key to argocd-cm that attaches a Kubernetes label selector to the list/watch calls Argo CD makes for a given API group/kind/cluster. Because the filtering happens server-side, non-matching objects are never transferred to Argo CD or loaded into the cluster cache at all, rather than being fetched and then discarded. For an Argo CD instance running against clusters with large volumes of irrelevant objects, this translates directly into a smaller cache footprint, lower memory use in the application controller, and reduced pressure on the Kubernetes API server
Configurable Serialization and Compression for Cached Manifests
@adityaraj178 contributed configurable serialization and compression for cached manifests in the gitops-engine cluster cache. When enabled, each manifest is serialized and compressed into a single byte array before being stored and replacing the deeply nested map[string]interface{} tree with a single heap allocation. Benchmarks on a live Argo CD instance showed controller memory dropping from 3.27 GB to under 1 GB with the default JSON plus fast-gzip combination — roughly a 70% reduction — for about 2.5% additional CPU, and another test on a large production instance measured around 30% memory savings with no observable CPU or performance cost.
Memory and CPU Usage Improvements
- Partial JSON Unmarshalling (@blakepettersson): Argo CD now uses lightweight partial JSON unmarshalling when full unmarshalling of resources is not needed, extracting only identifying metadata (apiVersion/kind/metadata.name/metadata.namespace) to reduce CPU/memory on hot paths.
- Dropping redundant application copies from the refresh path (@rumstead — Black Rock): eliminates redundant copying of Application objects from the Informer cache while performing Application refresh.
- Counting Applications per Cluster in a Single Pass (@rumstead): eliminates redundant list and the destination resolution by saving per-cluster Applications counts.
Repository Server Enhancements
Custom File Extensions for Directory-Type Sources
New feature developed by @nitishfy: until now, a directory source in Argo CD would only consider files ending in .yaml, .yml, .json, or .jsonnet as candidate manifests, and that filter was applied before the include/exclude globs were evaluated. This feature lets applications disable the implicit extension filter, so include becomes the single source of truth for which files the repo-server renders, allowing creating custom naming conventions.
Observability Enhancements
Support for OpenTelemetry tracing with configurable sampling
Thanks to an important contribution of @blakepettersson Argo CD’s OpenTelemetry support now goes beyond the auto-instrumented gRPC and HTTP calls it supported before. This release adds explicit spans across the application controller, the repo-server, and gitops-engine, so a single reconciliation can be followed end to end — from the controller through manifest generation in the repo-server, to the diff and sync work done by gitops-engine — with Argo CD -specific attributes attached to the spans.
The change also introduces an otlp-sample-ratio setting (configurable per component via <COMPONENT_NAME>_OTLP_SAMPLE_RATIO env. variable), letting you export just a part of traces and keep telemetry volume under control on large installations.
Hydration status label added to argocd_app_info
Until now, hydration health was only visible per-Application in the UI or API which made hydrator problems hard to diagnose and configure alerts on. Thanks to @crenshaw-dev (Intuit) Argo CD’s source hydrator now reports its state through Prometheus. The argocd_app_info metric got a new hydrator_status label that reflects the Source Hydrator phase for any Application with spec.sourceHydrator configured.
Custom Health Checks Improvements
This release brings about 20 new custom healthchecks, way too many to describe them in this blog. Thanks to everyone who contributes and helps to maintain custom health checks.
Support for Custom Deletion Messages
In the past, Argo CD always reported “Progressing” / “Pending deletion” for resources being deleted (i.e. metadata.deletionTimestamp set). Thanks to @crenshaw-dev Argo CD 3.6 brings two improvements:
- The default health check now lists any pending finalizers
- Health checks may return an optional deletionMessage value — if set, the default behavior is overridden
This new field allows health checks to report more detailed information about the resource being deleted and the pending finalizers. If a resource is pending deletion for a long time, the health message can explain to the user the purpose of the pending finalizers and warn about any risks of manually removing them.
Notable Bug Fixes
This release contains more than 170 bugfixes, Here we selected several of them that we consider most notable, addressing issues affecting reliability, performance, and user experience:
- clean the repo checkout on revision change (@emil-ep — Red Hat): After switching revisions, leftover charts/*.tgz files made the repo-server render the wrong Helm sub-chart version. Only a restart fixed it.
- controller dropped refresh requests during reconcile (@dudinea — Octopus Deploy): A webhook refresh or hydrate request that arrived during an ongoing refresh was lost until the next periodic poll.
- server-side diff respects impersonation (@pjiang-dev — Intuit): The diff dry-run used the controller’s own credentials instead of the impersonated ServiceAccount, bypassing the least-privilege boundary.
- ApplicationSet: restore ignoreApplicationDifferences after normalization (@pjiang-dev): Regression from the concurrency refactor: the controller overwrote fields that ignore rules should protect. This broke the argocd-image-updater write-back pattern.
- repo type normalization regressions (@peikk0, @Churi12): oci:// sources without credentials failed with “unsupported scheme oci”. Multi-source apps with an untyped Helm chart plus manifest-generate-paths failed with “repository not found”.
- reduce unnecessary sync operations caused by refreshes (@agaudreault): Every refresh during a sync re-ran the operation, with live GETs and repeated etcd writes. This is a large controller and API server load reduction for busy installs.
- fix of OIDC Refresh Flow (@ris-tlp): Argo CD now correctly renews expired user sessions using OIDC refresh tokens. Previously token refresh was broken and the user was redirected again to the login screen.
Thank You
Thank you to everyone who contributed code, reviews, docs, bug reports
and enhancements proposals. Welcome to all the new participants!
Where Can I Get the New Release?
For installation instructions and the full changelog, check out the
github release page.
Please see detailed upgrade instructions in the upgrading guide
We’d love to hear your feedback! Find us on the #argo-cd channel in
CNCF Slack to share your experience, report issues, or just say hi.
Argo CD v3.6 Less Memory, Faster UI, Better Rollout Visibility was originally published in Argo Project on Medium, where people are continuing the conversation by highlighting and responding to this story.
High-quality, ubiquitous, and portable telemetry to enable effective observability
Kubernetes attributes processor reaches v1.0.0 milestone
The Kubernetes attributes processor, which enriches your telemetry with Kubernetes metadata, has officially moved to v1.0.0! You can try it out on your custom distro, and it is also available as part of the latest opentelemetry-collector-contrib and ope