Keycloak incubating

Keycloak is an open-source identity and access management solution for modern applications and services, built on top of industry security standard protocols.

nightly

Admin v2: Mask client secrets for view-only users (#52731)

* Admin v2: Mask client secrets for view-only users

When a user lacks manage permission on a client, strip client secrets
from GET /clients/{client} and GET /clients responses to match v1 behavior.

Closes #52197

Signed-off-by: Shatrughan Rai <polyglot.dev@outlook.com>

# Conflicts:
# rest/admin-v2/services/src/main/java/org/keycloak/services/client/DefaultClientService.java

* Refactor getClientListSecretMaskedForViewOnly test

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Signed-off-by: Shatrughan Rai <polyglot.dev@outlook.com>

* Fix: test scope

fixes an unclosed Stream in the authorization test by wrapping it
in try-with-resources, matching the pattern used in tests.

Fixes: #52197

Signed-off-by: Shatrughan Rai <polyglot.dev@outlook.com>

* moving the fix to the new client resource type

Signed-off-by: Steve Hawkins <shawkins@redhat.com>

---------

Signed-off-by: Shatrughan Rai <polyglot.dev@outlook.com>
Signed-off-by: Steve Hawkins <shawkins@redhat.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Steve Hawkins <shawkins@redhat.com>

Fluentd graduated

Fluentd is an open source data collector for unified logging layer

Fluentd v1.19.4

Bug Fix

  • #5504 buffer: enforce decompression_size_limit on the chunk IO path
  • #5488 buffer: clamp exported buffer size metrics to non-negative values
  • #5487 buffer: fix spurious BufferOverflowError caused by queue_size leaking when a chunk purge fails
  • #5456 buffer: fix stage_byte_size leak on staged to unstaged chunk demotion that could eventually raise spurious BufferOverflowError in plugins implementing #format
  • #5502 in_syslog: enforce message_length_limit on TCP/TLS transport
    • The default value of message_length_limit is changed from 2048 to 8192 to match rsyslog's default MaxMessageSize.
  • #5501 output: fix incomplete path traversal check in extract_placeholders
  • #5423 output: treat JSON::GeneratorError as unrecoverable error
  • #5500 parser_syslog: fix NameError when RFC5424 timestamp has repeated spaces
  • #5497 parser_syslog: fix NameError when RFC3164 timestamp has repeated spaces
  • #5444 parser_syslog: avoid excessive backtracking when parsing malformed RFC5424 structured data
  • #5486 chunk: ensure to close the Tempfile for decompressed data
  • #5471 plugin base: bound the number of worker lock files by hashing the path into a fixed set of buckets
  • #5445 out_forward: stop the endless "ack in response and chunk id in sent data are different" warning storm by discarding (instead of reusing) a keepalive socket whose ack failed or came back with a mismatched chunk id
  • #5478 config: accept empty lines in quoted strings
  • #5433 config: fix a config error when a single scalar value is given to an array option in YAML config syntax (for example retryable_response_codes: 503)
  • #5472 supervisor: reduce memory usage of cleanup_lock_dir with huge number of lock files

Enhancement

  • #5503 in_http: add <auth> for basic authentication and <security> for client network allowlisting

Misc

Contributors to this release (Alphabetical order)

  • Akash Kumar
  • Alexander Urieles
  • Arpit Jain
  • JSap0914
  • Kentaro Hayashi
  • Mehrdad Biukian
  • onice
  • Shizuo Fujita
  • Vedant Madane
LoxiLB sandbox

eBPF based cloud-native load-balancer. Powering Kubernetes|Edge|5G|IoT|XaaS Apps.

vlatest

Merge pull request #874 from TrekkieCoder/main

gh-868 Generate packages runnable with systemd

Porter sandbox

Porter enables you to package your application artifact, client tools, configuration and deployment logic together as a versioned bundle that you can distribute, and install with a single command

canary

This is a "canary" release of the most recent commits on our main branch. It is not a tagged release and should not be used unless you are helping us test out a new feature. Canary is not stable. If you want a stable release, use latest.

kagent sandbox

Kagent is an open source programming framework designed for DevOps and platform engineers to run AI agents in Kubernetes

v1.0.0-alpha5

What's Changed

Features

Bug Fixes

  • fix: flush buffered OTEL logs by @yashrajshuklaaa in #2772
  • fix(byo): centralize ADK runtime configuration by @EItanya in #2965
  • fix(kagent-adk): compact MCP App tool results only when the result carries a UI resource by @AmirF194 in #2579
  • fix: wire Helm authentication settings into controller startup by @EItanya in #2971

Documentation

  • docs: add API design guidelines and refresh OIDC architecture by @EItanya in #2958
  • refactor: standardize environment settings and generate their reference by @EItanya in #2967

Other Changes

  • chore: keep planning documents in gitignored .plans by @EItanya in #2955
  • chore: update Substrate to v0.3.0-alpha1 by @EItanya in #2963
  • ci: overlap slow E2E harness scenarios by @EItanya in #2969

New Contributors

Full Changelog: v1.0.0-alpha4...v1.0.0-alpha5

Open Policy Containers sandbox

A docker-inspired CLI for building, tagging, pushing, pulling, and signing OPA policies to and from OCI-compliant registries.

policy v0.4.2

Changelog

xRegistry sandbox

The xRegistry project defines an abstract model for managing metadata about resources and provides a REST-based interface to discover, create, modify and delete those resources.

dev

Latest development build of the 'xr(server)' executables. The commit pointer and zip/tar files are old, do not use them.

Cortex incubating

A multitenant, horizontally scalable Prometheus as a Service

v1.22.0-rc.2

This is the third release candidate for v1.22.0.

It contains one fix on top of v1.22.0-rc.1: #7858, which fixes the new cortex_ingester_ingestion_delay_seconds histogram losing most of its observations. It also moves the integration tests to a project-hosted MinIO image (#7864); that change does not affect the release artifacts.

Images are published for cortex, query-tee, test-exporter and thanosconvert:

docker pull quay.io/cortexproject/cortex:v1.22.0-rc.2

Changelog

  • [CHANGE] Ruler: Remove the deprecated -ruler.evaluation-delay-duration flag and its ruler_evaluation_delay_duration per-tenant limit. Use -ruler.query-offset / ruler_query_offset, which no longer takes the higher of the two values. Cortex decodes the runtime config strictly, so a leftover ruler_evaluation_delay_duration override makes the runtime config fail to load: Cortex exits at startup (module failed, module=runtime-config), and on an already-running process every reload fails, pinning the last good overrides and dropping cortex_runtime_config_last_reload_successful to 0. Run grep -r ruler_evaluation_delay_duration over your runtime configs before upgrading. #7792
  • [CHANGE] Remove the deprecated -<prefix>.fifocache.size flag and its size YAML field (deprecated in 1.1.0). Use -<prefix>.fifocache.max-size-items or -<prefix>.fifocache.max-size-bytes; a cache configured only via size now starts with no capacity. #7791
  • [CHANGE] Querier: Remove the deprecated -querier.ingester-metadata-streaming flag and its ingester_metadata_streaming YAML field (deprecated in 1.18.0, default true). Streaming RPCs are now always used for the metadata APIs. Also removes the dead hidden ingester_streaming YAML field left over from -querier.ingester-streaming. #7791
  • [CHANGE] Remove deprecated CLI flags that have been no-ops for at least two minor releases. All of them were flag-only (no YAML config option) and already had no effect, so the only impact is that passing them now fails at startup. Remove them from your command lines before upgrading. #7790
    • -querier.ingester-streaming (deprecated in 1.17.0)
    • -querier.iterators (deprecated in 1.17.0)
    • -querier.batch-iterators (deprecated in 1.17.0)
    • -querier.query-store-for-labels-enabled (deprecated in 1.18.0)
    • -querier.max-outstanding-requests-per-tenant (deprecated in 1.18.0; use -frontend.max-outstanding-requests-per-tenant)
    • -query-scheduler.max-outstanding-requests-per-tenant (deprecated in 1.18.0; use -frontend.max-outstanding-requests-per-tenant)
    • -blocks-storage.tsdb.wal-compression-enabled (deprecated in 1.19.0; use -blocks-storage.tsdb.wal-compression-type)
    • -ingester.max-series-per-query (a chunks-storage limit, ignored since blocks storage; use -querier.max-fetched-series-per-query)
  • [CHANGE] Ingester: Formally deprecate -blocks-storage.tsdb.max-exemplars, scheduled for removal in v1.24.0. Use the per-tenant max_exemplars limit instead. The flag still works as the global fallback when max_exemplars is 0, but setting it now logs a warning and increments deprecated_flags_inuse_total. #7793
  • [CHANGE] Ingester: Graduate native histogram ingestion (-blocks-storage.tsdb.enable-native-histograms) from experimental. #7789
  • [CHANGE] Querier: Make query time range configurations per-tenant: query_ingesters_within, query_store_after, and shuffle_sharding_ingesters_lookback_period. Uses model.Duration instead of time.Duration to support serialization but has minimum unit of 1ms (nanoseconds/microseconds not supported). #7160 #7323
  • [CHANGE] Alertmanager: Remove the obsolete startup migration of local state files into per-tenant directories (scheduled for removal in 1.11.0). Upgrading from a release older than 1.9.0 with a persisted local state directory now requires upgrading to an intermediate release first, so the migration can run. #7513
  • [CHANGE] Cache: Setting -blocks-storage.bucket-store.metadata-cache.bucket-index-content-ttl to 0 will disable the bucket-index cache. #7446
  • [CHANGE] HA Tracker: Move -distributor.ha-tracker.failover-timeout from a global config to a per-tenant runtime config. The flag name and default value (30s) remain the same. #7481
  • [FEATURE] Parquet: Support sharded parquet file conversion and querying. #7610
  • [FEATURE] Parquet Converter: Add experimental -parquet-converter.max-num-columns flag to automatically shard parquet files when the number of columns exceeds the configured limit. This prevents failures when a TSDB block has more unique label names than the parquet library's column limit (32767). #7624
  • [FEATURE] Distributor: Add experimental -distributor.num-query-workers flag to use a goroutine worker pool for query fan-out calls to ingesters. Reuses pre-grown goroutine stacks to eliminate the runtime.copystack overhead (~8% CPU) observed on rulers with wide ingester fan-out. Falls back to spawning a new goroutine when no worker is available. #7623
  • [FEATURE] Ingester: Add experimental active series tracker that counts active series by configurable label matchers (including regex) per tenant and exposes cortex_ingester_active_series_per_tracker metric. Configured via active_series_trackers in runtime config overrides. #7476
  • [FEATURE] Ingester: Add experimental head-only queried series metric. cortex_ingester_queried_head_series tracks unique series queried from head via HLL. Enabled via -ingester.head-queried-series-metrics-enabled. #7500
  • [FEATURE] Ruler: Add per-tenant ruler_alert_generator_url_template runtime config option to customize alert generator URLs using Go templates. Includes a jsonEscape template function for safely embedding expressions in JSON-encoded URL parameters (e.g., Grafana Explore panes). Supports Grafana Explore, Perses, and other UIs. #7302 #7458
  • [FEATURE] Distributor: Add experimental -distributor.enable-start-timestamp flag for Prometheus Remote Write 2.0. When enabled, StartTimestamp (ST) is ingested. #7371
  • [FEATURE] Memberlist: Add -memberlist.cluster-label and -memberlist.cluster-label-verification-disabled to prevent accidental cross-cluster gossip joins and support rolling label rollout. #7385
  • [FEATURE] Querier: Add timeout classification to classify query timeouts as 4XX (user error) or 5XX (system error) based on phase timing. When enabled, queries that spend most of their time in PromQL evaluation return 422 Unprocessable Entity instead of 503 Service Unavailable. #7374
  • [FEATURE] Querier: Implement Resource Based Throttling in Querier. #7442
  • [FEATURE] Querier: Add resource-based query eviction that automatically cancels the heaviest running query when CPU or heap utilization exceeds configured thresholds. #7488
  • [FEATURE] Storage: Add support for Oracle Cloud Infrastructure (OCI) Object Storage as a backend for blocks, ruler, and alertmanager storage. Configured via -<prefix>.oci.* flags with backend: oci. #7718
  • [FEATURE] StoreGateway: Add experimental optional limit blocks-storage.bucket-store.max-concurrent-data-bytes on the data bytes (postings, series and chunks) fetched via the Series() API call and processed concurrently across all queries per store gateway to protect from oomkill. This returns an error that is retryable at querier level. #7271
  • [FEATURE] Engine: Add -querier.selector-batch-size and -ruler.selector-batch-size flags to configure series batching in the Thanos promQL engine. 0 disables batching. #7763
  • [ENHANCEMENT] Upgrade prometheus alertmanager version to v0.32.1. #7462
  • [ENHANCEMENT] Tenant Federation: Avoid purging the regex resolver LRU cache on user-sync ticks when the set of known users has not changed. #7489
  • [ENHANCEMENT] Parquet Converter: Add parquet-converter.max-block-label-names limit to skip conversion of TSDB blocks with too many label names. #7625
  • [ENHANCEMENT] Parquet Converter: Add a ring status page to expose the ring status. #7455
  • [ENHANCEMENT] Parquet: Add -blocks-storage.bucket-store.parquet-query-concurrency flag to configure the maximum number of concurrent goroutines applied at each level of parquet query processing in store-gateway: shard querying, row group processing, and column materialization. #7613
  • [ENHANCEMENT] Parquet: Add a row ranges cache for parquet query filtering in querier and store-gateway. #7478
  • [ENHANCEMENT] Ingester: Add cortex_ingester_ingestion_delay_seconds native histogram metric to track the delay between sample ingestion time and sample timestamp. #7443
  • [ENHANCEMENT] Ingester: Add WAL record metrics to help evaluate the effectiveness of WAL compression type (e.g. snappy, zstd): cortex_ingester_tsdb_wal_record_part_writes_total, cortex_ingester_tsdb_wal_record_parts_bytes_written_total, and cortex_ingester_tsdb_wal_record_bytes_saved_total. #7420
  • [ENHANCEMENT] Distributor: Introduce dynamic Symbols slice capacity pooling. #7398 #7401
  • [ENHANCEMENT] Metrics Helper: Add native histogram support for aggregating and merging, including dual-format histogram handling that exposes both native and classic bucket formats. #7359
  • [ENHANCEMENT] Cache: Add per-tenant TTL configuration for query results cache to control cache expiration on a per-tenant basis with separate TTLs for regular and out-of-order data. -frontend.out-of-order-results-cache-ttl falls back to -frontend.results-cache-ttl when unset, and then to the global cache backend TTL. #7357 #7775
  • [ENHANCEMENT] Update build image and Go version to 1.26. #7434 #7437 #7716 #7726
  • [ENHANCEMENT] Upgraded container base images from alpine:3.23 to gcr.io/distroless/static-debian12, reducing image size and attack surface. #7637
  • [ENHANCEMENT] Query Scheduler: Add cortex_query_scheduler_tracked_requests metric to track the current number of requests held by the scheduler. #7355
  • [ENHANCEMENT] Compactor: Prevent partition compaction to compact any blocks marked for deletion. #7391
  • [ENHANCEMENT] Distributor: Optimize memory allocations by reusing the existing capacity of these pooled slices in the Prometheus Remote Write 2.0 path. #7392
  • [ENHANCEMENT] Upgrade gRPC from v1.71.2 to v1.79.3 to address CVE-2026-33186. #7460 #7463
  • [ENHANCEMENT] Query Frontend: Add query_too_expensive reason to QFE and reason field to query stats. #7479
  • [ENHANCEMENT] Instrument Ingester CPU profile with source for read APIs. #7494
  • [ENHANCEMENT] Ingester: Convert expanded postings cache from FIFO to LRU eviction to retain frequently-queried entries under memory pressure. #7510
  • [ENHANCEMENT] Querier: Detach series label and chunk data from gRPC unmarshal buffers in store-gateway streaming path, allowing the Go GC to reclaim receive buffers. #7519
  • [ENHANCEMENT] Distributor: Added cortex_distributor_received_histogram_buckets metric to track number of buckets in received native histogram samples before validation, per user. #7569
  • [ENHANCEMENT] Ingester: Add lazy regex evaluation on head postings cache miss. Defers expensive regex matchers on high-cardinality labels to per-series filtering when a selective equality matcher already narrows the result set. Configured via -blocks-storage.expanded_postings_cache.head.lazy-matcher-max-cardinality (disabled by default). #7553
  • [ENHANCEMENT] Store Gateway: Resolve the parquet shard count from the bucket index instead of reading the converter mark for each block, reducing object storage calls when the bucket index is enabled. A component label is added to the bucket index loader metrics to distinguish store-queryable and store-gateway. #7648
  • [ENHANCEMENT] Query Frontend: Improve the slow query log with source, user_agent, engine_type, block_store_type, and query stats fields to aid slow query diagnosis. #7601
  • [ENHANCEMENT] Ring: Add ring metric to count number of duplicate tokens. #7626
  • [ENHANCEMENT] Upgrade prometheus version to v3.9.1. #7535
  • [ENHANCEMENT] Metrics: Add native histogram support to all remaining production histograms, enabling dual-format (classic + native) exposition across all Cortex components. #7636
  • [ENHANCEMENT] Ring: Cache ShuffleShardWithLookback subrings. The cached entry is invalidated on topology change or once now reaches the earliest RegisteredTimestamp + lookbackPeriod of any included instance. #7628
  • [ENHANCEMENT] Query Frontend: Rename time_taken field to time_taken_ms and make it return millisecond count. #7649
  • [ENHANCEMENT] Update prometheus alertmanager version to v0.33.0. #7647
  • [ENHANCEMENT] Querier/Ingester: Detach ingester series from gRPC buffers to reduce heap. #7670
  • [ENHANCEMENT] Ingester: Add cortex_ingester_tsdb_head_max_timestamp metric that re-exports the TSDB head max timestamp (prometheus_tsdb_head_max_time) per user, to help investigate ingestion issues like out-of-bounds (too old sample) errors. #7694
  • [ENHANCEMENT] Ingester: Include the TSDB head max time in the out of bounds and too old sample error messages, so that users can see how far behind the accepted time range a rejected sample is. #7695
  • [ENHANCEMENT] Compactor: Reduce object storage GET calls when updating the bucket index by skipping re-reading parquet converter markers for blocks that already have a valid-version parquet entry in the previous index. #7669
  • [ENHANCEMENT] Upgrade Thanos and promql-engine to latest. #7505 #7691 #7740 #7788
  • [ENHANCEMENT] Ruler: Adjust ruler frontend decoder to not wrap query error messages with execution prefix, this makes error responses consistent between internal and external ruler paths. #7741
  • [ENHANCEMENT] Distributor: Deduplicate metric metadata when converting PRW 2.0 requests. PRW 2.0 attaches metadata to every series, so a metric family was previously expanded into one MetricMetadata per series. #7760
  • [ENHANCEMENT] Ingester: Add cortex_ingester_head_metric_names gauge exposing the number of unique metric names in the TSDB head per tenant. Registered when -ingester.active-series-metrics-enabled is true. #7514
  • [ENHANCEMENT] Query Frontend: Log X-Grafana-User header in query stats, slow query, and query request logs when Grafana's send_user_header is enabled. #7799
  • [ENHANCEMENT] Querier: Use non-pointer HistogramBucket slice in response codec. #7809
  • [ENHANCEMENT] Update build image and Go version to 1.27.0. #7807 #7814
  • [ENHANCEMENT] Querier: Reduce merge iterator BatchSize from 12 to 8. #7823
  • [BUGFIX] Querier: Fix queryWithRetry and labelsWithRetry returning (nil, nil) on cancelled context by propagating ctx.Err(). #7370
  • [BUGFIX] Metrics Helper: Fix non-deterministic bucket order in merged histograms by sorting buckets after map iteration, matching Prometheus client library behavior. #7380
  • [BUGFIX] Distributor: Return HTTP 401 Unauthorized when tenant ID resolution fails in the Prometheus Remote Write 2.0 path. #7389
  • [BUGFIX] Packaging: Fix RPM and deb packages to install the binary to /usr/bin, install the systemd unit to the correct system path (/usr/lib/systemd/system for RPM, /lib/systemd/system for deb), and mark the sysconfig/default env file as a config file so it is not overwritten on upgrade. #7445
  • [BUGFIX] Compactor: Handle not-found and access-denied errors from Attributes() in bucket index updater, preventing a stale cached Get() from causing the entire cleanup cycle to fail when meta.json has been deleted from object storage. #7454
  • [BUGFIX] Compactor: Fix stale cortex_bucket_index_last_successful_update_timestamp_seconds metric not being cleaned up when tenant ownership changes due to ring rebalancing. This caused false alarms on bucket index update rate when a tenant moved between compactors. #7485 #7487
  • [BUGFIX] Compactor: Fix flake in TestCompactor_DeleteLocalSyncFiles and TestPartitionCompactor_DeleteLocalSyncFiles by polling on user ownership rather than just the CompactionRunsCompleted counter, which increments even when the second compactor sees zero owned users due to a transient ring-view skew at startup. #7565
  • [BUGFIX] Ingester: Close TSDB when compaction fails during createTSDB, preventing resource leaks (file descriptors, mmap handles) that could lead to ingester instability. #7560
  • [BUGFIX] Tenant Federation: Fix result cache returning stale data after a new tenant is added when -tenant-federation.regex-matcher-enabled=true. The resolved tenant set is now hashed and included in the cache key so that any change to the matched tenant list automatically invalidates cached entries. Non-regex users are unaffected. #7562
  • [BUGFIX] Tenant Federation: Fix regex resolver clearing known users list when user scan fails. #7534
  • [BUGFIX] Ingester: Fix inflight query counter leak when resource-based query protection rejects a request. #7539
  • [BUGFIX] Ingester: Release the TSDB appender on every early-return path in Push (e.g. out-of-order label set) by deferring Rollback. Previously such requests leaked TSDB head series references, mmap'd chunks and pending state per request, causing the cortex_ingester_tsdb_head_active_appenders gauge to grow unbounded. #7528
  • [BUGFIX] Ingester: Fix panic: send on closed channel in ActiveQueriedSeriesService on shutdown by removing the redundant channel close in stopping() and relying on ctx.Done() to signal worker exit. #7533
  • [BUGFIX] Ring: Fix ring token conflict resolution only applied to updated instance and make constantly token conflict check during instance observe period. #7554
  • [BUGFIX] Ring: Fix DoBatch never running its cleanup callback when a per-instance callback panics. wg.Done() is now deferred, so wg.Wait() no longer blocks forever and the context timers and request buffers owned by the cleanup function are released. #7559
  • [BUGFIX] Query Frontend: Fix native histogram responses not being handled correctly in minTime() sort ordering for split_by_interval merge. #7555
  • [BUGFIX] Compactor: Ensure visit marker heartbeat goroutine completes before blocks cleaner returns. #7386
  • [BUGFIX] Querier: Fix unbounded resource leak in the bucket-scan blocks finder (used when the bucket index is disabled). Per-tenant metadata fetchers, their Prometheus registries, and on-disk meta caches are now evicted once a tenant is no longer active, instead of being retained for the lifetime of the process. #7573
  • [BUGFIX] Alertmanager: Fix data race between ApplyConfig's dispatcher/inhibitor startup and Stop during config reload and shutdown, and reject lazy tenant creation after shutdown begins. #7618
  • [BUGFIX] Distributor: Release the push worker pool goroutines on shutdown by stopping the async executor during the stopping phase when -distributor.num-push-workers is set. #7602
  • [BUGFIX] Query Frontend: Track data selection min and max time of query requests in query stats regardless of whether query priority and query rejection are enabled. #7724
  • [BUGFIX] Querier: Fix flake in integration tests TestQuerierWithStoreGatewayDataBytesLimits and TestQuerierWithBlocksStorageLimits by waiting for the querier to see the store-gateway ACTIVE in the ring before querying. #7614
  • [BUGFIX] Ruler: Register xfunctions (xincrease, xrate, xdelta) in the global parser before loading rule files. #7621
  • [BUGFIX] Security: Reject empty entries in -distributor.sign-write-requests-keys caused by stray or trailing commas (e.g. newkey,). Previously these were silently accepted and produced an empty signing key, which downgraded HMAC stream-push authentication to a forgeable signature. Misconfigured flags now fail at process startup; audit your configs before upgrading. #7587
  • [BUGFIX] Config: Fix validation of explicit zero values in non-empty YAML root sections, allowing configs such as flusher: { exit_after_flush: false } while continuing to reject empty root sections. #7700
  • [BUGFIX] Querier: Fix panic due to request tracker truncating multi-byte UTF-8 character #7640
  • [BUGFIX] Ingester: Fix panic (HistogramProtoToHistogram called with a float histogram) when ingesting a float native histogram with a zero count (e.g. a staleness marker or empty histogram). The decoder is now selected by histogram type via IsFloatHistogram() instead of by count value. #7645
  • [BUGFIX] Querier: Fix parquet queryable fallback returning a nil error instead of the actual query error in LabelValues and LabelNames. #7638
  • [BUGFIX] Storage: Default the Azure endpoint_suffix to blob.core.windows.net instead of empty. Cortex builds the Thanos Azure config directly and bypasses Thanos' default, so an unset suffix produced an invalid FQDN (<account>.) and components hung on startup with DNS errors. #5449 #7687
  • [BUGFIX] Store Gateway: Fix misleading "no index cache backend addresses" validation error being reported for chunks-cache, metadata-cache, and parquet caches when their memcached or redis backend is configured without addresses. The message is now the cache-type-agnostic "no cache backend addresses". #7675
  • [BUGFIX] Querier/Query Frontend: Fix DNS watcher dropping all query-frontend/scheduler worker connections on a transient DNS lookup failure. #7698
  • [BUGFIX] Ring: Fix DynamoDB KV CAS not retrying on transactional conditional check failures. TransactWriteItems reports condition failures as TransactionCanceledException with a ConditionalCheckFailed cancellation reason, which was not recognized as retryable, so any concurrent ring update conflict (e.g. many ingesters joining during a rolling update) failed immediately instead of re-reading and retrying. TransactionConflict cancellation reasons are also treated as retryable. #7706
  • [BUGFIX] Distributor: Return HTTP 499 (Client Closed Request) instead of 500 when a remote-write or OTLP push is canceled by the client, so client-side cancellations are no longer counted as server-side errors. #7717
  • [BUGFIX] Querier: Fix gRPC codes.Canceled errors being mapped to HTTP 500 instead of 499 when a client cancels a query. #7738
  • [BUGFIX] Fix the gRPC DNS watcher's SRV record path deleting all known endpoints when the SRV query succeeds but every target's A record lookup fails. #7745
  • [BUGFIX] Compactor: Fix spurious bucket operation fail after retries error logs emitted during partial block cleanup. #7749
  • [BUGFIX] Alertmanager: Fix panic in validateAlertmanagerConfig when receiver config traversal encounters nil interface values. #7751
  • [BUGFIX] Parquet Converter: Fix auto_forget_delay having no effect. The ring lifecycler was created without the auto-forget delegate, so unhealthy instances were never automatically removed from the ring. #7752
  • [BUGFIX] Compactor: Properly handle error from ReadPartitionedGroupInfo in UpdatePartitionedGroupInfo. #7766
  • [BUGFIX] Alertmanager: Reject the global mattermost_webhook_url_file setting in per-tenant configs, consistent with every other global *_file setting. #7768
  • [BUGFIX] Alertmanager: Tighten per-tenant config validation to reject additional file-based settings. #7767
  • [BUGFIX] Querier: Fix panic (index out of range [-1]) in the active request tracker when truncating a match[]/query value made entirely of invalid UTF-8 continuation bytes. The backwards scan for a rune boundary now stops at index 0 instead of underflowing. #7743
  • [BUGFIX] Config: Fix CSV-list flags/YAML fields (e.g. -compactor.enabled-tenants) treating an explicitly empty string as a one-element list containing an empty tenant name instead of an empty list. #7714
  • [BUGFIX] Tenant Federation: Fix regex tenant federation dropping tenants when -blocks-storage.users-scanner.cache-ttl is set. The regex resolver sorted the user list returned by the users scanner in place, corrupting the scanner cache and progressively losing tenants on every sync until the cache expired. #7812
  • [BUGFIX] Tenant Federation: Fix regex tenant federation resolving to an empty user list right after startup. #7811
  • [BUGFIX] Ingester: Don't count a forced head compaction skipped because blocks shipping is in progress as a failure. Previously such skips incremented cortex_ingester_tsdb_compactions_failed_total, producing spurious alerts. #7842
Spin sandbox

Spin is a framework for building and deploying serverless applications in WebAssembly.

canary

This is a "canary" release of the most recent commits on our main branch. Canary is not stable.
It is only intended for developers wishing to try out the latest features in Spin, some of which may not be fully implemented.

Submariner sandbox

Submariner enables direct networking between Pods and Services in different Kubernetes clusters, either on-premises or in the cloud.

0.23.4

Update Submariner dependencies to v0.23.4

Signed-off-by: Automated Release <release@submariner.io>

Radius sandbox

Radius is a cloud-native application platform that enables developers and the platform engineers that support them to collaborate on delivering and managing cloud-native applications that follow organizational best practices for cost, operations and security, by default.

Radius v0.61.0

Announcing Radius v0.61.0

Today we're happy to announce the release of Radius v0.61.0. Check out the highlights below, along with the full changelog for more details.

We would like to extend our thanks to all the new and existing contributors who helped make this release possible!

Intro to Radius

If you're new to Radius, check out our website, radapp.io, for more information. Also visit our getting started guide to learn how to install Radius and create your first app.

Highlights

Updates to extensible Radius resources

This release adds more support for managing Radius.Core resources: rad resource list --preview can now list preview environments and applications; rad env create and rad env update --preview can reference a Recipe Pack in another resource group; and each Radius.Core environment must use a unique Kubernetes namespace. Deleting a preview environment now also deletes its applications and resources, rather than leaving them behind.

rad deploy now warns when a template uses the older Applications.* resource types and suggests the corresponding Radius.* replacements where available. Support for the old types will be removed in an upcoming release, so use the warning to migrate your templates. The deployment continues normally in this release.

Deploy templates directly from HTTPS URLs

rad deploy now supports deploying Bicep and ARM JSON templates hosted at HTTPS URLs. Radius validates the URL, resolves relative Bicep module references from the remote template’s location, and redacts URL credentials from command output. See #12676 for details.

Share secrets between connected resources

Connections to Recipe-backed resources can now provide both ordinary values and references to producer-managed secrets. Radius stores Recipe secret outputs in managed Radius.Security/secrets resources and makes references available to consuming Recipes without decrypting or exposing secret values. See #12709 for details.

More services in the default Kubernetes Recipe Pack

The default Kubernetes Recipe Pack now includes PostgreSQL, RabbitMQ, and Redis, expanding the backing services that can be deployed without configuring a custom Recipe Pack. You can download the pack manually and use it. Automatic install of this recipe pack via rad init is still in progress.

Live deployment progress for Radius Canvas

Radius deploy workflows now publish rotating deployment-progress snapshots that Radius Canvas can use to display resource status while a deployment is running. The snapshots redact environment-variable values and are retained for one day. See #12727 for details.

Prevent resource-name collisions between applications

Recipe-created cloud resources are now named using their application and environment, preventing separate applications or environments from accidentally sharing the same resource. This avoids conflicts during deployment and helps protect resources from being changed or deleted by another application. See [#12987](#12987) for details.

Breaking changes

This change affects you only if you used Radius’s older Kubernetes onboarding features for your application workloads. It does not remove support for running Radius on Kubernetes or affect workloads that don’t rely on these features.

If you applied a Kubernetes manifest that creates a Recipe custom resource in your cluster, Radius will no longer act on that object to provision backing services or refresh generated connection Secrets. This is the legacy Kubernetes Recipe resource—not a Radius Recipe definition or a Recipe Pack. If you added radapp.io/enabled or related annotations to a Kubernetes Deployment so Radius would provide connection settings as environment variables, Radius will no longer inject or update those settings.

To keep using Radius to deploy and connect application resources, migrate these workloads to Bicep/ARM deployments or Flux GitOps before upgrading. Both remain supported. Helm may leave old Recipe custom resource definitions and objects in the cluster; the upgrade does not migrate or clean them up.See #12952 for details.

New contributors

Welcome to our new contributors who have merged their first PR in this release!

Upgrading to Radius v0.61.0

You can upgrade to this release by upgrading your Radius CLI then running rad upgrade kubernetes. Only incremental version upgrades are supported. Consult the upgrade documentation for full details.

Review the breaking changes before upgrading if you use the legacy Kubernetes Recipe custom resource or annotation-based Deployment onboarding.

Full changelog

Full Changelog: v0.60.0...v0.61.0

KCL sandbox

A constraint-based record & functional language mainly used in configuration and policy scenarios.

v0.13.0

What's Changed

  • fix: the release of linux arm64 artifact is missing by @zong-zhe in #1923
  • fix: correct IP address type tests by @johngmyers in #1925
  • chore: let ci for macos arm not failfast by @liangyuanpeng in #1931
  • chore: change oci-distribution to new crate name:oci-client by @liangyuanpeng in #1932
  • feat: add a deepwiki doc reference by @zong-zhe in #1933
  • chore: bump wasm32-wasi to wasm32-wasip1 by @zong-zhe in #1934
  • add some tests for net by @liangyuanpeng in #1936
  • chore: add some unit tests for net by @liangyuanpeng in #1942
  • Fix: Ensure math.floor Returns an Integer Instead of a Float by @priyansh-saxena1 in #1946
  • fix: enforce type checking for list of schema elements by @priyansh-saxena1 in #1949
  • fix: invalid override spec error message by @Peefy in #1951
  • perf: use rustc hash for the indexmap in the whole proj by @Peefy in #1952
  • fix net to_ip4 and to_ip6 by @liangyuanpeng in #1937
  • fix: rename ip16 to ip6 by @liangyuanpeng in #1962
  • Docs: Be specific about UUID version by @jfharden in #1996
  • Fix typos in language server documentation by @aliazlan4 in #1994
  • fix: (lsp formatter) preserve trailing comments within collection exp… by @aliazlan4 in #1993
  • fix: replace serde_yaml with serde_yaml_ng by @vlada-dudr in #1987
  • chore: update Rust dependencies to latest compatible versions by @turner-hemmer in #1997
  • chore: remove un-used vars and flags by @Peefy in #1998
  • refactor: kcl api protobuf build by @Peefy in #2001
  • chore(deps): bump quinn-proto from 0.11.3 to 0.11.13 in /kclvm/tools by @dependabot[bot] in #2002
  • refactor: doc badges by @Peefy in #2003
  • refactor: cargo workspace for the entire project by @Peefy in #2004
  • chore: remove un-used compiler crate by @Peefy in #2005
  • ci: kcl lib linux glibc 2.17 release by @Peefy in #2006
  • refactor: rename kcl api function args and result by @Peefy in #2007
  • chore: bump rust 1.91 and edition 2024 & fix all lint issues in compiler_base by @Peefy in #2008
  • refactor: remove un-used runtime components by @Peefy in #2009
  • fix: kcl code lint errors by @Peefy in #2010
  • fix: wasm build issues by @Peefy in #2011
  • fix: serde yaml 1.1 on and yes string value by @Peefy in #2012
  • chore: use kcl cli downloads badge by @Peefy in #2015
  • fix: kcl linux release in CI by @Peefy in #2014
  • feat: add index signature for the KCL schema type API by @Peefy in #2016
  • feat: add function type for the KCL schema type API by @Peefy in #2017
  • fix: musl lib release by @Peefy in #2019
  • fix: vet failed on lambda calling by @Peefy in #2025
  • fix: lambda para wrong scope by @Peefy in #2028
  • fix: lazy eval for the if stmt by @Peefy in #2029
  • feat: parse file preserve symlink path by @Peefy in #2030
  • fix: template exec yaml encode obj by @Peefy in #2031
  • fix: json array literal stream output by @Peefy in #2032
  • fix: schema parameters config set val eval by @Peefy in #2033
  • fix: schema parameter validation by @Peefy in #2035
  • fix: schema dup instances by @Peefy in #2036
  • fix: option builtin func parameter types by @Peefy in #2037
  • feat: lambda lazy eval in the schema internal scope by @Peefy in #2038
  • fix: eval lambda forward reference by @Peefy in #2039
  • fix: format style on 2 empty lines at EOF by @Peefy in #2040
  • fix: config list parse by @Peefy in #2044
  • feat(json): add merge function for RFC 7396 JSON Merge Patch by @Arpit529Srivastava in #2045
  • fix: runtime int parse error by @Peefy in #2055
  • feat: add selector expr parse in the quant expr target by @Peefy in #2058
  • perf: sema resolver for source map and homogeneous schema array by @Peefy in #2060
  • docs: add building and testing chapters (fixes #1935) by @DCchoudhury15 in #2062
  • refactor: remove unwrap() calls in advanced resolver (#1166) by @jschoone in #2064
  • Add Upbound to ADOPTERS.md by @ytsarev in #2066
  • chore: fix compiler warnings by @johngmyers in #2071
  • feat(builtin): add reduce function by @johngmyers in #2070
  • fix(evaluator): update lazy scope cache when variable is reassigned by @johngmyers in #2075
  • fix(evaluator): update lazy scope cache for += list assignment by @johngmyers in #2077
  • fix(evaluator): ensure assign-by-value semantics for schema assignment by @Peefy in #2081
  • feat(parser): support braces for lambda arguments and multiline support by @f4z3r in #2082
  • feat(format): format lambda signatures based on line length by @f4z3r in #2085
  • Update Cargo.toml for performance by @rtainaan in #2088
  • perf(ast): use monotonic counter for AstIndex instead of Uuid::new_v4 by @rtainaan in #2091
  • perf(runtime,evaluator): use SmolStr for dict and schema config keys by @rtainaan in #2093
  • perf(runtime,sema): avoid per-call regex compile and quadratic scans in type-string helpers by @rtainaan in #2092
  • perf(parser): cache path canonicalization and package resolution during load by @rtainaan in #2094
  • perf(sema): cache parsed schema and rule doc strings during resolver run by @rtainaan in #2095
  • fix(evaluator): keep reassigned globals order-independent in lazy scope by @rtainaan in #2097
  • fix(evaluator): keep config keys literal when shadowed by a lambda param by @rtainaan in #2096
  • Add is_IPv6 to net and validations to datetime by @31puneet in #2098
  • fix(sema): prevent swallowed errors for unresolved schema expressions… by @DCchoudhury15 in #2099
  • Adding multi-line string dedenting by @31puneet in #2100
  • fix(evaluator): avoid recursive global aliases in evaluation by @f4z3r in #2103
  • fix(driver): guard current_dir() panic on wasm32 (#2104) by @Viscous106 in #2105
  • fix(runtime): release plugin dispatch lock before invoking handler to… by @Viscous106 in #2106
  • Core implementation for dry-run formatting by @f4z3r in #2108
  • chore: bump workspace version to 0.12.4 by @Peefy in #2110
  • fix(lsp): perform inherited attribute lookup on schema resolution by @f4z3r in #2111
  • fix(resolver): allow != operator on dicts, schemas, and lists by @f4z3r in #2112
  • fix(runtime): avoid 2^d exponential walk in schema_check_attr_optional by @Peefy in #2114
  • test(error-format): add C-API and stderr coverage for ExecProgramArgs.error_format by @Peefy in #2115
  • test(evaluator): add regression tests for issue 1918 by @Peefy in #2116
  • fix(evaluator): avoid usize underflow in lazy scope backtrack by @Peefy in #2117
  • chore(evaluator): apply cargo fmt by @Peefy in #2122
  • fix(vet): load sibling .k files when validating inside a kcl.mod package by @Peefy in #2121
  • feat(fmt): honor .editorconfig for indentation (#1930) by @Peefy in #2119
  • fix(lsp): respect .gitignore in workspace file watcher and surface MaxFilesWatch by @Peefy in #2118
  • fix(evaluator): skip eager mixin application duplicated by lazy replay (#1772) by @Peefy in #2120
  • fix(resolver): pre-scan top-level lambdas for forward reference type checking by @Peefy in #2125
  • fix(evaluator, runtime): improve schema error locations (fixes #1898, #1845) by @Peefy in #2128
  • fix(api): call malloc_trim(0) after exec_program on Linux glibc by @Peefy in #2129
  • fix(evaluator): avoid double evaluation of forward-referenced top-level fields (fixes #1759) by @56steve in #2131
  • fix(resolver): type-check config literals in list/dict schema attribute defaults (fixes #1737) by @56steve in #2132
  • fix(parser): surface both pkg paths in duplicate-package diagnostic (#1620) by @Peefy in #2133
  • fix(lsp): react to kcl.mod/kcl.yaml modify events (#1623) by @Peefy in #2134
  • fix(lsp): recover from missing schema_ty + drop wrong .k-file import completions (fixes #1736) by @Peefy in #2135
  • fix(evaluator): let config entries shadow enclosing schema attrs (fixes #1769) by @Peefy in #2136
  • fix(lsp): recover from VFS Create event for files outside opened_files (fixes #1830) by @Peefy in #2137
  • fix(evaluator): don't replay setter statements whose effect already applied (fixes #1833) by @Peefy in #2139
  • feat(runtime): allow datetime.now to format an explicit epoch via optional ticks (fixes #2057) by @56steve in #2141
  • fix(format): leave source unchanged when parsing fails (fixes #1882) by @Peefy in #2140
  • fix(evaluator): handle no-op setter replay for if/elif mixin branches (fixes #1835) by @Peefy in #2142
  • fix(evaluator): don't run nested schema check on default value (fixes #1979) by @Peefy in #2144
  • test(grammar): end-to-end test for empty schema-typed default merge (#1979) by @56steve in #2145
  • feat(api): honour error_format on the C API error envelope by @Peefy in #2146
  • feat(yaml): opt into multi-line block-scalar strings via new multiline_string flag (fixes #1954) by @Peefy in #2149
  • fix(sema): collapse single-file import to its containing directory when they refer to the same files (fixes #1970) by @Peefy in #2150
  • fix(runtime): honour JSON/YAML null as override when merging into schema (fixes #1763) by @Peefy in #2151
  • fix(runtime): preserve YAML block-scalar trailing newline (fixes #1894) by @Peefy in #2152
  • feat(runtime): add file.readbase64 for binary-safe file reads (fixes #1874) by @Peefy in #2155
  • fix(format): keep inline trailing comments on the same line as their node (fixes #1756) by @Peefy in #2154
  • fix(parser): allow exec_program with only k_code_list (fixes lib#217) by @Peefy in #2157
  • test(lsp): give rename_test enough wall time for the LSP to compile the file by @Peefy in #2158
  • fix(query): list_variables handles Binary (union) expressions (fixes #2026) by @Peefy in #2159
  • feat(api): skip un-requested format encoder when --format is set (fixes #2051) by @Peefy in #2160
  • fix(evaluator): honour child schema attribute overrides (fixes #1707) by @Peefy in #2164
  • feat(sourcemap): emit Source Map v3 for generated YAML (#1630) by @Peefy in #2165
  • perf(sema): memoize get_fully_qualified_name and skip rebuild on no-op compiles (fixes #1545) by @Peefy in #2161
  • feat(runtime): emit kcl_info_meta marker for @info(type="attr") (fixes #2047) by @Peefy in #2168
  • refactor(sema): add file-scope infrastructure for issue #1740 by @Peefy in #2171
  • feat(test): add line-level coverage for kcl test (fixes #1481) by @Peefy in #2162
  • perf(sema): incremental FQN map updates + lookup memoization (fixes #1237) by @Peefy in #2167
  • feat(parser): support ES6-style shorthand entries in config/schema literals (fixes #1275) by @Peefy in #2163
  • perf(evaluator): skip pass-3 for unused package bodies (fixes #1758) by @Peefy in #2169
  • feat(lsp): run KCL unit tests by clicking a CodeLens (fixes #1482) by @Peefy in #2170
  • fix(sema): wire file scope into check() and resolve_var for #1740 by @Peefy in #2172
  • feat(lint): support ./... to lint all packages by @Peefy in #2173
  • feat(tools): bundle a program into a single file by @Peefy in #2174
  • perf(lsp): speed up import completion by @Peefy in #2175
  • fix(runtime): apply width, alignment and fill when formatting string values (fixes #2176) by @56steve in #2177
  • refactor(compiler): unify pkgpath to filesystem path conversion by @Peefy in #2178
  • fix(driver): restore get_pkg_list match binding + dedupe LSP path conversion by @Peefy in #2179
  • test(lsp): add Server helpers for e2e tests (part of #1374) by @56steve in #2181
  • test(lsp): move hover tests to insta snapshots (part of #1374) by @56steve in #2180
  • docs: add KubeStellar Console guided install reference (#2073) by @neo0007777 in #2182
  • feat(lsp): auto-update dependencies through the kcl mod toolchain (fixes #1428) by @Peefy in #2183
  • feat(api): declare the GetSchemaTypeMappingUnderPath rpc in spec.proto by @Peefy in #2184
  • feat(lsp): walk up to find kcl.mod/kcl.yaml/kcl.work workspace (fixes #1510) by @Peefy in #2185
  • chore(release): bump workspace version to 0.13.0 by @Peefy in #2186
  • ci: add tag-triggered release workflow for kcl-lib and kcl-language-server by @Peefy in #2187

New Contributors

Full Changelog: v0.11.2...v0.13.0

kagent sandbox

Kagent is an open source programming framework designed for DevOps and platform engineers to run AI agents in Kubernetes

v1.0.0-alpha4

What's Changed

Features

Bug Fixes

  • fix: honor kagent Harness command and args by @marosset in #2934
  • fix(ui): wait for schedule deletion redirect in live test by @Dragonzz27 in #2879
  • fix(hitl): compose the same pause text in every runtime and always declare the extension by @QuentinBisson in #2472
  • fix(harness): hold the Claude prompt until Claude Code can trace the turn by @dhaifley in #2943

Other Changes

  • chore(deps): bump the go-minor-patch group across 1 directory with 39 updates by @dependabot[bot] in #2913
  • Bump kagent-tools to 0.3.0 by @inFocus7 in #2930

Full Changelog: v1.0.0-alpha3...v1.0.0-alpha4

KServe incubating

Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes

v0.21.0

What's Changed

  • refactor(e2e): auto-assign llminferenceservice & consolidate markers by @vivekk16 in #5837
  • fix(e2e): make KEDA operator pod label configurable via env var by @vivekk16 in #5842
  • feat(servingruntime): add resourceClaims support for DRA by @VedantMahabaleshwarkar in #5828
  • feat: restructure CRD management for independent installation by @Jooho in #5843
  • fix(llmisvc): handling stopped members in group by @bartoszmajsak in #5847
  • fix(e2e): per-test namespace isolation for llmisvc e2e by @jlost in #5832
  • fix(ci): apply HF env fixes to s3-init job and improve failure diagnostics by @lizzzcai in #5846
  • fix(e2e): migrate vLLM CPU image from ECR to Docker Hub and cache in CI by @lizzzcai in #5845
  • chore(llmisvc): migrate EndpointPickerConfig apiversion by @zdtsw in #5774
  • feat(llmisvc): replace UDS tokenizer sidecar with vLLM render deployment by @vivekk16 in #5712
  • feat(llmisvc): migrate WVA from VA CRD to annotation-based discovery by @vivekk16 in #5722
  • test(e2e): add canary deployment lifecycle tests by @maskarb in #5833
  • docs: add measured startup benchmarks for OCI model delivery paths by @kliukovkin in #5852
  • test(llmisvc): e2e tests for canary deployment by @bartoszmajsak in #5849
  • fix(e2e): use shared security context for canary lifecycle tests by @bartoszmajsak in #5858
  • fix(e2e): add diagnostics to canary wait_ready by @bartoszmajsak in #5857
  • fix(e2e): enable TLS for canary inference-sim when configured by @bartoszmajsak in #5859
  • fix(e2e): inject default LLMISVC annotations via env var by @bartoszmajsak in #5860
  • fix(e2e): preserve test namespaces on failure by @bartoszmajsak in #5861
  • chore: bump dulwich to >=1.2.5 in kserve-storage by @mholder6 in #5840
  • test(llmisvc): remove IstioShadowService fixture from envtests by @bartoszmajsak in #5870
  • docs: update llmisvc samples from v1alpha1 to v1alpha2 by @aneeshkp in #5804
  • fix(e2e): add gateway convergence wait to canary llmisvc test by @bartoszmajsak in #5876
  • fix(e2e): poll for CanaryPredictorReady condition by @maskarb in #5884
  • fix(samples): remove NixlConnector from prefix cache routing e2e sample by @vivekk16 in #5883
  • fix(e2e): improve test stability with resource cleanup and polling by @mwaykole in #5863
  • feat(llmisvc): migrate deprecated flag to metrics-data-source plugin by @zdtsw in #5877
  • ci(manifests): add kustomize validation workflow by @bartoszmajsak in #5878
  • fix(llmisvc): skip v1alpha2 InfPool reconcile when CRD absent by @bartoszmajsak in #5890
  • fix(e2e): update stale apiVersion in fixture by @jlost in #5872
  • fix(e2e): remove duplicate body field in ResponseRecord by @bartoszmajsak in #5865
  • chore: add license headers to e2e test files by @bartoszmajsak in #5864
  • feat: add pytest --collect-only for e2e tests to precommit by @jlost in #5787
  • ci(e2e): isolate conversion tests and fix result file clobbering by @jlost in #5879
  • feat: integrate with cluster TLS security profile by @ugiordan in #5791
  • fix: prevent agent crash when model spec unmarshal fails by @unichronic in #5700
  • fix(llmisvc): escape quotes in generated --kv-transfer-config by @iasthc in #5880
  • feat(llmisvc): drop old flag to support new TLS flag for disagg-sidecar by @zdtsw in #5875
  • fix(llmisvc): migrate --model-server-metrics-scheme when set on Args by @zdtsw in #5892
  • fix(utils): eliminate data race in API resource discovery cache by @bartoszmajsak in #5891
  • fix(python): map numpy unicode string dtype to BYTES in from_np_dtype by @harshhh817 in #5889
  • fix(llmisvc): strip validation from LLMIServiceConfig CRD by @bartoszmajsak in #5888
  • fix(e2e): stabilize canary traffic tests by @bartoszmajsak in #5896
  • docs(llmisvc): agentic tool calling guide and samples by @tessapham in #5906
  • fix(llmisvc): skip pod spec build when tokenizer disabled by @vivekk16 in #5913
  • fix(ci): reduce max_tokens for LoRA llmisvc tests by @pierDipi in #5736
  • refactor(e2e): failure diagnostics for canary lifecycle tests by @bartoszmajsak in #5919
  • fix(e2e): default GITHUB_SHA to 'latest' for test collection by @bartoszmajsak in #5911
  • fix(router): resolve flaky TestServerTimeout nil pointer panic by @bartoszmajsak in #5887
  • fix: template storage resources in inferenceservice-config ConfigMap by @mgonzalezo in #5697
  • fix(rbac): harden controller manager ClusterRoles by @mholder6 in #5785
  • fix(test): assert worker count via exec not log PIDs by @jlost in #5893
  • feat: add canary traffic splitting to RawDeployment HTTPRoutes by @maskarb in #5912
  • fix(llmisvc): detect scheduler plugins from template args by @vivekk16 in #5931
  • feat(llmisvc): migrate remaining --model-server-metrics-* flags for llm-d by @zdtsw in #5930
  • ci: cache openapi-generator JAR to avoid Maven rate limiting by @bartoszmajsak in #5939
  • fix(llmisvc): close v1alpha1 validation gaps for LoRA and DRA annotations by @bartoszmajsak in #5938
  • fix(llmisvc): detect latency-producer plugin from template args by @vivekk16 in #5937
  • fix(e2e): stabilize flaky HuggingFace LLM assertions by @jlost in #5925
  • chore(llmisvc): aligns additional url discovery by @bartoszmajsak in #5940
  • fix(llmisvc): re-enqueue group peers when status model names change by @jlost in #5943
  • feat(llmisvc): add direct KEDA scaling by @0-Zaid-0 in #5839
  • fix(reconcilers): set owner refs inside RawKubeReconciler.Reconcile by @Jooho in #5934
  • ci: add timeout-minutes to all GHA workflow jobs by @jlost in #5933
  • deps: bump cryptography to >=48.0.1 by @bartoszmajsak in #5957
  • ci: work around missing /etc/cni/net.d on current Ubuntu runners by @bartoszmajsak in #5961
  • release: prepare release v0.20.0 by @cjohannsen-cloudera in #5955
  • docs: fix DataPlane.explain return docstring to match tuple return by @anxkhn in #5760
  • feat(transfor): inject the ca bundle env for httpx by @spolti in #5932
  • feat: add support for Python 3.13 by @noah4477 in #5897
  • chore: update black to 26.3.1 to fix CVE-2026-32274 by @spolti with @Copilot in #5953
  • fix(llmisvc): fully qualify opt-125m-cpu sample image and bump to v0.23.0 by @KillianGolds in #5970
  • feat(llmisvc): add rolloutStrategy for controlling rolling update beh… by @dagrayvid in #5916
  • docs: remove star history chart in README.md by @terrytangyuan in #5979
  • fix: bump go-jose/go-jose/v4 to v4.1.4 (GHSA-78h2-9frx-2jm8) by @spolti with @Copilot in #5985
  • fix(isvc): isvc headless status.address.url has incorrect port by @spolti in #5968
  • fix(agent): migrate AWS SDK from v1 to v2 by @spolti in #5438
  • fix(isvc): gate canary traffic on readiness to prevent transient 503s by @yuzisun in #5984
  • feat(llmisvc): validate rendered config templates by @bartoszmajsak in #5918
  • fix: bump aiohttp to 3.13.3 in python/storage to fix zip bomb DoS by @spolti with @Copilot in #5971
  • fix(pmmlserver): bump pyjnius to 1.7.0 for Python 3.13 support by @kliukovkin in #5978
  • fix: upgrade Pillow to >=12.2.0 (CVE-2026-40192 / GHSA-whj4-6x5x-4v2j) by @spolti with @Copilot in #5959
  • test(llmisvc): add reverse direct KEDA/WVA actuator case to v1alpha2 by @vivekk16 in #5995
  • test(llmisvc): rename test-llmisvc-autoscaling-keda e2e job by @vivekk16 in #5969
  • feat(llmisvc): allow idleReplicaCount=0 for true KEDA scale-to-zero by @vivekk16 in #5996
  • fix(deps): bump golang.org/x/crypto to v0.52.0 in bgtest (CVE-2026-46595) by @spolti with @Copilot in #5993
  • feat(storage): add oci+fetch:// KServe-side OCI image pull (Step 3, #4083) by @kliukovkin in #5739
  • fix: bump Pillow to >=12.3.0 to patch CVE-2026-54058 by @spolti with @Copilot in #5998
  • fix(deps): kin-openapi nil-pointer panic in request validation by @spolti with @Copilot in #5992
  • ci: scope publisher PR path filters and drop redundant amd64 builds by @mateenali66 in #5963
  • ci(llmisvc): e2e test job for weighted InferencePool scenarios by @bartoszmajsak in #5886
  • test(modelcar): add uidModelcar regression test by @andresllh in #5954
  • fix(deps): bump cryptography to >=50.0.0 (CVE-2026-69247) by @spolti with @Copilot in #5999
  • fix(deps): update Pillow to 12.3.0 in alibiexplainer (CVE-2026-59200) by @spolti with @Copilot in #6010
  • feat: add ModelScope storage download support (ms:// URI scheme) by @xrwang8 in #5330
  • fix: reconcile orderings with consistent keys by @bartoszmajsak in #6021
  • fix(deps): bump pyasn1 to 0.6.4 to fix CVE-2026-59885 by @spolti with @Copilot in #6011
  • chore: membership promotion by @Jooho in #6025
  • test(llmisvc): extend e2e coverage for standalone KEDA autoscaling by @vivekk16 in #6018
  • fix(llmisvc): grant configmap write RBAC for CA bundle sync by @cabrinha in #5965
  • fix: support both full JWK and raw key formats in JWE decryption by @maskarb in #6034
  • chore: membership promotion update by @vivekk16 in #6040
  • deps(llmisvc): bump llm-d images to the latest 0.10 release by @zdtsw in #6026
  • feat(ci): detect CLI flag drift when image tags are bumped by @bartoszmajsak in #6035
  • fix: apply ${TAG} to docker image build/push targets by @Jooho in #6057
  • fix(deps): pin urllib3>=2.6.3 to address CVE-2026-21441 by @spolti with @Copilot in #6023
  • fix(qpext): bump opentelemetry-go to v1.43.0 to address CVE-2026-39883 by @spolti with @Copilot in #6019
  • refactor: kustomize image overrides and add deploy-dev-kind-localmodel. by @Jooho in #6060
  • fix: update kserve install script for old version/add a new install script by @Jooho in #6061
  • fix(ci): build arm images natively by @maskarb in #6070
  • fix(llmisvc): detect removals when reconciling generated routing rules by @bartoszmajsak in #6046
  • fix: don't assume LoRA served names are valid object names by @bartoszmajsak in #6077
  • fix(autoscaler): skip KEDA ScaledObject for components without autoScaling by @spolti in #6029
  • feat(llmisvc): localModelCache support for LoRA adapters by @mholder6 in #5690
  • chore(llmisvc): remove deprecated plugin and rename flag by @zdtsw in #6067
  • feat: improve autogluon artifact loading security by @ZabinskiMichal in #5803
  • fix(httproute): restrict canary weights to predictor rules only by @maskarb in #5944
  • fix(deps): bump setuptools to 78.1.1 in python/kserve (CVE-2025-47273) by @spolti with @Copilot in #6055
  • fix(llmisvc): validate LoRA adapters declared in a config by @bartoszmajsak in #6092
  • fix(llmisvc): disambiguate colliding LoRA adapter mount paths by @bartoszmajsak in #6085
  • feat(kernelcache): add kernelcache/mcv initial code donation by @spolti in #5590
  • feat(autogluonserver): inject synthetic item ID for two-column ts by @Wojciech-Rebisz in #6009
  • fix(llmisvc): give P/D engines a NixlConnector so KV actually transfers by @nickaggarwal in #6027
  • test(llmisvc): verify doc samples against the real admission chain by @bartoszmajsak in #6039
  • build(make): lint shell scripts with shellcheck by @bartoszmajsak in #6038
  • fix(quick-install): remove stray generated script by @Jooho in #6101
  • fix(llmisvc): drop vLLM flag removed upstream from gpt-oss samples by @bartoszmajsak in #6037
  • feat(transformer): allow ssl configuration on transformer by @spolti in #6045
  • refactor: split Makefile into fragment files by @Jooho in #6058
  • feat: reload rotated TLS certificates by @maskarb in #6106
  • fix(llmisvc): allow triggerAuthName without authModes for KEDA by @ThatsMrTalbot in #6004
  • refactor(llmisvc): default EPP scheduler config as presets, latest plugins by @zdtsw in #5921
  • fix: clean up orphaned autoscaler and OTel resources on canary promotion by @maskarb in #5899
  • feat(kernelcache): add security verification framework by @Jooho in #6099
  • feat(kernelcache): cert-mode signature verification by @Jooho in #6103
  • chore(deps): upgrade Envoy AI Gateway to v1.1.0 by @cjohannsen-cloudera in #6122
  • fix(llmisvc): template base llmisvcconfig files in the correct namespace by @NoOverflow in #6088
  • feat(llmisvc): automatically label workloads with InferencePool ref by @pierDipi in #5624
  • fix(rbac): add leader election role for localmodel controller by @mwaykole in #6120
  • fix(deps): allow pyarrow to use newer versions > 20 by @spolti in #6125
  • feat(kernelcache): add signing contract by @Jooho in #6118
  • refactor(kernelcache): extract security data contracts into types package by @Jooho in #6130
  • fix(llmisvc): render presets against the merged configuration by @bartoszmajsak in #6126
  • fix(ci): use kustomize binary from <project_root>/bin/ by @spolti in #6128
  • feat(deps): upgrade golang to 1.26 by @spolti in #6127
  • fix(llmisvc): grant events create/patch to the scheduler Role by @UgaTheDev in #6142
  • chore(build): stamp Go toolchain into tool cache key by @bartoszmajsak in #6146
  • fix(llmisvc): match group backends by their referenced pool by @bartoszmajsak in #6149
  • fix(llmisvc): release terminating members from owned routes by @bartoszmajsak in #6156
  • fix(autogluon): prevent path traversal in model loading by @Wojciech-Rebisz in #5802
  • feat: adding Power support for storage-initializer. by @Sunidhi-Gaonkar1 in #5235
  • fix(ci): rerun cancelled workflow runs by @jlost in #6135
  • feat(api): add InferenceService tracing configuration by @maskarb in #6108
  • fix(llmisvc): recognizes header case-insensitively by @bartoszmajsak in #6161
  • test: helpers for ConfigMap and HTTPRoute by @bartoszmajsak in #6163
  • test(llmisvc): wait for workload readiness by @bartoszmajsak in #6157
  • fix(llmisvc): report HTTPRoutesReady after a route is observed by @bartoszmajsak in #6164
  • chore(deps): upgrade KEDA to v2.20.2 by @vivekk16 in #6153
  • release: prepare release v0.21.0-rc0 by @cjohannsen-cloudera in #6160
  • release: prepare release v0.21.0-rc1 by @cjohannsen-cloudera in #6231
  • fix(release): generate missing install manifests for v0.21.0-rc1 by @cjohannsen-cloudera in #6273
  • release: prepare release v0.21.0 by @cjohannsen-cloudera in #6288

New Contributors

Full Changelog: v0.20.0...v0.21.0

Kuadrant sandbox

Kuadrant combines Gateway API and Istio-based gateway controllers to enhance application connectivity. It enables platform engineers and application developers to easily connect, secure, and protect their services and infrastructure across multiple clusters with policies for TLS, DNS, application authentication & authorization, and rate limiting.

v1.6.0-rc1

Release Candidate 1

This is a release candidate for v1.6.0. Pending QE validation before GA.

Component Versions

Component Version
Authorino Operator v0.27.0 (ships Authorino v0.28.0)
Limitador Operator v0.19.0 (ships Limitador v2.5.0)
DNS Operator v0.18.0
MCP Gateway v1.0.1
WASM Shim v0.15.0
Console Plugin v0.7.0
Developer Portal Controller v0.3.0

What's New in 1.6.0

Extensions SDK

Out-of-process (OOP) extensions are now fully supported: handshake and session management, authentication via TokenReview/SAR, protocol version validation on handshake, and automatic state rebuild on reconnect. The operator prunes extension resources on disconnect.

MCP Gateway Integration

MCP Gateway (controller + broker) and MCP Inspector plugin backend are now bundled with and fully managed by the kuadrant-operator lifecycle.

TokenRateLimitPolicy: AI Workload Support

New dataExtraction field for extracting response.totalTokens from LLM responses, enabling token-based rate limiting for AI workloads with built-in defaults and wasm config generation.

Default NetworkPolicy (Zero-Trust)

Kuadrant now ships deny-all NetworkPolicies for Authorino, Limitador, and the operator namespace out of the box, with targeted allow rules for metrics scraping, Console Plugin access, and WASM plugin traffic.

DNS Operator Lifecycle Management

DNS Operator is now fully managed by kuadrant-operator's KuadrantControlPlane lifecycle (alongside Authorino Operator and Limitador Operator), replacing the previous separate deployment model.

Egress Observability

Metrics and distributed tracing for egress gateways, documented end-to-end with span chain details (kuadrant-filter, Authorino, Limitador), access log field mappings, and custom metric label configuration.

OpenShift TLS Security Profile Propagation

Reads the cluster-wide OpenShift APIServer CR and propagates tlsMinVersion and tlsCipherSuites to the Authorino CR, so Authorino'''s TLS configuration matches the cluster-wide security profile. Falls back to Intermediate profile defaults on non-OpenShift clusters.

What's Changed

Bug Fixes

Security & Dependencies

Toolchain

Other Changes

New Contributors

Full Changelog: v1.5.3...v1.6.0-rc1

Cozystack sandbox

Cozystack is a free PaaS platform and framework for building private clouds and providing users/customers with managed Kubernetes, KubeVirt-based VMs, databases as a service, NATS, message brokers, etc. with GPU support in VMs and Kubernetes clusters.

v1.6.4-rc.1

[Backport release-1.6] feat(kubevirt): expose migration configuration…

… through platform values (#4402)

This reopens a fork backport from a branch in this repository so its CI
can run. The commit is the one from #4339 by @yankawai, unchanged. On
release-1.6 the pull-request CI pushes the images it builds, and a run
from a fork has no registry credentials, so #4339 stopped at the build
jobs and never reached e2e. The description below is the author's.

## What this PR does

Manual backport of #4254 to `release-1.6`.

The automated backport, #4322, stopped with conflict markers in four
files as its only commit, so it is red on DCO and unmergeable as it
stands.

Every conflict has the same cause: the `kubevirt.disabledFeatureGates`
work landed on main after 1.6 was cut, and the cherry-pick carried it
into hunks this line does not have. Four differences follow:

- platform values gains a `kubevirt` block holding `migrations` alone,
and the iaas bundle gains only the migrations hop;
- the wiring test takes the four migrations cases and the case for a
null `kubevirt` block, without the assertions that belong to
disabledFeatureGates cases this line does not have;
- `update_idempotency_test.sh` gains `run_update_logged`, a four-line
helper that arrived on main with that work and the new cases need;
- the kubevirt Makefile header counts six sed patches rather than five.

Verified on this branch: 9 kubevirt chart tests, 31 platform wiring
tests (102 across the platform suite), and `update_idempotency_test`
PASS under GNU make and sed. Removing the new awk guard from the
Makefile turns that test red naming the migrations block, so the ported
guard is doing work. On this line `(.Values.kubevirt).migrations` is the
only reader of `.Values.kubevirt`; rewriting it as
`.Values.kubevirt.migrations` turns the null-block case red with a nil
pointer.

Feature summary, unchanged from #4254: cluster-wide migration settings
could only be applied by patching the KubeVirt CR, because the kubevirt
chart did not render `migrations` and the platform bundle did not
forward it. `kubevirt.migrations` now passes through the generated
Package to `spec.configuration.migrations`. A hand patch of that Package
is undone on the next platform render, so this hop is the only supported
setter.

```yaml
kubevirt:
migrations:
bandwidthPerMigration: 625M
parallelMigrationsPerCluster: 2
parallelOutboundMigrationsPerNode: 1
```

### Screenshots

Not a UI change.

### Downstream repositories

Walked the trigger map against the diff: platform values, the iaas
bundle and the kubevirt chart. The v1.6 reference page of the platform
values had no `kubevirt` section, so the row goes there as a follow-up
that should land with this backport.

- [ ] No downstream repository is affected by this change
- [x] [cozystack/website](https://github.com/cozystack/website) -
follow-up: cozystack/website#705

### Release note

```release-note
feat(kubevirt): expose `kubevirt.migrations` in the platform values and forward it to `spec.configuration.migrations` on the KubeVirt CR, so live-migration bandwidth and parallelism can be set without patching a Package that the platform re-renders.
```

Backstage incubating

Backstage is an open platform for building developer portals

v1.55.2

This patch release contains the following commits:

  • fix: normalize TechDocs Markdown extension configuration
Backstage incubating

Backstage is an open platform for building developer portals

v1.54.9

This patch release contains the following commits:

  • fix: normalize TechDocs Markdown extension configuration
Backstage incubating

Backstage is an open platform for building developer portals

v1.50.7

This patch release contains the following commits:

  • fix: normalize TechDocs Markdown extension configuration
Cozystack sandbox

Cozystack is a free PaaS platform and framework for building private clouds and providing users/customers with managed Kubernetes, KubeVirt-based VMs, databases as a service, NATS, message brokers, etc. with GPU support in VMs and Kubernetes clusters.

v1.5.4

v1.5.4 (2026-08-19)

v1.5.4 is the final release of the 1.5 line. It is a stability release: it backports fixes for a webhook certificate-renewal outage that could block all pod creation, a KubeVirt VMI validation failure, several release-blocking crashloops (velero, cert-manager, SeaweedFS, flux-shard-operator), a silent PostgreSQL restore data-integrity bug, and a raft of smaller reliability and CI fixes accumulated on the branch. It also closes the SeaweedFS 4.31 rename fallout on the 1.5.x line and pins the CAPI kubeadm bootstrap objects that a later upgrade to 1.6 would otherwise prune — both of which need an operator to act, and both of which are covered in the section below.

These notes are measured against v1.5.2, not v1.5.3. v1.5.3 was tagged but its GitHub Release was left a draft and never published, so no user ever received it and every operator upgrading arrives from v1.5.2. Comparing against v1.5.3 would silently drop four commits — two user-facing fixes — that nobody has seen in a release. The range is v1.5.2..v1.5.4, 82 commits across 27 pull requests.

⚠️ Breaking Changes and Required Actions

There are no breaking API or values changes in v1.5.4. There are two things that need an operator, and both of them can cost data or wedge an upgrade if they are skipped. Read this section in full before applying the v1.5.4 Platform Package.

Pre-upgrade checks

Run these against the management cluster before upgrading.

1. SeaweedFS 4.31 rename — classify every instance, and re-run the audit even if you have run it before.

Cozystack v1.5.0 bumped the vendored SeaweedFS chart from 4.0.405 to 4.31.0. Before 4.31 the chart named its workloads after the chart (seaweedfs-master, seaweedfs-filer, seaweedfs-volume), ignoring the Helm release name. 4.31 names them after the release, and the data-plane release is <name>-system, so every StatefulSet wanted to become seaweedfs-system-*. StatefulSet names are immutable, so Helm could not rename in place — it stood up a second, duplicate set beside the running one.

What that duplicate does depends on the cluster. With as many nodes as master replicas the new masters cannot schedule (hard pod anti-affinity against the old ones), so the duplicate sits Pending/CrashLoopBackOff and the original keeps serving. With more nodes than masters the new, empty set comes up — and because both sets carry identical pod labels, the seaweedfs-s3 Service load-balances across them while both filers write to the same seaweedfs-db Postgres metadata store pointing at different volume servers. That is a data-integrity incident, not a cosmetic duplicate: reads of existing objects through the new endpoint miss, new writes land on empty volumes, and two master sets hand out volume IDs from independent sequences into one shared metadata table.

Separately, and landing on the same upgrade, the v1.5.0 database split moved the CNPG Cluster/seaweedfs-db — the filer metadata store, i.e. the index for every object in the tenant's S3 — out of the <name>-system release into its own <name>-db release. Migration 43 shipped comparing the owning release against the literal string seaweedfs-system, so it only ever fired for an instance named seaweedfs; an instance named anything else was skipped and had its Cluster pruned as a removed resource, with CNPG taking the PVC along with it. That prune is not a one-shot: Helm computes deletions by diffing the last deployed revision against the new manifest, so a tenant whose <name>-system last succeeded on a pre-split revision recomputes the same deletion on every upgrade attempt, including attempts that fail for unrelated reasons.

v1.5.4 pins fullnameOverride: seaweedfs in system/seaweedfs, so workloads are named after the chart exactly as they were before 4.31 and upgrading adopts the running set and its volumes in place. Migration 43 is fixed to match the -system suffix, and migration 45 re-runs the hand-over for clusters that already ran the hardcoded version. Two states cannot be adopted that way, and the chart fails the render rather than guess — the enforcing guard is packages/system/seaweedfs/templates/naming-guard.yaml, with a sibling copy in extra/seaweedfs so the refusal is visible on the SeaweedFS application itself.

Step 0 — seaweedfs-db ownership (read-only, do this first). This one destroys data rather than duplicating it, so clear it before anything else.

kubectl get cluster.postgresql.cnpg.io -A \
  -o custom-columns='NS:.metadata.namespace,NAME:.metadata.name,OWNER:.metadata.annotations.meta\.helm\.sh/release-name,KEEP:.metadata.annotations.helm\.sh/resource-policy'

Read the rows where NAME is seaweedfs-db:

OWNER KEEP Meaning
<name>-db keep Handed over. Nothing to do.
<name>-db (none) Installed fresh on ≥ v1.5.0. Safe — <name>-system never rendered the Cluster, so it is not in that release's prune baseline.
<name>-system (none) At risk. Migration 45 hands it over on the next platform upgrade. Do not reconcile <name>-system before the migration runs.
(no row at all) Already lost. The metadata index is gone: that tenant's S3 returns 500 and its objects are unreachable even though the volume PVCs still hold the bytes. No migration can rebuild it — restore the seaweedfs-db Postgres from a backup, or treat that tenant's object storage as lost. Note that <name>-db may still report Ready while this is true; trust the kubectl get cluster output, not the HelmRelease status.

Step 1 — classify every SeaweedFS instance (read-only, mutates nothing).

hack/seaweedfs-naming-audit.sh                 # whole cluster
hack/seaweedfs-naming-audit.sh tenant-foo      # or named namespaces
CLASS State Action
L Only the chart-named generation is present. None. The upgrade adopts it in place.
S Only the release-named generation is present — the instance was installed fresh on 1.5.x, and its data lives on data1-seaweedfs-system-volume-* PVCs. Re-bind those volumes onto the chart-named PVC names before upgrading. Pinning the chart name without that renames the workloads away from the data, and Helm cannot move data between PVCs. Runbook Step 2.
MIXED Both generations are present. One is an empty duplicate and one holds the data, and nothing durable in the object graph says which — so the chart refuses. Classify the tenant and delete the empty generation so exactly one remains; the render then adopts the survivor with no further action. Runbook Step 1, then 2a or 3.

Read the exit code, not just the table. The audit fails closed: a kubectl call that fails, or a Helm release payload it cannot decode, prints FATAL and exits non-zero, and the partial table must not be trusted. Only exit 0 means the table is the whole answer — and an empty table with exit 0 is a genuinely clean fleet.

If you have already run this audit, run it again on v1.5.4. The version of the script that shipped in v1.6.0 silenced every kubectl failure with 2>/dev/null, so a timeout or an RBAC denial produced an empty, "all clean" table byte-identical to an honestly clean fleet — a false clean, on the script whose output gates a runbook step that deletes PVCs. v1.5.4 is the first release on the 1.5 line to carry the audit at all, and it carries the fail-closed version (v1.6.1 and later carry it too). A clean result from a v1.6.0 checkout, or from main between 2026-07-20 and 2026-07-28, is not evidence of anything.

One unrelated filer change lands on the same upgrade and is worth knowing about while you are looking at this: the filer's postgres2 connection pool to that same seaweedfs-db metadata store was unconfigured, so every metadata lookup opened a fresh PostgreSQL connection and added seconds of latency to every S3 request. That is fixed here too (see the postgres2 connection pool entry below), and it needs no operator action — but if you have been treating slow S3 as a symptom of the rename, it may well have been this instead.

Recovery for S and MIXED is docs/operations/seaweedfs-431-rename-recovery.md. Do not guess which generation holds the data — the runbook exists because a duplicate that briefly served writes and later crashed is indistinguishable, on every durable signal, from one that never scheduled. A tenant that went 1.4.x straight to 1.6 never renamed and is unaffected by any of this; duplicates exist only on tenants that passed through 1.5.x.

2. The platform migration targetVersion moves from 45 to 46 — which changes what a later upgrade to 1.6 runs.

v1.5.4 is the first 1.5.x release stamped targetVersion: 46; v1.5.0 through v1.5.3 were all stamped 45. run-migrations.sh loops seq CURRENT (TARGET - 1), so a cluster that reaches 46 and later upgrades to v1.6 (targetVersion: 54) runs slots 46 through 53 — and never executes 1.6's own slot 45. This is a skip, not an ordering problem.

Slot 45 does not hold the same thing on both branches. On release-1.5 it is the SeaweedFS seaweedfs-db hand-over repair described above. On main and release-1.6 it is the pin that stamps helm.sh/resource-policy: keep onto the CAPI KubeadmConfigTemplate objects. 1.6 drops KubeadmConfigTemplate from the tenant kubernetes chart entirely — workers move to TalosConfigTemplate — so on that upgrade Helm sees the object in the previous release manifest, absent from the new one, and deletes it while the kubeadm-backed MachineSet is still mid-rollover with its bootstrap.configRef pointing at it. controller-manager then floods with reconcile errors, and where the Talos image fetch is slow or a MachineHealthCheck remediates, workers can hang pending with nothing to bootstrap from. Tenant Kubernetes only, and a noisy broken rollover rather than data loss — but it is not self-healing.

v1.5.4 closes this in two places, covering two disjoint populations, and dropping either would leave a real hole:

  • Existing clusters (stamped 45 or lower) pick the pin up from the migration: release-1.5's slot 45 now runs the keep-pin after the SeaweedFS repair, so the annotation is already in place by the time a later 1.6 upgrade skips 1.6's slot 45. The SeaweedFS half runs first deliberately — a missed hand-over loses a tenant's filer metadata, a missed pin is recoverable by hand — and both halves fail closed ahead of the version stamp, so a half-completed attempt is safe and the Job retries the whole slot. The pin selects on the app.kubernetes.io/managed-by=Helm label that Helm injects, so the KubeadmConfig children CAPI spawns from the template (owned by their Machine, never pruned by Helm) are correctly left alone. It is idempotent: an object already carrying keep is skipped without a write.

  • Fresh v1.5.4 installs cannot be reached by any migration, so the chart covers them instead. The platform chart renders the cozystack-version ConfigMap directly at targetVersion when it does not already exist, and the migration hook emits its Job only when that ConfigMap is already present — so a fresh install is stamped 46 having never run slot 45, or any other slot. The kubernetes chart therefore stamps helm.sh/resource-policy: keep on the KubeadmConfigTemplate at render time, so the object is born pinned no matter which path a cluster took to get there.

Verify after upgrading:

Every row should show keep.

kubectl get kubeadmconfigtemplates.bootstrap.cluster.x-k8s.io -A
-o custom-columns='NS:.metadata.namespace,NAME:.metadata.name,KEEP:.metadata.annotations.helm.sh/resource-policy'">

# Should print 46.
kubectl get configmap cozystack-version -n cozy-system -o jsonpath='{.data.version}{"\n"}'

# Every row should show keep.
kubectl get kubeadmconfigtemplates.bootstrap.cluster.x-k8s.io -A
-o custom-columns='NS:.metadata.namespace,NAME:.metadata.name,KEEP:.metadata.annotations.helm.sh/resource-policy'

A KubeadmConfigTemplate row showing <none> means the pin did not land on that object, and a later 1.6 upgrade will prune it out from under a live MachineSet. Recovering from that means recreating the template by hand with the right Helm ownership metadata, so fix it before upgrading to 1.6.

Manual actions required

  • BEFORE upgrading — recover SeaweedFS tenants classified S or MIXED. The chart refuses to render for them, which blocks the tenant's upgrade until an operator resolves it. Run hack/seaweedfs-naming-audit.sh and follow docs/operations/seaweedfs-431-rename-recovery.md.

  • BEFORE trusting an earlier audit result — re-run the audit from v1.5.4. The v1.6.0 copy of the script could report a false clean (see pre-upgrade check 1).

  • AFTER upgrading — confirm the migration stamp is 46 and every KubeadmConfigTemplate carries helm.sh/resource-policy: keep (see pre-upgrade check 2). This is what makes the later 1.6 upgrade safe.

  • No action, but expect one rollout each: the cert-manager webhook (it moves off port 10250) and cainjector (raised memory limit), the kube-ovn webhook, velero (new startupProbe), and the CNPG instances of every Postgres release with backup.enabled or bootstrap.enabled, whose rendered Cluster gains one spec.env entry for the S3 checksum fix.

Fixes

  • fix(kube-ovn): reload kubeovn-webhook serving certificate on cert-manager renewal: kube-ovn-webhook only loaded its TLS certificate once at startup. Once cert-manager rotated the backing Secret and the webhook pod outlived the old certificate's expiry (roughly a year after install), the pod kept serving the expired certificate, kube-apiserver rejected the TLS handshake, and — because the MutatingWebhookConfiguration uses failurePolicy: Fail — every pod creation in tenant namespaces was blocked, including virt-launcher, so no VM could start. The webhook now reloads its certificate from disk on renewal through a GetCertificate closure, with renewBefore widened to 720h. The reload path is hardened against the cases that would have made it unreliable in practice: the certificate file's mtime is observed before the key pair is read, so a Secret swap racing the read is retried on the next handshake rather than cached under a newer mtime and missed; change is detected by mtime inequality rather than strictly-newer, catching equal-mtime replacements and backward clock steps; a load attempt is recorded whatever the outcome, so a persistently unreadable file is retried once its mtime advances instead of on every handshake; and stat failures are logged rather than swallowed, so a broken mount is visible before expiry. The same change also bounds the webhook's http.Server with ReadHeaderTimeout, ReadTimeout, WriteTimeout and IdleTimeout — previously it set none, so a client that opened a connection and sent headers slowly could hold a goroutine and a file descriptor indefinitely (@IvanHunters in #3557, backport #3879).

  • fix(flux-shard-operator): repair sharded helm-controller crashloop behind an HTTP proxy: The cloned per-shard helm-controller inherited HTTP_PROXY/HTTPS_PROXY/NO_PROXY from the flux-aio container and lacked a startupProbe, so on clusters behind a corporate HTTP proxy the shard could not reach the (proxy-unreachable) in-cluster API server and crashlooped before ever becoming ready. Proxy environment variables are now stripped from the cloned container and a startupProbe is added so slow starts are tolerated instead of killed. The cloned probe inherits its timeoutSeconds from the liveness handler rather than being forced to 1s, so a future flux-aio shipping a larger liveness timeout cannot end up with a stricter startup probe and recreate the very crashloop this fixes (@IvanHunters in #3546, backport #3883).

  • fix(velero): add startupProbe so slow startup does not crashloop the install gate: Velero binds its health endpoint only after loading plugins and connecting to the API server, which can take minutes under a heavy parallel platform install — long enough to trip the liveness probe and crashloop the pod before the install-readiness gate ever passes. A startupProbe now defers liveness checks until the server has actually finished starting (@lexfrei in #3138, backport #3533).

  • fix(cert-manager): move the webhook off the kubelet's port: The cert-manager webhook listened on port 10250, the kubelet's own port. When the webhook Service resolved to a node IP instead of the pod IP, the connection landed on the kubelet, which answered with its own certificate, and the API server rejected every cert-manager admission call cluster-wide. The webhook now listens on a port the kubelet does not own (@lexfrei in #3359, backport #3366).

  • fix(cert-manager): raise cainjector memory limit to unblock caBundle injection: cainjector loads every webhook configuration, APIService, and CRD into its informer caches at startup; on a full Cozystack install that working set exceeded the 128Mi memory limit, so the single leader-elected cainjector pod was OOMKilled during cache population and crashlooped before it could inject any CA bundle, breaking every webhook that depends on cert-manager-issued CAs. The memory limit is raised so cache population can complete (@IvanHunters in #3199, backport #3202).

  • fix(mariadb): widen startup probe budget so bootstrap cannot be killed: The MariaDB operator built its startup probe without a failureThreshold, so it fell back to the Kubernetes default of 3 — giving a fresh bootstrap only ~40 seconds before being killed by the liveness probe the startup probe exists to defer. The budget is widened so a legitimately slow first bootstrap is no longer mistaken for a hung container (@lexfrei in #3344, backport #3364).

  • fix(seaweedfs): make naming audit fail closed on kubectl and payload errors: hack/seaweedfs-naming-audit.sh, used as the pre-upgrade gate before PVC deletion, silenced every kubectl failure with 2>/dev/null, so a transient API error produced an empty, "all clean" table indistinguishable from an honestly clean fleet. Every enumeration now routes through one fail-closed helper that names the failed query on stderr and exits non-zero, and a Helm release Secret whose payload cannot be decoded is fatal rather than a silent skip. The same change fixes a false clean reachable with no corruption at all: the chart-name extraction was a greedy (last-match) sed, and since Helm marshals the release's config after chart, a values subtree spelling chart.metadata.name shadowed the real chart name — the release then read as non-SeaweedFS and the tenant vanished from the report. It now takes the first match, folds newlines so a pretty-printed payload parses at all, and says which path and shape it expected when it cannot read one (@myasnikovdaniil in #3436, backport #3878).

  • fix(seaweedfs): close the 4.31 rename fallout on the 1.5.x→1.6 upgrade path: Cozystack v1.5.0 bumped the vendored SeaweedFS chart to 4.31.0, which renames workloads after the Helm release name; because StatefulSet names are immutable, upgrading through the 1.5.x line stood up a second, empty StatefulSet beside the running one instead of adopting it. This closes the remaining gaps in the adoption path: system/seaweedfs pins fullnameOverride: seaweedfs so the upgrade adopts the running workloads and volumes in place, the naming guard moves into the chart a platform upgrade actually re-renders and refuses when both naming generations exist, cluster-scoped COSI RBAC is named per namespace again (the 4.31 release-based names collided across tenants), and the seaweedfs-db hand-over now runs for every instance name — previously an instance not named seaweedfs had its filer metadata database pruned on upgrade. Migration 45 repairs clusters that already ran the old hand-over, and hack/seaweedfs-naming-audit.sh plus docs/operations/seaweedfs-431-rename-recovery.md guide classification and recovery. The refusal is expected for tenants that passed through 1.5.x — see the required-actions section above before upgrading (@myasnikovdaniil in #3339, backport #3370, building on the adopt-in-place pin from #3282).

  • fix(seaweedfs): configure postgres2 connection pool for the filer: The vendored SeaweedFS chart ships connection-pool settings only for the mysql store, which Cozystack does not use. The postgres2 store, which is enabled, had none — so the Go SQL pool kept zero idle connections and every filer metadata lookup opened a fresh PostgreSQL connection (TCP + TLS + SCRAM, roughly 300ms each). At around five path lookups per S3 operation that added about two seconds of latency to every S3 request, regardless of cluster load. Measured on a 12-node cluster with two filers: PUT of a 4KiB object went from 10s to 0.28s (p50) and HEAD from 2.6s to 0.16s. max_open is capped at 40 so two filer replicas stay under CNPG's default max_connections of 100. This fix was tagged in v1.5.3, which was never published, so v1.5.4 is the first release to carry it (@mattia-eleuteri in #2906, backport #3194).

  • fix(cozystack-basics): gate the hostname VAP policies on the VAP API: The route/gateway/ingress hostname ValidatingAdmissionPolicy templates rendered unconditionally. On a cluster missing the VAP API, an operator-generated HelmRelease with drift detection off would render the policies out on first install and never add them back. The templates are now gated on .Capabilities.APIVersions.Has, so installs on VAP-less clusters no longer silently drop hostname validation (@lexfrei in #3409, backport #3876).

  • fix(keycloak-configure): patch HelmRelease in release namespace on teardown: The keycloak-configure pre-delete teardown Job cleared the Flux HelmRelease finalizers with a kubectl patch aimed at a hardcoded namespace that didn't match the release's actual cozy-keycloak namespace, so the teardown Job's own RBAC could never reach the object it needed to patch and uninstalling the release could get stuck. The patch now targets the correct release namespace (@lexfrei in #3372, backport #3877).

  • fix(platform): forward backupStorage overrides to the backupstrategy-controller Package: The documented admin override for cozy-default backup S3 coordinates (spec.components.backupstrategy-controller.values on the platform Package) had no effect, because the cozystack-platform PackageSource exposes only a single platform component and silently ignored the override block, so admin-configured backup storage settings never reached the running backup controller. The platform chart now forwards the override through correctly (@androndo in #3333, backport #3402).

  • fix(backups): request S3 checksum only when required for barman-cloud (non-AWS S3 / Ceph RGW): Since botocore ~1.36 (early 2025) the default RequestChecksumCalculation is when_supported, so barman-cloud's boto3 attaches a flexible checksum to every PutObject. AWS S3 accepts it, but several S3-compatible backends — Ceph RADOS Gateway, the platform's own SeaweedFS system bucket, some MinIO and Cloudflare R2 builds — reject it with InvalidArgument: x-amz-content-sha256 must be UNSIGNED-PAYLOAD, failing every backup and WAL-archive upload. AWS_REQUEST_CHECKSUM_CALCULATION=when_required is now set through the CNPG Cluster's spec.env, which reaches the instance pods and therefore the barman-cloud subprocess the instance manager execs. On this branch that one setting covers all three S3 paths at once: the chart-rendered legacy spec.backup.barmanObjectStore, the same field SSA-patched by the CNPG backup driver in the useSystemBucket flow, and externalClusters recovery. Note this is a release-1.5-native equivalent of the fix on main, not a cherry-pick — main sets spec.instanceSidecarConfiguration.env on barmancloud.cnpg.io ObjectStore objects, and release-1.5 has no barman-cloud plugin and no ObjectStore CRD, so those objects would template cleanly and then apply to nothing. It is gated on backup.enabled or bootstrap.enabled, so a Postgres release with no S3 configured is untouched (@androndo in #3417, backport #3882).

  • fix(objectstorage-controller): converge BucketClaim readiness and speed up COSI provisioner failover: The COSI control plane (built from upstream container-object-storage-interface v0.2.2) never watched Bucket objects and dropped no-op resync deltas, so a BucketClaim whose backend Bucket became ready after the first reconcile stayed bucketReady=false forever, and the consuming application never got its BucketAccess. The controller now requeues until the backend Bucket is ready, emits a Warning event while it waits, and shortens the COSI leader-election lease so failover is faster (@lexfrei in #3034, backport #3532).

  • fix(tenant): inherit full ancestor label chain from parent namespace: Tenant namespace labelling only derived one ancestor level from .Release.Namespace, so a direct child of tenant-root never received the labels needed for its grandparent's <ancestor>-egress CiliumClusterwideNetworkPolicy to allow traffic, silently breaking cross-tenant egress for nested tenants more than one level deep. Namespaces now inherit the full ancestor label chain from their parent (@IvanHunters in #2912, backport #3191).

  • fix(backups/postgres): purge stale recovery cluster on repeat in-place restore: A repeat in-place PostgreSQL RestoreJob into a target that had already been restored once silently reported Succeeded while doing nothing — the PVC was untouched, no recovery pods appeared, and the database kept its pre-restore contents, a silent data-integrity failure. The restore path now purges the stale recovery cluster before restoring again, so a repeat restore actually recovers data instead of quietly no-opping. The freshness discriminator that decides whether a leftover Cluster belongs to this restore or a previous one is deliberately strict (creationTimestamp > StartedAt), so an exact timestamp tie resolves to "not fresh" and the caller purges — the conservative direction, since a genuinely fresh Cluster is always created well after StartedAt (@IvanHunters in #3318, backport #3321).

  • fix(fluxcd): omit empty distribution.artifact in FluxInstance for guest clusters: Guest Kubernetes clusters with the fluxcd addon enabled never became Ready, because the flux-instance template rendered spec.distribution.artifact unconditionally (unlike its sibling optional fields), and the empty string broke the intended air-gapped default of using the operator's embedded manifests. The field is now guarded like its siblings, so guest clusters with fluxcd enabled converge to Ready again (@IvanHunters in #3284, backport #3292).

  • fix(dashboard): unbreak CORS on expired session for k8s API calls: The dashboard SPA broke with CORS errors once the kc-access cookie expired: the Keycloak client's webOrigins was never set (so it stuck on a stale hostname after any rootUrl change), and oauth2-proxy mishandled the expired-session case for Kubernetes API calls. Both are fixed so an expired session no longer surfaces a broken, CORS-blocked SPA (@IvanHunters in #2788, backport #3291).

  • fix(apps/vpn): remove invalid foo field from urls Secret: A leftover debug foo field in the VPN app's <release>-urls Secret template is not part of the core/v1 Secret schema. Client-side apply tolerated it, but server-side apply's stricter validation rejected the object outright, so any VPN application upgrade failed and the release was left in a broken state. The stray field is removed (@IvanHunters in #3281, backport #3290).

  • fix(kubevirt): update KubeVirt to v1.8.4: Updates the vendored kubevirt-operator from v1.8.2 to v1.8.4, backporting the VMI checksum status-field validation fix (uint32 range) that Kubernetes 1.36 requires — without it, strict CRD numeric-format validation could leave VMIs stuck in Scheduled and unable to start. The update also carries assorted upstream bug fixes and CVE remediations; the KubeVirt custom resource itself is unchanged (@lexfrei in #2940, backport #3285).

  • fix(kubevirt-instancetypes): restore persistent EFI/TPM state: v1.5.1 stripped persistent: true from both preferredEfi and preferredTPM on the six shipped Windows preferences — windows.11, windows.2k22, windows.2k25 and their .virtio variants — to unblock live migration. Secure Boot and the vTPM stayed present, so the VMs still booted, but their state was no longer persisted: EFI NVRAM and vTPM contents were discarded on every VM restart, so anything the guest sealed to the vTPM (BitLocker being the common case) or wrote to NVRAM — Secure Boot key enrollment, boot entries — did not survive a power cycle. This affects anyone running Windows 11, Server 2022 or Server 2025 guests created from the shipped preferences on v1.5.1 or v1.5.2; v1.5.0 had persistence, and it is restored here. The strip was motivated by a real-looking concern that turned out to rest on an outdated premise — that the persistent-state-for-<vm> backend-storage PVC, ReadWriteOnce on the default replicated StorageClass, pins the VM to its node and blocks live migration and node drains, stalling cluster upgrades under evictionStrategy: LiveMigrate. KubeVirt has in fact migrated RWO-Filesystem backend storage since v1.4 (kubevirt/kubevirt#12629): it creates a fresh target state PVC and copies the small state blob during migration. Verified on the default replicated storage that a persistent-firmware VM reports LiveMigratable=True and gets its copy-on-target PVC, so persistence and live migration now hold at the same time. Existing VMs pick the change up on their next restart; already-running VMs are unaffected until then. This fix was tagged in v1.5.3, which was never published, so v1.5.4 is the first release to carry it (@kvaps in #3154, backport #3212, reverting #3006).

  • fix(migrations): pin KubeadmConfigTemplate on the 1.5 slot 45 migration and in the chart: Migration slot 45 holds different things on different branches — release-1.5 uses it for the SeaweedFS seaweedfs-db hand-over repair, while main and release-1.6 use it for the helm.sh/resource-policy: keep pin on the CAPI KubeadmConfigTemplate objects. Cutting v1.5.4 is what first ships targetVersion: 46, and run-migrations.sh loops seq CURRENT (TARGET - 1), so a cluster that reaches 46 later runs seq 46 53 on the way to 1.6 and never executes 1.6's slot 45 at all — the pin is skipped, and Helm then prunes the template while the kubeadm-backed MachineSet is still mid-rollover. 1.5's own slot 45 now applies the pin (after the SeaweedFS repair, and fail-closed) so an existing cluster picks it up on the way to 46, and the kubernetes chart stamps the annotation at render time so the object is born pinned — which is the only thing that covers a fresh v1.5.4 install, since such a cluster has its cozystack-version ConfigMap rendered straight at 46 and the migration hook never emits a Job at all. See the required-actions section above for the verification commands (@myasnikovdaniil in #3892).

  • fix(postgres-operator): align CNPG CRDs with 1.27.3: The vendored CloudNativePG CRDs were still pinned to 1.27.1 metadata while the operator image was already on 1.27.3, which prunes the new Backup.status.instanceID.sessionID field the newer instance-manager writes — causing false instance-manager restart errors and failed backups. The CRDs are aligned with the pinned 1.27.3 image so backups no longer fail on this mismatch (@myasnikovdaniil in #3526).

Development, Testing, and CI/CD

  • ci(release): move the v1.5.4 release path off the decommissioned runner pool: Every job on the release-1.5 release path was pinned to a [self-hosted] runner pool that no longer exists, so the release would silently sit Queued forever with nothing red to point at instead of failing loudly. The release, tag, and cache-warmer jobs are moved to the same runner shape release-1.6 already uses, unblocking future patch releases on this branch (@myasnikovdaniil in #3906).

  • ci: warm the build cache on release-1.5 pushes: release-1.5's build-cache warmer was dead code — it triggered on pushes to main (which never happens on this branch) and targeted the same decommissioned [self-hosted] runner — so every PR build on the branch ran cold and could exceed its 30-minute timeout. The warmer now triggers on release-1.5 pushes and targets a working runner, and the Build job is pointed at the branch-scoped cache namespace the warmer actually writes; without that the warmer would run and nothing would read it, because the cache-from refs still resolved to the shared namespace main writes, whose layers carry main's sources and miss on most of this branch's 28 packages. Builds only read the cache (WRITE_CACHE defaults to 0), so concurrent PR builds cannot race on a cache manifest and a miss degrades to a cold build rather than an error. The timeout is raised to 75 minutes at the same time: a warm build is around 15 minutes, but a cold one is around 45, and 30 minutes only ever fit the warm case — when it did not fit, the job reported a timeout instead of whatever actually went wrong (@myasnikovdaniil in #3469).

  • fix(ci, kubernetes): build container disks concurrently to unblock Build: Every release-1.5 PR Build job was dying at its timeout, and effectively all of the budget went to one target: image-ubuntu-container-disk builds one disk per supported Kubernetes minor — six of them — and each spends four to six minutes inside a libguestfs appliance installing that minor's kubelet and kubeadm into its own copy of the cloud image, roughly 28 minutes serially. None of it was cached and no warmer could fix that, because only build-main.yaml sets WRITE_CACHE=1 and main dropped this image from make image shortly before release-1.5 was cut, so the buildcache refs it reads had never been written — leaving release-1.5 as the only branch still building the image at all. The six builds share no per-version work, so they now run concurrently in a sub-make carrying its own -j (keeping the root make build serial), with buildkit deduplicating the shared guestfish and cloud-image stages across the concurrent solves and --output-sync=target keeping each version's log readable. The job also moves off the 4cpu/16gb runner shape, since six concurrent libguestfs appliances want roughly 9GB of RAM and a core each (@myasnikovdaniil in #3469).

  • ci: install flux in the release-1.5 cache warmer: The warmer's make build step shells out to flux push artifact, which the previous self-hosted runner had baked in but the new ephemeral runner shape does not; the cache warmer is updated to install the flux CLI so it can complete (@myasnikovdaniil in #3480).

  • test(seaweedfs): make the guard-parity and fake-kubectl assertions able to fail: The four assertions holding the invariant that neither SeaweedFS chart classifies on mutable claim timestamps or pod liveness were written as ! grep -qF X "$f", which cannot fail — POSIX and bash both exempt a !-negated pipeline from errexit, so the command ran, returned 1, and the test carried on reporting success. Those four were the entire body of the test, so it asserted nothing at all; and converting them naively fails, because all four names legitimately appear in the comment prose that records why each was rejected as a discriminator. They now assert against the template logic with comment blocks stripped, forbidding a live reference, and are proven in both directions. The fake kubectl used by the migration tests likewise now fails on unmodelled calls instead of returning success for anything it does not recognise (@myasnikovdaniil in #3436, backport #3878; #3892).

  • Regression coverage shipped alongside the fixes above: Each of the backported fixes carried its own test, and they are listed here rather than left silent because they are what stops these bugs coming back. A reconcile-level regression drives reconcileCNPGRestore end to end against a fake client, covering the actual call site of the repeat-restore bug rather than only its helpers (@IvanHunters in #3318, backport #3321); the e2e suite now asserts BucketAccess reaches accessGranted=true, asserts a BucketClaim converges to bucketReady promptly for both the bucket and Harbor paths, and dumps controller pods on a bucket failure (@lexfrei in #3034, backport #3532); keycloak-configure gains a run of the chart's helm-unittest suite (@lexfrei in #3372, backport #3877); the flux-shard-operator asserts the full startupProbe contract and rejects an alias regression, with the proxy-env drop and probe recorded in the sanitisation list (@IvanHunters in #3546, backport #3883); the kubeadm keep-pin folded into slot 45 has migration coverage (@myasnikovdaniil in #3892); and the PR labeler now maps the objectstorage-controller scope to area/storage (@lexfrei in #3034, backport #3532).

Documentation

  • [website] fix(backup): simplify UX with default backupclass and correct the admin override path: Rewrites the backup guides to match the corrected backupStorage override path shipped in #3402/#3333 above, so the documented admin override actually works and the default backup-class UX is simpler to follow (@androndo in cozystack/website#622).

  • [website] docs(backups): document PostgreSQL point-in-time recovery (PITR): Adds a guide for restoring a PostgreSQL cluster to a specific point in time, covering the required backup configuration and the restore procedure (@androndo in cozystack/website#629).

  • [website] docs(monitoring): add OIDC authentication guide for Grafana: New guide walking operators through enabling OIDC-based single sign-on for Grafana, including the admin-group and mode-toggle caveats (@IvanHunters in cozystack/website#597).

  • [website] docs(operations): document the gateway.http2 platform value: Documents the gateway.http2 platform value so operators know how to enable/disable HTTP/2 on the ingress gateway (@lllamnyp in cozystack/website#625).

  • [website] docs(talm): describe the preset value knobs for network and registries: Documents the new talm preset value knobs for tuning network sysctls and registry mirror configuration (@lexfrei in cozystack/website#633).

  • [website] docs: fix mermaid edge label and document wildcardSecretName for v1.5: Fixes an unrenderable mermaid diagram in the certificate documentation and documents the wildcardSecretName option for v1.5 (@lexfrei in cozystack/website#615).

  • [website] docs: correct the publishing reference and document the existingSecret cert mode: Corrects the certificate publishing reference and documents the third, existingSecret, certificate mode (@lexfrei in cozystack/website#619).

  • [website] docs(gpu): drop the manual KubeVirt patch step for GPU passthrough: The platform wires permittedHostDevices into the KubeVirt CR itself from packages/core/platform/files/gpu-passthrough-defaults.yaml, so the manual patch documented since v1.4 is obsolete; the guide now covers the pre-upgrade migration steps for hand-edited entries instead. Merged just after the v1.5.2 tag, so v1.5.4 is the first release these notes cover it in (@lexfrei in cozystack/website#556).

  • [website] docs: import the operator guides that lived in the cozystack repo: Moves the operator-facing guides that previously lived in the cozystack repository into the website's documentation, consolidating operator docs in one place (@myasnikovdaniil in cozystack/website#648).

  • [website] docs: regenerate the managed apps reference for v1.4.5 and v1.5.2: Refreshes the auto-generated managed-application reference pages so the published values tables match what those releases actually ship (@app/cozystack-ci in cozystack/website#600, cozystack/website#601).

Other repositories

talm (v0.32.0 → v0.34.0)

  • [talm] feat(charts): add preset value knobs: Adds configurable preset value knobs for network and registry settings, letting operators tune Talos machine config presets without hand-editing templates (@lexfrei in cozystack/talm#232).

  • [talm] chore(deps): migrate to Helm 4 and drop the cozystack/talos fork: Migrates talm's Helm library from v3 to v4 and drops the previously vendored cozystack/talos fork in favor of upstream Talos v1.13.7 (@lexfrei in cozystack/talm#231).

  • [talm] security: disable, then restore, KVM nested virtualization in Talos presets (CVE-2026-53359): talm first disabled KVM nested virtualization in its Talos presets to mitigate CVE-2026-53359, then reverted the blanket disable once the upstream fix landed, so presets keep nested virtualization available instead of being permanently disabled (@kvaps in cozystack/talm#224, @lexfrei in cozystack/talm#225).

  • [talm] feat(version): warn when project charts drift from the talm binary: talm now warns when a project's vendored charts have drifted from the version embedded in the talm binary, plus chart-drift-detection support in init --update, catching stale vendored charts before they cause a surprising render diff (@lexfrei in cozystack/talm#216).

ansible-cozystack (v1.5.2 → v1.6.1)

No tags were released in boot-to-talos (latest v0.7.1, 2026-03-19), cozyhr (latest v1.6.1, 2026-01-27), cozy-proxy (latest v0.3.0, 2026-04-28) or external-apps-example (no tags) during this release period.

Contributors

Thanks to everyone who contributed to this patch release:

Full Changelog: v1.5.2...v1.5.4

Download cozystack

Cozystack sandbox

Cozystack is a free PaaS platform and framework for building private clouds and providing users/customers with managed Kubernetes, KubeVirt-based VMs, databases as a service, NATS, message brokers, etc. with GPU support in VMs and Kubernetes clusters.

v1.6.2

v1.6.2 (2026-08-19)

A patch release with six fixes covering the backup-strategy controller, kube-ovn's webhook certificate, Velero CRD upgrades, CNPG barman-cloud backups, flux-shard-operator, and the published OpenAPI definitions, plus a release-pipeline reliability fix.

Fixes

  • fix(backupstrategy-controller): repair lookup-gated backup objects: The default backup Strategy CRs and the Velero BackupStorageLocation are gated on a Helm lookup performed while the referenced object is still being created; when that lookup came back empty the objects were skipped permanently, since helm-controller does not re-render a release whose chart and values are unchanged. The gate now resolves the default bucket credentials Secret through the RESTMapper, bounds each check, and tolerates an absent Secret instead of looping, so the default backup objects are created reliably instead of silently vanishing for months (@mattia-eleuteri in #3524, backport #3731).

  • fix(kube-ovn): reload kubeovn-webhook serving certificate on cert-manager renewal: kube-ovn-webhook loaded its TLS serving certificate once at startup and never re-read it; once cert-manager renewed the backing Secret and the old certificate expired, the apiserver's calls to the webhook failed verification and, because the MutatingWebhookConfiguration uses failurePolicy: Fail, every pod creation in tenant namespaces was rejected — including virt-launcher pods, blocking VMI startup. The webhook now serves its certificate through a reloading callback that re-reads the key pair when the mounted files change and widens renewBefore to 720h, so cert-manager renewals are honored without a pod restart (@IvanHunters in #3557, backport #3730).

  • fix(velero): apply CRD updates on upgrade via CreateReplace: Velero's CRDs stayed frozen at whatever version was first installed, since Helm never touches a chart's crds/ directory on upgrade; when the Velero image moved to a version that added new backup phases, the apiserver rejected phase transitions against the stale CRDs and backups silently stopped while the HelmRelease stayed green. The Velero package now opts into upgradeCRDs: CreateReplace, so CRDs are kept current on upgrade and backups keep working (@lexfrei in #3727, backport #3728).

  • fix(backups): request S3 checksum only when required for barman-cloud (non-AWS S3 / Ceph RGW): CNPG's barman-cloud plugin sidecar defaulted to computing a flexible checksum on every upload, which several S3-compatible backends (Ceph RGW, some MinIO / Cloudflare R2 builds) reject outright, so every backup and WAL-archive upload to those backends failed and ScheduledBackups never stored anything. Every barman-cloud ObjectStore Cozystack creates — Keycloak's system DB, the postgres app's backup and recovery stores, and the platform-managed system-bucket store — now sets AWS_REQUEST_CHECKSUM_CALCULATION=when_required, a safe default accepted by both AWS S3 and the affected backends (@androndo in #3417, backport #3767).

  • fix(flux-shard-operator): repair sharded helm-controller crashloop behind an HTTP proxy: The cloned helm-controller-shard<i> Deployment inherited HTTP_PROXY/HTTPS_PROXY/NO_PROXY from the flux-aio all-in-one wiring even though a standalone shard needs no external egress; behind an unreachable proxy the controller's blocking startup HTTPS call never completed, the manager never served /healthz, and every HelmRelease sharded to that controller was frozen. The sanitisation now also drops the inherited proxy env and adds a startupProbe derived from the liveness handler, so sharded HelmReleases keep reconciling in proxied environments instead of crashlooping forever (@IvanHunters in #3546, backport #3818).

  • fix(api): declare OpenAPIModelName for core and sdn types: The core and sdn API groups did not declare OpenAPIModelName the way the apps group already did, so their published OpenAPI definition names were Go import paths while every $ref pointing at them escaped each slash — the two spellings never matched, the reference dangled, and kubectl apply --validate failed on any resource against a cozystack-api built after the underlying Kubernetes 0.35 change. Declaring OpenAPIModelName for core and sdn too makes every published definition name the dotted Kubernetes model name, so client-side validation against the published OpenAPI works again (@myasnikovdaniil in #3808, backport #3812).

Development, Testing, and CI/CD

  • ci(release): complete the candidate-aware promotion pipeline on release-1.6: release-1.6 was missing the e2e and packages-verification jobs that Promote RC requires on its target base, so v1.6.1 was promoted with the rc e2e gate bypassed and the next patch release could not even be dispatched. Adds the rc-e2e job, the verify-release-candidate checks, hack/verify-promoted-packages.sh, hack/validate-changelog.sh and regression tests pinning the pipeline's contract, so future patch releases off release-1.6 run the same e2e and package-verification gates as main before promoting, and the tag-time changelog is validated and ported from the tag rather than regenerated (@myasnikovdaniil in #3893).

Documentation

  • [website] docs: import the operator guides that lived in the cozystack repo: Moves the operator-facing guides that used to live in the cozystack repo over to the documentation site, so operators find them alongside the rest of the docs instead of scattered across two repositories (@myasnikovdaniil in cozystack/website#648).

  • [website] docs(oidc): document private CA and staging trust: Documents how to configure tenant OIDC to trust a private certificate authority and staging certificates, closing a gap for operators running their own CA or testing with a staging issuer (@myasnikovdaniil in cozystack/website#650).

  • [website] feat(community): add a Community page and link it from the main menu: Adds a Community page linked from the site's main menu, giving visitors a single place to find how to get in touch with and contribute to the Cozystack community (@tym83 in cozystack/website#637).

  • [website] chore(telemetry): publish July 2026 and explain how the figures are derived: Publishes the July 2026 telemetry figures and documents how those figures are derived, giving the community visibility into adoption trends and how the numbers are calculated (@tym83 in cozystack/website#644).

  • [website] feat(blog): new Blockstor banner: Adds a new banner promoting Blockstor to the blog, improving the visibility of the storage control plane's announcement (@tym83 in cozystack/website#646).

Contributors

Thanks to everyone who contributed to this patch release:

Full Changelog: v1.6.1...v1.6.2

Download cozystack

werf sandbox

werf is a solution for implementing efficient and consistent software delivery to Kubernetes. It covers the entire CI/CD lifecycle and all related artifacts, glues commonly used tools (Git, Docker/Buildah, Helm, K8s) and facilitates best practices.

latest-signature

Notes added by 'git notes append'

Telepresence sandbox

Local development against a remote Kubernetes or OpenShift cluster

v2.32.1

Official Release Artifacts

Installers (with root daemon as a system service)

These installers include the option to run the root daemon as a system service, eliminating the need for elevated privileges when using Telepresence.

Platform Architecture Package
Linux (Debian/Ubuntu) amd64 telepresence-linux-amd64.deb
Linux (Debian/Ubuntu) arm64 telepresence-linux-arm64.deb
Linux (Fedora/RHEL) amd64 telepresence-linux-amd64.rpm
Linux (Fedora/RHEL) arm64 telepresence-linux-arm64.rpm
macOS amd64 telepresence-darwin-amd64.pkg
macOS arm64 telepresence-darwin-arm64.pkg
Windows amd64 telepresence-windows-amd64.msi
Windows arm64 telepresence-windows-arm64.msi

Standalone Binaries

Standalone binaries for manual installation. The root daemon runs on-demand with elevated privileges.

Platform Architecture Binary
Linux amd64 telepresence-linux-amd64
Linux arm64 telepresence-linux-arm64
macOS amd64 telepresence-darwin-amd64
macOS arm64 telepresence-darwin-arm64
Windows amd64 telepresence-windows-amd64.zip
Windows arm64 telepresence-windows-arm64.zip

Helm Chart

Code signing

Free code signing provided by SignPath.io, certificate by SignPath Foundation. See the code signing policy.

For more information, visit our installation docs.

Inspektor Gadget sandbox

Open source eBPF debugging and data collection tool for Kubernetes and Linux

Release v0.56.1

Welcome to the v0.55.1 bugfix release of Inspektor Gadget.

Bugfixes

  • [BACKPORT] USDT related fixes by @eiffel-fl in #5808:
    • uprobetracer: Parsing USDT notes in 32-bit ELF binaries no longer panics and crashes the process (dd9f3b7).
    • uprobetracer: Compressed USDT note sections are now rejected and reads from the section are capped, so a crafted binary can't exhaust memory (eab867e).
    • uprobetracer: USDT probe addresses are now mapped only through PT_LOAD file content, with overflow checks, so probes aren't attached to the wrong place (2a930bb).

Dependencies update

  • [BACKPORT] USDT related fixes by @eiffel-fl in #5808:
    • Bumped gRPC to v1.83.2 to pick up security fixes in gRPC (d2d8478).
    • Bumped otlploggrpc to v0.21.0 to pick up security fixes in OpenTelemetry, and updated the otel-logs operator for its API changes (bf45a79).
    • Ran go mod tidy on the trace_dns wasm module (69933dc).

Full Changelog: v0.56.0...v0.56.1

Prometheus graduated

The Prometheus monitoring system and time series database

3.15.0 / 2026-09-24

  • [CHANGE] PromQL: A range query whose end was not aligned to step caused subqueries inside it to evaluate past the parent's last actual step, inflating peakSamples in the query stats and against the query.max-samples limit, and wasting storage I/O reading samples that were never used in the result. Add tests to prevent regression of the fix made in #18081. #18598
  • [CHANGE] PromQL: Do not register a start timestamp reset if the start timestamp hasn't changed between subsequent samples. #19454
  • [CHANGE] Logging: Deprecate --log.level; use runtime.log_level configuration to supply the default level. #19511
  • [FEATURE] Configuration: Allow changing the process log level through runtime.log_level on configuration reload. #19511
  • [FEATURE] Prometheus: Add --auto-gomemlimit.refresh-interval flag to periodically re-detect the container or system memory limit and update GOMEMLIMIT at runtime. #18843
  • [FEATURE] Scraping: Add support for scraping targets via Unix Domain Sockets. #12024. #18091
  • [FEATURE] scrape: Implement OM2.0 scrape format. #18606
  • [ENHANCEMENT] Reduce TSDB head CPU utilization when initializing. #18001
  • [ENHANCEMENT] Docker SD: Add labels __meta_docker_container_image and __meta_docker_container_image_id. #19386
  • [ENHANCEMENT] Mixin: Add a p95/p99 remote-write send-batch latency panel to the remote-write dashboard. #19500
  • [ENHANCEMENT] Mixin: Support native histograms in the remote-write send-batch latency panel. #19522
  • [ENHANCEMENT] PromQL/TSDB: The --enable-feature=st-storage flag now automatically enables XOR2 float chunk encoding and ST-capable histogram chunk encoding, so you no longer need to pass xor2-encoding and histograms-st-encoding alongside it. #19518
  • [ENHANCEMENT] Remote write / Alertmanager: upgrade sigv4 to v0.5.0, adding session_name and tags fields for STS AssumeRole sessions. The previously undocumented service_name field is now also documented. #19569
  • [ENHANCEMENT] Scraping: Support zstd-compressed scrape responses, enabled via feature flag zstd-scrape. #19502
  • [ENHANCEMENT] TSDB: Stabilize the XOR2 float chunk encoding. --enable-feature=xor2-encoding is deprecated; use storage.tsdb.chunk_encoding.floats: xor2 instead. Check that other software reading the TSDB directly (e.g. Thanos sidecar) supports XOR2 before enabling. #19461
  • [ENHANCEMENT] TSDB: add prometheus_tsdb_head_appenders_created_total metric. #19411
  • [ENHANCEMENT] Tracing: add more spans to scrapes, API queries and rule evaluations. #19410
  • [ENHANCEMENT] UI: Show the effective configuration for each scrape pool on the Targets and Service Discovery pages. #19384
  • [ENHANCEMENT] scrape: Enable start time synthesis for summary _count and _sum series in scrape appender v2. #19323
  • [ENHANCEMENT] scrape: stop all pools in parallel for faster shutdowns. #19295
  • [ENHANCEMENT] storage/remote: Add undocumented failed_request_logging config field to debug log remote write V2 requests on send errors. #19249
  • [ENHANCEMENT] TSDB: Add fast path for XOR chunk decompression to speed up queries. #18049
  • [ENHANCEMENT] UI: Improve native histogram table formatting and add a background bar indicating the bucket count. #19332
  • [ENHANCEMENT] TSDB: Add prometheus_tsdb_head_series_pending_commit_underflow_total to report pending-sample reservation underflows. #19470
  • [PERF] AWS SD: Build RDS cluster labels once per cluster instead of once per instance. #19504
  • [PERF] AWS SD: Describe RDS instances of different clusters concurrently, bounded by request_concurrency. #19506
  • [PERF] AWS SD: Describe each ElastiCache resource once per refresh instead of twice. #19585
  • [PERF] Remote read: Avoid cloning labels for sampled reads when no external labels are configured. #19503
  • [PERF] Remote write: Reuse OTLP converter scratch state between requests. #19388
  • [PERF] scrape: conversion from classic to native histograms should only parse start times when enabled. #19446
  • [BUGFIX] PromQL: Fix info() enrichment for composite expressions with mixed @/offset references or selector-free vector branches, preventing metadata from being evaluated at an unrelated timestamp. #19387
  • [BUGFIX] PromQL: Fix info() enrichment when input series use different subsets of identifying labels. #19557
  • [BUGFIX] PromQL: Preserve metric-name dropping through the info function when delayed name removal is enabled. #19413
  • [BUGFIX] TSDB: Don't silently drop samples when head garbage collection removes a series while it is being appended to. #19272
  • [BUGFIX] TSDB: Keep series with uncommitted samples during selected- and stale-series compaction, including when appenders overlap. #19470
  • [BUGFIX] TSDB: Do not retain head series after a synthetic start-timestamp zero sample is rejected. #19470
  • [BUGFIX] TSDB: Prevent query panics during series eviction after WAL replay. #19664
  • [BUGFIX] TSDB: fix potential deadlock between mmapSeriesChunks and gcSeries. #19460
  • [BUGFIX] TSDB: fix default block reload interval for custom options. #19368
  • [BUGFIX] AWS SD: Do not crash on serverless MSK clusters or MSK clusters without Open Monitoring. #19194
  • [BUGFIX] AWS SD: Do not panic when the ElastiCache API omits optional fields of a serverless cache or cache cluster. #19435
  • [BUGFIX] AWS SD: Reject non-positive request_concurrency instead of hanging service discovery indefinitely. #19524
  • [BUGFIX] Discovery/AWS: Avoid a panic when discovering standalone ECS tasks with custom task groups. #19302
  • [BUGFIX] Discovery: Do not panic in AWS Lightsail service discovery when an instance is missing optional fields such as availability zone, blueprint, bundle, name, state or support code. #19324
  • [BUGFIX] Discovery: delete the stale prometheus_sd_last_update_timestamp_seconds series for a config that is removed on reload. #19131
  • [BUGFIX] HTTP: Avoid truncating compressed responses when handlers set Content-Length. #19661
  • [BUGFIX] IONOS SD: Do not panic when the API response omits the server, NIC or volume collections, or a server's properties. #19438
  • [BUGFIX] Metadata will not affect the number of Remote Write v2 shards. #19218
  • [BUGFIX] Mixins: Fix label mismatches that prevented PrometheusHAGroupNotIngestingSamples and PrometheusHAGroupCrashlooping from firing. #19444
  • [BUGFIX] Native histograms: DetectReset no longer misses a counter reset when a populated bucket behind an empty one disappears, which could make histogram rate()/increase() undercount. #19367
  • [BUGFIX] Never skip histogram buckets for histogram_stddev and histogram_stdvar functions. #19521
  • [BUGFIX] OTLP: Do not abort an entire OTLP payload ingestion if one metric has no datapoints. #19343
  • [BUGFIX] PromQL: Fix FastRegexMatcher false-positive match when a capturing group is directly adjacent to a literal (e.g. .*\|(foo)\|.*). #19516
  • [BUGFIX] PromQL: Fix a panic in range selectors using the experimental anchored or smoothed modifier when the selected series has no samples inside the query window, for example a query evaluated inside a scrape gap. #19431
  • [BUGFIX] PromQL: Fix empty results when a subquery with @ is used as the matrix argument of a call that is not step-invariant (for example quantile_over_time(scalar(x), metric[...:...] @ T)). #19187
  • [BUGFIX] PromQL: Make the "found duplicate series for the match group" many-to-many matching error message deterministic by sorting the two duplicate labels. #18810
  • [BUGFIX] PromQL: Preserve parentheses around duration literals on Expr.String() round-trip. #19403
  • [BUGFIX] PromQL: Reject duration-expression offsets and @ start() / @ end() before range selectors, matching the existing rejection of literal offsets and @ <timestamp>. #19406
  • [BUGFIX] PromQL: Report the position of the histogram argument rather than of a scalar argument in the native histogram NaN observation annotations of histogram_quantile and histogram_fraction. #19330
  • [BUGFIX] PromQL: info() now applies the @ modifier/offset when evaluating the info series, so info(v @ T) enriches as of T consistently instead of depending on the query start/eval time. #19266
  • [BUGFIX] Rules: Fix a panic when a rule manager created without a logger loads a rule file containing multiple YAML documents. #19433
  • [BUGFIX] Scrape: Do not append a stale marker for a series that is still exposed when the storage returns a new series reference for it. #19328
  • [BUGFIX] Scrape: JSON log formatter correctly format scrape target info. #19472
  • [BUGFIX] TSDB: Fix in-order chunk ID overflow by wrapping head chunk IDs modulo 2^23 so they never collide with the out-of-order flag bit. #19450
  • [BUGFIX] TSDB: Fix out-of-order chunk ID overflow by wrapping firstOOOChunkID modulo 2^23 instead of growing unbounded. #19216
  • [BUGFIX] UI: Remove an extraneous X-axis tick mark in the native histogram chart when using the "linear" display mode. #19326
  • [BUGFIX] discovery/aws: Avoid a panic when an ElastiCache ARN is missing its resource ID. #19333
  • [BUGFIX] discovery/aws: Do not panic on MSK clusters whose optional API fields are absent. #19584
  • [BUGFIX] discovery/aws: Do not panic when the EC2 API omits optional instance fields. #19512
  • [BUGFIX] discovery/aws: Don't panic on ECS tasks with absent optional fields. #19396
  • [BUGFIX] discovery/aws: MSK Optional Custom Configuration Fields #19422. #19422
  • [BUGFIX] discovery/ionos: Fix panic when the IONOS API omits optional server fields. #19418
  • [BUGFIX] discovery/kubernetes: Populate __meta_kubernetes_service_loadbalancer_ip from status.loadBalancer.ingress, falling back to deprecated spec.loadBalancerIP. #19404
  • [BUGFIX] histogram: Fix Compact moving buckets to wrong indices, and producing negative bucket counts for integer histograms, when more than one span is merged in the same pass. #19312
  • [BUGFIX] promtool: Fixed tsdb dump silently dropping native histogram samples. #18051
  • [BUGFIX] scrape: fix data race in Manager.TargetsDroppedCounts to avoid miscounting dropped targets. #19304
  • [BUGFIX] scrape: fix nil histogram when native and classic histograms are mixed in one metric family. #19452
  • [BUGFIX] TSDB: Fix WAL and GC log messages to emit human-readable duration strings instead of nanosecond integers. #19307
  • [BUGFIX] TSDB: Avoid WAL corruption after a failed WAL write in Agent mode. #19700
  • [BUGFIX] Agent: Ignore unknown WAL record types, to help users rolling back. #19814
werf sandbox

werf is a solution for implementing efficient and consistent software delivery to Kubernetes. It covers the entire CI/CD lifecycle and all related artifacts, glues commonly used tools (Git, Docker/Buildah, Helm, K8s) and facilitates best practices.

v2.79.2

Changelog

Bug Fixes

  • cleanup: stop import metadata cleanup leaking goroutines (#7864) (5b7332f), closes #7862
  • release: relabel only merged releases of the current branch (6e73257)

Installation

To install werf we strongly recommend following these instructions.

Alternatively, you can download werf binaries from here:

These binaries were signed with PGP and could be verified with the werf PGP public key. For example, werf binary can be downloaded and verified with gpg on Linux with these commands:

curl -sSLO "https://tuf.werf.io/targets/releases/2.79.2/linux-amd64/bin/werf" -O "https://tuf.werf.io/targets/signatures/2.79.2/linux-amd64/bin/werf.sig"
curl -sSL https://werf.io/werf.asc | gpg --import
gpg --verify werf.sig werf
werf sandbox

werf is a solution for implementing efficient and consistent software delivery to Kubernetes. It covers the entire CI/CD lifecycle and all related artifacts, glues commonly used tools (Git, Docker/Buildah, Helm, K8s) and facilitates best practices.

v3.5.0

Changelog

Features

  • build: hide no-op secondary images by default (#7894) (63371e3)
  • build: include werf version in the build report (#7859) (851eb02)
  • cleanup: report recent automatic host cleanup (81cf86f)
  • deploy: return rendered resources from ReleaseInstall (#7869) (9a6146b)
  • dev: publish JSON Schemas for werf configuration files (#7860) (25bfb25)

Bug Fixes

  • build, buildah: stop re-compressing parent layers on stage push (#7854) (9837b10)
  • build, cleanup: keep the parallel log streaming and its progress numbers in order (#7863) (35bb659)
  • build, stages: reuse an image whose dependencies are gone from the registry (#7876) (5ef3a11)
  • build: keep staged dockerfile targets apart by their base image (#7874) (349d55f)
  • build: show skipped images in build output (#7882) (145b39d)
  • build: synchronize legacy image metadata (#7892) (d82cd44)
  • ci: restore origin/3 daily coverage (#7890) (867e10d)
  • cleanup: avoid slow tag deletion in GitLab registries (#7866) (cdb0a19)
  • cleanup: preserve retained final images (#7867) (72162e5)
  • cleanup: register each temporary path with the GC once (#7875) (c49c9ae)
  • deploy: prevent panic when release deletion fails (#7861) (1015652)
  • deploy: render included charts at the project root (#7865) (83605cc)
  • deploy: report the inaccessible events feed through the nelm logger (#7857) (8faf82a)
  • registry: delete Harbor repositories through v2 API (#7891) (2f47707)

Installation

To install werf we strongly recommend following these instructions.

Alternatively, you can download werf binaries from here:

These binaries were signed with PGP and could be verified with the werf PGP public key. For example, werf binary can be downloaded and verified with gpg on Linux with these commands:

curl -sSLO "https://tuf.werf.io/targets/releases/3.5.0/linux-amd64/bin/werf" -O "https://tuf.werf.io/targets/signatures/3.5.0/linux-amd64/bin/werf.sig"
curl -sSL https://werf.io/werf.asc | gpg --import
gpg --verify werf.sig werf
werf sandbox

werf is a solution for implementing efficient and consistent software delivery to Kubernetes. It covers the entire CI/CD lifecycle and all related artifacts, glues commonly used tools (Git, Docker/Buildah, Helm, K8s) and facilitates best practices.

v3.6.0 [dev]

Changelog

Features

  • cleanup: scan several kubernetes namespaces for used images (59efb22)

Bug Fixes

  • build, ci: preserve embedded image tags and test embedded assets (#7898) (86eb31b)
  • build, deploy: fix stage reuse and release history regressions (#7897) (2411186)
  • build: avoid panics on missing or rejected stages (bbb08ad)
  • build: avoid redundant registry listings when publishing metadata (8d5fb5a)
  • build: avoid slow startup on image cache misses (1ba2e9f)
  • build: avoid slow startup on image cache misses (#7903) (905c0a5)
  • build: honor empty and partial git stage dependencies (5a83a7a)
  • build: honor empty and partial git stage dependencies (#7906) (dc9f5f6)
  • build: keep embedded stapel tags independent of runtime overrides (8194eed)
  • build: repair broken destination stages during registry copies (dfa8b3e)
  • build: validate internal base image cycles when reading config (d57b2fc)
  • cleanup: restore explicit namespace selection for Kubernetes scans (#7900) (2532c65)
  • cleanup: stop import metadata cleanup leaking goroutines (#7864) (5b7332f), closes #7862
  • deploy: apply .helmignore when reading the chart (#7830) (faf0672)
  • deploy: honor the release history limit environment variable (76a3c51)
  • host-cleanup: speed up project stage discovery (#7907) (b6c4a38)
  • release: relabel only merged releases of the current branch (12c5004)

Installation

To install werf we strongly recommend following these instructions.

Alternatively, you can download werf binaries from here:

These binaries were signed with PGP and could be verified with the werf PGP public key. For example, werf binary can be downloaded and verified with gpg on Linux with these commands:

curl -sSLO "https://tuf.werf.io/targets/releases/3.6.0/linux-amd64/bin/werf" -O "https://tuf.werf.io/targets/signatures/3.6.0/linux-amd64/bin/werf.sig"
curl -sSL https://werf.io/werf.asc | gpg --import
gpg --verify werf.sig werf
Hyperlight sandbox

A lightweight, secure container runtime solution designed for modern cloud-native workloads

Latest prerelease from main branch

What's Changed

Added

  • Per-direction virtqueue configuration through SandboxConfiguration and
    SandboxBuilder, with allocations included in scratch sizing.
  • Shared virtqueue framing with a 12-byte MsgHeader and external byte values.
  • ExternalValueSource implementations for RecvChain and Segments.
  • Producer batch completion without notification and segmented payload
    assembly and extraction without flattening.

Changed

  • Support overriding the guest log level when building or restoring initialized
    snapshots.
  • Snapshot::save now writes the guest memory blob sparsely, skipping all-zero
    blocks instead of writing them. A guest memory image is mostly untouched
    pages, so this cuts the bytes actually written by roughly the proportion of
    the guest's memory it never touched. The saved layout is byte-for-byte
    identical and its digest is unchanged, so this is transparent to readers and
    to previously saved snapshots. Filesystems that do not support sparse files
    store the blob as before.
  • Expose C guest ByteChunks values as pointer and length arrays.
  • Return typed hl_ReturnValue objects from C guest functions through
    hl_result_from_* constructors.
  • Guest tracing skips its per-call and per-callsite work while the guest log
    level is OFF. hyperlight_guest_tracing::is_trace_enabled reports whether
    the configured level is above OFF rather than whether the tracing state was
    allocated.
  • Breaking: Virtqueue rings and pools occupy host-owned scratch before page
    tables. Snapshots use ABI 5 and config schema v3. Existing snapshots must be
    regenerated.
  • Host virtqueue access uses checked copies and atomics across mapped scratch.
    Snapshot admission checks geometry, canonical rings, and distinct, aligned
    H2G pool slots.
    Consumers validate descriptors and payload accesses during use.
  • Virtqueue producers use concrete SlotPool allocation and BufferLease
    ownership. BufferMap supplies complete owners exposing initialized bytes.
  • VirtqProducer::reset and GuestContext::prepare_snapshot are unsafe.
    Peers must stay stopped with no live consumer-side chain handles until
    their consumers are reset or replaced.
  • ChainBuilder::build() allocates readable and writable requests.
    writable_avail() reserves available upper-tier slots within the descriptor
    budget. It may add zero slots to a nonempty chain.
  • Require guest logs and all host and guest function calls to use virtqueues.
  • Keep registered Rust guest return values typed until transport encoding so
    external byte results avoid intermediate FlatBuffer copies.
  • Store canonical virtqueue rings in versioned OCI transport layers. Config v3
    rejects snapshots without transport state.
  • Running snapshots checkpoint dirty virtqueues before capture. Ordinary calls
    keep their deferred result path.
  • Reject snapshot capture while guest-owned transport buffers are retained.
  • Use the reclaimed stack pages to raise the default G2H and H2G pools to 12
    and 8 pages.

Removed

  • RunPool and the run-specific AllocError::InvalidAlign variant.
  • Remove legacy stack I/O, its GuestHandle methods, and its sandbox
    configuration and builder options.
  • Embedded byte payload tables and their value-union variants.

Fixed

  • Linear-time segment consumption for fragmented virtqueue messages.
  • Allow reclaimed virtqueue completions to span multiple ring reuse cycles.
  • Virtqueue consumers return errors when payload copies or runtime bookkeeping
    cannot be allocated.
  • Use a 16 KiB-aligned default scratch size for Apple Silicon compatibility.
  • Guest virtqueue copies reject overlapping buffers before accessing memory.
  • Keep sandboxes usable after an H2G request exceeds available virtqueue capacity.
  • Snapshot checkpoints ignore idle cancellation and clear partial abort state.
  • Keep sandboxes usable when G2H calls exhaust reply capacity.
  • Reject incompatible transport snapshots before changing sandbox state.
  • Allow guest host-return conversions to call or log to the host.

Full Changelog (excl. dependencies)

Full Changelog (dependencies)

  • chore(deps): bump crate-ci/typos from 1.49.0 to 1.50.1 by @dependabot[bot] in #1797
  • chore(deps): bump smallvec from 1.15.2 to 1.16.0 by @dependabot[bot] in #1800
  • chore(deps): bump crossbeam-channel from 0.5.16 to 0.5.17 by @dependabot[bot] in #1808
  • chore(deps): bump cc from 1.4.4 to 1.4.5 by @dependabot[bot] in #1806
  • chore(deps): bump syn from 3.0.4 to 3.0.5 by @dependabot[bot] in #1805
  • chore(deps): bump Swatinem/rust-cache from 2.9.1 to 2.9.2 by @dependabot[bot] in #1710
  • chore(deps): bump crossbeam-queue from 0.3.13 to 0.3.14 by @dependabot[bot] in #1809
  • chore(deps): update and group mshv crates together by @simongdavies in #1813
  • chore(deps): bump actions-rust-lang/setup-rust-toolchain from 1.17.0 to 2.0.0 by @dependabot[bot] in #1818
  • chore(deps): bump the wasm-tools group across 1 directory with 3 updates by @dependabot[bot] in #1780
  • Configure Dependabot cooldown and Windows crate groups by @simongdavies in #1821
  • chore(deps): bump bitflags from 2.13.1 to 2.13.2 by @dependabot[bot] in #1830
  • chore(deps): bump wat from 1.258.0 to 1.259.0 by @dependabot[bot] in #1828
  • chore(deps): bump uuid from 1.26.0 to 1.26.1 by @dependabot[bot] in #1827
  • chore(deps): bump docker/build-push-action from 7.3.0 to 7.4.0 by @dependabot[bot] in #1845
  • chore(deps): bump cfg-if from 1.0.4 to 1.0.5 by @dependabot[bot] in #1853
  • chore(deps): bump syn from 3.0.5 to 3.0.6 by @dependabot[bot] in #1852

New Contributors

Full Changelog: v0.17.0...dev-latest

kpt sandbox

Automate Kubernetes Configuration Editing

v1.0.1

What's Changed

Features

Behavioral changes

  • Feature: Implement comprehensive field validation for Kptfile (#4735) by @adi-coderr in #4742
  • Accept scp-style and local upstream.git.repo values in Kptfile validation by @chiliec in #4750

Fixes

  • Fix copy-merge to preserve locally-added files in upstream-deleted directories by @kushnaidu in #4720
  • Fix Kptfile pipeline not updating on pkg update with fast-forward/force-delete-replace by @OisinJohnston2005 in #4704
  • Fix incorrect handling when unfetched packages are deleted from upstream by @aravindtga in #4695
  • run: add authoritative gitCommit ldflag override for version command (#4752) by @adi-coderr in #4757

Dependencies

  • Bump github.com/google/cel-go from 0.28.1 to 0.29.0 by @dependabot[bot] in #4771

Docs & internal

New Contributors

Full Changelog: v1.0.0...v1.1.0