Netdata

Infrastructure & DevOpsGPL-3.0

Real-time performance monitoring for systems and applications.

Latest v2.11.0 · by NetdataWritten in GoWebsitenetdata/netdataRSS

Release activity

Release activity — 11 releases across 11 days since Dec 3, 2025. Each cell is one day; darker means more releases that day. Nothing is recorded before Dec 3, 2025. Older weeks are hidden at this screen width.
JunJulAugSep
SundayNo releases on May 24, 2026No releases on May 31, 2026No releases on Jun 7, 2026No releases on Jun 14, 2026No releases on Jun 21, 2026No releases on Jun 28, 2026No releases on Jul 5, 2026No releases on Jul 12, 2026No releases on Jul 19, 2026No releases on Jul 26, 2026No releases on Aug 2, 2026No releases on Aug 9, 2026No releases on Aug 16, 2026No releases on Aug 23, 2026No releases on Aug 30, 2026No releases on Sep 6, 2026
MondayNo releases on May 25, 2026No releases on Jun 1, 2026No releases on Jun 8, 2026No releases on Jun 15, 2026No releases on Jun 22, 2026No releases on Jun 29, 2026No releases on Jul 6, 2026No releases on Jul 13, 2026No releases on Jul 20, 2026No releases on Jul 27, 2026No releases on Aug 3, 2026No releases on Aug 10, 2026No releases on Aug 17, 2026No releases on Aug 24, 2026No releases on Aug 31, 2026No releases on Sep 7, 2026
TuesdayNo releases on May 26, 2026No releases on Jun 2, 2026No releases on Jun 9, 2026No releases on Jun 16, 2026No releases on Jun 23, 2026No releases on Jun 30, 2026No releases on Jul 7, 2026No releases on Jul 14, 2026No releases on Jul 21, 2026No releases on Jul 28, 2026No releases on Aug 4, 2026No releases on Aug 11, 2026No releases on Aug 18, 2026No releases on Aug 25, 2026No releases on Sep 1, 2026No releases on Sep 8, 2026
WednesdayNo releases on May 27, 2026No releases on Jun 3, 2026No releases on Jun 10, 2026No releases on Jun 17, 2026No releases on Jun 24, 2026No releases on Jul 1, 2026No releases on Jul 8, 20261 release on Jul 15, 2026No releases on Jul 22, 2026No releases on Jul 29, 2026No releases on Aug 5, 20261 release on Aug 12, 2026No releases on Aug 19, 2026No releases on Aug 26, 2026No releases on Sep 2, 2026
ThursdayNo releases on May 28, 2026No releases on Jun 4, 2026No releases on Jun 11, 2026No releases on Jun 18, 2026No releases on Jun 25, 2026No releases on Jul 2, 2026No releases on Jul 9, 2026No releases on Jul 16, 2026No releases on Jul 23, 2026No releases on Jul 30, 2026No releases on Aug 6, 2026No releases on Aug 13, 2026No releases on Aug 20, 2026No releases on Aug 27, 2026No releases on Sep 3, 2026
FridayNo releases on May 29, 2026No releases on Jun 5, 2026No releases on Jun 12, 2026No releases on Jun 19, 2026No releases on Jun 26, 2026No releases on Jul 3, 2026No releases on Jul 10, 2026No releases on Jul 17, 2026No releases on Jul 24, 2026No releases on Jul 31, 2026No releases on Aug 7, 2026No releases on Aug 14, 2026No releases on Aug 21, 2026No releases on Aug 28, 2026No releases on Sep 4, 2026
SaturdayNo releases on May 30, 2026No releases on Jun 6, 2026No releases on Jun 13, 2026No releases on Jun 20, 2026No releases on Jun 27, 2026No releases on Jul 4, 2026No releases on Jul 11, 2026No releases on Jul 18, 2026No releases on Jul 25, 2026No releases on Aug 1, 2026No releases on Aug 8, 2026No releases on Aug 15, 2026No releases on Aug 22, 2026No releases on Aug 29, 2026No releases on Sep 5, 2026

11 releases since Dec 3, 2025

Changelog

v2.11.0

Latest
Added 15
  • Network Monitor Dashboard (Technical Preview) with device inventory, metrics, topology, flows, traps and alerts
  • Network Flows collection for NetFlow, IPFIX and sFlow from routers, switches and firewalls (Technical Preview)
  • Network Topology mapping of Layer 2 and Layer 3 network structure via SNMP with LLDP neighbours, forwarding databases, OSPF and BGP adjacencies (Technical Preview)
  • Application Dependency Mapping from Linux kernel to map processes, containers and Kubernetes workloads without instrumentation (Technical Preview)
  • SNMP Trap Listener to decode traps from 800+ vendor profiles into charts and readable messages
  • OpenTelemetry logs ingestion with storage and querying through the Logs interface
Changed 4
  • SNMP device coverage expanded from 229 to 273 profiles
  • Prometheus collector rebuilt with real relabeling engine and application profiles
  • Netdata Cloud Nodes and Alerts views rebuilt for 50,000-node scale
  • Netdata Cloud silencing rules gained timezone-anchored scheduling
Deprecated 1
  • Full support ended for RHEL 7.x, CentOS 7.x, and Amazon Linux 2

From Netdata

Table of Contents
Release Summary

Netdata has always been exceptionally good at one particular thing: telling you, per second and with no configuration, what is happening inside a machine. v2.11.0 is the release where that stops being the boundary.

Two things happen here, and together they change what Netdata is.

First, Netdata learns the network. A new Network Monitor dashboard brings device inventory, metrics, topology, flows, traps and alerts into one place. Behind it, Network Flows receives NetFlow, IPFIX and sFlow from your routers, switches and firewalls and turns raw records into a faceted view of who is talking to whom. Network Topology walks your devices over SNMP and draws the Layer 2 and Layer 3 map — LLDP neighbours, forwarding databases, OSPF and BGP adjacencies — while Application Dependency Mapping does the same for your software, reading your Linux hosts' kernels to map what each process, container and Kubernetes workload talks to, with nothing to instrument. The SNMP Trap Listener makes Netdata a trap receiver that decodes traps from 800+ vendor profiles into charts and readable messages instead of raw OIDs. Feeding all of it, SNMP device coverage grew from 229 to 273 profiles. The usual Netdata rules still apply: the profiles are already written, collection and storage stay on your own infrastructure, and there is little to configure beyond telling Netdata where the devices are. Network Monitor, Network Flows and Network Topology ship as technical previews — complete enough to run today, open to every Netdata user while in preview, and still moving fast.

Second, and just as important, logs become a first-class pillar of the platform. This release adds OpenTelemetry logs — OTLP log records are ingested, stored and queried from the same Logs tab that already serves systemd journal and Windows events, with journald, filelog and syslog receivers, configurable retention, log-to-metric conversion and parsing for unstructured lines. The deeper change is architectural: the same faceted, high-cardinality log engine now backs systemd journal, Windows events, OpenTelemetry logs, SNMP traps and network flows alike. Five very different signal types, one query surface, all stored on your own infrastructure.

Around those two, this release completes the major-cloud set with a new AWS CloudWatch collector — the counterpart to the Azure Monitor collector added in v2.10.0 — and rebuilds the Prometheus collector around a real relabeling engine with application profiles. Underneath all of it sits the largest stability and memory-safety effort we have ever shipped in a single release: 87 incremental parts across the database engine, streaming, ACLK, health, ML and libnetdata.

Netdata's strong endpoint monitoring capabilities have been enhanced. macOS monitoring was rebuilt natively: unified logs through Apple's OSLog framework, GPU, SMC and IOHID sensors, fans, power and battery, thermal pressure and NVMe health, none of which was reachable before without dropping to log show and powermetrics by hand. Windows added Active Directory and SMB monitoring on a reworked perflib layer, FreeBSD added system-call monitoring, and Linux added audit subsystem monitoring. One agent, four endpoint operating systems, per-second resolution on each.

In Netdata Cloud, AI troubleshooting just became more accurate and more relevant: you can now connect your own MCP servers to your Netdata Space — GitHub, PagerDuty and Atlassian, and your own custom servers — and Netdata AI will use them while investigating. Alongside that, new Fleet Management views render nodes as status hexagons, the Nodes and Alerts views were rebuilt to stay fluent in 50,000-node rooms, silencing rules gained timezone-anchored scheduling, and the mobile app opened up to everyone on free plans.

FeatureHighlightsDetails
Network Monitor DashboardOne place for the whole networkTechnical Preview• Device inventory, vendors, health and top talkers at a glance• Overview, Devices, Metrics, Topology, NetFlow, Traps and Alerts sub-tabs• Per-device and per-interface traffic, speed and error rates
Network FlowsNetFlow v5/v9, IPFIX, sFlowTechnical Preview• New high-throughput flow-analysis plugin with its own four-tier journal• Cisco ASA NSEL accounting• Enrichment with cloud provider IP ranges, BGP/RIPE RIS data, and private-IP labelling
Network TopologyL2 and L3 topology discoveryTechnical Preview• LLDP, FDB and STP based Layer 2 topology• L3 subnet segments plus OSPF and BGP adjacencies• MAC OUI vendor lookup and reverse DNS enrichment
Application Dependency MappingPer-host service maps, no instrumentationTechnical Preview• What each process talks to, read from the kernel socket table• Attributed per container, image, systemd unit and Kubernetes pod/workload (Linux)• Regroup the map by process name, container or PID
SNMP Trap Listener800+ vendor trap profiles• Profile-based trap decoding into charts and log entries• Trap enrichment and forwarding to SIEM• Configurable via the Dynamic Configuration UI
Expanded SNMP Device Coverage273 device profiles, up from 229• Cisco Catalyst, Cisco Nexus, FortiGate and Check Point improvements• BGP session monitoring and network-device licence monitoring• SNMPv3 context names, bitmask value mappings, ping_only mode
LogsOne engine, five signal types• OpenTelemetry (OTLP) log ingestion, query and retention• journald, filelog and syslog receivers; log-to-metric conversion• Same faceted engine now backs journal, Windows events, traps and flows
AWS CloudWatch Collector47 service profiles, 800+ metrics• Automatic resource discovery with tag filtering and multi-account support• Stock health alerts, incl. load-balancer target health and MSK• Query policies to control API cost on sparse metrics
New CollectorsCato Networks, PAN-OS, and more• SASE and firewall monitoring for Cato Networks and Palo Alto PAN-OS• New endpoint collectors for macOS, Windows, Linux and FreeBSD• HTTP service discovery, range heatmaps and chart series aggregation
Endpoint MonitoringmacOS, Windows, FreeBSD, Linux• Native macOS logs, GPU, sensors, fans, power, thermal and NVMe health• Windows Active Directory and SMB on a reworked perflib layer• FreeBSD system calls; Linux audit subsystem and nfds monitoring
Prometheus Collector OverhaulRelabeling and app profiles• Full relabeling engine with job-level metric relabeling• Profile-driven application detection for chart contexts• Single-pass stream parser, migrated to the V2 collector framework
Query Engine and APITop-N queries and correctness• New limit parameter for top-N dimension queries• New latest time-grouping with a collector-cache fast path• Deterministic cardinality limiting and group-by aggregation fixes
Netdata Cloud: AI Troubleshooting and MCPBring your own MCP servers• Connect GitHub, PagerDuty and Atlassian Cloud to a Space over OAuth• Netdata AI uses connected tools in conversations and reports• Read-only tools only, enforced by Netdata• Multi-algorithm capacity forecasting with automatic selection
Netdata Cloud: Alerts and NotificationsSilencing, labels, misconfiguration• Timezone-anchored, DST-safe silencing rules with auto-expiry• Host labels included in every notification channel• Space admins notified about misconfigured alerts• New Group Email integration for shared mailing lists• Mobile app now free for all
Netdata Cloud: Fleet Management and Scale50,000-node rooms• New Nodes Fleet Management views with node-group hexagons• Rebuilt home page with a geographic fleet map• Nodes and Alerts stay fluent at very high cardinality
Netdata Cloud: Dashboards and ChartsNew cards and aggregations• State Timeline, heatmap, alert-count, node-breakdown and inventory cards• Dashboard playlists, plus import and export• Percentage-of-group and latest time aggregation
Stability, Security and PerformanceProduction reliability• Faster agent startup through optimized host-context loading• MCP endpoints protected by a dedicated ACL and bearer tokens• Large memory-safety effort across the database engine, streaming and ACLK
Release Highlights
Network Monitor Dashboard — Technical Preview

[!NOTE] Technical Preview. The capabilities that follow — Network Monitor, Network Flows, Network Topology and Application Dependency Mapping — ship as technical previews. Everything described here works today and is open to every Netdata user while in preview. What is still moving is the shape around them: we are expanding coverage, refining the interfaces, and working out how these capabilities are ultimately packaged and supported, so expect their scope and availability to keep evolving over the next few releases. Preview feedback carries real weight, please tell us what you need through any of the support channels below.

Everything below is surfaced through one new place in the interface: a dedicated Network Monitor tab that brings device inventory, metrics, topology, flows and traps together instead of leaving them as separate tools.

Network Monitor dashboard showing device counts by type, network health, errors and drops, devices by vendor, top devices by traffic, and device inventory

Sub-TabWhat It Shows
OverviewDevice counts by type, network health, errors and drops, devices by vendor, top devices by traffic, and a full device inventory
DevicesEvery discovered network device, with vendor, type and address
MetricsPer-device and per-interface metrics: traffic in/out, interface speed, error and discard rates
TopologyThe discovered Layer 2 and Layer 3 map
NetFlowFlow analysis, with the full faceted search described below
TrapsDecoded SNMP traps as charts and searchable log entries
AlertsAlerts raised on network devices

The result is that an engineer investigating a network problem starts from "which devices do I have and are they healthy" and drills into flows, topology or traps without leaving the page.

Network Flows: NetFlow, IPFIX and sFlow — Technical Preview

Netdata can now receive and analyze network flow records. The new Network Flows plugin accepts NetFlow v5/v9, IPFIX and sFlow from routers, switches and firewalls, and turns them into a searchable, faceted view of who is talking to whom on your network.

Network Flows Sankey view grouping flows by source AS, protocol and destination AS, with the faceted filter sidebar

Flows can be grouped and sorted on any field and rendered as a table, Sankey diagram, time series, or country, state and city maps.

What You Get
FeatureDetails
Multi-Protocol IngestionNetFlow v5 and v9, IPFIX, and sFlow on configurable listeners
Cisco ASA NSELFirewall accounting records, including connection and NAT events
Four-Tier Local JournalRaw records plus 1-minute, 5-minute and 1-hour rollups, stored on the agent
EnrichmentCloud-provider IP ranges, BGP/RIPE RIS routing data, private-IP labelling, and reverse DNS
Faceted SearchFilter and group flows interactively from the Live tab
Accounting ChartsPer-listener and per-source metrics on records received, processed and dropped
Where Data Lives

Flow data stays on the agent. Collection and the four-tier journal remain local under the configured journal_dir — by default ${NETDATA_CACHE_DIR}/flows, typically /var/cache/netdata/flows/ for native packages.

[!IMPORTANT] During the Tech Preview, Network Flows is available also on the free Community tier and requires an agent connected to Netdata Cloud and a > signed-in user whose Space role permits sensitive functions. If you see Connect this agent to Netdata to use this function, the plugin is > working — the block is access control, not a failure. Connecting the agent does not move or offload its flow storage. See the access requirement for details.

The plugin ships in the static and Docker builds.

Network Topology and Application Dependencies — Technical Preview

Netdata now draws your infrastructure — twice. A new SNMP topology engine builds Layer 2 and Layer 3 maps of your network from the devices themselves. On your Linux hosts, the Network Viewer builds a second map: what each process, container and Kubernetes workload is talking to, read straight from the kernel.

SNMP network topology map showing discovered Cisco, Juniper and Aruba devices and their adjacencies, with map modes and inference strategies

Topology Discovery
LayerDiscovered FromProduces
Layer 2LLDP neighbours, forwarding databases (FDB), STPSwitch-to-switch and switch-to-host adjacencies
Layer 3Interface addressing and subnet dataL3 subnet segments and subnet adjacencies
RoutingOSPF neighbour tables, BGP peer tablesOSPF and BGP adjacency links between routers

Discovered devices are enriched with MAC OUI vendor lookup and reverse DNS, so the map is labelled with real vendor and host names rather than raw addresses.

Application Dependency Mapping

The same topology view that draws your switches also draws your software. On a monitored host, Netdata reads the kernel's live socket table and turns it into a dependency map: what each process is talking to, over which port — attributed, on Linux, to the container, image, systemd unit, or Kubernetes pod, namespace and workload that owns it.

There is nothing to instrument. No language agents, no sidecars, no code changes, no service mesh. These are the sockets your kernel already has, enumerated fresh each time you open the map and drawn as a graph — on hosts you are already monitoring, with no setup.

  • Attribution down to the workload — every process in the map carries its command line and the user it runs as, plus the container, image, systemd unit or Kubernetes pod, namespace and workload behind it.
  • Group the map the way you think — collapse by process name, container, or PID. A service running eight worker processes is one box when you want the architecture, and eight when you need the one misbehaving worker.
  • Aggregated or detailed — the dependency graph by default, or the same graph plus the individual socket evidence behind every link.
  • Linux, FreeBSD and macOS — container and Kubernetes attribution is Linux-only; on FreeBSD and macOS the map is drawn from processes and endpoints. On Windows, the Network Connections and Network Protocols tables remain available, the latter carrying SMB.

Today each host draws its own map: processes on the same host are linked to each other directly, and peers elsewhere appear as the addresses they are.

[!TIP] See the Network Topology documentation for setup and supported devices, and Application Dependency Mapping for what the connections map covers and on which platforms.

SNMP Trap Listener

Netdata can now act as an SNMP trap receiver. The new SNMP Trap Listener accepts traps from network devices, decodes them using profiles, and turns them into charts and searchable log entries.

SNMP traps decoded into an events distribution chart and readable log messages, with facets for trap name, severity, vendor and source IP

Instead of an opaque OID, a trap arrives as IF-MIB::linkDown with the message "Link down on interface 30 on MikroTik-router" — charted by trap name, filterable, and searchable alongside your other logs.

FeatureDetails
800+ Vendor ProfilesTrap definitions covering a broad range of network-equipment vendors
Profile-Based DecodingTraps are matched to profiles and rendered as meaningful charts and fields, not raw OIDs
EnrichmentTrap sources are enriched and deduplicated
SIEM ForwardingForward decoded traps onward to a SIEM
UI ConfigurationConfigure listeners through the Dynamic Configuration UI

[!TIP] See the SNMP Traps documentation for listener configuration, the field reference, enrichment, and SIEM forwarding.

Expanded SNMP Device Coverage

SNMP device profiles grew from 229 to 273 in this release, and the profile engine gained several capabilities.

  • Vendor improvements: Cisco Catalyst, Cisco Nexus, Fortinet FortiGate, Check Point, and Netgear (topology discovery).
  • BGP session monitoring for network devices.
  • Network-device licence monitoring, so you can alert before a licence lapses.
  • SNMPv3 context name support.
  • Structured mapping configuration and bitmask value mappings, so vendor status bitfields become readable states.
  • ping_only mode for devices you want reachability for without a full SNMP walk.
  • More reliable collection: snmpEngineTime as the primary uptime source, a lower default MaxOIDs of 20, corrected retry serialization, and tolerance for invalid optional typed row values.
Logs: OpenTelemetry Ingestion and a Unified Log Engine

Netdata already ingested OpenTelemetry metrics over OTLP. v2.11.0 adds logs: OTLP log records are received, stored through a dedicated storage path, and queried from the Logs tab alongside systemd journal and Windows event logs. Retention and storage defaults are configurable, and the otel-plugin is now also built and enabled on macOS.

In practice, an OpenTelemetry Collector can forward logs to Netdata from the journald, filelog and syslog receivers, and you can shape them on the way in — converting logs to metrics, or parsing unstructured lines into fields you can then filter on.

One Engine, Five Signal Types

The more consequential change is what now sits underneath. Netdata's faceted, high-cardinality log engine — journal files, facet extraction, and interactive filtering — became the common substrate for everything log-shaped:

SourceStatus in v2.11.0
systemd journalEstablished; this release fixes memory retention after queries and journal access validation
Windows eventsEstablished; row rendering fixes in this release
OpenTelemetry logsNew — OTLP ingestion with its own storage and query path
SNMP trapsNew — decoded traps become searchable log entries, not just charts
Network flowsNew — flow records land in a four-tier journal with the same faceted search

This is why a NetFlow record, a decoded IF-MIB::linkDown trap and an application log line all filter, group and search the same way. It also means improvements to the engine — indexing, retention accounting, memory behaviour — benefit all five at once. Groundwork for virtual log functions landed in this release as well.

AWS CloudWatch Collector

Monitor your AWS estate from Netdata with the new CloudWatch collector. It reads platform metrics from the CloudWatch API with automatic resource discovery: point it at your accounts, and it finds your resources, matches them to built-in profiles, and starts collecting.

This completes the major-cloud set — Netdata v2.10.0 added Azure Monitor, and this release adds its AWS counterpart.

What You Get
FeatureDetails
47 Built-in Service Profiles800+ metrics across compute, containers, databases, networking, storage, messaging, AI/ML and billing
Automatic DiscoveryDiscovers resources and enables matching profiles without per-resource configuration
Multi-AccountCollect from several AWS accounts in a single collector job
Resource Tag FilteringScope discovery by resource tags, and optionally attach tags as chart labels
Explicit TargetsDeclare exact targets, collection rules, and exact metric selections when you do not want discovery
Query PoliciesPer-rule and per-selection policies control how sparse metrics are queried, keeping CloudWatch API cost down
Built-in AlertsStock health alerts, including load-balancer target health and MSK cluster conditions
Collector Activity MetricsProfile-aware metrics on the collector's own API usage, so you can see what it is costing you
UI ConfigurationConfigure through the Netdata Dynamic Configuration UI without editing files
Supported Services
CategoryServices via Profiles
ComputeEC2, Auto Scaling, Lambda, Step Functions
ContainersECS, EKS (incl. control plane)
DatabasesRDS, DocumentDB, DynamoDB (incl. per-operation), ElastiCache, Redshift, OpenSearch
NetworkingALB (incl. target and target health), NLB, ELB, NAT Gateway, Site-to-Site VPN, CloudFront, API Gateway
PrivateLinkEndpoints and endpoint subnets, services with per-AZ, per-load-balancer and per-VPC-endpoint breakdowns
StorageS3 (incl. request metrics), EBS (incl. stalled I/O), EFS
MessagingSQS, SNS, EventBridge, Kinesis, Firehose, MSK (incl. per-cluster)
AI & MLBedrock
BillingTotal, per-service, per-linked-account and per-linked-account-per-service cost metrics

[!TIP] For setup and authentication options, see the CloudWatch collector documentation. To write your own profiles, see the AWS CloudWatch Profile Format.

New Collectors
CollectorMonitors
Cato NetworksCato SASE sites, tunnels and traffic
Palo Alto PAN-OSPAN-OS firewall health, sessions and throughput
Linux audit subsystemAudit events from the Linux kernel audit subsystem
Active Directory (windows.plugin)Domain controller health and replication
SMB protocol (windows.plugin)SMB server and client activity
FreeBSD system callsSystem-call activity on FreeBSD
macOS hardware and logsGPU, SMC and IOHID sensors and fans, power sources, thermal pressure, NVMe health and macOS unified logs — see Endpoint Monitoring below

Also new in the collector framework: HTTP service discovery, single-instance collector support, histogram buckets rendered as range heatmaps, chart series aggregation, and mutable chart labels.

Endpoint Monitoring: macOS, Windows, FreeBSD and Linux

Several of the collectors above belong to a larger push that runs through this whole cycle: bringing every endpoint operating system up to the depth Netdata has always had on Linux. Same agent, same per-second resolution, same zero-configuration behaviour — whether the endpoint is a Linux server, a Mac in a build farm, a Windows domain controller or a FreeBSD host.

PlatformWhat Landed in This Cycle
macOSA native overhaul: unified logs, GPU, sensors and fans, power and battery, thermal pressure, NVMe health and per-application metrics — detailed below
WindowsActive Directory and SMB protocol monitoring, a reworked perflib layer with missing memory metrics and path translation, and Windows connection and protocol tables in the Network Viewer
FreeBSDSystem-call monitoring, connection-map support in the Network Viewer, and build and counter-accuracy fixes
LinuxLinux audit subsystem monitoring, per-application nfds (open file descriptor) monitoring, and the groundwork for a new eBPF Go plugin
macOS: Native Logs, Sensors, GPU and Hardware Health

macOS is the largest single piece. Apple silicon put Macs into build farms, CI fleets and render pipelines, but macOS monitoring stayed shallow — for logs and hardware telemetry you dropped to log show and powermetrics by hand and read the output yourself. Netdata now collects that natively, with the same charts, alerts and retention as everything else.

AreaWhat Netdata Now Collects
Unified LogsRead through Apple's OSLog framework, not by parsing log show output. Severity histograms, faceted filtering on level, process, PID, sender, subsystem, category, thread, activity and signpost fields, full-text search and live tail — in the same Logs tab as journald and Windows events
GPU (Apple Silicon)Active-residency utilization, time-weighted clock frequency, performance-state residency, power draw and die temperature
Sensors and FansTemperature, voltage, current, power and fan speed from SMC and IOHID, summarised per subsystem with optional per-sensor detail
Thermal PressureNominal, moderate, heavy, sleeping and trapping states, per-subsystem thermal levels for CPU, GPU and I/O, and processor-hot assertions
Power and BatteryCharge, voltage, current, cycle count and temperature per battery or UPS — cycle count doubling as a wear indicator across an ageing laptop fleet
Storage HealthNVMe health read natively through IOKit — estimated endurance, available spare, composite temperature, power-on and power-cycle counts, unsafe shutdowns, data read and written, and media errors — plus smartctl through ndsudo
Per-Application MetricsCPU, memory and disk I/O, grouped the way macOS itself groups things: app bundles, framework helpers, daemons, driver extensions and third-party binaries
NetworkLive connections with their owning processes, TCP and UDP statistics, and the process-to-endpoint topology view, on a Darwin libproc backend

Also on macOS in this release: the otel-plugin is now built and enabled, so a Mac can ingest OpenTelemetry metrics and logs like any other host; apps.plugin gained bounded-cardinality process grouping for machines with heavy process churn; disk.util percent scaling was corrected; and an nd-run setenv SIGSEGV and a mach_smi stack buffer overflow were both fixed.

Because this work was backported to 2.10.4, many macOS users already have it.

[!TIP] Native macOS Monitoring: Logs, Sensors, GPU & Hardware Health walks through the entire macOS overhaul, chart by chart, with the reasoning behind each metric.

Windows
  • Active Directory monitoring in windows.plugin — domain controller health and replication.
  • SMB protocol monitoring — SMB server and client activity.
  • Perflib reworked, with missing memory metrics and path translation added.
  • Network Viewer on Windows — the Network Connections and Network Protocols tables, the latter carrying SMB data.
  • Carried in from the 2.10.x patch line: MSI installer fixes, Windows hardware detection, and Windows event row rendering.
FreeBSD
  • System-call monitoring, contributed by @DavidMarec.
  • Network Viewer support, so FreeBSD hosts draw a connection map from their processes and endpoints, as macOS does.
  • Agent compilation on FreeBSD fixed, plus counter size mismatches and a freebsd_ipfw off-by-one.
Linux
  • Linux audit subsystem monitoring through debugfs.plugin — audit events from the kernel audit subsystem.
  • nfds monitoring in apps.plugin, for per-application open file descriptors, contributed by @arch-yunus.
  • Groundwork for a new eBPF Go plugin (ebpfgo.plugin), and the cgroups–eBPF transport moved from shared memory to netipc IPC.
  • The Go sensors collector was removed; Linux hardware sensors have been collected by the C libsensors module in debugfs.plugin since v2.2.0, so nothing changes for users.

Cutting across platforms: a shared cross-OS sensors function and a common temperature-histogram context, so hardware sensor readings look and query the same way regardless of which platform reported them.

Prometheus Collector Overhaul

The Prometheus collector was rebuilt on the V2 collector framework and gained the pieces users kept asking for:

  • Relabeling engine, at both job level and profile level. The old label_prefix option is dropped in favour of it.
  • Application profiles (promprofiles) that detect the exporter and give its metrics proper chart contexts, instead of one undifferentiated pile of charts.
  • Profile-owned relabeling and fallback types, so a profile carries its own relabeling rules and decides how metrics it does not cover are charted — no per-job configuration needed to get a curated result.
  • Profile autogen selector for generating chart contexts from a profile.
  • Single-pass stream parser, replacing the previous multi-pass scrape parsing.

[!TIP] See the Prometheus Profile Format for writing your own profiles, and Prometheus Metric Relabeling for the relabeling rules.

Query Engine and API
  • limit parameter on data queries, for efficient top-N dimension queries.
  • New latest time-grouping, with a collector-cache fast path so "what is it right now" queries do not touch the database engine.
  • Deterministic cardinality limiting: cardinality_limit no longer crashes, limits only the result dataset by default, breaks ties deterministically, and reports partial results on mixed folds. The fold cut is exposed in jsonwrap v2.
  • Group-by correctness: fixes to percentage grouping at the second grouping level with a dimensions filter, to view.dimensions.sts.avg across sum/min/max/extremes/percentage aggregations, and per-row anomaly rate with a zero-denominator guard.
  • Accurate PARTIAL annotations: false PARTIAL annotations on complete data are fixed, raw partial-trimming metadata is exposed, and the pre-window storage point is excluded from the first bucket of tier queries.
  • Correct tier selection for sub-resolution windows.
  • csvjsonarray now always emits numeric timestamps and valid JSON when label-quotes is passed.
Netdata Cloud: AI Troubleshooting and MCP

[!NOTE] Netdata Cloud ships continuously rather than on the Agent's release cadence. The Cloud sections below cover everything that shipped between v2.10.0 (8 April 2026) and this release. If you use Netdata Cloud, you already have all of it — nothing here requires an upgrade. Items marked Beta are available to everyone; the label reflects how new they are, and means their interface and scope will keep moving over the next few releases.

MCP client capabilities — Beta. An alert tells you what changed. It rarely tells you why — that answer usually lives in the pull request that shipped minutes earlier, or the incident already open in PagerDuty. With MCP Connections, Netdata Cloud acts as an MCP client and reads those systems directly while it investigates.

MCP Connections settings, offering GitHub, PagerDuty, Atlassian Cloud and custom MCP server integrations

Note this is the reverse of the integration you may already use. Connecting Claude or Cursor to Netdata's own MCP server has been possible for a while; here, Netdata reaches out to your servers.

CapabilityDetails
Predefined IntegrationsGitHub, PagerDuty, and Atlassian Cloud for Jira, Confluence and Bitbucket — connected over OAuth, no tokens to paste
Custom MCP ServersPoint at any HTTPS MCP endpoint, with encrypted credentials, connection testing, health checks and tool selection
Read-Only, EnforcedNetdata enables read-only tools only. Tools that create, modify or delete are discovered but cannot be enabled from the UI
Used by Reports and ChatsConnected tools are available to Netdata AI in both AI conversations and generated reports
Admin-GatedConfigured per Space under Settings → AI → MCP Connections, behind an admin-only permission
On-PremisesAI and Insights features are now enabled on on-prem installations that have them configured

Why this matters: an investigation that used to stop at "CPU saturation started at 14:02" can now continue into "…which coincides with the deploy in this GitHub PR, and there is already a PagerDuty incident open for it."

[!IMPORTANT] MCP Connections require a Netdata Cloud paid plan and Space admin access to configure.

Alongside MCP, capacity forecasting was rewritten to run several forecasting algorithms and automatically select the best-performing one per metric, rather than applying a single model everywhere. Investigation prompts grew to 10,000 characters, and reports gained dedicated forecasting tooling.

Netdata Cloud: Alerts and Notifications

Silencing rules received the largest single set of changes:

  • Timezone-anchored scheduling. A rule is now bound to a timezone, so recurring maintenance windows behave correctly across DST transitions and multi-day windows no longer drift.
  • Notification options instead of severities, giving finer control over exactly what a rule suppresses.
  • Automatic deletion on expiry, so temporary silences do not accumulate, with a "rule expired" entry in the activity feed.
  • Rule authorship is recorded and shown, so you can tell who silenced what.
  • Node selection by room and multi-select, instead of picking nodes one at a time.

Notifications now carry host labels — across email, Slack, Mattermost, Rocket.Chat and mobile push (up to 100 labels per email). An alert arrives already carrying the environment, region, cluster or team that the node belongs to, so routing and triage no longer require a lookup. Slack and Mattermost push notifications also render message previews.

Misconfigured alerts are no longer silent. Space administrators are notified when alerts are misconfigured — for example alerts stuck in a raised state — with a direct link to a dedicated tab listing them.

A new Group Email notification integration sends Space-wide alerts to a shared group or mailing-list address, independently of each member's personal email notification settings — so an on-call alias or team distribution list receives alerts without every member having to configure their own. Configure it per Room and per notification type under Space settings → Alerts & Notifications. The address is verified with a token from a test email, which means you need access to that mailbox; the integration requires a paid plan and Space admin access.

Alert evaluation — testing an alert definition against historical data before deploying it, introduced in v2.10.0 — now evaluates several definitions in a single request, so tuning a set of related alerts no longer means one round-trip each. Netdata's AND/OR/NOT expressions are translated correctly, and time-range errors are clearer. The Alerts view also gained status filters and configurable alert cards.

Finally, the Netdata mobile app is now available to everyone on free plans. It was previously restricted to paid plans.

Netdata Cloud: Fleet Management and Scale

Cloud's node views were rebuilt around getting to the right node quickly in very large infrastructures.

Fleet Management view showing 23,495 servers as status hexagons, grouped by fleet state with per-location cards

  • Nodes Fleet Management views, with nodes rendered as group hexagons that reflect status and link straight through to the relevant alerts or the Nodes tab.
  • Rebuilt home page, including a geographic map card for distributed fleets.
  • Group nodes by alert status, in addition to the existing grouping dimensions.
  • SNMP overview tab, surfacing SNMP devices — including per-interface breakdowns on a single node.
  • Metric taxonomy kept in step with the Agent, so this release's new collectors are browsable in Cloud from day one — AWS CloudWatch, MSSQL transaction logs, and vSphere clusters, datastores and resource pools all have taxonomy entries.

On scale: the Nodes and Alerts views were rebuilt to stay fluent in rooms of 50,000 nodes, backed by a large query-path overhaul in the charts service — batched per-node and per-context routing, covering indexes, memory-optimized aggregation, and top-N pushdown with distributed top-k across agents. Large spaces load and interact noticeably faster.

On administration: Space settings gained a dedicated User invitations tab, SCIM provisioning was corrected to match the RFC (case-insensitive matching for filters and sorting, filtered emails[...].value PATCH paths, and inactive accounts included in user searches), and the Dynamic Configuration UI now masks passwords when configuring collectors from Cloud.

Netdata Cloud: Dashboards and Charts

New Dashboards — Beta. The custom dashboards experience was rebuilt: a substantially larger card library, ready-made layouts to start from, and the tools to manage dashboards at scale.

AdditionWhat It Does
State Timeline cardRenders any single context as a state timeline — good for up/down, health and status over time
Heatmap chartsHeatmap rendering with caching, pairing with the Agent's new histogram range heatmaps
Alert count cardA stat card sourced from live alert counts
Node breakdown and inventory cardsFleet composition and inventory as first-class dashboard components
Fleet map cardGeographic distribution, with configurable controls and persisted gauge thresholds
PlaylistsRotate through a set of dashboards automatically — for wall displays and NOC screens
Import and exportMove dashboards between spaces, or keep them in version control
Templates and layoutsStart from a ready-made layout — Infrastructure, Containers, Network, Applications, Operations or Nodes — instead of an empty canvas

Two new aggregations landed on the query side: percentage-of-group aggregation, which shows each dimension as its share of the group rather than an absolute value, and latest time aggregation, the Cloud counterpart to the Agent's new latest time-grouping. Live-tail accuracy was fixed so live charts no longer trim to empty or show partial tail windows, and the threshold at which dimensions collapse into OTHERS was raised tenfold, so high-cardinality charts keep showing real dimension names for much longer.

Security Hardening
ChangeImpact
Dedicated MCP ACL with bearer-token enforcementMCP endpoints are no longer reachable without a token
Anonymous MCP metadata limitedSigned-out callers see a reduced metadata surface
WebSocket decompression-bomb guard (CWE-409)A malicious compressed frame can no longer exhaust memory
ndsudo privilege-escalation checkEscalation attempts through ndsudo are rejected
NETDATA_HOST_PREFIX validationValidated in local_listeners, and format-checked to prevent % injection into printf strings
/api/v3/settings behind HTTP_ACL_DASHBOARDConnection allowlists are enforced; anonymous state-changing requests blocked
gRPC upgraded for CVE-2026-33186Removes a known vulnerability from the Go dependency tree
kickstart.sh argument sanitizationInstaller arguments that could reach a privileged context are validated or sanitized; claiming config is no longer written from user input as root, and uses a secure temporary file
CodeQL security-extended suiteBroader static-analysis coverage in CI
End of Full Support for RHEL 7.x, CentOS 7.x, and Amazon Linux 2

The 2.11.x release series will be the last stable releases of the Netdata Agent to provide pre-built RPM packages for Red Hat Enterprise Linux 7.x, CentOS 7.x, Amazon Linux 2, and other compatible platforms. Officially, per our usual platform support policy, our support for RHEL 7.x and CentOS 7.x should have ended more than two years ago when Red Hat ended the Maintenance Support 2 support phase for the platform. We chose at that time to continue our support to a limited extent as we had a very large number of customers still using the platform and didn’t have a good way for them to switch to a different installation type. Today, however, we have decent support for switching installation types, we have a much lower percentage of users using the platform, and we’re starting to run into technical issues due to continuing support for a platform that is now more than a decade old. Given these factors and the upstream end of life for Amazon Linux 2 (which is causing similar technical issues) earlier this year, we have decided to finally end full support for these platforms.

Native RPM packages for these platforms will not be published for stable releases starting with version 2.12.0 of the Netdata Agent, and are expected to stop being published in nightly builds at some point within the next few weeks after the release of version 2.11.0. Existing packages will continue to be available for the foreseeable future, but will eventually be removed as well. New installs will automatically switch to using static builds at the same time that we stop publishing native RPM packages for nightly builds. Users with existing installs are encouraged to switch those systems to static builds as well, instructions on how to do so can be found here.

Local builds for these platforms should continue to work at least until the release of version 2.12.0, but after that point we will start removing the platform-specific support code for builds on RHEL 7.x, CentOS 7.x, and Amazon Linux 2.

This does not affect users using newer versions of Red Hat Enterprise Linux, CentOS, Amazon Linux, or equivalent platforms.

This does not affect users who are using our static builds or Docker images on these platforms. Those will continue to work as-is without any need for user intervention, and our static builds are the recommended approach for using Netdata on these platforms going forwards.

Acknowledgments

We would like to thank our dedicated, talented contributors that make up this amazing community. The time and expertise that you volunteer is essential to our success.

  • @jmestwa-coder for bounding v2 journal header offsets before CRC reads and avoiding an out-of-bounds read in url_percent_escape_decode.
  • @DavidMarec for adding FreeBSD system-call monitoring and fixing a FreeBSD plugin counter size mismatch.
  • @Kelpy2004 for fixing the health sum lookup on incremental dimensions and /proc/interrupts name parsing.
  • @artem for fixing a stack buffer overflow in the macOS mach_smi collector.
  • @arch-yunus for adding nfds monitoring and fixing a file descriptor bug in apps.plugin.
  • @AJCxZ0 for silencing curl output in netdata-updater.sh.
  • @lavr for allowing pg_ls_dir execute privilege for replication slot files in the PostgreSQL collector.
  • @wangtsingx for fixing a misleading error message on PEM decode failure in x509check.
  • @Hashim1999164 for fixing the python.d DEB dependency typo and a duplicate ndsudo entry.
  • @Func86 for reducing the Docker image size by setting permission bits in the builder stage.
  • @kkzhsh for fixing a NetdataYAML.cmake compilation error.
  • @Zhao73 for fixing documentation typos.
  • OrbisAI Security for reporting and upgrading gRPC to address CVE-2026-33186.
Contributions
Collectors
  • Added NetFlow/IPFIX/sFlow flow-analysis plugin with a four-tier local journal, Cisco ASA NSEL accounting, enrichment from cloud-provider IP ranges and BGP/RIPE RIS data, faceted search, and accounting charts (netflow.plugin). (#22111, #23201, #22665, #22703, #22719, #22925, #23241, #23246, #23313, @ktsaou)
  • Added AWS CloudWatch collector with 47 service profiles, automatic resource discovery, multi-account support, resource tag filtering and labels, explicit targets and collection rules, exact metric selection, per-rule and per-selection query policies, declarative opt-in metrics, stock health alerts, and collector activity metrics (go.d/cloudwatch). (#22874, #22944, #22948, #22949, #23004, #23011, #23024, #23028, #23031, #23086, #23090, #23092, #23094, #23110, #23113, #23119, #23122, #23123, #23125, #23338, #23340, @ilyam8)
  • Added SNMP L2/L3 topology engine and collector, discovering Layer 2 adjacencies from LLDP/FDB/STP, L3 subnet segments and adjacency, and OSPF and BGP adjacency links, with MAC OUI vendor lookup and reverse DNS enrichment (go.d/snmp_topology). (#22109, #22215, #22780, #22783, #22790, #22822, #22988, @ktsaou, @ilyam8)
  • Added SNMP trap listener with profile-based trap ingestion, 800+ vendor trap profiles, enrichment, SIEM forwarding, and dynamic configuration support (go.d/snmp_traps). (#22652, #22693, #22702, #22849, @ktsaou)
  • Added network-viewer and streaming topology functions with a documented topology v1 payload contract, container grouping in connection topology, and a role field on presentation actor types (topology). (#22110, #22217, #22496, #22601, @ktsaou)
  • Added Cato Networks collector for SASE site, tunnel and traffic monitoring (go.d/cato_networks). (#22373, @ktsaou)
  • Added Palo Alto PAN-OS collector (go.d/panos). (#22389, @ktsaou)
  • Added Linux audit subsystem monitoring (debugfs.plugin). (#22077, @ktsaou)
  • Added Active Directory monitoring to the Windows plugin (windows.plugin). (#22093, @thiagoftsm)
  • Added SMB protocol monitoring to the Windows plugin (windows.plugin). (#22236, @thiagoftsm)
  • Added system-call monitoring on FreeBSD (freebsd.plugin). (#23082, @DavidMarec)
  • Added macOS hardware sensor and log collectors: GPU power, clock and temperatures, SMC and IOHID sensors and fans, power sources and thermal pressure, NVMe SMART, and macOS logs, with a powermetrics fallback via ndsudo, a cross-OS sensors function, a shared temperature-histogram context, and bounded-cardinality process grouping for apps.plugin on hosts with high process churn (macos.plugin, macos-logs.plugin). Backported to 2.10.4. (#22475, #23085, @ktsaou)
  • Added the basis of a new eBPF Go plugin (ebpfgo.plugin). (#22469, @thiagoftsm)
  • Overhauled the Prometheus collector: migrated to the V2 framework, added a relabeling engine with job-level and profile-owned metric relabeling, a promprofiles catalog, profile-driven app detection for chart contexts, profile-owned fallback types, a profile autogen selector, and a unified single-pass stream parser; also fixed summaries without quantiles and preserved typed counters whose names end in info (go.d/prometheus). (#22640, #22651, #22660, #22664, #22668, #22682, #22694, #23016, #23255, #23405, #23410, #23419, #23441, @ilyam8)
  • Expanded SNMP device coverage from 229 to 273 profiles, with improvements for Cisco Catalyst, Cisco Nexus, Fortinet FortiGate and Check Point, plus BGP session monitoring and network-device licence monitoring (go.d/snmp). (#22122, #22170, #22190, #22191, #22192, #22193, @ktsaou)
  • Added SNMPv3 context name support, structured mapping config with bitmask value mappings, a ping_only option, and reusable profile engine helpers (go.d/snmp). (#22175, #22177, #22180, #22181, #22200, @ktsaou, @ilyam8)
  • Added OpenTelemetry logs ingestion and query subsystem, enabled the otel-plugin on macOS, and added optional OTLP log export of accepted agent events (otel.plugin). (#22720, #23172, #23184, #23185, #23250, @vkalintiris)
  • Extended the Network Viewer to Windows connections with SMB data and to FreeBSD, stabilized topology around self-listeners, and integrated the eBPF plugin with the network viewer (network-viewer, ebpf.plugin). (#22253, #22470, #22585, #22608, #22632, #22715, @thiagoftsm, @ktsaou)
  • Added HTTP service discovery and single-instance collector framework support, with operational tests for HTTP, DynCfg and secretstore configurations (go.d). (#22256, #22752, #23317, #23324, #23327, #23333, #23349, @ilyam8)
  • Added histogram buckets charted as range heatmaps, chart series aggregation, mutable chart labels, optional instance labels, and context_namespace support for autogenerated chart contexts; fixed unlabeled contributor intersection and algorithm resolution from runtime metric kinds (go.d/chartengine). (#22642, #23013, #23087, #23341, #23414, #23424, #23425, @ilyam8, @ktsaou)
  • Migrated the vSphere collector to framework v2 and expanded its coverage, with topology overlay helpers (go.d/vsphere). (#22458, #22625, @ktsaou, @ilyam8)
  • Added SQL Agent job execution metrics and transaction log monitoring to the MSSQL collector (go.d/mssql). (#22319, #22730, @ilyam8)
  • Added an opt-in urienc modifier to percent-encode secret references, for credentials containing URI-reserved characters (go.d). (#22750, @ilyam8)
  • Added vnode-scoped metrics for Azure Monitor workloads (go.d/azure_monitor). (#22402, @ilyam8)
  • Reworked and improved perflib on Windows, and added missing memory metrics and path translation (windows.plugin). (#22164, #22194, #23189, #23205, @thiagoftsm)
  • Replaced the cgroups-eBPF shared-memory transport with netipc IPC, and added a cgroup-name Go helper (cgroups.plugin, ebpf.plugin). (#22221, #22685, @ktsaou)
  • Made stock alert overrides removable without a restart (dyncfg/health). (#22511, @stelfrag)
  • Added a FUNCTION_DEL protocol command to plugins.d (plugins.d). (#21685, @ktsaou)
  • Added offline Function test modes for NetFlow and systemd journal (netflow.plugin, systemd-journal.plugin). (#22638, @ktsaou)
  • Reconciled instance function availability and withdrew unavailable shared functions, so the dashboard no longer offers functions that cannot run (go.d). (#22870, #22871, #22853, @ilyam8)
  • Ran collector dyncfg commands concurrently on per-key lanes and moved jobmgr dyncfg domains onto a claimed executor, reducing configuration latency (go.d). (#22964, #22966, #23299, @ilyam8)
  • Added a collector taxonomy framework POC for integrations (integrations). (#22489, @ilyam8)
  • Fixed SNMP topology reverse DNS warming, index-derived endpoints, a snapshot data race on device labels, capability actor type preservation, and Netgear switch topology discovery via LLDP/FDB/STP (go.d/snmp_topology). (#22366, #22371, #22714, #22826, #22839, @ktsaou, @ilyam8)
  • Fixed streaming topology graph output and namespace-relative cgroup lookup for topology (topology). (#22432, #22727, @ktsaou)
  • Fixed jobmgr lifecycle and configuration reconciliation, pre-activation retirement races, and stopping-rejection classification (go.d/jobmgr). (#23238, #23279, #23362, @ilyam8)
  • Fixed the CloudWatch collector not being configurable from the dyncfg UI (go.d/cloudwatch). (#23314, @ilyam8)
  • Fixed a scale factor for arubaWiredTempSensorTemperature and tolerance for invalid optional typed row values (go.d/snmp). (#22846, #23376, @ilyam8)
  • Fixed macOS disk.util percent scaling, and stopped logging an error every cycle for macOS zombies (macos.plugin, apps.plugin). (#23292, #23297, @ktsaou, @vkalintiris)
  • Fixed macOS NVMe SMART collection: corrected the IOKit object lifecycle so interfaces and registry entries are released on every path, bounded discovery before device reads, admitted replacement devices after pruning, kept samples collected before a teardown error, and allowed the IOService:/ device paths that smartctl needs through ndsudo via a purpose-built validator. ndsudo is now also installed with native macOS builds, so SMART collection works when the Go and script plugins are disabled (macos.plugin, daemon). (#23445, #23447, @ktsaou)
  • Fixed the Cato collector to support Cato SDK v0.3.2 snapshots (go.d/cato_networks). (#23280, @ilyam8)
  • Fixed function timeouts over the limit being clamped instead of returning 400 (go.d/functions). (#23258, @ilyam8)
  • Prevented plugins from overwriting reserved dyncfg functions (plugins.d). (#23224, @stelfrag)
  • Unified histogram bucket dimension names (go.d). (#23021, @ilyam8)
  • Fixed a DBEngine retention accounting underflow (netflow.plugin). (#22914, @ktsaou)
  • Fixed the debugfs audit capability service limit (debugfs.plugin). (#22831, @ktsaou)
  • Fixed static journal facet filters and journal file access error handling and validation (systemd-journal.plugin). (#22310, #22456, @stelfrag, @ktsaou)
  • Excluded ND_REMAPPING bookkeeping entries from the journal index, avoided ValueGuardInUse when indexing otel journals, and required absolute journal directory paths (journal-index, journal-log-writer). (#22513, #22621, #22771, @vkalintiris)
  • Fixed eBPF and cgroup shutdown handling, exit paths and log messages (ebpf.plugin, cgroups.plugin). (#22242, #22414, #22574, @stelfrag, @Copilot, @thiagoftsm)
  • Fixed a cgroup-name timeout environment race (cgroups.plugin). (#23156, @ktsaou)
  • Fixed ZFS bugs in the diskspace plugin (diskspace.plugin). (#22188, @thiagoftsm)
  • Fixed non-Hard-drive entries being parsed as physical disks (go.d/adaptecraid). (#22355, @ilyam8)
  • Fixed handling of SSIDs containing whitespace (go.d/ap). (#22472, @Copilot)
  • Fixed /proc/interrupts name parsing (proc.plugin). (#22556, @Kelpy2004)
  • Allowed pg_ls_dir execute privilege for replication slot files (go.d/postgres). (#22488, @lavr)
  • Fixed a misleading error message on PEM decode failure (go.d/x509check). (#23134, @wangtsingx)
  • Fixed agent compilation on FreeBSD (freebsd.plugin). (#23221, @stelfrag)
  • Fixed the systemd journal plugin OTEL log directory discovery, then removed OTel journal discovery from it entirely now that the otel-plugin owns that path (systemd-journal.plugin). (#22569, #22572, @stelfrag)
  • Also fixed in the 2.10.x patch line and included here: nvidia_smi temperature and power collection with driver 580 XML variants, macOS mach_smi stack buffer overflow, apps.plugin file-descriptor accounting with new nfds monitoring, FreeBSD counter size mismatches and freebsd_ipfw/claim off-by-one, freeipmi.plugin watchdog underflow at low uptime, systemd-journal.plugin memory retention after queries, eBPF PID accounting shared-memory pool leak and 100% CPU spin, eBPF FD PID map iteration, pluginsd cleanup race and slot bounds check, diskspace mountpoint initialization, SNMP snmpEngineTime uptime source and MaxOIDs default, dyncfg transient 503s and handoff behavior, powerstore extra_details, v2 journal header offset bounds before CRC reads, and the fail2ban socket path move into ndsudo. (#22179, #22183, #22201, #22203, #22207, #22231, #22232, #22291, #22298, #22436, #22447, #22490, #22553, #22598, #22710, #22666, #22745, #23044, #23047, #23089)
  • Rebuilt the go.d job manager around a single-owner command kernel, routing scheduling and commands through an executor seam, running blocking collector work as supervised effects, moving effect lifecycle routing into a transition kernel, and pulling vnode snapshots at runtime boundaries (go.d/jobmgr). (#22951, #22952, #22979, #22986, #22990, #23203, @ilyam8)
  • Restructured the SNMP topology collector: extracted the graph model, shape package, v1 renderer, enrichment and Function packages, typed internal details, migrated it to V2 single-instance, and removed dead code paths (go.d/snmp_topology). (#22611, #22614, #22753, #22762, #22773, #22774, #22775, #22776, #22791, #22792, #22795, #22796, #22797, #22798, #22801, #22802, #22803, #22804, #22806, #22807, #22827, #22829, #22832, #22833, #22834, @ilyam8)
  • Restructured the SNMP traps collector to establish ownership boundaries: internalized the receiver, profile catalog, profile metric runtime and output backends, isolated the job runtime and trap dedup/telemetry, and reduced the profile surface (go.d/snmp_traps). (#23358, #23359, #23360, #23361, #23363, #23375, #23377, #23387, #23389, @ilyam8)
  • Reworked the metrix metric layer with a transactional, bounded descriptor lifecycle, split source layout by ownership, and eliminated repeated multiscope flattening and structured flatten allocations (go.d/metrix). (#23053, #23075, #23242, #23278, @ilyam8, @ktsaou)
  • Reworked go.d function declaration and publication: split the declaration API, replaced AgentWide with MethodScope, funnelled reconcile through the job manager, and bound single-instance functions to the runtime job (go.d). (#22767, #22856, #22857, #22868, #22869, @ilyam8)
  • Shared reverse DNS and SNMP state across SNMP collectors, and shared ping probing between the ping and snmp collectors (go.d/snmp). (#22189, #22770, #23385, @ilyam8)
  • Skipped disabled lifecycle cap work in the chart engine (go.d/chartengine). (#23282, @ktsaou)
  • Migrated the MongoDB Go driver to v2 and Docker API usage to Moby modules (go.d/mongodb, go.d). (#22738, #22742, @ilyam8)
  • Updated the vendored systemd journal SDK through 0.7.8 and the vendored NetIPC library (netflow.plugin, libnetdata). (#22649, #22680, #22713, #22729, #22731, #22785, #22859, #22936, #22993, #23033, #23093, @ktsaou)
  • Added test coverage for the job manager, dyncfg, SNMP topology scenarios, Prometheus V1 compatibility manifests, Azure Monitor schema behavior, and chart series (go.d). (#22641, #22794, #22843, #22851, #22950, #22991, #23316, #23319, #23320, #23328, #23346, @ilyam8)
  • Removed the Go sensors collector, superseded since v2.2.0 by the C libsensors module in debugfs.plugin (go.d/sensors). (#23237, @ilyam8)
  • Rendered non-finite summary quantile values as a gap, and let template dimensions inherit float from the series metric meta (go.d/framework). (#22643, #22681, @ilyam8)
  • Adjusted vnodes dyncfg availability and split dyncfg job-name validation per domain (go.d). (#22247, #22415, @ilyam8, @stelfrag)
  • Improved service rules description position in service discovery (go.d/sd). (#22413, @ilyam8)
  • Added charttpl Group.Clone() and Spec.MarshalTemplate() (go.d). (#22882, @ilyam8)
  • Improved journal file handling and logging in DBENGINE (dbengine). (#22152, @stelfrag)
  • Prepared the groundwork for virtual logs functions (systemd-journal.plugin). (#22584, @ktsaou)
Packaging/Installation
  • Migrated RPM package builds from netdata.spec.in to CPack, and built RPM packages through CPack in CI (packaging). (#23194, #23290, @vkalintiris)
  • Enabled the netflow plugin and the scripts.d plugin in static and Docker builds (build). (#22852, #22892, @ilyam8)
  • Replaced the vendored SQLite amalgamation with build-time generation, and bumped SQLite to 3.53.3 (build). (#21779, #23218, @vkalintiris, @stelfrag)
  • Raised the minimum Go version to 1.26.2 and updated the Go toolchain (build). (#23204, #23210, #23215, @ilyam8, @thiagoftsm)
  • Moved the Rust workspace to edition 2024 with MSRV 1.91, centralized workspace dependencies, and kept line tables instead of full debuginfo (build). (#22907, #23148, @ilyam8, @vkalintiris)
  • Added an ENABLE_ND_MCP cmake flag to build nd-mcp independently of go.d.plugin (build). (#22313, @Copilot)
  • Stopped static packages shipping builder runtime state, repaired otel directory ownership, and removed obsolete otel-signal-viewer artifacts on static upgrade (packaging). (#23251, #23440, @vkalintiris)
  • Removed Fedora 42 and openSUSE Leap 15.6 from CI and package builds, and synced CI with the officially supported Alpine versions (CI). (#22211, #22239, #22240, @Ferroin)
  • Correctly handled EPEL on RHEL 7, and preserved PWD while ensuring the temporary directory is set in kickstart.sh (packaging). (#23222, #23235, @Ferroin)
  • Explicitly triggered systemctl daemon-reload from RPM packages when needed, since some RPM-based systems have no file watcher to queue it (packaging). (#23300, @Ferroin)
  • Hardened kickstart.sh: arguments that could be evaluated in a privileged context are now strictly validated where possible and sanitized otherwise, claiming configuration is no longer written from user input while running as root, and it is written through a secure temporary file (packaging). (#23223, @Ferroin)
  • Improved error handling in the kickstart script when fetching files, and bumped the repository config package version it uses (packaging). (#22420, #22424, @Ferroin)
  • Made warnings and fatal errors more visible during installation (packaging). (#23252, @Ferroin)
  • Added a Markdown copy of the Windows package EULA (packaging). (#22421, @Ferroin)
  • Updated the bundled static curl to 8.20.0 (build). (#22473, @Copilot)
  • Made libnetdata standalone-linkable (libnetdata). (#22528, @vkalintiris)
  • Fixed the IBM MQ FetchContent check breaking incremental rebuilds, and assorted CI fixes for IBM MQ library handling (build, CI). (#22223, #22491, @ktsaou, @Ferroin)
  • Fixed the python.d DEB dependency typo and a duplicate ndsudo entry (packaging). (#23289, @Hashim1999164)
  • Fixed a NetdataYAML.cmake compilation error (build). (#21295, @kkzhsh)
  • Reduced the Docker image size by setting permission bits in the builder stage (packaging). (#21902, @Func86)
  • Corrected the Rust macro name in the spec _have_rust gate (packaging). (#22515, @vkalintiris)
  • Excluded go.mod and go.sum from the packaging workflow, disabled compression for GHA artifact uploads of already-compressed files, forced the updater to actually update in CI checks, and added a go fix check plus a reworked SNMP fixture workflow covering topology engine tests and assorted CI updates and fixes (CI). (#22172, #22426, #22427, #22441, #22661, #22816, #23272, #23422, @ilyam8, @Ferroin, @stelfrag)
  • Enabled the CodeQL security-extended suite, aligned its ignore paths, and filtered build results out of CodeQL analysis (CI). (#22245, #23336, @ktsaou, @stelfrag)
  • Added the SQLite version to build details and the startup log message (daemon). (#22208, @stelfrag)
  • Also fixed in the 2.10.x patch line and included here: Windows MSI installer issues, claiming only when both token and rooms are provided, the search for Visual Studio tooling across multiple versions, netdata-updater fetch error handling, and silenced curl output. (#22422, #22639, #22723, #22751, #22754)
Documentation
  • Added a Network Performance Monitoring documentation section with an SNMP integration catalog, capability overviews, an Integrations submenu, screenshots, and corrected catalog brand icons and SNMP-trap tiles (docs/npm). (#22840, #22854, #22860, #22863, @ktsaou, @shyamvalsan)

View originalPermalink
How v2.11.0 went

v2.10.4

Added 3
  • macOS hardware sensor collectors for GPU power, clock, and temperatures, SMC and IOHID sensors and fans, power sources and thermal pressure, NVMe SMART, with a powermetrics fallback via ndsudo
  • Cross-OS sensors function and shared temperature-histogram context
  • Bounded-cardinality process grouping for apps.plugin on macOS to keep per-process monitoring meaningful on hosts with high process churn
Fixed 15
  • DBENGINE hardening against corrupted or truncated data files with extent disk-size and uncompressed-page-bounds validation, SIGBUS protection and header-offset bounds checks on v2 journal walks
  • Page-cache races in pgc_page_add and pgc_queue_del, deadlock on dimension creation, and crash on virtual node takeover
  • Data validation from SQLite to prevent crashes on corrupted databases and improved UUID handling in SQLite functions
  • Chart indexes remain allocated while a host is archived and guarded against null root index in chart lookups
  • rrdcontext metadata leak on non-dbengine hosts
  • Database engine now rotates active datafile and requeues pending extent on unrecoverable write errors so a failing disk no longer stalls the engine

From Netdata

Release notes

Netdata v2.10.4 is a patch release to address issues discovered since v2.10.3.

This is a larger-than-usual patch release focused on stability, memory safety, and crash resilience across the database engine, streaming, Cloud connectivity, and collectors. It also backports the new macOS hardware sensor collectors to the 2.10.x line.

New in this release
  • macOS hardware sensor collectors: GPU power, clock, and temperatures, SMC and IOHID sensors and fans, power sources and thermal pressure, NVMe SMART, with a powermetrics fallback via ndsudo, plus a cross-OS sensors function and a shared temperature-histogram context (#22475, #23085, @ktsaou)
  • Bounded-cardinality process grouping for apps.plugin on macOS, keeping per-process monitoring meaningful on hosts with high process churn (#23085, @ktsaou)
Database engine and metadata
  • Hardened DBENGINE against corrupted or truncated data files: extent disk-size and uncompressed-page-bounds validation, SIGBUS protection and header-offset bounds checks on v2 journal walks, improved journal file access error handling, and spinlock-holder identification for datafile/journal deadlock diagnostics (#22324, #22514, #22310, #22725, @stelfrag; #22666, @jmestwa-coder)
  • Fixed page-cache races in pgc_page_add and pgc_queue_del, a deadlock on dimension creation, and a crash on virtual node takeover (#22466, #22400, #22417, @stelfrag)
  • Validated data loaded from SQLite to prevent crashes on corrupted databases, and improved UUID handling and error reporting in SQLite functions (#22679, #22233, @stelfrag)
  • Kept chart indexes allocated while a host is archived and guarded against a null root index in chart lookups, avoiding null dereferences during queries (#23074, #23056, @stelfrag)
  • Fixed an rrdcontext metadata leak on non-dbengine hosts (#22438, @stelfrag)
  • Rotated the active datafile and requeued the pending extent on unrecoverable write errors, so a failing disk no longer stalls the database engine (#23048, @ktsaou)
Streaming and replication
  • Tracked and accounted memory allocation size in replication queries, and fixed a sender replication counter leak on obsolete charts (#22756, #22428, @stelfrag)
  • Rejected oversized ZSTD frames and handled decompression errors gracefully (log and fail the connection instead of fatal()) (#22830, @stelfrag)
  • Fixed the streaming receiver discarding already-delivered data when a child disconnects (#23118, @ktsaou)
Cloud connectivity (ACLK)
  • Prevented rare unbounded one-core CPU spins in the cloud-connection loops, and avoided returning an uninitialized packet_id on publish failure (#22879, @ktsaou; #22504, @stelfrag)
Alerts and machine learning
  • Bounded the alert notification execution wait so a stuck notification script cannot block progress, and fixed a shutdown race when restoring alert information from the database (#22626, #22448, @stelfrag)
  • Added safeguards against ML database corruption and streamlined the recovery process (#22478, @stelfrag)
Collectors
  • Fixed systemd-journal.plugin memory retention after queries and its apps.plugin accounting (#23089, @ktsaou)
  • Fixed file descriptor accounting in apps.plugin (adding nfds monitoring) and eBPF FD PID map iteration (#22447, @arch-yunus; #22436, @stelfrag)
  • Fixed buffer overflows and counter-size issues in platform collectors: a stack buffer overflow in macOS mach_smi, FreeBSD counter size mismatches and an off-by-one in freebsd_ipfw and the claim code, and a freeipmi.plugin watchdog underflow at low system uptime (#22553, @artem; #23044, @DavidMarec; #22710, @vkalintiris; #22490, @ktsaou)
  • Fixed Windows Events row rendering crashes with stricter size checks, safer XML parsing, and robust variant handling (#22872, @ktsaou)
  • Restored stable temperature and power collection in go.d/nvidia_smi with NVIDIA driver 580 XML output variants (#23047, @copilot-swe-agent)
  • Fixed Windows hardware detection so virtual machines are no longer reported as bare metal (#22942, @thiagoftsm)
  • Moved the fail2ban socket path into ndsudo for go.d/fail2ban (#22745, @ilyam8)
  • Fixed a pluginsd cleanup race with an active collector, added a slot bounds check to the pluginsd parser, and initialized mountpoint state before the slow worker in the diskspace plugin (#22207, #22598, #22298, @stelfrag)
Security and API
  • The /api/v3/settings endpoint is now gated behind HTTP_ACL_DASHBOARD instead of HTTP_ACL_NOCHECK, enforcing connection allowlists and blocking anonymous state-changing requests (#22896, @stelfrag)
  • Fixed the csvjsonarray output format emitting invalid JSON when label-quotes is passed (#23115, @ktsaou)
Core stability and memory safety
Installation and packaging
  • Windows: fixed MSI installer issues, adjusted the installer to run claiming only when both the token and rooms are provided (with safer wide-character handling and a capped HTTP claim response size), and fixed the search for Visual Studio tooling to check for multiple versions (#22751, #22754, @thiagoftsm; #22723, @Ferroin)
  • Improved error handling in netdata-updater when fetching files and silenced curl output (#22422, @Ferroin; #22639, @AJCxZ0)
Support options

As we grow, we stay committed to providing the best support ever seen from an open-source solution. Should you encounter an issue with any of the changes made in this release or any feature in the Netdata Agent, feel free to contact us through one of the following channels:

  • Netdata Learn: Find documentation, guides, and reference material for monitoring and troubleshooting your systems with Netdata.
  • GitHub Issues: Make use of the Netdata repository to report bugs or open a new feature request.
  • GitHub Discussions: Join the conversation around the Netdata development process and be a part of it.
  • Community Forums: Visit the Community Forums and contribute to the collaborative knowledge base.
  • Discord Server: Jump into the Netdata Discord and hang out with like-minded sysadmins, DevOps, SREs, and other troubleshooters. More than 2000 engineers are already using it!
View originalPermalink
How v2.10.4 went

v2.10.3

Changed 2
  • Switched the SNMP collector's primary uptime source to SNMP-FRAMEWORK-MIB::snmpEngineTime in seconds to avoid the ~497-day TimeTicks wrap of hrSystemUptime/sysUpTime while keeping the existing systemUptime metric name and HR-MIB fallback
  • Split dynamic configuration job-name validation per domain so service discovery, vnode, and secret-store names can include dots such as FQDNs, while collectors retain strict naming rules
Fixed 2
  • Fixed a per-PID shared-memory pool leak in ebpf.plugin that filled the 32,768-slot pool within ~15 hours and pegged a CPU core at 100% in an infinite map-iteration loop by using non-allocating lookup for aggregation paths, zeroing freshly-allocated slots, sweeping module bits on exit, and properly cleaning up stale shared-memory and semaphore objects on init
  • Removed the unused extra_details field from the go.d/powerstore Hardware struct to fix /hardware response decoding errors that caused job check failures

From Netdata

Release notes

Netdata v2.10.3 is a patch release to address issues discovered since v2.10.2.

This patch release provides the following bug fixes and updates:

  • Fixed a per-PID shared-memory pool leak in ebpf.plugin that, on hosts with normal process churn, filled the 32,768-slot pool within ~15 hours and then pegged a CPU core at 100% in an infinite map-iteration loop; aggregation paths now use a non-allocating lookup, freshly-allocated slots are zeroed, module bits are swept on exit, and stale shared-memory and semaphore objects are properly cleaned up on init (#22232, @ktsaou)
  • Switched the SNMP collector's primary uptime source to SNMP-FRAMEWORK-MIB::snmpEngineTime (in seconds) to avoid the ~497-day TimeTicks wrap of hrSystemUptime/sysUpTime, while keeping the existing systemUptime metric name and HR-MIB fallback (#22231, @ilyam8)
  • Split dynamic configuration job-name validation per domain so service discovery, vnode, and secret-store names can include dots (e.g., FQDNs), while collectors retain strict naming rules (#22247, @ilyam8)
  • Removed the unused extra_details field from the go.d/powerstore Hardware struct to fix /hardware response decoding errors that caused job check failures (#22291, @ilyam8)
Support options

As we grow, we stay committed to providing the best support ever seen from an open-source solution. Should you encounter an issue with any of the changes made in this release or any feature in the Netdata Agent, feel free to contact us through one of the following channels:

  • Netdata Learn: Find documentation, guides, and reference material for monitoring and troubleshooting your systems with Netdata.
  • GitHub Issues: Make use of the Netdata repository to report bugs or open a new feature request.
  • GitHub Discussions: Join the conversation around the Netdata development process and be a part of it.
  • Community Forums: Visit the Community Forums and contribute to the collaborative knowledge base.
  • Discord Server: Jump into the Netdata Discord and hang out with like-minded sysadmins, DevOps, SREs, and other troubleshooters. More than 2000 engineers are already using it!
View originalPermalink
How v2.10.3 went

v2.10.2

Changed 1
  • Reduced default SNMP MaxOIDs from 60 to 20 for more reliable polling
Fixed 2
  • Fixed ZFS-related crashes in diskspace.plugin by guarding a NULL filesystem field and replacing the blocking pool-capacity collector with a lightweight cache fed by existing statvfs calls, preventing coredumps on degraded or exporting ZFS pools
  • Eliminated false-positive "timed out waiting for enable/disable decision" warnings by removing the wait-decision timeout and making the dyncfg command path non-droppable
Removed 1
  • Removed 32-bit counter fallbacks from the IF-MIB profile to prevent counter type switching and overflow between collection cycles

From Netdata

Release notes

Netdata v2.10.2 is a patch release to address issues discovered since v2.10.1.

This patch release provides the following bug fixes and updates:

  • Fixed ZFS-related crashes in diskspace.plugin by guarding a NULL filesystem field and replacing the blocking pool-capacity collector with a lightweight cache fed by existing statvfs calls, preventing coredumps on degraded or exporting ZFS pools (#22188, @thiagoftsm)
  • Reduced default SNMP MaxOIDs from 60 to 20 for more reliable polling, and removed 32-bit counter fallbacks from the IF-MIB profile to prevent counter type switching and overflow between collection cycles (#22203, @ilyam8)
  • Eliminated false-positive "timed out waiting for enable/disable decision" warnings by removing the wait-decision timeout and making the dyncfg command path non-droppable, so back-pressure flows upstream instead of producing 503 errors (#22201, @ilyam8)
Support options

As we grow, we stay committed to providing the best support ever seen from an open-source solution. Should you encounter an issue with any of the changes made in this release or any feature in the Netdata Agent, feel free to contact us through one of the following channels:

  • Netdata Learn: Find documentation, guides, and reference material for monitoring and troubleshooting your systems with Netdata.
  • GitHub Issues: Make use of the Netdata repository to report bugs or open a new feature request.
  • GitHub Discussions: Join the conversation around the Netdata development process and be a part of it.
  • Community Forums: Visit the Community Forums and contribute to the collaborative knowledge base.
  • Discord Server: Jump into the Netdata Discord and hang out with like-minded sysadmins, DevOps, SREs, and other troubleshooters. More than 2000 engineers are already using it!
View originalPermalink
How v2.10.2 went

v2.10.1

Added 1
  • Add ping_only option to the SNMP collector, allowing collection of only ICMP round-trip time metrics while still using SNMP at startup for device identification and labeling
Fixed 2
  • Fix SNMP collector retries serialization so a zero value is correctly preserved, and disable bulk walk detection when MaxRepetitions is set to 0
  • Reduce transient 503 errors on dynamic configuration enable/disable by buffering the command channel and extending the handoff timeout in the go.d plugin

From Netdata

Release notes

Netdata v2.10.1 is a patch release to address issues discovered since v2.10.0.

This patch release provides the following bug fixes and updates:

  • Added ping_only option to the SNMP collector, allowing collection of only ICMP round-trip time metrics while still using SNMP at startup for device identification and labeling (#22180, @ilyam8)
  • Fixed SNMP collector retries serialization so a zero value is correctly preserved, and disabled bulk walk detection when MaxRepetitions is set to 0 (#22179, @ilyam8)
  • Reduced transient 503 errors on dynamic configuration enable/disable by buffering the command channel and extending the handoff timeout in the go.d plugin (#22183, @ilyam8)
Support options

As we grow, we stay committed to providing the best support ever seen from an open-source solution. Should you encounter an issue with any of the changes made in this release or any feature in the Netdata Agent, feel free to contact us through one of the following channels:

  • Netdata Learn: Find documentation, guides, and reference material for monitoring and troubleshooting your systems with Netdata.
  • GitHub Issues: Make use of the Netdata repository to report bugs or open a new feature request.
  • GitHub Discussions: Join the conversation around the Netdata development process and be a part of it.
  • Community Forums: Visit the Community Forums and contribute to the collaborative knowledge base.
  • Discord Server: Jump into the Netdata Discord and hang out with like-minded sysadmins, DevOps, SREs, and other troubleshooters. More than 2000 engineers are already using it!
View originalPermalink
How v2.10.1 went

v2.10.0

Added 10
  • Secrets management with 4 resolver types (environment variables, files, commands, and secretstore references) and 4 backends (AWS Secrets Manager, Azure Key Vault, GCP Secret Manager, HashiCorp Vault)
  • Nagios plugins collector for running Nagios-compatible checks with automatic performance data charts, threshold-based alerting with soft/hard state logic, and execution metrics
  • Azure Monitor support
  • Enhanced AI conversations with ability to resume and expand generated AI reports
  • Ask AI about active alerts or charts, their root causes, and recommended remediation steps
  • Complete support for recurrence rules for alerts silencing

From Netdata

Table of Contents
Release Summary

Netdata v2.10.0 introduces secrets management, a Nagios plugins collector, Azure Monitor support, and broad stability improvements.

FeatureHighlightsDetails
Secrets Management4 resolver types, 4 backends• Environment variables, files, commands, and secretstore references• AWS Secrets Manager, Azure Key Vault, GCP Secret Manager, HashiCorp Vault
Enhanced AI ConversationsReports and Alerts integration• Resume and expand generated AI reports • Ask AI about active alerts or charts, their root causes, and recommended remediation steps
Advanced AlertingSilencing, Evaluation & Acknowledge• Complete support for recurrence rules for Alerts Silencing • Test alert definitions against historical data before deploying • Acknowledge alerts you're working on
Improved Custom DashboardsComplete Customisation• Add Metrics, Logs, Events, Node List, Alerts, Live Functions and Text components • Duplicate Custom Dashboards
Nagios Plugins CollectorRun any Nagios-compatible check• Automatic performance data charts• Threshold-based alerting with soft/hard state logic• Execution metrics (duration, CPU, memory)
Azure Monitor Collector38 service profiles, 1,300+ metrics• Automatic resource discovery via Azure Resource Graph• Multi-subscription, flexible scoping by resource groups/regions/tags
Azure AD for Database CollectorsMSSQL, PostgreSQL, Generic SQL• Service principal, managed identity, and default credential chain• Passwordless auth for Azure-hosted databases
New Storage CollectorsDell PowerStore, Dell PowerVault• Hardware health, performance, capacity, and sensor monitoring• Built on V2 collector framework
Expanded Collector CoveragevSphere, MSSQL, SNMP, Docker• vSphere datastores/clusters/resource pools• MSSQL Always On AG monitoring• SNMP IPSec/VPN profiles for FortiGate, Juniper, MikroTik, Check Point• Docker container listing function
OpenTelemetry ImprovementsMetrics pipeline overhaul• Proper slot-based aggregation with configurable intervals• Multi-slot ingestion with out-of-order support
Stability & PerformanceProduction reliability• Faster agent startup• ML prediction optimization• Alerts API speedup• Multiple crash and race condition fixes
Release Highlights
Secrets Management

Keep collector credentials out of plain-text configuration files. Netdata now lets you reference secrets in collector configurations instead of storing them directly. Passwords, tokens, and API keys are resolved at runtime from the source you choose.

Resolvers

Netdata provides four resolver types, from simple to enterprise-grade:

ResolverSyntaxBest For
Environment variable${env:VAR_NAME}Secrets already injected into the Netdata service environment
File${file:/absolute/path}Secrets stored in local files on disk (Docker Secrets, Kubernetes Secrets mounted as volumes)
Command${cmd:/absolute/path args}Secrets returned by a trusted local command (e.g., 1Password CLI, custom scripts)
Secretstore${store:<kind>:<name>:<operand>}Secrets managed centrally in cloud providers or Vault
Supported Secretstore Backends

For organizations managing secrets centrally, Netdata integrates with four secretstore backends:

BackendKindOperand FormatAuthentication Modes
AWS Secrets Manageraws-smsecret-name[#key]Environment credentials, ECS task role, EC2 instance profile
Azure Key Vaultazure-kvvault-name/secret-nameService principal, Managed identity, Default credential chain
Google Secret Managergcp-smproject/secret[/version]Metadata server, Service account file
HashiCorp Vaultvaultpath#keyToken, Token file

Configure secretstores through the Netdata Dynamic Configuration UI (Collectors → go.d → SecretStores) or via configuration files under /etc/netdata/go.d/ss/.

Example

A MySQL collector referencing its password from HashiCorp Vault:

# /etc/netdata/go.d/mysql.conf
jobs:
  - name: mysql_prod
    dsn: "netdata:${store:vault:vault_prod:secret/data/netdata/mysql#password}@tcp(127.0.0.1:3306)/"

You can mix resolver types freely. For example, read the username from an environment variable and the password from a secretstore in the same config value:

jobs:
  - name: mysql_prod
    dsn: "${env:MYSQL_USER}:${store:vault:vault_prod:secret/data/netdata/mysql#password}@tcp(127.0.0.1:3306)/"
How It Works
  • Secrets are resolved each time a collector job starts or restarts.
  • Updating a secretstore automatically restarts collector jobs that use it so they pick up new credentials.
  • Secretstore configuration values themselves also support ${env:...}, ${file:...}, and ${cmd:...} resolvers, so backend credentials don't need to be stored in plain text either.

[!TIP] For detailed setup, backend-specific authentication guides, and troubleshooting, see the Secrets Management documentation.

Enhanced AI Conversations

Netdata AI Conversations are now more powerful. You can now add Active Alerts, Charts or AI-generated reports to start an AI conversation to provide better, more meaningful responses to your conversations.

Advanced Alerting

With this release, Netdata provides our users with the ability to:

  • Graphically Evaluate Alerts on historical data before they are deployed

  • Create complex recurrence rules on Alerts Silencing to manage your maintenance windows

  • Acknowledge Alerts that your teams are working on

Improved Custom Dashboards

With this release, we have improved our custom dashboards support to:

  • Add metrics charts, logs, live functions, events feeds, active alerts counts, node list, etc.
  • Duplicate Dashboards to help create new custom dashboards from existing ones
Nagios Plugins Collector

Run Nagios-compatible plugins and custom scripts directly inside Netdata with the new Nagios Plugins collector. Any executable that follows the Nagios plugin output format works out of the box: packaged Nagios plugins, custom shell scripts, or compiled binaries.

What You Get
FeatureDetails
Check State MonitoringTracks OK, WARNING, CRITICAL, and UNKNOWN states from the process exit code
Automatic Performance Data ChartsNagios performance data values are parsed and charted automatically
Threshold-Based AlertingWhen perfdata includes warning/critical thresholds, Netdata derives a threshold state chart (ok/warning/critical) and creates built-in alerts
Execution MetricsMeasures run duration, CPU time, and memory usage of each check
Nagios-Style SchedulingConfigurable check intervals, retry intervals, max check attempts, and time periods
Soft/Hard State LogicNon-OK results retry at retry_interval up to max_check_attempts before triggering alerts, preventing false alarms from transient failures
Built-in Alerts

The collector ships with four stock alerts:

AlertTriggers When
nagios_job_execution_state_warnA check is in WARNING hard state
nagios_job_execution_state_critA check is in CRITICAL hard state
nagios_job_perfdata_threshold_state_warnA perfdata metric exceeds its warning threshold
nagios_job_perfdata_threshold_state_critA perfdata metric exceeds its critical threshold

All stock alerts suppress soft retry states. Alerts fire only after the configured max_check_attempts consecutive failures.

Example
# /usr/local/lib/netdata/checks/check_api.sh
#!/bin/sh
resp=$(curl -s -o /dev/null -w "%{http_code} %{time_total}" --max-time 5 "$1" 2>/dev/null)
code=$(echo "$resp" | cut -d' ' -f1)
time=$(echo "$resp" | cut -d' ' -f2)

[ "$code" -ge 500 ] && echo "CRITICAL - HTTP $code | response_time=${time}s;~:2;~:5" && exit 2
[ "$code" -ne 200 ] && echo "WARNING - HTTP $code | response_time=${time}s;~:2;~:5" && exit 1
echo "OK - HTTP $code | response_time=${time}s;~:2;~:5"
exit 0
# /etc/netdata/scripts.d/nagios.conf
jobs:
  - name: api_health
    plugin: /bin/sh
    args: ["/usr/local/lib/netdata/checks/check_api.sh", "http://localhost:8080/health"]
    check_interval: 1m
    retry_interval: 15s
    max_check_attempts: 3

This produces check state charts, execution metrics, and because the script emits perfdata with thresholds (response_time with warn >2s / crit >5s), an automatic response time chart with a threshold state chart and built-in alerts.

[!TIP] For the full setup guide, custom scripts walkthrough, and time period configuration, see the Nagios Plugins collector documentation.

Azure Monitor Collector

Monitor your entire Azure infrastructure from Netdata with the new Azure Monitor collector. It collects platform metrics from the Azure Monitor Metrics API with automatic resource discovery: point it at your subscriptions, and it finds your resources, matches them to built-in profiles (preconfigured sets of metrics), and starts collecting metrics.

[!IMPORTANT] This is an alpha feature. The configuration format may change in upcoming releases.

What You Get
FeatureDetails
38 Built-in Service Profiles1,300+ metrics across compute, databases, networking, storage, AI/ML, analytics, containers, and more
Automatic DiscoveryDiscovers resources via Azure Resource Graph and enables matching profiles without per-resource configuration
Multi-SubscriptionMonitor resources across multiple Azure subscriptions in a single collector job
Flexible Discovery ScopingNarrow discovery by resource groups, regions, or tags, or use a custom Azure Resource Graph KQL query
Built-in AlertsPre-configured alerts for critical Azure resource conditions
UI ConfigurationConfigure through the Netdata Dynamic Configuration UI without file editing
Supported Services
CategoryServices via Profiles
ComputeVirtual Machine, Virtual Machine Scale Set, App Service (incl. Functions), Container Apps, Container Instances
KubernetesAKS
DatabasesSQL Database, SQL Elastic Pool, SQL Managed Instance, MySQL Flexible, PostgreSQL Flexible, Cosmos DB, Cache for Redis
NetworkingLoad Balancer, Application Gateway, Front Door, Firewall, NAT Gateway, VPN Gateway, ExpressRoute Circuit, ExpressRoute Gateway
Storage & AnalyticsStorage Account, Data Explorer, Data Factory, Stream Analytics, Synapse Analytics, Log Analytics
MessagingEvent Hubs, Service Bus, Event Grid, IoT Hub
AI & MLCognitive Services, Machine Learning, Application Insights
IntegrationAPI Management, Logic Apps, Key Vault, Container Registry

[!TIP] For detailed setup instructions, authentication options, custom profiles, and tuning, see the Azure Monitor collector setup guide.

Acknowledgments
  • @kiwixz for increasing statsd UDP buffer size to match localhost MTU, preventing packet truncation.
  • @DarkByteZero for fixing MongoDB top-queries timeout, netdev_mutex deadlock, and apps.plugin use-after-free crash.
Contributions
Collectors
  • Overhauled OpenTelemetry metrics pipeline with proper slot-based aggregation, multi-slot ingestion with out-of-order support, configurable intervals, and optimized JSON handling (otel.plugin). (#21771, #21893, #21896, #22085, @vkalintiris)
  • Added secrets management support, allowing collector credentials to be stored in external secret stores like AWS Secrets Manager, Azure Key Vault, GCP Secret Manager, and HashiCorp Vault (go.d). (#21951, #22081, #22083, @ilyam8)
  • Rewrote scripts.d collector to V2 framework with alertable state support for Nagios-style custom scripts (go.d/scripts.d). (#21908, #22008, @ilyam8)
  • Added Azure Monitor collector for monitoring Azure cloud resources with built-in alerts (go.d/azure_monitor). (#21993, #22007, #22095, @ilyam8)
  • Added Azure AD authentication support for MSSQL, PostgreSQL, and generic SQL collectors (go.d). (#21905, #21995, @ktsaou, @ilyam8)
  • Added Dell PowerVault ME4/ME5 storage array collector with system health, hardware status, performance, and capacity monitoring (go.d/powervault). (#21936, @ktsaou)
  • Added Dell PowerStore storage array collector with cluster, appliance, volume, node, and hardware monitoring (go.d/powerstore). (#21929, @ktsaou)
  • Added Always On Availability Group monitoring to MSSQL collector with AG health, replica states, sync metrics, and WSFC cluster status (go.d/mssql). (#21927, @ktsaou)
  • Added IPSec/VPN monitoring profiles for FortiGate, Juniper, MikroTik, and Check Point (go.d/snmp). (#21926, @ktsaou)
  • Added per-Windows-service process tree grouping to apps.plugin with auto-discovery via Service Control Manager (apps.plugin). (#21925, @ktsaou)
  • Added datastore, cluster, and resource pool monitoring to vSphere collector with capacity, IOPS, and latency metrics (go.d/vsphere). (#21924, @ktsaou)
  • Added Docker container listi1ng function for interactive container overview in the dashboard (go.d/docker). (#21868, @ilyam8)
  • Increased statsd UDP buffer size to match localhost MTU, preventing packet truncation (statsd). (#21822, @kiwixz)
  • Improved Windows plugin with CPU temperature chart visibility and Control Panel enhancements (windows.plugin). (#21797, @thiagoftsm)
  • Fixed dynamic configuration rollback for non-disruptive Service Discovery update failures (go.d). (#21861, @ilyam8)
  • Fixed Service Discovery pipeline name derivation from source context (go.d). (#22105, @ilyam8)
  • Fixed collector static functions not registering on first job start (go.d). (#21894, @ilyam8)
  • Fixed MongoDB top-queries function timeout due to incorrect default timeout (go.d/mongodb). (#21860, @DarkByteZero)
  • Fixed smartctl collector treating non-fatal exit codes as errors when disk health data is still valid (go.d/smartctl). (#21858, @ilyam8)
  • Fixed netdev_mutex deadlock when /proc/net/dev reading fails, permanently blocking network interfaces (proc.plugin). (#21839, @DarkByteZero)
  • Fixed apps.plugin use-after-free crash on parent pointer dereference during process collection (apps.plugin). (#21838, @DarkByteZero)
  • Fixed service discovery to skip unsupported discoverer configurations instead of failing (go.d). (#21818, @ilyam8)
  • Fixed DCGM exporter service discovery (go.d). (#21800, @ilyam8)
  • Refactored Azure Monitor discovery and profile selection with proper configuration modes (go.d/azure_monitor). (#22088, #22095, @ilyam8)
  • Improved Windows plugin codebase with internal improvements and fixes (windows.plugin). (#22039, @thiagoftsm)
  • Introduced V2 metrics collection framework with float dimension support, structured family handling, and redesigned function management (go.d). (#21769, #21825, #21850, #21851, #21859, #21909, #21979, @ilyam8)
  • Changed "no instances configured" function response code to 422 (go.d). (#21903, @ilyam8)
  • Removed legacy MSSQL collector from Windows plugin, completing migration to go.d (windows.plugin). (#21876, #21886, @thiagoftsm)
  • Restructured go.d agent internals with improved dyncfg lifecycle, logger attribution, and codebase organization (go.d). (#21803, #21808, #21817, #21821, #21830, #21833, #21840, #21842, @ilyam8)
  • Added conditional and rate-limited logging to go.d logger (go.d). (#21813, @ilyam8)
Packaging/Installation
  • Optimized release build profile for runtime speed with fat LTO and opt-level=3 (build). (#22046, @vkalintiris)
  • Fixed Windows config editor default in installer (packaging). (#21957, @stelfrag)
  • Fixed systemd-journal.plugin permissions on offline installs (packaging). (#21953, @ralphm)
  • Added Fedora 44 to CI and package builds (CI). (#21943, @Ferroin)
  • Fixed control files for DEB packages causing uninstallation issues (packaging). (#21940, @Ferroin)
  • Added Ubuntu 26.04 to CI and package builds (CI). (#21939, @Ferroin)
  • Fixed Windows installer driver installation across different Windows versions (packaging). (#21911, @thiagoftsm)
  • Updated Go toolchain to v1.26.0 (build). (#21866, @ilyam8)
  • Dropped Ubuntu 20.04 from CI and package builds after upstream EOL (CI). (#21647, @Ferroin)
  • Increased minimum language standards to C17 and C++17 with updated Protobuf and Abseil (build). (#21574, @Ferroin)
  • Updated libbpf and synced eBPF repositories (build). (#21982, @thiagoftsm)
Documentation
Other Notable Changes
  • Optimized ML prediction performance with circular buffer, improved model loading, and safer cluster center handling (ML). (#21795, #22042, #22073, #22104, @stelfrag)
  • Sped up alerts API filtering with host status snapshots for efficient prefiltering (health). (#21984, @stelfrag)
  • Improved streaming stability by reducing shutdown timing issues (streaming). (#21992, @thiagoftsm)
  • Improved netdatacli ping to properly report agent readiness with distinct exit codes (0=ready, 1=initializing, 255=unreachable) (daemon). (#21965, @stelfrag)
  • Added periodic timezone refresh for correct DST handling across all components (daemon). (#21944, @stelfrag)
  • Improved logger by removing spinlock during I/O to prevent deadlocks (daemon). (#21928, @stelfrag)
  • Improved agent startup and restart time with optimized mmap settings and prefetching (dbengine). (#21891, @stelfrag)
  • Added rejection of incoming streaming connections for locally collected vnodes (streaming). (#21889, @stelfrag)
  • Enforced Pushover API field length limits to prevent notification failures on long URLs (health/notifications). (#21882, @Copilot)
  • Added SOCKS5 and SOCKS5H proxy support to ACLK for environments requiring SOCKS proxies (ACLK). (#21831, @stelfrag)
  • Added environment variable expansion in host labels with ${VAR} and ${VAR:-default} syntax (daemon). (#21796, @ktsaou)
  • Added SNI support for streaming SSL/TLS connections, enabling reverse proxy compatibility (streaming). (#21715, @Copilot)
  • Preserved UTF-8 characters in RRD string fields instead of replacing them with underscores (daemon). (#21694, @ktsaou)
  • Added YAML support to libnetdata and migrated log2journal to shared YAML parser (libnetdata). (#20544, @ktsaou)
  • Fixed eBPF PID cleanup to avoid use-after-free during iteration (ebpf.plugin). (#22098, @stelfrag)
  • Fixed undefined behavior in timezone TZif 64-bit parsing (daemon). (#22097, @stelfrag)
  • Fixed ACLK SIGABRT during shutdown that could lead to database corruption (ACLK). (#22051, @thiagoftsm)
  • Fixed cache SIGSEGV in aral during journal v2 processing (dbengine). (#22026, @thiagoftsm)
  • Fixed SIGSEGV in context dimension entry lookup (dbengine). (#22025, @thiagoftsm)
  • Fixed clock handling in eBPF and other plugins to prevent crashes (ebpf.plugin). (#22009, @thiagoftsm)
  • Fixed lock order during context processing to avoid deadlock (dbengine). (#21996, @stelfrag)
  • Fixed uninitialized vnode stale timeout leaking between host definitions (streaming). (#21983, @stelfrag)
  • Fixed thread safety and initialization handling in Windows hardware info collection (windows.plugin). (#21885, #21958, @stelfrag)
  • Fixed health API to properly acquire and release RRD instances (health). (#21952, @stelfrag)
  • Fixed data race in ML training during host stop (ML). (#21844, @stelfrag)
  • Fixed context hub cleanup with queue bounding and stale entry handling (dbengine). (#21832, @stelfrag)
  • Fixed potential use-after-free in RAM mode metric release (dbengine). (#21809, @stelfrag)
  • Fixed URL validation in cloud config that incorrectly rejected valid URLs (ACLK). (#21805, @stelfrag)
  • Fixed crash when processing corrupted journal files (dbengine). (#21794, @stelfrag)
  • Fixed cache page queue race condition with revalidation under lock (dbengine). (#21793, @stelfrag)
  • Fixed FreeBSD 15.0 build failure due to ipfw structure changes (daemon). (#21843, @Copilot)
  • Fixed race condition during pluginsd dimension array operations (plugins.d). (#21628, @stelfrag)
  • Removed compilation warnings across Linux, Windows, BSD platforms and removed variable-length arrays for C11+ compatibility (build). (#21827, #21961, #22024, #22066, #22091, #22103, @thiagoftsm, @stelfrag)
  • Improved DBENGINE log messages with structured identifiers, tier prefixes, and correct file names (dbengine). (#22045, #22047, #22053, #22086, @stelfrag)
  • Optimized eBPF memory handling with inline Judy storage and reduced fragmentation (ebpf.plugin). (#22050, @stelfrag)
  • Added ML unit tests for anomaly detection features (ML). (#22043, @stelfrag)
  • Added dictionary benchmark tool (libnetdata). (#22041, @stelfrag)
  • Created journal directory if missing before inotify watch setup (otel.plugin). (#22019, @vkalintiris)
  • Fixed various Coverity issues across timezone, ML, and health modules (various). (#21777, #21878, #21971, #21986, #22054, #22094, @thiagoftsm, @stelfrag)
  • Changed missing user configuration directory log to info level (daemon). (#21969, @vkalintiris)
  • Refactored stream connection definition parsing for reusability (streaming). (#21921, @stelfrag)
  • Optimized UID/GID cache updates in apps plugin to once per cycle (apps.plugin). (#21864, @stelfrag)
  • Added input validation for socket connection definitions (libnetdata). (#21881, @stelfrag)
  • Reduced log noise for journal indexing limit warnings on online files (journal-viewer). (#21816, @vkalintiris)
  • Improved ACLK proxy logging with protocol details and credential status (ACLK). (#21789, @stelfrag)
  • Fixed potential crash in expression evaluation print function (libnetdata). (#21790, @stelfrag)
  • Improved datafile deletion with retry mechanism and synchronization (dbengine). (#21781, @stelfrag)
  • Adjusted eBPF user ring buffer handling (ebpf.plugin). (#21676, @thiagoftsm)
  • Improved metadata storage with prepared statements for chart and host labels (dbengine). (#21492, @stelfrag)
  • Added float dimension support to plugins.d protocol and go.d framework (plugins.d, go.d). (#21349, #21825, @ktsaou, @ilyam8)
Deprecation notice
Changed in this release
Important Changes in Next Minor Release
Important Changes in Next Major Release

Deprecated Components

Component TypeVersions Being Deprecated
APIsv1, v2

What This Means

Only the v3 API and v3 Dashboard will be supported starting with the next major release. These newer versions offer improved performance, enhanced features, and better security.

Support options

As we grow, we stay committed to providing the best support ever seen from an open-source solution. Should you encounter an issue with any of the changes made in this release or any feature in the Netdata Agent, feel free to contact us through one of the following channels:

  • Premium Support: Customers who wish to have a direct channel with Netdata and prioritized support with defined SLAs can contact us.
  • Netdata Learn: Find documentation, guides, and reference material for monitoring and troubleshooting your systems with Netdata.
  • GitHub Issues: Use the Netdata repository to report bugs or open a new feature request.
  • GitHub Discussions: Join the conversation around the Netdata development process and be a part of it.
  • Community Forums: Visit the Community Forums and contribute to the collaborative knowledge base.
  • Discord Server: Jump into the Netdata Discord and hang out with like-minded sysadmins, DevOps, SREs, and other troubleshooters. More than 2000 engineers are already using it!
View originalPermalink
How v2.10.0 went

v2.9.0

Added 10
  • Interactive query analysis functions for 14+ databases including ClickHouse, CockroachDB, Couchbase, Elasticsearch, OpenSearch, MongoDB, MS SQL Server, and MySQL/MariaDB to identify slow queries and performance bottlenecks from the Netdata dashboard
  • Top Queries function for ClickHouse to show aggregated query statistics from system.query_log with execution time, memory usage, and rows read/written
  • Top Queries and Running Queries functions for CockroachDB to display statement statistics and currently executing statements via SHOW CLUSTER STATEMENTS
  • Top Queries function for Couchbase to show completed N1QL requests from system:completed_requests with service time and result count
  • Top Queries function for Elasticsearch and OpenSearch to display active search tasks with running time and search type information
  • Top Queries function for MongoDB to show slow operations from system.profile with execution time and document examination details
Changed 1
  • SQL Server collector now provides UI-based configuration

From Netdata

Table of Contents
Release Summary

Netdata v2.9.0 brings powerful database observability and expanded OpenTelemetry support.

FeatureHighlightsDetails
Interactive Query Analysis14+ databases supported• Identify slow queries• Debug bottlenecks• Monitor operations• No manual database connections required
Rebuilt SQL Server CollectorRewritten in Go• Comprehensive metrics• Query Store integration• UI-based configuration
OpenTelemetry Log IngestionLogs via otel plugin• Stored in systemd-compatible journal files• Configurable retention policies
Stability ImprovementsProduction reliability• Under-the-hood fixes for more robust operation
Release Highlights
New Top Tab Functions: 14+ Databases and SNMP

We've added interactive query analysis functions to database collectors, letting you identify slow queries, long-running operations, and performance bottlenecks directly from the Netdata dashboard—no need to connect to the database and run diagnostic queries manually.

DatabaseFunctionsDescription
ClickHouseTop Queries• Aggregated query stats from system.query_log• Execution time, memory usage, rows read/written
CockroachDBTop Queries• Statement statistics from crdb_internal• Execution counts, latency, rows processed
Running Queries• Currently executing statements via SHOW CLUSTER STATEMENTS• Client info, duration, distributed execution status
CouchbaseTop Queries• Completed N1QL requests from system:completed_requests• Service time, result count, error tracking
Elasticsearch/OpenSearchTop Queries• Active search tasks from Tasks API• Running time, search type, node distribution
MongoDBTop Queries• Slow operations from system.profile• Execution time, docs examined, keys examined, plan summary
MS SQL ServerTop Queries• Query Store statistics• CPU time, logical reads/writes, memory grants, parallelism
Deadlock Info• Latest deadlock from system_health Extended Events• Victim process, lock mode, wait resource
Error Info• Recent SQL errors from Extended Events session• Error number, message, query text
MySQL/MariaDBTop Queries• Digest statistics from performance_schema• Execution time, lock time, rows examined/sent
Deadlock Info• Latest InnoDB deadlock from SHOW ENGINE INNODB STATUS• Victim transaction, lock mode, wait resource
Error Info• Recent SQL errors from Performance Schema history• Error number, SQLSTATE, message per query digest
Oracle DBTop Queries• SQL statistics from V$SQLSTATS• CPU time, elapsed time, buffer gets, disk reads
Running Queries• Active sessions from V$SESSION• Wait events, blocking sessions, SQL text
PostgreSQLTop Queries• Statement statistics from pg_stat_statements• Total/mean time, shared blocks hit/read, temp blocks
Running Queries• Active queries from pg_stat_activity• Duration, wait events, client info, backend state
ProxySQLTop Queries• Query digest from stats_mysql_query_digest• Execution time, rows affected/sent, errors
RedisTop Queries• Slow commands from SLOWLOG• Command name, duration, client info
RethinkDBRunning Queries• Active jobs from rethinkdb.jobs• Query text, duration, involved servers
YugabyteDBTop Queries• YSQL statistics from pg_stat_statements• Execution time, calls, rows processed
Running Queries• Active backends from pg_stat_activity• Query state, wait events, elapsed time
Generic SQLUser-Defined• Custom SQL functions defined in job configuration• Interactive table views in Top tab for any SQL database

What You Can Do:

  • Identify resource hogs: Sort by total execution time, CPU time, or I/O
  • Spot frequent queries: Find high-call-count queries that may benefit from caching
  • Debug in real-time: See currently running queries and which client initiated them
  • No extra tools needed: All analysis happens in the Netdata dashboard's Top tab

[!IMPORTANT] Some functions require database-specific configuration (e.g., enabling Query Store, pg_stat_statements, or profiling).

See each collector's documentation for prerequisites.

Also Added: SNMP Network Interfaces Function

CollectorFunctionDescription
SNMPNetwork Interfaces• Real-time interface status and traffic• Packets in/out, errors, discards• Filterable by type (Ethernet, Aggregation, Virtual)
Microsoft SQL Server Collector (Go)

We've rebuilt the Microsoft SQL Server collector from the ground up in Go, replacing the legacy C implementation. This brings comprehensive SQL Server monitoring with real-time query analysis capabilities.

What You Get:

FeatureDetails
Full Metric CoverageConnections, batch requests, buffer manager, memory, wait statistics, locks, per-database transactions, I/O latency, file sizes, and SQL Agent job status
Top Queries FunctionReal-time view of your most resource-intensive queries directly in the dashboard—see execution time, CPU, I/O, memory grants, and parallelism metrics
UI ConfigurationCreate and manage monitoring jobs through the Netdata dashboard—no manual config file editing required
Replication MonitoringTrack replication status, latency, and warnings for transactional/merge publications
Flexible AuthenticationSQL authentication, Windows integrated auth, named instances, and remote servers

Top Queries: Identify Performance Bottlenecks

The new Top Queries function pulls data from SQL Server's Query Store to show:

  • Queries consuming the most total execution time or CPU
  • I/O patterns (logical reads, physical reads, writes)
  • Memory grants and tempdb usage
  • Parallelism (DOP) across executions

[!IMPORTANT] Requires Query Store enabled on monitored databases (SQL Server 2016+).

Query text may contain sensitive data—ensure proper dashboard access controls.

OpenTelemetry Logs Ingestion

We're expanding our OpenTelemetry support! Following the metrics support added in v2.7.0, the otel plugin now ingests OpenTelemetry logs through the same OTLP gRPC endpoint.

What's New
  • Log Ingestion: Send logs from your OpenTelemetry Collector or instrumented applications directly to Netdata.
  • Custom Storage: Logs are stored in systemd-compatible journal files using a Rust implementation (no libsystemd dependency required).
  • Unified Interface: View OpenTelemetry logs alongside system logs in the Logs tab with full search and filtering capabilities.
Configuration & Retention

Control your log storage with configurable retention policies:

  • Set limits by disk usage, file count, or age
  • Define custom storage locations
  • Tune file size and time span settings
Availability

Available on Linux agents (beta), with support for additional platforms coming soon.


[!NOTE]
This feature is in beta. We welcome your feedback!

Contributions
Collectors
  • Added dynamic configuration support to go.d Service Discovery with UI-driven templates (#21680, #21718, #21730, @ilyam8)
  • Added SNMP profile metadata mapping for “Westermo Teleindustri AB” to “Westermo” (#21727, @ilyam8)
  • Added a dcgm collector for NVIDIA dcgm-exporter with full field classification (#21721, @ktsaou)
  • Added jitter and variance metrics to the go.d ping collector (#21683, @ktsaou)
  • Added a Kubernetes API Server collector (#21682, @ktsaou)
  • Added function support to the go.d SQL collector to enable interactive, filterable table views in the Top tab (#21666, #21671, #21678, @ilyam8)
  • Added PostgreSQL running-queries and enhanced top-queries support with pg_stat_monitor detection (#21656, @ktsaou)
  • Added Windows support to hardware monitoring collectors with direct vendor CLI execution and improved reliability (#21635, @ktsaou)
  • Added new cloud functions for MySQL and MSSQL to provide error and deadlock information (#21645, #21650, #21652, @ktsaou, @ilyam8)
  • Added smart ZFS dataset deduplication to the diskspace plugin to reduce duplicate reporting (#21643, @ilyam8)
  • Added top-queries and running-queries functions for nine additional database collectors: Redis, ClickHouse, Elasticsearch, Couchbase, ProxySQL, OracleDB, CockroachDB, YugabyteDB, and RethinkDB (go.d). (#21607, @ktsaou)
  • Added interface type categorization to SNMP collector for easier filtering (go.d/snmp). (#21605, @ilyam8)
  • Added interfaces function for SNMP collector with per-interface traffic, packet, error/discard, and status metrics with computed rates (go.d/snmp). (#21604, @ilyam8)
  • Implemented Top Queries Functions framework for PostgreSQL, MySQL, MSSQL, and MongoDB with a reusable Functions API for building Netdata Functions in Go (go.d). (#21595, @ktsaou)
  • Added a new Microsoft SQL Server (MSSQL) collector to go.d.plugin with comprehensive monitoring of connections, transactions, replication, and more (go.d/mssql). (#21583, @ktsaou)
  • Added hostgroup-level summary metrics and hostgroup label to all ProxySQL backend charts (go.d/proxysql). (#21549, @ilyam8)
  • Added ProxySQL health alerts for SHUNNED/OFFLINE_HARD backends and hostgroups with no ONLINE backends (go.d/proxysql). (#21548, @ilyam8)
  • Added IPC mutexes chart for Windows systems (windows.plugin). (#21474, @thiagoftsm)
  • Added journal-viewer plugin replacing the existing systemd-journal plugin with improved OpenTelemetry logs ingestion in a systemd-compatible way (journal-viewer-plugin). (#21356, #21561, #21568, #21578, #21581, #21625, @vkalintiris)
  • Added EATON UPS SNMP profile for automatic discovery and collection of battery, power, and status metrics (go.d/snmp). (#21355, @ilyam8)
  • Added four new Windows sensors: human presence, human proximity, electrical capacitance, and inductance (windows.plugin). (#21319, @thiagoftsm)
  • Added a scripts.d plugin to run unmodified Nagios plugins inside Netdata (#21294, #21759, #21774, @ktsaou)
  • Fixed systemd-units.plugin response content type to application/json (systemd-units.plugin). (#21621, @ilyam8)
  • Fixed Windows plugin metric names, error logging, and cleanup hooks (windows.plugin). (#21597, @thiagoftsm)
  • Fixed ProxySQL collector to use timeout when pinging instances (go.d/proxysql). (#21573, @ilyam8)
  • Fixed go.d logger to disable terminal color on Windows (go.d). (#21562, @ilyam8)
  • Fixed cgroup duplicate detection to only treat a match as a duplicate when the existing cgroup is enabled and available (cgroups.plugin). (#21559, @ilyam8)
  • Fixed SNMP collector nil dereference when no profiles are found (go.d/snmp). (#21557, @ilyam8)
  • Fixed Windows plugin driver lifecycle and CPU temperature collection thread safety (windows.plugin). (#21523, @thiagoftsm)
  • Fixed eBPF cleanup by properly cleaning up shared memory and semaphores (ebpf.plugin). (#21501, @stelfrag)
  • Fixed Windows storage collector memory leaks with dictionary delete callbacks (windows.plugin). (#21500, @stelfrag)
  • Fixed MSSQL ODBC bindings and cursor handling for improved collection stability (windows.plugin). (#21478, @thiagoftsm)
  • Fixed Windows plugin metrics, error logging, and cleanup across multiple collectors (windows.plugin). (#21474, @thiagoftsm)
  • Fixed Windows plugin service enumeration and chart correctness (windows.plugin). (#21466, @thiagoftsm)
  • Fixed incorrect perflib mappings, labels, and priorities in Windows plugin collectors (windows.plugin). (#21458, @thiagoftsm)
  • Fixed MSSQL NULL result handling to prevent bad metrics in replication scenarios (windows.plugin). (#21417, @thiagoftsm)
  • Improved panic handling by always capturing and logging backtraces in agent logs (journal-viewer) (#21681, @vkalintiris)
  • Added function-only mode to the go.d plugin, allowing modules and jobs to expose functions without collecting metrics (#21646, @ilyam8)
  • Refactored the go.d function API for cleaner, more consistent, and more maintainable collector code (#21633, #21655, @ilyam8)
  • Added RequireCloud flag for database functions to restrict sensitive query data to cloud-only access (go.d). (#21629, @ilyam8)
  • Set default update interval to 10 seconds for all go.d collector functions (go.d). (#21616, @ilyam8)
  • Skipped internal SQL Server databases (model_*) in MSSQL go.d collector (go.d/mssql). (#21588, @ilyam8)
  • Removed per-metric collect_* options from MSSQL go.d collector - all metrics now always collected (go.d/mssql). (#21586, @ilyam8)
  • Added option to trust MSSQL server certificates in Windows plugin for easier TLS setup (windows.plugin). (#21582, @thiagoftsm)
  • Migrated pgbouncer collector to jackc/pgx/v5 (go.d/pgbouncer). (#21510, @ilyam8)
  • Merged MSSQL queries to reduce runtime needed for metrics collection (windows.plugin). (#21491, @thiagoftsm)
  • Added MSSQL blocked processes collection using SQL query with configurable option (windows.plugin). (#21426, @thiagoftsm)
  • Formatted MSSQL Windows collector code to match project style (windows.plugin). (#21383, @thiagoftsm)
  • Added support for float-type dimensions in go.d collectors (go.d). (#21362, #21363, @ilyam8)
  • Added MSSQL session connections chart alongside user connections (windows.plugin). (#21350, @thiagoftsm)
  • Added MSSQL user connections collection via SQL query for improved reliability over Perflib (windows.plugin). (#21348, @thiagoftsm)
  • Reorganized MSSQL perflib code for improved readability (windows.plugin). (#21334, @thiagoftsm)
Packaging/Installation
  • Restored the C-based systemd-journal.plugin (#21729, @vkalintiris)
  • Updated bundled libbpf to v1.6.3.1p_netdata (#21717, @thiagoftsm)
  • Simplified LZ4 detection by requiring liblz4 ≥ 1.7.1 and enabling LZ4 builds consistently (#21669, @Ferroin)
  • Fixed $releasever_major usage to only apply on RHEL 9+, keeping RHEL 8 on $releasever. (#21610, #21649, @Ferroin)
  • Fixed repo URL to use only major version for RHEL 9/10. (#21596, @Ferroin)
  • Updated Azure signing action after rebranding for Windows builds. (#21591, @ilyam8)
  • Fixed 32-bit builds of the journal-viewer plugin. (#21580, @vkalintiris)
  • Enabled ARM package builds in PR CI runs for better testing coverage. (#21570, @Ferroin)
  • Restored correct permissions and capabilities for systemd-journal.plugin during source installs. (#21569, @ilyam8)
  • Removed Ubuntu 25.04 from CI and package builds after upstream EOL. (#21519, @Ferroin)
  • Fixed major version detection for native packages in updater. (#21485, @ilyam8)
  • Removed Fedora 41 from CI and package builds after upstream EOL. (#21475, @Ferroin)
  • Added automatic repair for systems affected by the /dev/null deletion bug in the updater. (#21432, @Ferroin)
  • Fixed RPM updater check URL generation for RHEL-family distros. (#21431, @Ferroin)
  • Fixed updater bug that could delete /dev/null when URL check fails. (#21428, @ilyam8)
  • Fixed EPEL installation order to prevent failed installs on EL systems. (#21410, @ilyam8)
  • Fixed incorrect find command patterns in kickstart, causing installation failures. (#21399, @ilyam8)
  • Added concurrency to the check-markdown workflow to prevent duplicate runs. (#21393, @ilyam8)
  • Added a check to report when updates fail due to native packages no longer being published. (#21288, @Ferroin)
  • Added intelligent fallback behavior in the kickstart script based on native package availability detection. (#21262, @Ferroin)
  • Implemented systemd-tmpfiles for handling required directories at runtime. (#21243, @Ferroin)
  • Minor improvements to CMake code following best practices. (#21146, @Ferroin)
  • Suppressed known build warnings for cleaner build output. (#20102, @Ferroin)
Documentation
  • Fixed incorrect icon filenames in integration metadata so icons render correctly (#21767, @Ancairon)
  • Fixed metadata validation for ibm.d, network-viewer, and netdata-otel (#21756, @ktsaou)
  • Updated documentation to reference Cloud MCP alongside Agent and Parent MCP (#21746, @ktsaou)
  • Restructured integration categories, fixed related_resources, added guides (#21742, #21749, #21751, #21753, @ktsaou)
  • Improved cgroups.plugin discoverability by organizing it into per-technology modules (#21737, #21740, @ktsaou)
  • Added the missing type = api to streaming parent configuration examples (#21739, @ktsaou)
  • Added documentation for the Netdata Cloud MCP server (#21736, @juacker)
  • Added clarification to Plan & Billing docs about when cancellations take effect (#21728, @kanelatechnical)
  • Added documentation recommending bearer token protection as the preferred security method (#21712, @ktsaou)
  • Added documentation for opting out of Windows telemetry (#21710, @kanelatechnical)
  • Added documentation for logs support in the OpenTelemetry plugin (#21705, @vkalintiris)
  • Added documentation for access control and feature availability (#21703, @ktsaou)
  • Fixed relative Markdown links to ensure proper ingestion by the Learn site (#21688, #21689, #21690, @ktsaou)
  • Improved Netdata architecture description (#21642, @ktsaou)
  • Added node-states-and-transitions to learn map (#21631, @ktsaou)
  • Added reporting documentation page (#21520, @ktsaou)
  • Added comprehensive node states and transitions documentation covering Live, Stale, Offline, and Unseen states with timing details. (#21630, @ktsaou)
  • Improved go.d collector functions metadata with version-aware descriptions. (#21626, @ilyam8)
  • Updated collector function docs, schema, and templates for improved clarity. (#21622, @ilyam8)
  • Added collector function documentation with metadata for Top Queries/Running Queries across 14 collectors. (#21618, @ilyam8)
  • Clarified funcapi types and enums with explicit documentation comments. (#21617, @ilyam8)
  • Removed obsolete TODO documentation files. (#21614, @ilyam8)
  • Added alert configuration ordering and override documentation with the conceptual model. (#21613, @ktsaou)
  • Updated kickstart.md note formatting. (#21600, @kanelatechnical)
  • Added documentation for Real-Time Conversations with live exhibits and updated AI docs structure. (#21594, @shyamvalsan)
  • Added On-Prem Release Notes to documentation map. (#21571, @Ancairon)
  • Added ProxySQL alerts to metadata.yaml for documentation. (#21565, @ilyam8)
  • Added SOC 2 Type 2 compliance badge to Security and Privacy documentation. (#21554, @Ancairon)
  • Added reciprocal links between ephemerality, identity, and template documentation. (#21552, @ktsaou)
  • Fixed broken links in vm-templates.md and various documentation files. (#21546, @ilyam8)
  • Added unclaim/reclaim node guide for moving agents between Spaces. (#21539, @kanelatechnical)
  • Added VM templates and cloning guide for preparing Netdata installations for cloning. (#21527, @ktsaou)
  • Cleaned up go.d/prometheus metadata by removing archived and unmaintained repositories. (#21498, @ilyam8)
  • Added AI-Powered Alert Configuration documentation for creating alerts in plain English. (#21482, @kanelatechnical)
  • Added automatic node instance cleanup documentation to nodes-ephemerality page. (#21471, @kanelatechnical)
  • Added "docker" keyword to the cgroups plugin for search discoverability. (#21469, @Ancairon)
  • Clarified Docker image tags: netdata/netdata defaults to latest, use :stable for stable releases. (#21467, @ilyam8)
  • Added troubleshooting guide for SNMP chart gaps with timing diagnostics. (#21452, @ilyam8)
  • Fixed admonition syntax in SNMP documentation. (#21445, @ilyam8)
  • Removed third-party network device Prometheus exporters from documentation. (#21443, @ilyam8)
  • Added keywords and improved the overview for SNMP documentation. (#21442, @ilyam8)
  • Updated PLATFORM_SUPPORT.md with current distro versions and static build details. (#21422, @kanelatechnical)
  • Removed Docker prerequisite to install smartmontools as it's now included in helper images. (#21413, @ilyam8)
  • Added keyword support to the documentation map for improved search relevance. (#21402, @Ancairon)
  • Updated Home tab documentation to reflect the current Netdata Cloud UI. (#21386, @kanelatechnical)
  • Removed snippet with broken link from go_expvar metadata. (#21382, @Ancairon)
  • Fixed dangling link to sizing documentation. (#21372, @ralphm)
  • Updated security and privacy design links in map.csv. (#21369, @ilyam8)
Other Notable Changes
  • Ensured dyncfg-generated alerts fully sync to Cloud on creation and prevented duplicate configuration sends (#21706, @stelfrag)
  • Improved agent event posting stability by preventing unintended stdout writes and disabling signal-based timeouts (#21696, @stelfrag)
  • Improved statement preparation with spinlock protection to avoid shutdown crashes. (#21608, @stelfrag)
  • Used preallocated static buffer when writing daemon status during shutdown. (#21593, @stelfrag)
  • Added health database check before processing alerts and saving logs. (#21592, @stelfrag)
  • Fixed MCP test client prompts/resources support and schema validation. (#21521, @ktsaou)
  • Improved SSL certificate verification error logging with better memory handling. (#21513, @stelfrag)
  • Improved ACLK query execution with better error handling and resource cleanup. (#21479, @stelfrag)
  • Improved metadata storage stability with better JudyL error handling and vacuum throttling. (#21468, @stelfrag)
  • Sent complete node info when node is marked ephemeral for accurate Cloud data. (#21456, @stelfrag)
  • Improved streaming connection loss detection with TCP keepalive and dead connection detection. (#21430, @stelfrag)
  • Improved thread join handling to prevent race conditions during shutdown. (#21421, @stelfrag)
  • Skipped retention check during datafile initialization to speed up startup. (#21387, @stelfrag)
  • Added option to configure custom API URL for Telegram notifications. (#21376, @ilyam8)
  • Fixed a timeout race that could prematurely close long-running initial web requests (#21722, @ktsaou)
  • Fixed shutdown crash by safely handling destroyed dictionaries during function execution (#21713, @stelfrag)
  • Fixed ML model loading to correctly map columns and prevent incorrect k-means contexts (#21700, @stelfrag)
  • Fixed buffer overflow in the UTF-8 text sanitizer to prevent crashes on invalid or partial input (#21698, @stelfrag)
  • Fixed a segmentation fault during datafile extent flush by correcting UUID reference handling (#21667, @stelfrag)
  • Fixed claim script newline handling to ensure valid configuration files across different shells (#21665, @ktsaou)
  • Fixed Matrix notifications to comply with the Matrix API and prevent request errors (#21659, @ktsaou)
  • Fixed CPU frequency detection on FreeBSD to prevent NaN values in system info (#21658, @ktsaou)
  • Fixed race condition during thread exit cleanup. (#21603, @stelfrag)
  • Fixed function error message key to use camelCase for UI compatibility. (#21602, @ilyam8)
  • Fixed potential NULL dereference when releasing metric buffers. (#21615, @stelfrag)
  • Fixed race condition on shutdown by properly initializing discovery_thread. (#21563, @stelfrag)
  • Fixed crash during ML predictions by acquiring lock before accessing dimension data. (#21555, @stelfrag)
  • Fixed incorrect column index for old_value in SQLite health query. (#21533, @stelfrag)
  • Fixed health alert db lookup parser with correct operator mappings (< to LESS, > to GREATER) and added == operator support. (#21529, @ktsaou)
  • Fixed metric page list retrieval with UUID null check and bounds validation. (#21575, @stelfrag)
  • Fixed netdata_mutex_destroy macro to call correct debug function. (#21579, @stelfrag)
  • Fixed memory leaks in ACLK error paths for batch processing and alert configuration. (#21541, @stelfrag)
  • Fixed journal v2 migration to safely handle deleted metrics and null handles. (#21514, @stelfrag)
  • Fixed race condition during journal file deletion by unmapping before deleting. (#21512, @stelfrag)
  • Fixed ACLK MQTT send path to skip garbage-collected fragments preventing crashes. (#21483, @stelfrag)
  • Fixed functions event loop to propagate exit code instead of calling exit() directly. (#21455, @stelfrag)
  • Fixed MQTT packet ACK handling to prevent use-after-free. (#21416, @stelfrag)
  • Fixed service shutdown detection during health initialization for faster shutdown. (#21329, @stelfrag)
  • Fixed log queries to skip internal field-remapping entries, preventing incorrect field names from appearing in results (#21755, @vkalintiris)
  • Improved function execution by centralizing timeout and progress handling in the agent (#21723, @vkalintiris)
  • Improved the “deferred response too big” log with detailed context including size, limit, plugin, and transaction ID (#21720, @vkalintiris)
  • Downgraded inotify name-modify logs to info to reduce false error noise during systemd file rotation (#21719, @vkalintiris)
  • Added configurable indexing limits to prevent memory spikes from large journals and reduced log noise (#21716, @vkalintiris)
  • Ensured journal files are marked as archived on shutdown and added graceful signal handling to prevent unnecessary re-indexing (#21707, @vkalintiris)
  • Added strict validation for journal object reads to prevent out-of-bounds access and improved writer updates for consistency (#21697, @vkalintiris)
  • Fixed journal handling to avoid misclassifying active systemd files as archived during transient state changes (#21695, @vkalintiris)
  • Improved journal window management by handling mmap failures safely and tightening error handling (#21693, @vkalintiris)
  • Improved panic handling by always capturing and logging backtraces in agent logs (#21681, @vkalintiris)
  • Removed redundant libuv loop calls to improve event loop stability and reduce overhead (#21657, @stelfrag)
  • Removed Sentry error reporting from journal-viewer plugin (#21612, @vkalintiris)
Deprecation notice
Changed in this release
Important Changes in Next Minor Release
MS SQL Server Collector: C Version Removal

The legacy C-based MS SQL Server collector (part of windows.plugin) will be removed in the next minor release. The Go-based collector (go.d/mssql) is now the default and recommended option.

The Go version offers significant advantages:

  • More features — includes the new Top Queries function for real-time query analysis.
  • Multi-job support — monitor multiple SQL Server instances from a single agent.
  • Remote connections — connect to SQL Server instances over the network, not just localhost.
  • UI configuration — create and manage monitoring jobs directly from the Netdata dashboard.
  • Easier maintenance — benefits from the go.d framework's shared infrastructure and faster development cycles.

If you're currently using the C collector, see the Go collector setup documentation for configuration instructions.

Important Changes in Next Major Release

Deprecated Components

Component TypeVersions Being Deprecated
APIsv1, v2

What This Means

Only the v3 API and v3 Dashboard will be supported starting with the next major release. These newer versions offer improved performance, enhanced features, and better security.

Support options

As we grow, we stay committed to providing the best support ever seen from an open-source solution. Should you encounter an issue with any of the changes made in this release or any feature in the Netdata Agent, feel free to contact us through one of the following channels:

  • Premium Support: Customers who wish to have a direct channel with Netdata and prioritized support with defined SLAs can contact us.
  • Netdata Learn: Find documentation, guides, and reference material for monitoring and troubleshooting your systems with Netdata.
  • GitHub Issues: Use the Netdata repository to report bugs or open a new feature request.
  • GitHub Discussions: Join the conversation around the Netdata development process and be a part of it.
  • Community Forums: Visit the Community Forums and contribute to the collaborative knowledge base.
  • Discord Server: Jump into the Netdata Discord and hang out with like-minded sysadmins, DevOps, SREs, and other troubleshooters. More than 2000 engineers are already using it!
View originalPermalink
How v2.9.0 went

v2.8.5

Changed 2
  • Update the Netdata Docker image to Debian 13.3 with the latest upstream updates and security fixes
  • Standardize alert configuration to use a consistent type for time group values
Fixed 5
  • Fix CSV parser configuration serialization so empty settings are correctly omitted from logs configuration output
  • Fix ProxySQL backend status metrics in the go.d collector to correctly reflect reported backend states
  • Fix proc file parsing to support non-seekable files, restoring network stats collection on affected kernels
  • Fix edit-config to ignore inherited container environment values, preventing incorrect command handling
  • Fix vnode label configuration by allowing arbitrary label keys in the go.d schema, enabling labels to be set via the UI

From Netdata

Netdata v2.8.5 is a patch release to address issues discovered since v2.8.4.

This patch release provides the following bug fixes and updates:

  • Updated the Netdata Docker image to Debian 13.3 with the latest upstream updates and security fixes
  • Standardized alert configuration to use a consistent type for time group values (#21528, @ilyam8)
  • Fixed CSV parser configuration serialization so empty settings are correctly omitted from logs configuration output (#21526, @ilyam8)
  • Fixed ProxySQL backend status metrics in the go.d collector to correctly reflect reported backend states (#21524, @ilyam8)
  • Fixed proc file parsing to support non-seekable files, restoring network stats collection on affected kernels (#21507, @ilyam8)
  • Fixed edit-config to ignore inherited container environment values, preventing incorrect command handling (#21505, @ilyam8)
  • Fixed vnode label configuration by allowing arbitrary label keys in the go.d schema, enabling labels to be set via the UI (#21503, @ilyam8)
Support options

As we grow, we stay committed to providing the best support ever seen from an open-source solution. Should you encounter an issue with any of the changes made in this release or any feature in the Netdata Agent, feel free to contact us through one of the following channels:

  • Netdata Learn: Find documentation, guides, and reference material for monitoring and troubleshooting your systems with Netdata.
  • GitHub Issues: Make use of the Netdata repository to report bugs or open a new feature request.
  • GitHub Discussions: Join the conversation around the Netdata development process and be a part of it.
  • Community Forums: Visit the Community Forums and contribute to the collaborative knowledge base.
  • Discord Server: Jump into the Netdata Discord and hang out with like-minded sysadmins, DevOps, SREs, and other troubleshooters. More than 2000 engineers are already using it!
View originalPermalink
How v2.8.5 went

v2.8.4

Added 1
  • Include Go version in build info
Changed 1
  • Update go toolchain to v1.25.5

From Netdata

Netdata v2.8.4 is a patch release to address issues discovered since v2.8.3.

This patch release provides the following bug fixes and updates:

Support options

As we grow, we stay committed to providing the best support ever seen from an open-source solution. Should you encounter an issue with any of the changes made in this release or any feature in the Netdata Agent, feel free to contact us through one of the following channels:

  • Netdata Learn: Find documentation, guides, and reference material for monitoring and troubleshooting your systems with Netdata.
  • GitHub Issues: Make use of the Netdata repository to report bugs or open a new feature request.
  • GitHub Discussions: Join the conversation around the Netdata development process and be a part of it.
  • Community Forums: Visit the Community Forums and contribute to the collaborative knowledge base.
  • Discord Server: Jump into the Netdata Discord and hang out with like-minded sysadmins, DevOps, SREs, and other troubleshooters. More than 2000 engineers are already using it!
View originalPermalink
How v2.8.4 went

v2.8.3

Added 1
  • Added collection statistics to the SNMP go.d collector and exposed them as per-profile charts
Changed 5
  • Improved go.d job management by avoiding blocking during job shutdown, keeping other jobs responsive
  • Logged data collection duration when go.d collections are skipped due to a previous run still in progress
  • Updated ndexec runner to return captured stdout on command failures
  • Updated bundled components used in static builds
  • Updated DEB package builds to use XZ or zstd compression instead of gzip
Fixed 9
  • Fixed the go.d AP collector to handle unknown station statistics gracefully
  • Fixed incorrect plugin and binary paths on Windows for the go.d plugin by correctly detecting the installation prefix at startup
  • Fixed and standardized Windows AD, ADCS, and ADFS charts, and improved stability and accuracy of Windows hardware metrics collection
  • Fixed SNMP service discovery by accepting version "2" and using the default polling interval instead of overriding it
  • Fixed the MSSQL errors chart in the Windows plugin by using the correct metric
  • Disabled the MongoDB exporter on affected Ubuntu versions due to insecure libbson dependencies
Removed 1
  • Removed the strict rabbitmq_version check to improve RabbitMQ collector compatibility with older brokers

From Netdata

Netdata v2.8.3 is a patch release to address issues discovered since v2.8.2.

This patch release provides the following bug fixes and updates:

  • Fixed the go.d AP collector to handle unknown station statistics gracefully (#21461, @ilyam8)
  • Fixed incorrect plugin and binary paths on Windows for the go.d plugin by correctly detecting the installation prefix at startup (#21451, @ilyam8)
  • Improved go.d job management by avoiding blocking during job shutdown, keeping other jobs responsive (#21448, @ilyam8)
  • Fixed and standardized Windows AD, ADCS, and ADFS charts, and improved stability and accuracy of Windows hardware metrics collection (#21433, #21454, @thiagoftsm)
  • Fixed SNMP service discovery by accepting version “2” and using the default polling interval instead of overriding it (#21424, @ilyam8)
  • Logged data collection duration when go.d collections are skipped due to a previous run still in progress (#21423, #21425, @ilyam8)
  • Fixed the MSSQL errors chart in the Windows plugin by using the correct metric (#21412, @thiagoftsm)
  • Removed the strict rabbitmq_version check to improve RabbitMQ collector compatibility with older brokers (#21411, @ilyam8)
  • Added collection statistics to the SNMP go.d collector and exposed them as per-profile charts (#21409, @ilyam8)
  • Updated ndexec runner to return captured stdout on command failures (#21405, @ilyam8)
  • Disabled the MongoDB exporter on affected Ubuntu versions due to insecure libbson dependencies (#21403, @Ferroin)
  • Updated bundled components used in static builds (#21401, @Ferroin)
  • Fixed off-by-one validation in Journal v2 by treating extent_index == extent_entries as invalid, preventing out-of-bounds access (#21400, @stelfrag)
  • Fixed timed waits in completion to handle spurious wakeups and honor shutdown timeouts, preventing premature timeouts and hangs (#21395, @stelfrag)
  • Prevented datafiles from exceeding their max size by making size checks atomic and accounting for incoming writes (#21390, @stelfrag)
  • Updated DEB package builds to use XZ or zstd compression instead of gzip (#21310, @Ferroin)
Support options

As we grow, we stay committed to providing the best support ever seen from an open-source solution. Should you encounter an issue with any of the changes made in this release or any feature in the Netdata Agent, feel free to contact us through one of the following channels:

  • Netdata Learn: Find documentation, guides, and reference material for monitoring and troubleshooting your systems with Netdata.
  • GitHub Issues: Make use of the Netdata repository to report bugs or open a new feature request.
  • GitHub Discussions: Join the conversation around the Netdata development process and be a part of it.
  • Community Forums: Visit the Community Forums and contribute to the collaborative knowledge base.
  • Discord Server: Jump into the Netdata Discord and hang out with like-minded sysadmins, DevOps, SREs, and other troubleshooters. More than 2000 engineers are already using it!
View originalPermalink
How v2.8.3 went

v2.8.2

Changed 2
  • Replaced dots with slashes in OTEL metric families to enable hierarchical grouping in dashboards
  • Prioritized environment-provided config directories in go.d for runtime overrides
Fixed 6
  • Adjusted Windows sensors initialization by moving COM and Sensor API setup into the sensors thread for better stability
  • Fixed user group updates to ensure proper Proxmox group assignment in Docker entrypoint
  • Ensured Netdata has access to NVIDIA device files by adding the netdata user to the appropriate group, fixing GPU monitoring on non-Debian systems
  • Prevented replication from stalling by detecting empty-response loops and safely completing the process
  • Stopped replication when the parent is already caught up, preventing stalls and unnecessary gap-filling
  • Silenced Redis client library logs in the go.d Redis collector to reduce noise

From Netdata

Netdata v2.8.2 is a patch release to address issues discovered since v2.8.1.

This patch release provides the following bug fixes and updates:

  • Adjusted Windows sensors initialization by moving COM and Sensor API setup into the sensors thread for better stability (#21374, @stelfrag)
  • Replaced dots with slashes in OTEL metric families to enable hierarchical grouping in dashboards (#21371, @ilyam8)
  • Fixed user group updates to ensure proper Proxmox group assignment in Docker entrypoint (#21364, @ilyam8)
  • Ensured Netdata has access to NVIDIA device files by adding the netdata user to the appropriate group, fixing GPU monitoring on non-Debian systems (#21358, #21359, @ilyam8)
  • Prevented replication from stalling by detecting empty-response loops and safely completing the process (#21357, @stelfrag)
  • Stopped replication when the parent is already caught up, preventing stalls and unnecessary gap-filling (#21352, @stelfrag)
  • Prioritized environment-provided config directories in go.d for runtime overrides (#21345, @ilyam8)
  • Silenced Redis client library logs in the go.d Redis collector to reduce noise (#21344, @ilyam8)
Support options

As we grow, we stay committed to providing the best support ever seen from an open-source solution. Should you encounter an issue with any of the changes made in this release or any feature in the Netdata Agent, feel free to contact us through one of the following channels:

  • Netdata Learn: Find documentation, guides, and reference material for monitoring and troubleshooting your systems with Netdata.
  • GitHub Issues: Make use of the Netdata repository to report bugs or open a new feature request.
  • GitHub Discussions: Join the conversation around the Netdata development process and be a part of it.
  • Community Forums: Visit the Community Forums and contribute to the collaborative knowledge base.
  • Discord Server: Jump into the Netdata Discord and hang out with like-minded sysadmins, DevOps, SREs, and other troubleshooters. More than 2000 engineers are already using it!
View originalPermalink
How v2.8.2 went
View all

Discussion

If you publish Netdata, you can claim this product by proving you administer its repository.