High-severity NVIDIA vulnerability lets unauthenticated attackers crash GPU monitoring

Hundreds of internet-exposed graphics processing unit (GPU) servers were open to a high-severity flaw in NVIDIA’s DCGM Exporter (CVE-2026-47483) that lets unauthenticated attackers crash the monitoring service and may disrupt AI workloads, according to Lava.

Lava reported the flaw to NVIDIA, which rated it 8.2 on the CVSS scale and published a security bulletin on July 28, 2026.

GPU servers are computers built around GPUs and are used for AI, machine learning and scientific computing workloads. DCGM Exporter is a monitoring tool that reads data directly from the GPUs on a host, including temperature, utilization, memory usage, power consumption and error events, and publishes it over HTTP, typically on port 9400, for systems like Prometheus to collect.

The data includes the exact GPU model and a unique ID (UUID) for each GPU. According to Lava researcher Michael Katchinskiy, this is enough for a potential attacker to see what hardware a server runs, how heavily it is used and whether it is experiencing errors.

After realizing how much these endpoints revealed, Katchinskiy asked, “How many of them are exposed to the internet?”

Who was exposed

Katchinskiy and his team found more than 2,000 servers exposing the exporter to the internet. Over four scans conducted between March and May 2026, these servers reported more than 12,000 unique GPUs, with an estimated value of $100 million.

“Every host returned metrics over plaintext HTTP, and none required authentication,” Katchinskiy wrote.

Among the exposed hardware were NVIDIA’s Blackwell Ultra B300, H200 and H100 GPUs, built for large AI workloads, along with consumer RTX 5090 and 4090 cards.

NVIDIA DCGM Exporter vulnerability CVE-2026-47483

(Source: Lava)

The exposed exporters belonged to nearly 300 organizations. The US hosted 5,274 of the exposed GPUs (44%), followed by Romania with 2,054 and China with 1,967.

Some of the exposed services ran on customer infrastructure at GPU cloud providers, including Voltage Park, Lambda, Northern Data and DigitalOcean.

The largest group was associated with Voltage Park, with 672 public Node Exporter hosts and 71 public DCGM Exporter hosts. The company confirmed that most of the Node Exporter instances and all of the DCGM Exporter instances were deployed by customers, and its security team contacted the affected customers.

Debug endpoints let attackers crash the exporter

About a quarter of the exposed hosts also served Go’s /debug/pprof/ profiling endpoints alongside /metrics. Some of these endpoints keep a request open for a duration set by the caller, and a large number of concurrent unauthenticated requests increases memory consumption.

Katchinskiy wrote that the team first assumed the exposed profiling endpoints were the result of operator misconfiguration. They then reproduced the behavior using NVIDIA’s official DCGM Exporter container without modifying it.

Deployments exposing the exporter on a routable interface could expose /debug/pprof/ as well.

“With enough concurrent unauthenticated requests, the exporter could run out of memory and crash, cutting off visibility into GPU health and activity,” he explained.

The CPU and RAM pressure could also slow the host and interfere with training or inference workloads running on it, particularly without strict resource limits. Hundreds of the public DCGM Exporters the team observed exposed this attack surface.

Node Exporter leaks server and network details

Researchers also examined Prometheus Node Exporter, which monitors a server’s hardware and operating system. They found 12,096 public Node Exporter hosts reporting metrics from NVIDIA/Mellanox InfiniBand and RoCE adapters, exposing adapter models, firmware versions, link state and fabric activity. Some revealed hostnames, OS and kernel versions, and BIOS details.

One endpoint returned enough data to identify a Dell PowerEdge XE9680 running Ubuntu 22.04.5 LTS, its BIOS and adapter firmware, and an active 400 Gb/s port.

“Exact versions let an attacker go straight to matching known vulnerabilities. Where Node Exporter and DCGM Exporter overlapped, we could tie GPU activity to the server, software stack and network around it, all without authentication,” Katchinskiy noted.

Fixes and mitigations

DCGM Exporter users should upgrade to version 4.8.2 or later and make sure the –enable-pprof flag is off unless profiling is needed. In current versions, the profiling endpoint is opt-in.

Katchinskiy advises keeping Node Exporter, DCGM Exporter and Prometheus off the public internet unless there is a specific need and access control in place. Exporters should be bound to a loopback or private interface, with firewall rules or security groups limiting access to the monitoring infrastructure. Prometheus query APIs and target pages need the same protection.

“Our research showed that publicly exposed monitoring services can reveal far more than basic telemetry,” Katchinskiy said. “Across thousands of systems, we could see what GPUs were running, how they were being used, and details about the servers and networks around them.”

In the case of DCGM Exporter, he added, the exposure “went beyond visibility” and allowed an attacker to exhaust resources, crash GPU monitoring and potentially affect workloads running on the same host.

“As AI infrastructure scales, organizations need to protect every layer of the stack and know exactly what they are responsible for versus what their provider is responsible for,” Katchinskiy concluded.

More about

Don't miss