The gap isn't a missing dashboard. It's a missing closed loop.
DCGM for telemetry, Base Command Manager for orchestration, Run:ai for AI-driven scheduling and ops. NVIDIA valued that operational layer enough to acquire Run:ai for roughly $700M.
No equivalent. MI300X/MI355X fleets run on borrowed tooling, spreadsheets, and whoever remembers what happened last time.
| Category | Examples | What it covers |
|---|---|---|
| Provisioning / LCM | Tinkerbell, xCAT, MAAS, Bright, Rocks | Bare-metal imaging — not ongoing health. |
| Observability | Prometheus, Grafana, Elastic | Dashboards — not remediation. |
| HPC schedulers | Slurm, PBS, LSF | Job placement — not fleet health. |
| Fabric telemetry | NVIDIA UFM | Deep, but NVIDIA/InfiniBand-only. |
MI300X/MI355X deployments are ramping as neoclouds and hyperscalers diversify GPU supply beyond a single vendor. Every fleet that crosses from a few racks to thousands of nodes hits the same wall: manual operations stops scaling long before the hardware does. The window is open now — before an incumbent, or NVIDIA's own tooling, extends sideways into AMD the way the Run:ai acquisition extended NVIDIA's reach.