NVIDIA NIM Fleet Overview
Real-time status and reliability for every free NVIDIA NIM endpoint, probed continuously.
- 24
- Endpoints
- 3
- Healthy
- 10
- Incidents
- 1m ago
- Last probe
Partial degradation
3 of 24 endpoints healthy
Avg latency
1647ms
Avg throughput
27.9tok/s
Incidents 24h
10critical
Updated 1m ago
Loading…
Recommended Now
nemotron-3-super-120b-a12b
NVIDIA
Why this model
TTFT
824ms
Throughput
36.7 tok/s
Uptime
100%
Volatility
highly unstable
Queue
low
Reliability
99
Active Incidents
20laguna-xs-2.1 partially recovered
9h ago
INFOlaguna-xs-2.1 degraded (timeout)
9h ago
CRITICALnemotron-3.5-content-safety degraded (timeout)
9h ago
CRITICALnemotron-3-ultra-550b-a55b partially recovered
9h ago
INFOnemotron-3-ultra-550b-a55b degraded (timeout)
9h ago
CRITICALnemotron-3-nano-omni-30b-a3b-reasoning recovered to healthy
9h ago
INFOModel Fleet
Search, filter, pin & export · 24 endpoints
Reliability & SLA
Uptime history · time-of-day latency · error budgets
Loading reliability history…
Loading latency history…
About this dashboard
Measured, not reported
What you are looking at
NIM Stats is an independent operational dashboard for the free NVIDIA NIM API endpoints. A collector sends a real chat-completion request to every tracked model on a fixed cadence — roughly every ten minutes — and records what actually happened: time to first token, end-to-end latency, sustained tokens per second, whether the call succeeded or timed out, and how congested the endpoint appeared. Every figure above is one of those measurements, taken from outside NVIDIA's network. None of it is a vendor-published status claim.
Healthy means the endpoint is serving normally. Busy means it is serving but with elevated latency or congestion. Jammed means it is failing or timing out on probe. TTFT is milliseconds until the first token arrives, which is what an interactive chat feels; throughput is sustained tokens per second, which is what a batch job feels. The two rarely rank endpoints the same way, so the fleet table reports both.
How to use it
Pick the highest-reliability endpoint with a healthy status, or read the recommendation at the top of the page. If a call you were already making starts failing, check whether that endpoint is jammed here before you go debugging your own client. Measurements come from a single vantage point on a fixed cadence, so treat them as a strong prior for which endpoint to try first rather than a service-level guarantee — your own latency will vary with geography, network path, and prompt size.
Agents and scripts can read every page on this site as clean Markdown at the same URL by sending Accept: text/markdown, or by appending .md to the path. Start at /llms.txt for what this site covers and when to reach for it. For a time series or per-endpoint history rather than a summary, the NIM Stats API reference documents the public read-only JSON API — no key, no rate limit — with an OpenAPI 3.1 spec at /openapi.json. NIM Stats is not affiliated with, endorsed by, or operated by NVIDIA Corporation.