NIM Stats · Command Center

NVIDIA NIM Fleet Overview

Real-time status and reliability for every free NVIDIA NIM endpoint, probed continuously.

24
Endpoints
3
Healthy
10
Incidents
1m ago
Last probe
Fleet status

Partial degradation

3 of 24 endpoints healthy

Healthy 3Busy 8Jammed 13

Avg latency

1647ms

Avg throughput

27.9tok/s

Incidents 24h

10critical

Updated 1m ago

1629ms
▲ 61.0% · 12h
avg2021min779max4724

Loading…

nemotron-3-super-120b-a12b

NVIDIA

busy

Why this model

Low congestion (3%)
High session reliability
Low queue pressure
Excellent uptime (100%)
Minimal timeout rate

TTFT

824ms

Throughput

36.7 tok/s

Uptime

100%

Volatility

highly unstable

Queue

low

Reliability

99

20

laguna-xs-2.1 partially recovered

9h ago

INFO

laguna-xs-2.1 degraded (timeout)

9h ago

CRITICAL

nemotron-3.5-content-safety degraded (timeout)

9h ago

CRITICAL

nemotron-3-ultra-550b-a55b partially recovered

9h ago

INFO

nemotron-3-ultra-550b-a55b degraded (timeout)

9h ago

CRITICAL

nemotron-3-nano-omni-30b-a3b-reasoning recovered to healthy

9h ago

INFO

Search, filter, pin & export · 24 endpoints

Uptime history · time-of-day latency · error budgets

0Meeting SLA
0Breaching
10mAllowed downtime

Loading SLA data…

Loading reliability history…

0 models
036912151821

Loading latency history…

FasterSlower · TTFT by hour (UTC)

Measured, not reported

What you are looking at

NIM Stats is an independent operational dashboard for the free NVIDIA NIM API endpoints. A collector sends a real chat-completion request to every tracked model on a fixed cadence — roughly every ten minutes — and records what actually happened: time to first token, end-to-end latency, sustained tokens per second, whether the call succeeded or timed out, and how congested the endpoint appeared. Every figure above is one of those measurements, taken from outside NVIDIA's network. None of it is a vendor-published status claim.

Healthy means the endpoint is serving normally. Busy means it is serving but with elevated latency or congestion. Jammed means it is failing or timing out on probe. TTFT is milliseconds until the first token arrives, which is what an interactive chat feels; throughput is sustained tokens per second, which is what a batch job feels. The two rarely rank endpoints the same way, so the fleet table reports both.

How to use it

Pick the highest-reliability endpoint with a healthy status, or read the recommendation at the top of the page. If a call you were already making starts failing, check whether that endpoint is jammed here before you go debugging your own client. Measurements come from a single vantage point on a fixed cadence, so treat them as a strong prior for which endpoint to try first rather than a service-level guarantee — your own latency will vary with geography, network path, and prompt size.

Agents and scripts can read every page on this site as clean Markdown at the same URL by sending Accept: text/markdown, or by appending .md to the path. Start at /llms.txt for what this site covers and when to reach for it. For a time series or per-endpoint history rather than a summary, the NIM Stats API reference documents the public read-only JSON API — no key, no rate limit — with an OpenAPI 3.1 spec at /openapi.json. NIM Stats is not affiliated with, endorsed by, or operated by NVIDIA Corporation.