Fleet Operations
Discover
Live operational intelligence across every free NVIDIA NIM endpoint — ranked by throughput, reliability, and congestion, refreshed continuously.
- 24
- Endpoints
- 6
- Providers
- 23
- Degraded
- 2m ago
- Last probe
Fleet Degraded1 of 24 endpoints operational · 23 degraded
1815
LIVEThroughput▲18%
27.7tok/s
P95 Latency▼2.7%
7,155ms
Reliability▲2.2pt
65.4%
Congestion▼2.2pt
45%
Model Registry
24 endpoints
Fleet Performance · 12h
2274ms
▲ 4.1% · 12havg2161min779max4266
Loading…
Provider Intelligence
6 providers · 24 models · 23 degraded
ProviderMCongestionUptimeDegRouteRecovery
DOther5
53%
66.00%
50/5
DNVIDIA11
27%
75.46%
100/11
DMistral1
100%
56.67%
10/1
DMeta4
56%
46.67%
40/4
DGoogle2
55%
83.33%
20/2
DDeepSeek1
100%
0.00%
10/1
Overall Rankings
24 ranked · throughput · reliability · congestion
Runner Up
riva-translate-4b-instruct-v1.1
NVIDIA
520 pts
Best Overall
nemotron-3-super-120b-a12b
NVIDIA
547 pts
Third
nemotron-3-nano-omni-30b-a3b-reasoning
NVIDIA
519 pts
1547A
nemotron-3-super-120b-a12b
NVIDIA
2520A
riva-translate-4b-instruct-v1.1
NVIDIA
3519A
nemotron-3-nano-omni-30b-a3b-reasoning
NVIDIA
4508A
nemotron-3.5-content-safety
NVIDIA
5475A
gpt-oss-20b
Other
6475A
llama-3.2-11b-vision-instruct
Meta
7466A
nemotron-3-ultra-550b-a55b
NVIDIA
8456A
diffusiongemma-26b-a4b-it
Google
9446A
muse-glimmer-30b
Meta
10428A
riva-translate-4b-instruct-v2
NVIDIA
11405A
kimi-k3
Other
12396A
laguna-xs-2.1
Other
13384A
nemotron-3.5-lightning-30b-a3b
NVIDIA
14344A
llama-3.1-nemotron-safety-guard-8b-v3
NVIDIA
15295A
ising-calibration-1.5-31b
NVIDIA
16278B
llama-3.1-nemoguard-8b-topic-control
NVIDIA
17219D
glm-5.3
Other
18219D
gemma-4-31b-it
Google
19171D
mistral-nemotron
Mistral
209D
glm-5.3-flash
Other
210D
llama-3.1-nemoguard-8b-content-safety
NVIDIA
220D
llama-guard-4-12b
Meta
230D
llama-3.2-90b-vision-instruct
Meta
240D
deepseek-v4.1-flash
DeepSeek
Trend Analysis
24 models tracked
Rising1
riva-translate-4b-instruct-v1.1100%
Improving13
glm-5.373%
laguna-xs-2.173%
gpt-oss-20b90%
nemotron-3.5-lightning-30b-a3b70%
llama-3.1-nemotron-safety-guard-8b-v377%
Stable10
glm-5.3-flash3%
riva-translate-4b-instruct-v2100%
nemotron-3.5-content-safety97%
nemotron-3-ultra-550b-a55b87%
nemotron-3-super-120b-a12b100%
Declining0
All clear
Reliability Matrix
recent · by model
riva-translate-4b-instruct-v2
82
82
82
83
82
82
83
83
84
84
84
84
riva-translate-4b-instruct-v1.1
93
93
93
94
94
94
97
97
97
97
97
97
nemotron-3-super-120b-a12b
97
97
97
97
93
93
93
93
93
93
95
95
nemotron-3-nano-omni-30b-a3b-reasoning
96
96
96
97
97
97
96
96
96
93
93
92
nemotron-3.5-content-safety
96
96
96
96
96
96
96
94
94
94
0
86
llama-3.2-11b-vision-instruct
0
91
91
91
91
91
92
92
92
92
93
93
diffusiongemma-26b-a4b-it
86
86
86
87
88
92
92
92
91
91
91
91
gpt-oss-20b
91
0
89
0
85
85
85
84
0
80
80
80
kimi-k3
0
0
83
83
83
84
88
88
88
88
88
87
muse-glimmer-30b
86
86
86
86
86
82
81
84
0
83
83
83
nemotron-3-ultra-550b-a55b
85
81
81
81
81
0
88
88
88
88
0
84
llama-3.1-nemotron-safety-guard-8b-v3
0
81
81
0
82
82
0
0
0
74
73
73
glm-5.3
74
75
76
0
0
70
70
71
0
73
73
0
laguna-xs-2.1
0
75
75
75
75
75
75
75
0
0
68
68
gemma-4-31b-it
70
74
74
73
73
73
74
73
0
74
0
0
nemotron-3.5-lightning-30b-a3b
73
73
0
68
0
0
0
63
67
68
71
73
mistral-nemotron
63
64
68
67
67
0
64
0
0
0
0
0
llama-3.1-nemoguard-8b-topic-control
0
0
66
68
70
0
67
0
64
66
68
68
ising-calibration-1.5-31b
0
0
0
55
0
0
55
0
53
0
53
56
glm-5.3-flash
0
0
0
0
0
0
0
0
0
0
0
0
llama-3.1-nemoguard-8b-content-safety
0
0
0
0
0
0
0
0
0
0
0
0
llama-guard-4-12b
0
0
0
0
0
0
0
0
0
0
0
0
llama-3.2-90b-vision-instruct
0
0
0
0
0
0
0
0
0
0
0
0
deepseek-v4.1-flash
0
0
0
0
0
0
0
0
0
0
0
0
Incident Timeline
23 active
Congestion spike detected — consider failover
glm-5.3-flash · Other
Congestion spike detected — consider failover
glm-5.3 · Other
Congestion spike detected — consider failover
laguna-xs-2.1 · Other
Congestion spike detected — consider failover
gpt-oss-20b · Other
Congestion spike detected — consider failover
nemotron-3.5-lightning-30b-a3b · NVIDIA
Congestion spike detected — consider failover
nemotron-3-ultra-550b-a55b · NVIDIA
Congestion spike detected — consider failover
llama-3.1-nemotron-safety-guard-8b-v3 · NVIDIA
Congestion spike detected — consider failover
llama-3.1-nemoguard-8b-topic-control · NVIDIA
Congestion spike detected — consider failover
llama-3.1-nemoguard-8b-content-safety · NVIDIA
Congestion spike detected — consider failover
ising-calibration-1.5-31b · NVIDIA
Congestion spike detected — consider failover
mistral-nemotron · Mistral
Congestion spike detected — consider failover
llama-guard-4-12b · Meta
Congestion spike detected — consider failover
llama-3.2-90b-vision-instruct · Meta
Congestion spike detected — consider failover
gemma-4-31b-it · Google
Congestion spike detected — consider failover
deepseek-v4.1-flash · DeepSeek
Throughput degradation on primary endpoint
riva-translate-4b-instruct-v2 · NVIDIA
Throughput degradation on primary endpoint
riva-translate-4b-instruct-v1.1 · NVIDIA
Throughput degradation on primary endpoint
nemotron-3-super-120b-a12b · NVIDIA
Throughput degradation on primary endpoint
nemotron-3-nano-omni-30b-a3b-reasoning · NVIDIA
Throughput degradation on primary endpoint
kimi-k3 · Other
Throughput degradation on primary endpoint
muse-glimmer-30b · Meta
Throughput degradation on primary endpoint
llama-3.2-11b-vision-instruct · Meta
Throughput degradation on primary endpoint
diffusiongemma-26b-a4b-it · Google