Fleet Operations

Discover

Live operational intelligence across every free NVIDIA NIM endpoint — ranked by throughput, reliability, and congestion, refreshed continuously.

24
Endpoints
6
Providers
23
Degraded
2m ago
Last probe
Fleet Degraded1 of 24 endpoints operational · 23 degraded
Throughput▲18%
27.7tok/s
P95 Latency▼2.7%
7,155ms
Reliability▲2.2pt
65.4%
Congestion▼2.2pt
45%
24 endpoints
2274ms
▲ 4.1% · 12h
avg2161min779max4266

Loading…

6 providers · 24 models · 23 degraded
ProviderMCongestionUptimeDegRouteRecovery
DOther5
53%
66.00%
5
0/5
DNVIDIA11
27%
75.46%
10
0/11
DMistral1
100%
56.67%
1
0/1
DMeta4
56%
46.67%
4
0/4
DGoogle2
55%
83.33%
2
0/2
DDeepSeek1
100%
0.00%
1
0/1
24 ranked · throughput · reliability · congestion
Runner Up
riva-translate-4b-instruct-v1.1
NVIDIA
520 pts
Best Overall
nemotron-3-super-120b-a12b
NVIDIA
547 pts
Third
nemotron-3-nano-omni-30b-a3b-reasoning
NVIDIA
519 pts
1
nemotron-3-super-120b-a12b
NVIDIA
547A
2
riva-translate-4b-instruct-v1.1
NVIDIA
520A
3
nemotron-3-nano-omni-30b-a3b-reasoning
NVIDIA
519A
4
nemotron-3.5-content-safety
NVIDIA
508A
5
gpt-oss-20b
Other
475A
6
llama-3.2-11b-vision-instruct
Meta
475A
7
nemotron-3-ultra-550b-a55b
NVIDIA
466A
8
diffusiongemma-26b-a4b-it
Google
456A
9
muse-glimmer-30b
Meta
446A
10
riva-translate-4b-instruct-v2
NVIDIA
428A
11
kimi-k3
Other
405A
12
laguna-xs-2.1
Other
396A
13
nemotron-3.5-lightning-30b-a3b
NVIDIA
384A
14
llama-3.1-nemotron-safety-guard-8b-v3
NVIDIA
344A
15
ising-calibration-1.5-31b
NVIDIA
295A
16
llama-3.1-nemoguard-8b-topic-control
NVIDIA
278B
17
glm-5.3
Other
219D
18
gemma-4-31b-it
Google
219D
19
mistral-nemotron
Mistral
171D
20
glm-5.3-flash
Other
9D
21
llama-3.1-nemoguard-8b-content-safety
NVIDIA
0D
22
llama-guard-4-12b
Meta
0D
23
llama-3.2-90b-vision-instruct
Meta
0D
24
deepseek-v4.1-flash
DeepSeek
0D
24 models tracked
Rising1
riva-translate-4b-instruct-v1.1100%
Improving13
glm-5.373%
laguna-xs-2.173%
gpt-oss-20b90%
nemotron-3.5-lightning-30b-a3b70%
llama-3.1-nemotron-safety-guard-8b-v377%
Stable10
glm-5.3-flash3%
riva-translate-4b-instruct-v2100%
nemotron-3.5-content-safety97%
nemotron-3-ultra-550b-a55b87%
nemotron-3-super-120b-a12b100%
Declining0
All clear
recent · by model
riva-translate-4b-instruct-v2
82
82
82
83
82
82
83
83
84
84
84
84
riva-translate-4b-instruct-v1.1
93
93
93
94
94
94
97
97
97
97
97
97
nemotron-3-super-120b-a12b
97
97
97
97
93
93
93
93
93
93
95
95
nemotron-3-nano-omni-30b-a3b-reasoning
96
96
96
97
97
97
96
96
96
93
93
92
nemotron-3.5-content-safety
96
96
96
96
96
96
96
94
94
94
0
86
llama-3.2-11b-vision-instruct
0
91
91
91
91
91
92
92
92
92
93
93
diffusiongemma-26b-a4b-it
86
86
86
87
88
92
92
92
91
91
91
91
gpt-oss-20b
91
0
89
0
85
85
85
84
0
80
80
80
kimi-k3
0
0
83
83
83
84
88
88
88
88
88
87
muse-glimmer-30b
86
86
86
86
86
82
81
84
0
83
83
83
nemotron-3-ultra-550b-a55b
85
81
81
81
81
0
88
88
88
88
0
84
llama-3.1-nemotron-safety-guard-8b-v3
0
81
81
0
82
82
0
0
0
74
73
73
glm-5.3
74
75
76
0
0
70
70
71
0
73
73
0
laguna-xs-2.1
0
75
75
75
75
75
75
75
0
0
68
68
gemma-4-31b-it
70
74
74
73
73
73
74
73
0
74
0
0
nemotron-3.5-lightning-30b-a3b
73
73
0
68
0
0
0
63
67
68
71
73
mistral-nemotron
63
64
68
67
67
0
64
0
0
0
0
0
llama-3.1-nemoguard-8b-topic-control
0
0
66
68
70
0
67
0
64
66
68
68
ising-calibration-1.5-31b
0
0
0
55
0
0
55
0
53
0
53
56
glm-5.3-flash
0
0
0
0
0
0
0
0
0
0
0
0
llama-3.1-nemoguard-8b-content-safety
0
0
0
0
0
0
0
0
0
0
0
0
llama-guard-4-12b
0
0
0
0
0
0
0
0
0
0
0
0
llama-3.2-90b-vision-instruct
0
0
0
0
0
0
0
0
0
0
0
0
deepseek-v4.1-flash
0
0
0
0
0
0
0
0
0
0
0
0
23 active
Congestion spike detected — consider failover
glm-5.3-flash · Other
2m ago
Congestion spike detected — consider failover
glm-5.3 · Other
2m ago
Congestion spike detected — consider failover
laguna-xs-2.1 · Other
3m ago
Congestion spike detected — consider failover
gpt-oss-20b · Other
2m ago
Congestion spike detected — consider failover
nemotron-3.5-lightning-30b-a3b · NVIDIA
2m ago
Congestion spike detected — consider failover
nemotron-3-ultra-550b-a55b · NVIDIA
2m ago
Congestion spike detected — consider failover
llama-3.1-nemotron-safety-guard-8b-v3 · NVIDIA
2m ago
Congestion spike detected — consider failover
llama-3.1-nemoguard-8b-topic-control · NVIDIA
2m ago
Congestion spike detected — consider failover
llama-3.1-nemoguard-8b-content-safety · NVIDIA
2m ago
Congestion spike detected — consider failover
ising-calibration-1.5-31b · NVIDIA
2m ago
Congestion spike detected — consider failover
mistral-nemotron · Mistral
2m ago
Congestion spike detected — consider failover
llama-guard-4-12b · Meta
2m ago
Congestion spike detected — consider failover
llama-3.2-90b-vision-instruct · Meta
2m ago
Congestion spike detected — consider failover
gemma-4-31b-it · Google
2m ago
Congestion spike detected — consider failover
deepseek-v4.1-flash · DeepSeek
2m ago
Throughput degradation on primary endpoint
riva-translate-4b-instruct-v2 · NVIDIA
2m ago
Throughput degradation on primary endpoint
riva-translate-4b-instruct-v1.1 · NVIDIA
2m ago
Throughput degradation on primary endpoint
nemotron-3-super-120b-a12b · NVIDIA
2m ago
Throughput degradation on primary endpoint
nemotron-3-nano-omni-30b-a3b-reasoning · NVIDIA
2m ago
Throughput degradation on primary endpoint
kimi-k3 · Other
2m ago
Throughput degradation on primary endpoint
muse-glimmer-30b · Meta
2m ago
Throughput degradation on primary endpoint
llama-3.2-11b-vision-instruct · Meta
2m ago
Throughput degradation on primary endpoint
diffusiongemma-26b-a4b-it · Google
2m ago