inclusionAI Ling-3.0-flash — 124B total / 5.1B active MoE, hybrid Mamba+MLA attention, 256K context. Requires a vendor vLLM fork (inclusionAI/vllm, branch ling_3_0).
Results
Comparing 3 setups — 3 config differences
| hardware | vllm-ling-3p0-flash-8xh200 (spot) | vllm-ling-3p0-flash-8xh200 (on-demand) | vllm-ling-3p0-flash-2xb200-runpod |
|---|---|---|---|
| engine | vllm | vllm | vllm |
| accelerators | H200:8 | H200:8 | B200:2 |
| parameter | vllm-ling-3p0-flash-8xh200 (spot) | vllm-ling-3p0-flash-8xh200 (on-demand) | vllm-ling-3p0-flash-2xb200-runpod |
|---|---|---|---|
max-num-seqs | — | — | 384 |
moe-backend | — | — | triton |
tensor-parallel-size | 4 | 4 | 2 |
- vllm-ling-3p0-flash-8xh200 (spot)
- vllm-ling-3p0-flash-8xh200 (on-demand)
- vllm-ling-3p0-flash-2xb200-runpod
Leaderboard
| # | chart | setup | hardware | config | tok/s | $/1M completion |
|---|---|---|---|---|---|---|
| 1 | vllm-ling-3p0-flash-8xh200 (spot) | vllm · H200:8 | TP=4+8KV cache
Model
Parallelism
Other
| 10,950.6 | $0.8570 | |
| 2 | vllm-ling-3p0-flash-8xh200 (on-demand) | vllm · H200:8 | TP=4+8KV cache
Model
Parallelism
Other
| 10,950.6 | $2.0202 | |
| 3 | vllm-ling-3p0-flash-2xb200-runpod | vllm · B200:2 | TP=2 · max_seqs=384+9Kernels
KV cache
Model
Parallelism
Scheduler
Other
| 5,352.2 | $0.6684 |