inclusionAI Ling-3.0-flash — 124B total / 5.1B active MoE, hybrid Mamba+MLA attention, 256K context. Requires a vendor vLLM fork (inclusionAI/vllm, branch ling_3_0).

Results

Comparing 3 setups 3 config differences

hardwarevllm-ling-3p0-flash-8xh200 (spot)vllm-ling-3p0-flash-8xh200 (on-demand)vllm-ling-3p0-flash-2xb200-runpod
enginevllmvllmvllm
acceleratorsH200:8H200:8B200:2
parametervllm-ling-3p0-flash-8xh200 (spot)vllm-ling-3p0-flash-8xh200 (on-demand)vllm-ling-3p0-flash-2xb200-runpod
max-num-seqs384
moe-backendtriton
tensor-parallel-size442
  • vllm-ling-3p0-flash-8xh200 (spot)
  • vllm-ling-3p0-flash-8xh200 (on-demand)
  • vllm-ling-3p0-flash-2xb200-runpod

Leaderboard

#chartsetuphardwareconfigtok/s$/1M completion
1vllm-ling-3p0-flash-8xh200 (spot)vllm · H200:8
TP=4+8

KV cache

enable-prefix-caching
true

Enable prefix caching for shared prompt prefixes.

gpu-memory-utilization
0.85

Fraction of GPU memory for the model executor (0–1). Per-instance limit.

Model

trust-remote-code
true

Trust remote code when downloading model and tokenizer.

Parallelism

tensor-parallel-size
4

Number of tensor-parallel groups.

Other

enable_auto_tool_choice
true
mamba_cache_mode
align
reasoning_parser
ling3
speculative_config
{"method": "mtp", "num_speculative_tokens": 3}
tool_call_parser
ling3
10,950.6$0.8570
2vllm-ling-3p0-flash-8xh200 (on-demand)vllm · H200:8
TP=4+8

KV cache

enable-prefix-caching
true

Enable prefix caching for shared prompt prefixes.

gpu-memory-utilization
0.85

Fraction of GPU memory for the model executor (0–1). Per-instance limit.

Model

trust-remote-code
true

Trust remote code when downloading model and tokenizer.

Parallelism

tensor-parallel-size
4

Number of tensor-parallel groups.

Other

enable_auto_tool_choice
true
mamba_cache_mode
align
reasoning_parser
ling3
speculative_config
{"method": "mtp", "num_speculative_tokens": 3}
tool_call_parser
ling3
10,950.6$2.0202
3vllm-ling-3p0-flash-2xb200-runpodvllm · B200:2
TP=2 · max_seqs=384+9

Kernels

moe-backend
triton

MoE expert computation kernel backend.

KV cache

enable-prefix-caching
true

Enable prefix caching for shared prompt prefixes.

gpu-memory-utilization
0.85

Fraction of GPU memory for the model executor (0–1). Per-instance limit.

Model

trust-remote-code
true

Trust remote code when downloading model and tokenizer.

Parallelism

tensor-parallel-size
2

Number of tensor-parallel groups.

Scheduler

max-num-seqs
384

Max sequences processed in a single iteration.

Other

enable_auto_tool_choice
true
mamba_cache_mode
align
reasoning_parser
ling3
speculative_config
{"method": "mtp", "num_speculative_tokens": 3}
tool_call_parser
ling3
5,352.2$0.6684