StepFun Step-3.7-Flash — hybrid SWA/global-attention thinking model with MTP speculative decoding. Requires a full 8-GPU node (TP=8).

Results

Comparing 2 setups — same config and hardware.

  • vllm-step3p7-flash-8xh200 (on-demand)
  • vllm-step3p7-flash-8xh200 (spot)

Leaderboard

#chartsetuphardwareconfigtok/s$/1M completion
1vllm-step3p7-flash-8xh200 (on-demand)vllm · H200:8
TP=8+7

Model

trust-remote-code
true

Trust remote code when downloading model and tokenizer.

Parallelism

enable-expert-parallel
true

Use expert parallelism instead of tensor parallelism for MoE layers.

tensor-parallel-size
8

Number of tensor-parallel groups.

Other

disable_cascade_attn
true
enable_auto_tool_choice
true
reasoning_parser
step3p5
speculative_config
{"method": "mtp", "num_speculative_tokens": 3}
tool_call_parser
step3p5
24,184.7$0.9277
2vllm-step3p7-flash-8xh200 (spot)vllm · H200:8
TP=8+7

Model

trust-remote-code
true

Trust remote code when downloading model and tokenizer.

Parallelism

enable-expert-parallel
true

Use expert parallelism instead of tensor parallelism for MoE layers.

tensor-parallel-size
8

Number of tensor-parallel groups.

Other

disable_cascade_attn
true
enable_auto_tool_choice
true
reasoning_parser
step3p5
speculative_config
{"method": "mtp", "num_speculative_tokens": 3}
tool_call_parser
step3p5
19,452.6$0.4372