StepFun Step-3.7-Flash — hybrid SWA/global-attention thinking model with MTP speculative decoding. Requires a full 8-GPU node (TP=8).
Results
Comparing 2 setups — same config and hardware.
- vllm-step3p7-flash-8xh200 (on-demand)
- vllm-step3p7-flash-8xh200 (spot)
Leaderboard
| # | chart | setup | hardware | config | tok/s | $/1M completion |
|---|---|---|---|---|---|---|
| 1 | vllm-step3p7-flash-8xh200 (on-demand) | vllm · H200:8 | TP=8+7Model
Parallelism
Other
| 24,184.7 | $0.9277 | |
| 2 | vllm-step3p7-flash-8xh200 (spot) | vllm · H200:8 | TP=8+7Model
Parallelism
Other
| 19,452.6 | $0.4372 |