Tok/s vs cost of self-hosting inclusionAI/Ling-3.0-flash-fp8 on H200:8.
Run it yourself
ling-3.0-flash-fp8@v1 · H200:8
$ iwant launch ling-3.0-flash-fp8Results
Dashed lines estimate the same throughput at another cloud's price for the same GPUs.
| setup | $/hour | peak tok/s | $/1M at peak | |
|---|---|---|---|---|
| ling-3.0-flash-fp8@v1 · H200:8 | $52.03 | 17,522 | $0.8249 | |
| ling-3.0-flash-fp8@v1 · AWS ap-northeast-2 spot (estimated) | $12.18 | 17,522 | $0.1932 | |
| ling-3.0-flash-fp8@v1 · AWS us-east-1 on-demand (estimated) | $63.30 | 17,522 | $1.0035 |