Tok/s vs cost of self-hosting inclusionAI/Ling-3.0-flash-fp8 on H200:8.

Run it yourself

ling-3.0-flash-fp8@v1 · H200:8
$ iwant launch ling-3.0-flash-fp8

Results

Dashed lines estimate the same throughput at another cloud's price for the same GPUs.

setup$/hourpeak tok/s$/1M at peak
ling-3.0-flash-fp8@v1 · H200:8$52.0317,522$0.8249
ling-3.0-flash-fp8@v1 · AWS ap-northeast-2 spot (estimated)$12.1817,522$0.1932
ling-3.0-flash-fp8@v1 · AWS us-east-1 on-demand (estimated)$63.3017,522$1.0035