Skip to main content
Model shopping fails when you change the prompt and the model at the same time. Fix the usage block first, then compare totals.

Set a fixed usage block

Price the baseline

Price the candidate

Repeat with the cheaper model and the provider that lists it. Search first when you do not know which host carries the model:
Then call POST /v1/traces/cost with the same usage block.

Compare totals

If the candidate total is lower, pull live rates confirmed the savings before you switch.

Full case study

For accuracy and latency tradeoffs on a real routing workload, see Reduce LLM spend by switching models.