Find your crossover#
Dedicated monthly cost divided by effective per-million-token API cost gives the volume crossover before engineering and reliability costs. An H100 at $1,690 and an effective $0.16 per million tokens gives about 10.56 billion tokens per month. This does not prove a single H100 can serve that workload.
Add utilization#
A reserved GPU is billed while idle. Include token throughput at your concurrency, peak capacity, load time, failover, observability and on-call effort.
Make the decision#
Use shared API capacity for variable demand. Evaluate a dedicated endpoint when you have stable load, a compatible model and measured throughput.