weightsapi.INFERENCEConsole
Navigation
Self-hosting versus API: do the arithmetic

Self-hosting versus API: do the arithmetic

Dedicated monthly cost divided by effective per-million-token API cost gives the volume crossover before engineering and reliability costs. An H100 at $1,690 and an effective $0.16 per million tokens gives about 10.56 billion tokens per month. This does not prove a single H100 can serve that workload.

Find your crossover#

Dedicated monthly cost divided by effective per-million-token API cost gives the volume crossover before engineering and reliability costs. An H100 at $1,690 and an effective $0.16 per million tokens gives about 10.56 billion tokens per month. This does not prove a single H100 can serve that workload.

Add utilization#

A reserved GPU is billed while idle. Include token throughput at your concurrency, peak capacity, load time, failover, observability and on-call effort.

Make the decision#

Use shared API capacity for variable demand. Evaluate a dedicated endpoint when you have stable load, a compatible model and measured throughput.