ALWAYS-ON GPU HOSTING
Reserved capacity
A GPU stays allocated around the clock, whether requests arrive or not.
$2,900/ month
- Full GPU rented continuously
- Idle capacity remains billable
- Reference H100 economics
SOVEREIGN ENDPOINTS PRICING
Keep dedicated model capacity available when requests arrive—without keeping the underlying GPU rented around the clock.
Beta offer: $10/month, includes $50 in usage credits. Pricing depends on model, GPU type, GPU count, and required capacity.SERVERLESS ECONOMICS
Pay for active inference compute instead of idle capacity.
A GPU stays allocated around the clock, whether requests arrive or not.
Attach GPU capacity for inference, then release it when the workload is idle.
Illustrative comparison based on the original InferX economics reference. Actual cost varies by model, GPU class, active compute time, and deployment configuration.
HOW INFERX MAKES IT POSSIBLE
DEPLOYMENT