SOVEREIGN ENDPOINTS PRICING

    You should not pay for idle GPUs.

    Keep dedicated model capacity available when requests arrive—without keeping the underlying GPU rented around the clock.

    Beta offer: $10/month, includes $50 in usage credits. Pricing depends on model, GPU type, GPU count, and required capacity.

    SERVERLESS ECONOMICS

    Serverless compute vs always-on GPU hosting.

    Pay for active inference compute instead of idle capacity.

    ALWAYS-ON GPU HOSTING

    Reserved capacity

    A GPU stays allocated around the clock, whether requests arrive or not.

    $2,900/ month
    • Full GPU rented continuously
    • Idle capacity remains billable
    • Reference H100 economics
    INFERX SERVERLESS

    Active compute only

    Attach GPU capacity for inference, then release it when the workload is idle.

    ~$220/ month
    • Pay for active GPU time
    • Scale to zero between requests
    • Reference serverless workload

    Illustrative comparison based on the original InferX economics reference. Actual cost varies by model, GPU class, active compute time, and deployment configuration.

    HOW INFERX MAKES IT POSSIBLE

    Why the economics are different.

    InferX retains initialized runtime state independently from active GPU compute. The model remains ready while capacity is attached only for the request lifecycle.Explore InferX Architecture
    01Snapshot ReadyRETAINED
    02GPU RestoreATTACH
    03Burst ComputeACTIVE
    04Release GPUZERO
    05Snapshot RetainedREADY

    DEPLOYMENT

    Configure capacity around the workload—not idle time.

    Deploy a Sovereign Endpoint™