PRICING

    Pricing built around how you run inference.

    Choose a path: start with InferX Free, add usage credits as needed, or deploy the InferX Platform with your infrastructure team.

    INFERX FREE

    InferX Free

    Start building with free models on InferX. Run inference on enterprise-grade GPUs with OpenAI-compatible APIs and Zero Data Retention.

    $0 / month
    • Free models available
    • Runs on enterprise-grade GPUs
    • OpenAI-compatible API
    • Zero Data Retention
    • Good for trying InferX and testing endpoints
    Start Free
    PAY AS YOU GO

    Pay As You Go

    Add prepaid credits anytime

    No subscriptionprepaid usage credits
    • Use credits across models and endpoints
    • No monthly commitment
    • Good for testing, demos, or occasional usage
    • Upgrade from InferX Free anytime
    Sign in to dashboard to add usage credits
    ENTERPRISE

    Enterprise Platform

    End-to-end deployment of your inference stack.

    Customplatform and on-prem deployment
    • Secure container runtime
    • Sub-second cold starts
    • Run multiple models on the same GPU
    • Plug-and-play deployment
    • Runs on your existing Kubernetes stack
    • 50-70% infrastructure savings*
    • Private cloud, on-prem, bare metal, and GPU cloud support
    • OpenAI-compatible APIs

    *Savings depend on workload patterns, model size, traffic shape, and existing GPU utilization.

    SERVERLESS ECONOMICS

    Serverless compute vs always-on GPU hosting.

    Pay for active inference compute instead of idle capacity.
    THE OLD WAY

    Dedicated H100

    Always on. Always billing.

    Hourly rate
    $4.00 / hr
    Hours / month
    730 hrs
    Calculation
    $4.00 × 730
    $2,900per month

    Paying for idle capacity.

    THE INFERX WAY

    InferX Serverless

    Active compute only. Scale to zero between requests.

    Active GPU time
    $3.95 / GPU hr
    Standby cost
    Scale to zero
    Reference workload
    Bursty usage
    ~$220per month

    Pay only for active compute.

    ~13×lower reference monthly cost$2,900 → ~$220for the same production workload pattern
    Illustrative comparison based on the original InferX economics diagram. Costs vary by model, GPU class, active compute time, and deployment configuration.
    SERVERLESS COMPUTE

    True Serverless Compute

    Deploy your custom model with no idle GPU time. InferX bills active compute by the second and uses GPU Snapshot Restore for sub-second cold starts when requests arrive.

    H100$3.95 / GPU-hour about $0.00110 / GPU-second
    H200$3.95 / GPU-hour about $0.00110 / GPU-second
    B300$6.95 / GPU-hour about $0.00193 / GPU-second
    This is for custom deployments / dedicated inference endpoints. Hosted pay-per-token API pricing is separate.

    PRICING FAQ

    Simple access, clear billing.

    What is included in InferX Free?

    InferX Free gives developers access to free models running on enterprise-grade GPUs, with OpenAI-compatible APIs and Zero Data Retention.

    Can I use InferX without a subscription?

    Yes. Pay As You Go lets you add usage credits when you need them without a monthly subscription.

    What is available for Enterprise Platform customers?

    Enterprise Platform includes end-to-end deployment of the InferX runtime across private cloud, on-prem, bare metal, GPU cloud, or existing Kubernetes infrastructure.

    Do unused credits roll over?

    Unused credits do not currently roll over.