PRICING

    Pricing built around how you run inference.

    Choose a path: subscribe to the Developer Plan, add usage credits as needed, or design a production deployment with the InferX team.

    FREE TIER

    Free Tier

    Free models available

    $0 / month
    • Great for testing and iterating
    • OpenAI-compatible APIs
    • Zero data retention
    • No credit card required
    Start free
    PAY AS YOU GO

    Pay As You Go

    Add prepaid credits anytime

    No subscriptionprepaid usage credits
    • Use credits across models and endpoints
    • No monthly commitment
    • Good for testing, demos, or occasional usage
    • Upgrade to Developer Plan anytime
    Add usage credits
    Need enterprise solutions?Private deployments, dedicated capacity, and custom pricing for production teams.

    INTRODUCTORY MODEL PRICING

    Introductory pricing compared with common market API rates.

    Representative hosted endpoint rates for popular ready-now models.
    ModelMarket rateInferX rateCached inputSavings
    DeepSeek V4 FlashInput $0.14 / 1M
    Output $0.28 / 1M
    Input $0.07 / 1M
    Output $0.11 / 1M
    $0.01 / 1M50% input61% output

    InferX supports cached-input pricing for repeated context and cache-hit workloads.

    SERVERLESS ECONOMICS

    Serverless compute vs always-on GPU hosting.

    Pay for active inference compute instead of idle capacity.
    THE OLD WAY

    Dedicated H100

    Always on. Always billing.

    Hourly rate
    $4.00 / hr
    Hours / month
    730 hrs
    Calculation
    $4.00 × 730
    $2,900per month

    Paying for idle capacity.

    THE INFERX WAY

    InferX Serverless

    Active compute only. Scale to zero between requests.

    Active GPU time
    $3.95 / GPU hr
    Standby cost
    Scale to zero
    Reference workload
    Bursty usage
    ~$220per month

    Pay only for active compute.

    ~13×lower reference monthly cost$2,900 → ~$220for the same production workload pattern
    Illustrative comparison based on the original InferX economics diagram. Costs vary by model, GPU class, active compute time, and deployment configuration.
    SERVERLESS COMPUTE

    True Serverless Compute

    Deploy your custom model with no idle GPU time. InferX bills active compute by the second and uses GPU Snapshot Restore for sub-second cold starts when requests arrive.

    H100$3.95 / GPU-hour about $0.00110 / GPU-second
    H200$3.95 / GPU-hour about $0.00110 / GPU-second
    B300$6.95 / GPU-hour about $0.00193 / GPU-second
    This is for custom deployments / dedicated inference endpoints. Hosted pay-per-token API pricing is separate.

    PRICING FAQ

    Simple access, clear billing.

    What is included in the Developer Plan?

    The Developer Plan is $1/month and includes $5 in usage credits once per month, plus access to discounted InferX pricing.

    Can I use InferX without a subscription?

    Yes. Pay As You Go lets you add usage credits when you need them without a monthly subscription.

    What is available for Enterprise customers?

    Enterprise plans can include private deployments, higher capacity, support, and custom pricing.

    Do unused credits roll over?

    Unused credits do not currently roll over.