ENDPOINT PRICING

    Hosted models. Usage-based access.

    Compare confirmed production pricing for InferX hosted model endpoints.

    MONTHLY ACCESS

    An easy way to try InferX.

    Beta offer: $10/month, includes $50 in usage credits.

    Start using InferX with included monthly credits. Explore production endpoints, models, and APIs without committing to infrastructure. If you use all included credits, continue seamlessly with pay-per-usage pricing.

    MONTHLY ACCESS$20 / month $10 / month
    INCLUDED USAGE CREDITS$50 every billing cycle
    AFTER INCLUDED CREDITSPay per usage
    Start for $10
    Production-ready API modelsOpenAI-compatible hosted endpoints
    ModelInput / 1MOutput / 1MCached input / 1MContextAction
    AgentsAgents A1Independent
    Input / 1M
    $0.12
    Output / 1M
    $0.90
    Cached input / 1M
    $0.024
    Context
    262,000
    Use Endpoint
    Devstral 2 123BMistral AI
    Input / 1M
    $0.35
    Output / 1M
    $1.80
    Cached input / 1M
    $0.07
    Context
    128,000
    Use Endpoint
    OrnithOrnith 1.0 35BOrnith
    Input / 1M
    $0.14
    Output / 1M
    $1.00
    Cached input / 1M
    $0.028
    Context
    262,000
    Use Endpoint
    Qwen3 Coder NextAlibaba Cloud
    Input / 1M
    $0.18
    Output / 1M
    $0.90
    Cached input / 1M
    $0.036
    Context
    256,144
    Use Endpoint
    Qwen3 Coder Next — No ThinkingAlibaba Cloud
    Input / 1M
    $0.15
    Output / 1M
    $0.75
    Cached input / 1M
    $0.03
    Context
    260,000
    Use Endpoint
    Qwen3 Embedding 8BAlibaba Cloud
    Input / 1M
    $0.01
    Output / 1M
    $0.00
    Cached input / 1M
    $0.002
    Context
    Use Endpoint
    Qwen3.6 27BAlibaba Cloud
    Input / 1M
    $0.25
    Output / 1M
    $2.80
    Cached input / 1M
    $0.05
    Context
    262,144
    Use Endpoint
    Qwen3.6 35B A3BAlibaba Cloud
    Input / 1M
    $0.12
    Output / 1M
    $0.90
    Cached input / 1M
    $0.024
    Context
    262,000
    Use Endpoint
    Qwen3.6 35B A3B — No ThinkingAlibaba Cloud
    Input / 1M
    $0.10
    Output / 1M
    $0.75
    Cached input / 1M
    $0.02
    Context
    262,000
    Use Endpoint
    MimoMimo v2.5Xiaomi
    Input / 1M
    $0.09
    Output / 1M
    $0.19
    Cached input / 1M
    $0.015
    Context
    1,000,000
    Use Endpoint
    Hy3Tencent
    Input / 1M
    $0.14
    Output / 1M
    $0.58
    Cached input / 1M
    $0.035
    Context
    Use Endpoint
    DeepSeek V4 FlashDeepSeek
    Input / 1M
    $0.09
    Output / 1M
    $0.19
    Cached input / 1M
    $0.01
    Context
    1,000,000
    Use Endpoint

    Prices and model metadata are maintained in the InferX marketing catalog. Beta pricing: DeepSeek V4 currently reflects 50% off; Mimo v2.5 currently reflects 40% off. Dedicated endpoint pricing depends on the selected model and GPU configuration.