ENDPOINT PRICING

    Hosted models. Usage-based access.

    Compare confirmed production pricing for InferX hosted model endpoints.

    INFERX FREE

    Start with free models on InferX.

    InferX Free gives developers access to free models running on enterprise-grade GPUs, with OpenAI-compatible APIs and Zero Data Retention.

    Try InferX, test endpoints, and move to pay-per-token models or higher limits when your workload is ready.

    ENTRY TIERFree models available
    APIOpenAI-compatible
    DATA POLICYZero Data Retention
    Start Free
    Production-ready API modelsOpenAI-compatible hosted endpoints
    ModelInput / 1MOutput / 1MCached input / 1MContextAction
    GLM 5.3 FlashZhipu AI
    Input / 1M
    $0.30$0.09New · 70% off
    Output / 1M
    $1.00$0.30New · 70% off
    Cached input / 1M
    $0.03$0.009New · 70% off
    Context
    Use Endpoint
    Qwen 3.8 27BAlibaba Cloud
    Input / 1M
    $0.45$0.01896% discount
    Output / 1M
    $3.20$0.12896% discount
    Cached input / 1M
    Context
    Use Endpoint
    AgentsAgents A1Independent
    Input / 1M
    $0.12
    Output / 1M
    $0.90
    Cached input / 1M
    $0.024
    Context
    262,000
    Use Endpoint
    Devstral 2 123BMistral AI
    Input / 1M
    $0.35$0.1460.00% off
    Output / 1M
    $1.80$0.7260.00% off
    Cached input / 1M
    $0.035$0.01460.00% off
    Context
    128,000
    Use Endpoint
    OrnithOrnith 1.0 35BOrnith
    Input / 1M
    $0.14
    Output / 1M
    $1.00
    Cached input / 1M
    $0.028
    Context
    262,000
    Use Endpoint
    Qwen3 Coder NextAlibaba Cloud
    Input / 1M
    $0.18
    Output / 1M
    $0.90
    Cached input / 1M
    $0.036
    Context
    256,144
    Use Endpoint
    Qwen3 Coder Next — No ThinkingAlibaba Cloud
    Input / 1M
    $0.15
    Output / 1M
    $0.75
    Cached input / 1M
    $0.03
    Context
    260,000
    Use Endpoint
    Qwen3.6 27BAlibaba Cloud
    Input / 1M
    $0.25$0Free · 100% discount
    Output / 1M
    $2.80$0Free · 100% discount
    Cached input / 1M
    $0.05$0Free · 100% discount
    Context
    262,144
    Use Endpoint
    Qwen3.6 35B A3BAlibaba Cloud
    Input / 1M
    $0.12$0Free · 100% discount
    Output / 1M
    $0.90$0Free · 100% discount
    Cached input / 1M
    $0.024$0Free · 100% discount
    Context
    262,000
    Use Endpoint
    Qwen3.6 35B A3B — No ThinkingAlibaba Cloud
    Input / 1M
    $0.10
    Output / 1M
    $0.75
    Cached input / 1M
    $0.02
    Context
    262,000
    Use Endpoint
    Gemma 4 31B IT FP8Google
    Input / 1M
    $0.11$0Free · 100% off
    Output / 1M
    $0.35$0Free · 100% off
    Cached input / 1M
    $0.011$0Free · 100% off
    Context
    262,144
    Use Endpoint
    GPT-OSSGPT-OSS 20BOpenAI
    Input / 1M
    $0.07$0Free · 100% off
    Output / 1M
    $0.25$0Free · 100% off
    Cached input / 1M
    $0.03$0Free · 100% off
    Context
    20,000
    Use Endpoint
    DeepSeek V4 Flash 0731DeepSeek
    Input / 1M
    $0.14$0.042New · 70% off
    Output / 1M
    $0.28$0.084New · 70% off
    Cached input / 1M
    $0.028$0.0084New · 70% off
    Context
    1,000,000
    Use Endpoint
    DeepSeek V4.1 FlashDeepSeek
    Input / 1M
    $0.30$0.1260% off
    Output / 1M
    $1.20$0.4860% off
    Cached input / 1M
    $0.006$0.002460% off
    Context
    Use Endpoint

    Prices and model metadata are maintained in the InferX marketing catalog. Free promotional models show their original rate crossed out with the active InferX price. Dedicated endpoint pricing depends on the selected model and GPU configuration.