Free Tier
Free models available
- Great for testing and iterating
- OpenAI-compatible APIs
- Zero data retention
- No credit card required
PRICING
Choose a path: subscribe to the Developer Plan, add usage credits as needed, or design a production deployment with the InferX team.
Free models available
Includes $5 usage credits
Add prepaid credits anytime
INTRODUCTORY MODEL PRICING
| Model | Market rate | InferX rate | Cached input | Savings |
|---|---|---|---|---|
| DeepSeek V4 Flash | Input $0.14 / 1M Output $0.28 / 1M | Input $0.07 / 1M Output $0.11 / 1M | $0.01 / 1M | 50% input61% output |
InferX supports cached-input pricing for repeated context and cache-hit workloads.
SERVERLESS ECONOMICS
Always on. Always billing.
Paying for idle capacity.
Active compute only. Scale to zero between requests.
Pay only for active compute.
Deploy your custom model with no idle GPU time. InferX bills active compute by the second and uses GPU Snapshot Restore for sub-second cold starts when requests arrive.
Hosted production models with usage measured by input and output tokens.
Dedicated model capacity priced around model and GPU configuration.
Commercial licensing for customer-operated and provider infrastructure.
Confirmed information about billing, commitments, credits, and support.
PRICING FAQ
The Developer Plan is $1/month and includes $5 in usage credits once per month, plus access to discounted InferX pricing.
Yes. Pay As You Go lets you add usage credits when you need them without a monthly subscription.
Enterprise plans can include private deployments, higher capacity, support, and custom pricing.
Unused credits do not currently roll over.