Products
AI Infrastructure for Every Deployment Model.
From a single API request to a sovereign platform inside your own cluster, InferX gives every team a path to production.
01 / Product
Endpoints
Run production AI models instantly through OpenAI-compatible endpoints with pay-per-usage pricing.
- OpenAI compatible
- Pay per usage
- Featured open models
- Streaming responses
APIMODELTOKENS
02 / Product
Sovereign Endpoints™
Deploy from 200+ models or bring your own model with zero-ops dedicated inference.
- Dedicated infrastructure
- Private networking
- Bring your own model
- Sub-second cold starts
VPCPRIVATEMODEL
03 / Product
InferX On-Prem
Deploy the InferX Platform inside your own infrastructure.
- Deploy inside Kubernetes
- Private cloud
- On-prem
- Multi-tenant runtime
CLUSTERRUNTIMEGPU
04 / Product
Skill Function
Protected callable skills with isolated context and the right model for each skill.
- Composable AI capabilities
- Reusable skills
- Agent orchestration
- Schema-driven execution
AGENTSKILLACTION
Compare
Choose the right operating model.
The same InferX runtime, delivered with the level of control your workload requires.
| Endpoints | Sovereign Endpoints™ | InferX On-Prem | Skill Function | |
|---|---|---|---|---|
| Deployment | Shared cloud | Dedicated cloud | Your environment | Managed skills |
| Billing | Per token | Reserved capacity | Platform license | Per invocation |
| Networking | Public API | Private VPC | Customer network | Secure endpoint |
| Operations | InferX managed | InferX managed | Customer controlled | InferX managed |
| Best for | Fast iteration | Production isolation | Sovereign control | Agent workflows |
