All models
Alibaba CloudReady now
Qwen3.6 35B A3B
A mixture-of-experts model balancing large-model intelligence with small-model speed.
Context length262,000
GPU requirement1 GPU
Concurrency4.80x
API model IDQwen3.6-35B-A3B-FP8
Input pricing$0.12 / 1M
Output pricing$0.90 / 1M
Cached input pricing$0.024 / 1M
Recommended use cases
Where this model fits.
Supported features
Ready for production.
OpenAI compatible
Scale to zero
API example
Use the interface you already know.
Every published text endpoint uses an OpenAI-compatible request surface.
curl
curl https://api.inferx.com/v1/chat/completions \
-H "Authorization: Bearer $INFERX_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "Qwen3.6-35B-A3B-FP8", "messages": [...] }'