Xiaomi MiMo v2.5 TPM Ramp
Target TPM increased from 0.10M to 30.00M by enlarging prompt tokens per request.
BENCHMARKS
Every benchmark includes hardware, model, methodology, and reproducible assumptions.
Read methodologyEnd-to-end startup measured from request admission to a model-ready runtime.
Measured on production hardware. Full environmental fields remain listed below; unpublished fields are not inferred.
The time difference comes from changing the work performed before the first request can run.
Select a model class and accelerator to inspect published restore measurements.
GPU Snapshot Restore · ~70 GB FP16 model · NVIDIA H100
Model readiness is retained while active GPU capacity is released between bursts.
Published serving measurements include workload shape, request volume, failure behavior, and latency context.
Target TPM increased from 0.10M to 30.00M by enlarging prompt tokens per request.
Scaling efficiency requires matched model, interconnect, tensor-parallel, and request settings.
| GPU COUNT | TOPOLOGY | THROUGHPUT | SCALING EFFICIENCY | STATUS |
|---|---|---|---|---|
| 1 | Not published | — | — | Pending validated run |
| 2 | Not published | — | — | Pending validated run |
| 4 | Not published | — | — | Pending validated run |
| 8 | Not published | — | — | Pending validated run |
The current public record contains the fields below. Missing fields are shown explicitly.
| FIELD | VALUE | DISCLOSURE STATUS |
|---|---|---|
| Hardware | NVIDIA H100 | Published |
| Driver | Not published | Required for reproduction |
| CUDA | Not published | Required for reproduction |
| vLLM version | Not published | Required for reproduction |
| Precision | FP16 | Published |
| Model | Qwen3.6-35B-A3B | Published |
| Prompt | Not published | Required for reproduction |
| Batch size | Not published | Required for reproduction |
A benchmark is useful only when its boundaries and omissions are visible.
Measure startup separately from steady-state serving. Do not combine model initialization with token generation throughput.
Cold-start timing begins when a request enters the startup path and ends when the initialized model runtime is ready to serve.
Comparisons must use the same model, precision, accelerator class, and readiness definition.
Publish run count, warm-up policy, distribution, and outlier handling with every production dataset.
Unavailable driver, CUDA, framework, prompt, batch, and serving measurements remain explicitly unpublished.
Separate measured observations from illustrative scenarios and derived estimates.
The public harness and complete run manifest will be linked here when the remaining environment fields are published.
# Public benchmark harness status: pending publication # Required run manifest model: Qwen3.6-35B-A3B precision: FP16 hardware: NVIDIA H100