Property Qwen3_600MInt4
Qwen3_600MInt4
Arm/qwen3-0-6b-onnx-genai-int4-kquantlast-emb-int4. 495 MB decimal on disk across six
files; 461.25 MiB of resident weights; Apache-2.0 - the preset for an app that cannot take
the Llama terms, and the model the nightly lane uses.
Smaller on disk is not smaller in memory, and here are the numbers. 28 layers, 8 KV
heads, head size 128 - 112 KiB of KV per token, 3.5x this catalogue's larger preset,
against a third of the weights. At 4096 tokens its KV cache alone is 448 MiB. Its declared
context_length is 40960, so a generator built with no max_length would
allocate 4480 MiB of KV cache - which is exactly the jetsam scenario the memory budget
and the belt-and-braces search.max_length exist to prevent.
Its only published throughput is AWS Graviton (70.4 tok/s, peak 838 MB decimal), and that
figure is deliberately not in MeasuredPeakBytes: the budget reads that field
alone and has no way to discount a measurement by where it was taken, so a server peak
entering it would silently become a phone budget.
public static ChatPreset Qwen3_600MInt4 { get; }