Table of Contents

Property Qwen3_600MInt4

Namespace
Qavren.Edge.Chat
Assembly
Qavren.Edge.Chat.Onnx.dll

Qwen3_600MInt4

Arm/qwen3-0-6b-onnx-genai-int4-kquantlast-emb-int4. 495 MB decimal on disk across six files; 461.25 MiB of resident weights; Apache-2.0 - the preset for an app that cannot take the Llama terms, and the model the nightly lane uses.

Smaller on disk is not smaller in memory, and here are the numbers. 28 layers, 8 KV heads, head size 128 - 112 KiB of KV per token, 3.5x this catalogue's larger preset, against a third of the weights. At 4096 tokens its KV cache alone is 448 MiB. Its declared context_length is 40960, so a generator built with no max_length would allocate 4480 MiB of KV cache - which is exactly the jetsam scenario the memory budget and the belt-and-braces search.max_length exist to prevent.

Its only published throughput is AWS Graviton (70.4 tok/s, peak 838 MB decimal), and that figure is deliberately not in MeasuredPeakBytes: the budget reads that field alone and has no way to discount a measurement by where it was taken, so a server peak entering it would silently become a phone budget.

public static ChatPreset Qwen3_600MInt4 { get; }

Property Value

ChatPreset