Table of Contents

Property Llama32_1BInstructInt4

Namespace
Qavren.Edge.Chat
Assembly
Qavren.Edge.Chat.Onnx.dll

Llama32_1BInstructInt4

Arm/llama-3-2-1b-instruct-onnx-genai-int4-kquantlast-emb-int8-vivo-x300. 1.241 GB decimal on disk across six files; 1167.52 MiB of resident weights; context 4096 as shipped (the publisher patched it down from Llama's stock 131072 so the KV cache fits a phone). Geometry: 16 layers, 8 KV heads, head size 64, vocab 128256 - 32 KiB of KV per token. Measured on a vivo X300 (Android 16, ORT 1.27.0 CPU EP, 4 threads): 30.564 tok/s decode, TTFT 950 ms, peak RSS 1280 MiB, model load 1.61 s, MMLU 5-shot 44.2%. No iOS figure exists for this or any other ORT GenAI model.

public static ChatPreset Llama32_1BInstructInt4 { get; }

Property Value

ChatPreset