1bit.MONSTERDocs GitHub ↗

Llama — Meta Dense Transformers

Meta's dense transformers with broad backend coverage — GGUF runs on every GPU backend and CPU, with NPU support for the 1B/3B/8B tier.

Models

Model Params 1BP Size Backend(s) Perf
Llama-3.2-1B 1B 581 MB GGML-Vulkan / ZINC / NPU
Llama-3.2-3B 3B 1.7 GB GGML-Vulkan / ZINC / NPU / HIP
Llama-3.1-8B 8B 4.1 GB GGML-Vulkan / ZINC / NPU / HIP
TinyLlama-1.1B 1.1B 328 MB GGML-Vulkan / ZINC / NPU

Notes

See also: full model support detail · benchmarks SSOT · all families