1bit.MONSTERDocs GitHub ↗

DeepSeek — MoE with Multi-Head Latent Attention

DeepSeek's MoE family uses Multi-Head Latent Attention (MLA). The full V2/V3/R1 family runs through GPU HIP; distilled variants (Llama- and Qwen-based) add NPU + Vulkan coverage.

Models

Model Params 1BP Size Backend(s) Perf
DeepSeek-V2 / V3 / R1 up to 671B GPU HIP 20 tok/s
DeepSeek-R1-Distill-Llama-8B 8B 4.1 GB GGML-Vulkan / ZINC / NPU / HIP 44 tok/s
DeepSeek-R1-Distill-Qwen-7B 7B 3.8 GB ZINC / NPU / HIP

Notes

See also: full model support detail · benchmarks SSOT · all families