1bit.MONSTERDocs GitHub ↗

Zyphra — Complete End-to-End Pipeline

Zyphra's portfolio spans the entire AI stack: EEG → LLM (dense, MoE, Mamba) → TTS → voice cloning. 1bit supports all of it — 11 models in 1BP format, plus EEG and TTS pipelines documented for ecosystem completeness. This is the family the engine was tuned against, and the one that powers the JARVIS pipeline.

Legend: 🧠 LLM · 👁️ vision · 🗣️ voice · 🧬 EEG · 🏁 end-to-end validated

Models

Model Params 1BP Size Backend(s) Pipeline Perf
ZAYA1-8B 8.8B 6.6 GB¹ HIP / NPU 🧠🗣️ 64 tok/s HIP
ZAYA1-74B-preview 74B 45.8 GB² HIP 🧠🗣️ 16.7 tok/s HIP³ (measured 2026-08-05)
ZAYA1-VL-8B 8.8B HIP (vision) 👁️🧠🗣️
ZR1-1.5B 1.5B 781 MB ZINC / NPU 🧠🗣️ 26 tok/s ZINC
BlackMamba-1.5B 1.5B 970 MB Mamba1 HIP 🧠🗣️ 79.4 tok/s 🏁
BlackMamba-2.8B 2.8B 1.8 GB Mamba1 HIP 🧠🗣️ 46.0 tok/s 🏁
Zamba2-1.2B-v2 1.2B 1.1 GB HIP / CPU 🧠 30 tok/s HIP
Zamba2-2.7B-v2 2.7B 2.4 GB HIP / CPU 🧠
Zamba2-7B-v2 7B 6.6 GB HIP / CPU 🧠
Zamba-7B-v1 7B 4.3 GB Mamba1 HIP 🧠

¹ ZAYA1-8B 1BP is ~6.6 GB full-weight — the 149 MB entry on HF is MoE-expert-stripped; use the complete file. ² Earlier catalogs listed a 739 MB 1BP for the 74B — that is physically impossible for a full 74B model (even ternary ≈ 18 GB) and refers only to a stripped-expert variant; the runnable 74B Q4_K_M GGUF is 45.8 GB. ³ 74B measured on ROCm TheRock HIP 7.15a, zaya-llama.cpp/build-hip (Juste-Leo2 Zaya branch), full GPU offload — see benchmarks/RESULTS-zaya1-74b-benchmark-2026-08-05.md.

Architectures in this family

Full pipeline depth

1BP catalog

Model 1BP Size Verified
ZAYA1-8B, ZAYA1-74B-preview 6.6 GB / 45.8 GB ✅ loads
ZR1-1.5B 373 MB hosted
Zamba2-1.2B / 2.7B / 7B v2 1.15 – 7.25 GB hosted
BlackMamba-1.5B / 2.8B 1.0 / 1.9 GB ✅ loads

See also: full model support detail · benchmarks SSOT · all families