1bit.MONSTERDocs GitHub โ†—

๐Ÿ“œ Historical launch post โ€” This was written at project launch (July 2026). Binary size, model counts, and tok/s figures are from the early NPU-only phase. Current numbers are significantly different. See README for up-to-date data.

Launch Plan โ€” 1bit.MONSTER

The Show HN post that makes 1bit.MONSTER the reference standard for NPU inference.

One-Liner Pitch

120 KB binary. 94 tok/s on AMD's NPU. Zero Python. Zero dependencies. MIT license. I reverse-engineered AMD's proprietary NPU stack in 4 days. Here's what I built.

Target Audience

Timing

Day Action
T-7 Final polish โ€” docs, README, release
T-3 Post Show HN draft to Discord for feedback
T-1 Push final release with packages attached
T-0 7:00 AM PT / 14:00 UTC โ€” Post Show HN
T+0 Monitor comments, respond authentically
T+1 Follow-up: Post to r/LocalLLaMA and AMD community
T+3 Update docs based on feedback, fix first bugs

Show HN Draft

Title

I reverse-engineered AMD's NPU stack in 4 days. Here's a 120 KB binary that runs LLMs at 94 tok/s.

Alternative titles:

Body

Hi HN,

I reverse-engineered AMD's XDNA 2 NPU stack and built an open-source
inference engine in C++23. It's 120 KB. Zero Python. Zero dependencies.
One binary. Runs 5 LLMs on the same chip you already have.

Why this matters:

AMD's NPU is rated for 50 TOPS INT8. FastFlowLM (their proprietary
runtime) gets ~94 tok/s on Qwen3-0.6B. My engine matches that โ€” but
it's 120 KB, MIT-licensed, and compiles with one g++ command.

No Python. No pip. No Docker. No MLIR toolchain. Just g++ and run.

curl -sL https://1bit.monster/install.sh | bash


What I learned reverse-engineering the NPU:

1. AMD's MLIR toolchain is 2+ GB and requires a proprietary Chess
   compiler license. It generates xclbins (NPU instruction blobs)
   from C++ kernels. I bypassed all of it.

2. The XRT runtime is open-source but undocumented. I linked against
   libxrt_coreutil directly โ€” 3 functions is all you need to submit
   work to the NPU.

3. The real breakthrough was batch decoding. Single-token decode on
   the NPU is 244 ms/tok (4 tok/s). But at batch size 16, it's
   16 ms/tok (63 tok/s). Batch 32: 36 ms/tok for all models at once.

Current performance:

  | Engine           | Speed          | Models           |
  |------------------|----------------|------------------|
  | GGML-Vulkan      | 373 tok/s      | Qwen3-0.6B Q4_K  |
  | GGML-Vulkan      | 662 tok/s      | SmolLM2-135M Q4_K|
  | GPU (HIP kernels)| 157 GB/s GEMV  | 1-bit ternary     |
  | NPU (FLM proxy)  | 94 tok/s (hist.)| Qwen3-0.6B       |

The binary auto-detects which model you have and dispatches the
right xclbin. No recompilation per model.

Tech stack:
- C++23 NPU engine โ†’ XRT โ†’ XDNA 2 NPU
- C++ GPU engine โ†’ HIP + Vulkan โ†’ Radeon 8060S
- Single binary, MIT license

If you have a Ryzen AI Max+ 395 (Strix Halo), you can run this
right now:

curl -sL https://1bit.monster/npu-install.sh | bash 1bit-npu model.q4nx 16


The repo: https://github.com/1bit-MONSTER/1bit-MONSTER

I've been working on this for 2.5 months. Happy to answer questions
about the NPU, XRT, INT8 quantization, or why I chose C++ over
everything else.

โ€” bong-water-water-bong

Launch Checklist

Post-Launch Priorities

Timeframe Action
T+0โ€“2h Respond to every HN comment personally
T+2โ€“24h Fix first bugs from feedback, push hotfix release
T+24โ€“72h Post to Reddit communities
T+72h Write blog post: "What I learned reverse-engineering the AMD NPU"
T+1 week Add Windows support if there's demand
T+2 weeks Benchmarks from other Strix Halo users (community validation)
T+1 month Speculative decode (Phase 2 from roadmap)

Key Messages to Reinforce