Qwen 3.8 Distilled
Runs on your deviceempero-ai · Aug 2026 · Distilled from Qwen 3.8
Empero's full-parameter distillations of the Qwen 3.8 flagship into small, fast sizes - a big knowledge jump over same-size models (the 9B scores near the giants on broad knowledge), with step-by-step reasoning. Text only.
- Context
- 256K tokens
- Smallest download
- 1.3 GB
- Runs from
- 4 GB RAM
- License
- Apache 2.0
Will Qwen 3.8 Distilled run on your machine?
This is the same sizing the app uses - not a marketing estimate.
On Apple Silicon, system memory is also graphics memory - pick "None / integrated / Apple Silicon" and read the bars against your Mac's memory. The app checks your real hardware before every download.
Variants & sizes
| Variant | Download | Quantization | Needs | Apple Silicon |
|---|---|---|---|---|
| 2B | 1.3 GB | Q4_K_M | 4 GB RAM | Metal (GGUF) |
| 4B | 2.8 GB | Q4_K_M | 8 GB RAM | Metal (GGUF) |
| 9B | 5.8 GB | Q4_K_M | 16 GB RAM | Metal (GGUF) |
The app picks the right variant for your machine and verifies every download. Model files come straight from their official repositories, pinned to exact revisions.
Context window, compared
License
Qwen 3.8 Distilled is released under the Apache 2.0 - read the terms. The weights download to your machine and stay there; Your Own AI adds no accounts, tracking, or lock-in on top.
Frequently asked questions
- Can I run Qwen 3.8 Distilled offline?
- Yes. Qwen 3.8 Distilled is a downloadable model: with Your Own AI it runs entirely on your computer, works with no internet connection, and your conversations never leave your device.
- What hardware does Qwen 3.8 Distilled need?
- The smallest variant (2B, a 1.3 GB download) runs on a machine with 4 GB of memory. Larger variants want more memory or a graphics card - the interactive checker on this page grades every variant against your hardware, using the same logic as the app.
- Is Qwen 3.8 Distilled free to use?
- The model weights are released under the Apache 2.0 license and download at no charge. Your Own AI itself is free and open source for local use.
- How much context does Qwen 3.8 Distilled support?
- 256K tokens of context.
Run Qwen 3.8 Distilled in two clicks
Download Your Own AI, pick Qwen 3.8 Distilled, and the app sizes it to your hardware. Private by architecture - no account needed for local use.