GLM-4.7 Flash

Runs on your device

Zhipu (Z.ai) · Jan 2026

Zhipu's fast GLM model — a strong all-rounder tuned for agentic coding, reasoning, and tool use. The lighter "Flash" tier of the GLM family; the 30B MoE runs only 3B parameters per token.

Context
198K tokens
Smallest download
18.3 GB
Runs from
32 GB RAM
License
MIT
codingagenticreasoningMLX for Apple Silicon

Will GLM-4.7 Flash run on your machine?

This is the same sizing the app uses - not a marketing estimate.

30B-A3B (MoE)18.3 GB downloadRuns - experts in RAM, the rest on the GPU

Mixture-of-experts: the always-on layers sit on your card, the experts sit in system memory - bigger than your card, still quick.

On Apple Silicon, system memory is also graphics memory - pick "None / integrated / Apple Silicon" and read the bars against your Mac's memory. The app checks your real hardware before every download.

Variants & sizes

VariantDownloadQuantizationNeedsApple Silicon
30B-A3B (MoE)MoE18.3 GBQ4_K_M32 GB RAMMLX · 16.9 GB

The app picks the right variant for your machine and verifies every download. Model files come straight from their official repositories, pinned to exact revisions.

License

GLM-4.7 Flash is released under the MIT - read the terms. The weights download to your machine and stay there; Your Own AI adds no accounts, tracking, or lock-in on top.

Frequently asked questions

Can I run GLM-4.7 Flash offline?
Yes. GLM-4.7 Flash is a downloadable model: with Your Own AI it runs entirely on your computer, works with no internet connection, and your conversations never leave your device.
What hardware does GLM-4.7 Flash need?
The smallest variant (30B-A3B (MoE), a 18.3 GB download) runs on a machine with 32 GB of memory. Larger variants want more memory or a graphics card - the interactive checker on this page grades every variant against your hardware, using the same logic as the app.
Is GLM-4.7 Flash free to use?
The model weights are released under the MIT license and download at no charge. Your Own AI itself is free and open source for local use.
How much context does GLM-4.7 Flash support?
198K tokens of context.

Run GLM-4.7 Flash in two clicks

Download Your Own AI, pick GLM-4.7 Flash, and the app sizes it to your hardware. Private by architecture - no account needed for local use.