Models that run on your device

46 downloadable variants across 24 model families - from 2 GB laptop models to frontier-class mixtures of experts. Everything here runs fully offline, keeps your conversations on your machine, and needs no account for local use.

Muse GlimmerRecommended

Meta · Aug 2026

Meta's new agent-first model - built for reliable tool use, working in project folders, and recovering from its own mistakes. Sees images too. Brand-new: early support, expect rough edges.

30Bfrom 12.4 GBruns from 16 GB RAM128K contextsees imagesMLXApache 2.0

Qwen 3.8Recommended

Alibaba's Qwen team · Aug 2026

Alibaba's newest. Frontier coding, math, and reasoning with built-in thinking; a hybrid attention design keeps long documents fast. The strongest model here for 24GB-class graphics cards.

27Bfrom 16.5 GBruns from 32 GB RAM256K contextsees imagesMLXApache 2.0

Qwen 3.6Recommended

Alibaba's Qwen team · Apr 2026

Top-tier coding, math, and reasoning. The 35B MoE runs fast for its size — only 3B parameters active per token.

27B · 35B-A3B (MoE)from 12.3 GBruns from 24 GB RAM256K contextMLXApache 2.0

Gemma 4Recommended

Google · Apr 2026

Google's latest open model. Strong writing, analysis, and instruction-following. E2B/E4B run on modest machines; the 26B is a fast MoE.

E2B · E4B · 12B · 26B-A4B (MoE) · 31Bfrom 3.1 GBruns from 8 GB RAM128K contextsees imagesMLXApache 2.0 (Gemma 4 terms)

Ministral 3

Mistral AI · Dec 2025

Mistral's latest small model family. Fast inference, great for on-device use.

3B · 8B · 14Bfrom 2.2 GBruns from 8 GB RAM256K contextMLXApache 2.0

Phi-4 MiniRecommended

Microsoft · Feb 2025

Microsoft's compact model. Best option for machines with limited RAM.

3.8Bfrom 2.5 GBruns from 8 GB RAM128K contextMLXMIT

GPT-OSS (OpenAI)

OpenAI · Aug 2025

OpenAI's open-weight model. Strong reasoning, agentic tasks, and function calling.

20B (3.6B active) · 120B (5.1B active)from 11.6 GBruns from 16 GB RAM128K contextApache 2.0

LFM2.5Recommended

Liquid AI · May 2026

Liquid AI's small mixture-of-experts - 8B total, 1.5B active per token, so it runs fast on ordinary machines. Strong instruction following and tool use for its size; 10 languages.

8B-A1B (MoE)from 5.2 GBruns from 12 GB RAM125K contextLFM Open License v1.0

Granite 4.0 TinyRecommended

IBM · Sep 2025

IBM's small hybrid mixture-of-experts - 7B total, about 1B active per token, with a 1M-token context. Apache-2.0. A fast, permissive everyday model for long documents.

7B-A1B (MoE)from 4.2 GBruns from 8 GB RAM1M contextApache 2.0

Nemotron 3.5 LightningRecommended

NVIDIA · Aug 2026

NVIDIA's newest open reasoning model - a 30B mixture-of-experts with 3B active per token and a 256K context. Strong reasoning, coding, and tool use; runs on a 32 GB machine with the experts in main memory.

30B-A3B (MoE)from 18.9 GBruns from 32 GB RAM256K contextOpenMDW 1.1

Ling-mini 2.0

Ant Group (inclusionAI) · Sep 2025

Ant Group's mid-size mixture-of-experts - 16B total, 1.4B active per token. MIT-licensed, 128K context; a quick all-rounder for 24 GB machines.

16B-A1.4B (MoE)from 9.9 GBruns from 24 GB RAM128K contextMIT

DeepSeek V4 Flash

DeepSeek · Jul 2026

DeepSeek's V4 Flash - a 284B mixture-of-experts with 13B active per token and a 1M-token context. Frontier-class reasoning and agentic work on a workstation with a big graphics card and 128 GB or more of main memory.

284B-A13B (MoE)from 96.9 GBruns from 128 GB RAM1M contextMIT

GLM-5.2

Zhipu AI · Jun 2026

Zhipu's GLM-5.2 - a 753B mixture-of-experts with a 1M-token context. For very large workstations: a big graphics card and 384 GB or more of main memory, experts in main memory.

753B (MoE)from 253.9 GBruns from 384 GB RAM1M contextMIT

DeepSeek R1 0528

DeepSeek · May 2025

Powerful chain-of-thought reasoning. Excels at complex problem solving, math, and logic.

8Bfrom 5.0 GBruns from 16 GB RAM128K contextMIT

Devstral Small 2

Mistral AI · Dec 2025

Mistral's agentic coding model. Tops open models on SWE-bench at its size — purpose-built for software engineering and code agents.

24Bfrom 14.3 GBruns from 16 GB RAM384K contextApache 2.0

Ornith 1.5

DeepReinforce · Aug 2026

DeepReinforce's agentic coding model, self-improved with RL - 1.5 extends the loop to generating its own training tasks. State-of-the-art among open coders at its size, purpose-built for terminal coding agents and tool use. The 9B runs on modest machines; the 35B is a mixture-of-experts (3B active per token) that runs fast for its size and adds vision.

9B · 35B-A3B (MoE)from 5.6 GBruns from 16 GB RAM256K contextMLXMIT

GLM-4.7 Flash

Zhipu (Z.ai) · Jan 2026

Zhipu's fast GLM model — a strong all-rounder tuned for agentic coding, reasoning, and tool use. The lighter "Flash" tier of the GLM family; the 30B MoE runs only 3B parameters per token.

30B-A3B (MoE)from 18.3 GBruns from 32 GB RAM198K contextMLXMIT

Qwen3-Coder

Alibaba's Qwen team · Jul 2025

Alibaba's agentic coding model — a 30B MoE with only 3B active per token. Tuned for repository-scale work and tool use (Qwen Code, Cline).

30B-A3B (MoE)from 18.6 GBruns from 32 GB RAM256K contextMLXApache 2.0

Qwen 3.5 Opus Distilled

Jackrong · Mar 2026

Qwen 3.5 distilled from Claude Opus reasoning traces. Enhanced chain-of-thought capabilities.

4B · 9B · 27Bfrom 2.7 GBruns from 8 GB RAM256K contextApache 2.0

Qwen 3.5 Uncensored

HauhauCS · Mar 2026

Qwen 3.5 with safety guardrails removed. No content filtering or refusals.

2B · 4B · 9B · 27Bfrom 1.2 GBruns from 4 GB RAM256K contextApache 2.0

MedGemma

Google · Jan 2026

Google's open medical models - discuss your own health records, lab results, and medical images (X-rays, skin photos, scans) privately on your device. For understanding and preparing questions, not diagnosis.

4B (v1.5) · 27Bfrom 2.4 GBruns from 8 GB RAM128K contextsees imagesMLXHealth AI Developer Foundations terms

Qwen 3.8 Distilled

empero-ai · Aug 2026

Empero's full-parameter distillations of the Qwen 3.8 flagship into small, fast sizes - a big knowledge jump over same-size models (the 9B scores near the giants on broad knowledge), with step-by-step reasoning. Text only.

2B · 4B · 9Bfrom 1.3 GBruns from 4 GB RAM256K contextApache 2.0

Qwen 3.8 Uncensored

orcarouter · Aug 2026

The Qwen 3.8 flagship with refusal behavior substantially reduced (openly documented - reduced, not eliminated). Same frontier coding, math, and reasoning, and it can see images with its vision add-on. For 24GB-class graphics cards.

27Bfrom 16.5 GBruns from 32 GB RAM256K contextsees imagesApache 2.0

Qwythos 9B

empero-ai · Jul 2026

A community Qwen 3.5-based merge — uncensored and multimodal (it can see images you attach), with step-by-step reasoning and tool use. A creative, unfiltered generalist. v2 trains out the repetition loops of the original.

9Bfrom 5.4 GBruns from 16 GB RAM1M contextsees imagesMLXApache 2.0

Looking for frontier models instead? Browse the online catalog - one account, pay per use.