Models that run on your device
46 downloadable variants across 24 model families - from 2 GB laptop models to frontier-class mixtures of experts. Everything here runs fully offline, keeps your conversations on your machine, and needs no account for local use.
Muse GlimmerRecommended
Meta · Aug 2026Meta's new agent-first model - built for reliable tool use, working in project folders, and recovering from its own mistakes. Sees images too. Brand-new: early support, expect rough edges.
Qwen 3.8Recommended
Alibaba's Qwen team · Aug 2026Alibaba's newest. Frontier coding, math, and reasoning with built-in thinking; a hybrid attention design keeps long documents fast. The strongest model here for 24GB-class graphics cards.
Qwen 3.6Recommended
Alibaba's Qwen team · Apr 2026Top-tier coding, math, and reasoning. The 35B MoE runs fast for its size — only 3B parameters active per token.
Gemma 4Recommended
Google · Apr 2026Google's latest open model. Strong writing, analysis, and instruction-following. E2B/E4B run on modest machines; the 26B is a fast MoE.
Ministral 3
Mistral AI · Dec 2025Mistral's latest small model family. Fast inference, great for on-device use.
Phi-4 MiniRecommended
Microsoft · Feb 2025Microsoft's compact model. Best option for machines with limited RAM.
GPT-OSS (OpenAI)
OpenAI · Aug 2025OpenAI's open-weight model. Strong reasoning, agentic tasks, and function calling.
LFM2.5Recommended
Liquid AI · May 2026Liquid AI's small mixture-of-experts - 8B total, 1.5B active per token, so it runs fast on ordinary machines. Strong instruction following and tool use for its size; 10 languages.
Granite 4.0 TinyRecommended
IBM · Sep 2025IBM's small hybrid mixture-of-experts - 7B total, about 1B active per token, with a 1M-token context. Apache-2.0. A fast, permissive everyday model for long documents.
Nemotron 3.5 LightningRecommended
NVIDIA · Aug 2026NVIDIA's newest open reasoning model - a 30B mixture-of-experts with 3B active per token and a 256K context. Strong reasoning, coding, and tool use; runs on a 32 GB machine with the experts in main memory.
Ling-mini 2.0
Ant Group (inclusionAI) · Sep 2025Ant Group's mid-size mixture-of-experts - 16B total, 1.4B active per token. MIT-licensed, 128K context; a quick all-rounder for 24 GB machines.
DeepSeek V4 Flash
DeepSeek · Jul 2026DeepSeek's V4 Flash - a 284B mixture-of-experts with 13B active per token and a 1M-token context. Frontier-class reasoning and agentic work on a workstation with a big graphics card and 128 GB or more of main memory.
GLM-5.2
Zhipu AI · Jun 2026Zhipu's GLM-5.2 - a 753B mixture-of-experts with a 1M-token context. For very large workstations: a big graphics card and 384 GB or more of main memory, experts in main memory.
DeepSeek R1 0528
DeepSeek · May 2025Powerful chain-of-thought reasoning. Excels at complex problem solving, math, and logic.
Devstral Small 2
Mistral AI · Dec 2025Mistral's agentic coding model. Tops open models on SWE-bench at its size — purpose-built for software engineering and code agents.
Ornith 1.5
DeepReinforce · Aug 2026DeepReinforce's agentic coding model, self-improved with RL - 1.5 extends the loop to generating its own training tasks. State-of-the-art among open coders at its size, purpose-built for terminal coding agents and tool use. The 9B runs on modest machines; the 35B is a mixture-of-experts (3B active per token) that runs fast for its size and adds vision.
GLM-4.7 Flash
Zhipu (Z.ai) · Jan 2026Zhipu's fast GLM model — a strong all-rounder tuned for agentic coding, reasoning, and tool use. The lighter "Flash" tier of the GLM family; the 30B MoE runs only 3B parameters per token.
Qwen3-Coder
Alibaba's Qwen team · Jul 2025Alibaba's agentic coding model — a 30B MoE with only 3B active per token. Tuned for repository-scale work and tool use (Qwen Code, Cline).
Qwen 3.5 Opus Distilled
Jackrong · Mar 2026Qwen 3.5 distilled from Claude Opus reasoning traces. Enhanced chain-of-thought capabilities.
Qwen 3.5 Uncensored
HauhauCS · Mar 2026Qwen 3.5 with safety guardrails removed. No content filtering or refusals.
MedGemma
Google · Jan 2026Google's open medical models - discuss your own health records, lab results, and medical images (X-rays, skin photos, scans) privately on your device. For understanding and preparing questions, not diagnosis.
Qwen 3.8 Distilled
empero-ai · Aug 2026Empero's full-parameter distillations of the Qwen 3.8 flagship into small, fast sizes - a big knowledge jump over same-size models (the 9B scores near the giants on broad knowledge), with step-by-step reasoning. Text only.
Qwen 3.8 Uncensored
orcarouter · Aug 2026The Qwen 3.8 flagship with refusal behavior substantially reduced (openly documented - reduced, not eliminated). Same frontier coding, math, and reasoning, and it can see images with its vision add-on. For 24GB-class graphics cards.
Qwythos 9B
empero-ai · Jul 2026A community Qwen 3.5-based merge — uncensored and multimodal (it can see images you attach), with step-by-step reasoning and tool use. A creative, unfiltered generalist. v2 trains out the repetition loops of the original.
Looking for frontier models instead? Browse the online catalog - one account, pay per use.