Features

Everything below is free and works offline, except where a paid plan is named. Plans add online reach - they never take anything away.

Your AIs

One AI per job

A coding partner, a writing editor, a research assistant, a health specialist - each its own AI with its own model, memory, knowledge, and permissions. They share what they should (what they know about you, a project's notes) and keep the rest apart.

Characters for personal or work use

Build a character from 18 archetypes: a companion with a backstory, or a colleague who knows your internal docs cold. Either way, no server reads along. Export any character as one signed pack file, import others' and verify who made them - or start from the ready-made characters on this site.

Specialists follow their own rules

A health AI answers on your device by default and asks before anything goes online. A coding AI gets project tools and an agent-ready model. A research AI can search the web with cited sources when you allow it. The job shapes the rules, not the other way round.

Response styles and modes

Conversational to thorough per AI; chat, report, and code modes per message, with automatic mode detection you can turn off.

Doing work

Projects: agentic coding in chat

Open a project folder and your AI reads files, edits, and runs commands - always with your permission, every step recorded. A free add-on installs in one click.

Approvals, your way

Your AI can ask before every action, run ordinary project work unasked while risky or irreversible steps still ask, or approve everything - one simple choice, set as your default or per project, switchable mid-session. Whatever you choose, every action it takes is written to your records.

Run in your terminal

Shell commands an AI writes can hand off to your own terminal, pre-filled at your prompt in the project folder - you stay the one who presses enter.

Working with content

Files, images, and vision

Attach documents, spreadsheets, code, and images. Vision models read pictures; a one-time ~30 MB add-on reads scanned paper, fully on-device.

Verify sources

Answers about your documents can be checked claim-by-claim against the source, with the exact supporting quote - automatically or on demand.

Code answers, kept readable

Long code in a reply shows as a compact preview with one tap to see the whole thing - so answers stay readable even when the code runs long.

Web search with citations

When you allow online use, current-events questions can search the web and cite their sources - with live progress while they research.

Memory

A shared memory of you

Your AIs learn who you are from your chats - and Your Memory shows every fact with where it came from. Edit anything, forget anything, or pause learning entirely. Nothing is hidden from you.

Each AI remembers its own life

Every AI keeps its own memory of what you've done together and recalls the right moments when they matter - alongside memories you hand it directly.

They know you as a whole, not as a list

From what they have learned, your AIs keep a short summary of who you are - "How your AIs see you" - written on your device and rewritten as things change. Read it any time on Your Memory; if it is wrong, fix the fact it came from and it rewrites itself. Every other AI builds its picture of you on their servers and keeps it to itself. Yours never goes online.

Remember anything, anywhere

A Remember button under every reply, or highlight any text and keep just that. Choose the destination: one AI, all of them, or the project you're working in.

Project notes

Each project keeps notes all your AIs share - conventions, decisions, context. Working AIs can save notes as they go (you approve each one), and sessions distill their takeaways when you finish.

Knowledge you give them

Add documents an AI should always know - it keeps its own copy and uses the relevant parts when they help.

Memory that survives

All of it - profile, per-AI, and project memory - is stored encrypted on your device, backs up to your Flowsta Vault, and comes back on a new machine.

How memory works →

Smart routing

It measures, not assumes

On your device, routing picks by real numbers: as you use each model, the app times how fast it actually runs on this computer and how long it takes to load, and ranks by that - not a spec sheet, not someone else's benchmark. Online models follow the preferences you set. Two people with the same catalog get different picks, each right for their machine.

The mode is the consent

"Auto - Offline Only" never touches the internet. "Auto - Online and Offline" answers with a frontier online model by default and keeps a question on your device when a model here is better for it - one dial in Settings sets how much goes online. Pick a specific model and there's no routing at all.

The right model per question

One online model for everyday questions, another for the genuinely hard ones, a web-search model for anything that needs current information - each a pick you can change, with new models arriving without an app update. Health questions stay on your device by default - going online with one is a choice that asks first, every time.

It always says why

Every reply names the model that answered, where it ran, and why it was chosen - with one-click second opinions: "Redo on your device" or "Try this answer online". Settings keeps a live list of recent routing decisions.

Your levers

Prefer fastest or strongest on-device. Privacy-first to freshness-first online. Per-category picks for which online model handles what, prices shown up front - and project work can stay entirely on your device.

How routing works →

Models and performance

Best for this computer

The models page opens with a pick per activity - coding with you, everyday chat, seeing images, health questions - each chosen from what your machine can actually run. No guessing, and no downloads that were never going to fit.

Model management without a terminal

Browse, download, and switch open models in the app, with hardware-fit guidance - what runs fully on your GPU, what will be slower, what won't fit. Models that can drive project work carry an Agentic label, so you know before you download.

Every model says who made it

Model cards name the maker of the weights, who packaged the files you download, and for community builds, what they are based on and how they differ. You always know whose work you are running.

Engines for your hardware

Vulkan and Metal out of the box, and a one-click CUDA engine for NVIDIA machines - the app picks safe defaults and steps down gracefully if your hardware protests. More than one GPU? Bigger models spread across them.

Bigger than your graphics card

Mixture-of-experts models can run with their always-on layers on the graphics card and their experts in system memory - so a model well beyond your card's size still runs at a good pace. The app knows which models split well and grades them honestly.

Models live where you say

Choose where model files are stored - a second drive, an external disk - and downloads follow, with free space checked before every download and existing models moved safely.

Apple Silicon MLX engine (preview)

Macs with Apple Silicon can add an optional MLX engine, then fetch MLX versions of supported models - chats run on MLX while everything else stays on the standard engine. Whether it is faster depends on your Mac; nothing changes unless you install it.

Online frontier models (paid plans)

Frontier models from every major lab - an optional paid service you sign in to with your Flowsta Identity. A monthly allowance, fair metered pricing, every price shown up front. An add-on, never a dependency.

Browse every model →

The inference engine

An OpenAI-compatible endpoint, serving your AIs

While the app is open, your machine serves a standard endpoint. Point agent frameworks like Hermes Agent and OpenClaw, coding editors, or your own scripts at it - and they get YOUR AIs, not a bare model. No API key needed on your own machine.

Connect your own server

Point the app at any OpenAI-compatible endpoint - a bigger model on your homelab box, a llama.cpp server on the LAN. It health-checks, measures its real speed, and its models join your picker.

The whole stack rides along

In either direction, everything applies: the personality you authored answers, memories inform it, routing treats connected hardware as a candidate, and every conversation lands in your signed records - external apps show an API badge on the Memory page.

Trust and your data

Signed conversation records

Every conversation is written into a tamper-evident record on your device - built on Holochain, the peer-to-peer foundation under Flowsta - and encrypted with a key only you hold.

Proof of which model answered

Answers from models on your device are recorded with a fingerprint of the exact model file that produced them, alongside its name - your records do not just say what was said, they prove which weights said it. Online answers carry the provider's word; your own machine's answers carry proof.

Export and receipts

Export any conversation as a readable file with knowable contents - and optionally sign it with your Flowsta identity so anyone can verify it at flowsta.com.

Backup and device-loss recovery

Keys, conversations, AIs, and memories back up to your Flowsta Vault and come back on a new machine. Your data exports readable, keys included - no lock-in, by design.

Open source

AGPL-licensed, source on GitHub. Verify the privacy claims yourself - that's the point of making them verifiable.

Everyday comfort

Native on macOS, Windows, and Linux

One app, three platforms - the same features and the same privacy whichever machine you're on. Installs like any other app, no terminal required.

Light and dark themes

The whole app in light or dark - switch anytime from the header menu, and every page follows instantly.

Help where you need it

Dismissible help tips explain each surface as you first meet it. Turn them all off in Settings - or bring back the ones you dismissed.

Wondering what you'd actually do with all this? See the uses →

Download free