Local Mind for Private AI
Run LLMs privately, offline
Only for Mac
Free
Mac
**Your machine is now a private AI server.**
Local Mind runs state-of-the-art open-source large language models — GLM-5.2, DeepSeek-V4-Flash, Lacuna, Qwen — entirely on your machine. No cloud. No account. No subscription. Your prompts and your data never leave your machine.
**Automatically fit to your Mac.** It detects your chip, GPU, memory size and disk space, then picks the best model, quantization, and execution mode that actually fits the box.
Download, Serve, and Run. No config files, no guesswork.
**Private and offline.** Everything runs on-device through a GPU engine. Once a model is downloaded, you can go fully offline. Nothing is logged, uploaded, or shared.
**Pay once. Own it.** A single one-time purchase — no monthly fees, unlimited local use.
**Works with your tools.** Local Mind serves a local chat-completions API to developer tools on your local network so other machines can use its on-device model.
**Fits real hardware.** From 8 GiB (fast 4B coding models at ~50 tok/s decode) to 512 GiB machine (DeepSeek-V4-Flash at ~22 tok/s decode). Models larger than RAM run via SSD-streaming or hybrid CPU/GPU execution, at a slower but steady speed — the app never leaves you at a dead end.
**Features**
- Auto-fit hardware advisor: best model + execution mode for your exact Mac
- GLM-5.2, DeepSeek-V4-Flash, plus fast Qwen3.8, Lacuna-S1, Muse Glimmer coding models
- 100% private, offline-capable, on-device inference
- OpenAI + Anthropic compatible AI agent on this machine or LAN peers
- Resumable, integrity-verified downloads for categorized models (SHA-256 verified)
- Reuse GGUF models already on your disk — no re-download
- Live speed + memory/GPU monitor
Note: model weights are large, open-source downloads, from a few GB up to hundreds of GB depending on the model your Mac can run. The app shows the exact size and asks before downloading or deleting.
Ratings & Reviews
- This app hasn’t received enough ratings or reviews to display an overview.
• Two new models. GLM-5.3-Flash — 320B, and the first GLM that runs entirely in memory on a 128 GB Mac rather than streaming from disk. Qwen3.8-Flash-Next — 125B, the fastest large model in the app. Measured on an M1 Ultra: 24 and 31 tokens/sec respectively.
• GLM-5.2 retired. At 217 GB it never fit in memory on a 128 GB Mac and only ever streamed; GLM-5.3-Flash replaces it and runs resident on the same hardware.
The developer, Yijun Yu, indicated that the app’s privacy practices may include handling of data as described below. For more information, see the developer’s privacy policy .
Data Not Linked to You
The following data may be collected but it is not linked to your identity:
- Diagnostics
Accessibility
The developer has not yet indicated which accessibility features this app supports. Learn More
Information
- Seller
- Yijun Yu
- Size
- 14.2 MB
- Category
- Developer Tools
- Compatibility
Requires macOS 14.0 or later and a Mac with Apple M1 chip or later.
- Mac
Requires macOS 14.0 or later and a Mac with Apple M1 chip or later.
- Mac
- Languages
- English and Simplified Chinese
- Age Rating
4+
- 4+
- Copyright
- © 2026 Yijun Yu
