Local Mind for Private AI

Run LLMs privately, offline

Only for Mac

Free

Mac

**Your machine is now a private AI server.** Local Mind runs state-of-the-art open-source large language models — GLM-5.2, DeepSeek-V4-Flash, Lacuna, Qwen — entirely on your machine. No cloud. No account. No subscription. Your prompts and your data never leave your machine. **Automatically fit to your Mac.** It detects your chip, GPU, memory size and disk space, then picks the best model, quantization, and execution mode that actually fits the box. Download, Serve, and Run. No config files, no guesswork. **Private and offline.** Everything runs on-device through a GPU engine. Once a model is downloaded, you can go fully offline. Nothing is logged, uploaded, or shared. **Pay once. Own it.** A single one-time purchase — no monthly fees, unlimited local use. **Works with your tools.** Local Mind serves a local chat-completions API to developer tools on your local network so other machines can use its on-device model. **Fits real hardware.** From 8 GiB (fast 4B coding models at ~50 tok/s decode) to 512 GiB machine (DeepSeek-V4-Flash at ~22 tok/s decode). Models larger than RAM run via SSD-streaming or hybrid CPU/GPU execution, at a slower but steady speed — the app never leaves you at a dead end. **Features** - Auto-fit hardware advisor: best model + execution mode for your exact Mac - GLM-5.2, DeepSeek-V4-Flash, plus fast Qwen3.8, Lacuna-S1, Muse Glimmer coding models - 100% private, offline-capable, on-device inference - OpenAI + Anthropic compatible AI agent on this machine or LAN peers - Resumable, integrity-verified downloads for categorized models (SHA-256 verified) - Reuse GGUF models already on your disk — no re-download - Live speed + memory/GPU monitor Note: model weights are large, open-source downloads, from a few GB up to hundreds of GB depending on the model your Mac can run. The app shows the exact size and asks before downloading or deleting.

  • This app hasn’t received enough ratings or reviews to display an overview.

• Two new models. GLM-5.3-Flash — 320B, and the first GLM that runs entirely in memory on a 128 GB Mac rather than streaming from disk. Qwen3.8-Flash-Next — 125B, the fastest large model in the app. Measured on an M1 Ultra: 24 and 31 tokens/sec respectively. • GLM-5.2 retired. At 217 GB it never fit in memory on a 128 GB Mac and only ever streamed; GLM-5.3-Flash replaces it and runs resident on the same hardware.

The developer, Yijun Yu, indicated that the app’s privacy practices may include handling of data as described below. For more information, see the developer’s privacy policy .

  • Data Not Linked to You

    The following data may be collected but it is not linked to your identity:

    • Diagnostics

Privacy practices may vary, for example, based on the features you use or your age. Learn More

The developer has not yet indicated which accessibility features this app supports. Learn More

Seller
  • Yijun Yu
Size
  • 14.2 MB
Category
  • Developer Tools
Compatibility
Requires macOS 14.0 or later and a Mac with Apple M1 chip or later.
  • Mac
    Requires macOS 14.0 or later and a Mac with Apple M1 chip or later.
Languages
  • English and Simplified Chinese
Age Rating
4+
Copyright
  • © 2026 Yijun Yu