On-Device models are here, powered by MLX. Run AI models directly on your iPhone and iPad and share them on your local network with other devices if you want to.
Take the Reins of your AI.
Reins connects to your Ollama, LM Studio or any OpenAI-compatible server and runs local AI models directly on-device. It brings powerful AI models like Gemma, Qwen, Llama and more to your iPhone and iPad — with tools, web search, and full control over every conversation. No login. No data collection.
PLUG-n-PLAY
Either connect to your Ollama, LM Studio or OpenAI-compatible server, or run models on-device. Reins just works, keeps everything private, every chat stays yours.
TOOLS & WEB SEARCH
Go beyond chat. Reins gives models the ability to search the web, fetch pages, calculate and reason across multiple steps — no API key required for web search.
FULL MODEL CONTROL
Download, switch and manage your models directly from the app. Connect multiple servers, configure API keys and custom headers and switch between them instantly. Access Ollama Cloud when you need it.
POWERFUL CHAT
Attach images, PDFs, CSVs and text files to any conversation. Enable thinking mode for deeper reasoning when you need it. Branch chats to explore different paths without losing your original thread. Never lose a cut-off response, pick up exactly where it stopped. Export or back up your conversations anytime.
FINE-GRAINED CONTROL
Set a different system prompt for every conversation. Dial in temperature, context size, max tokens and more for precise control over model behavior.
ON DEVICE MODELS
Run AI models directly on your iPhone and iPad without any setup. It's powered by MLX to achieve the best performance on Apple Silicon. You can download any MLX model even if it's not listed in the library. Reins also works with Apple Intelligence's on-device Foundation Models.
BACKGROUND PROCESSING
No need to wait around for complex reasoning, long responses or model downloads. Switch apps or lock your screen and Reins keeps generating in the background. Come back and it'll be waiting for you. (Requires iOS or iPadOS 26.0+)
Reins works with your Ollama, LM Studio or OpenAI-compatible server, or runs AI models on-device with no server required. Self-hosted or Ollama Cloud both work too.
Privacy Policy: https://getreins.app/privacy
Terms of Use: https://getreins.app/terms
We have been utilizing this as both the front end and back in operations in AI and VI operations analytics within our infrastructure we have found that this has been a very exciting platform and tool that allows us to use existing infrastructure resources that are not to the general public, but at the same time balancing out public access in our processes operations in our research. I do recommend that if you’re going to utilize this in an environment that is in a corporate structure or research center that you allow this to run in your infrastructure that has the suitable configuration, including TPU and GPU included in the configuration inside the servers itself and have high speed storage interaction and connection connections that allow the process to collect the information necessary and then provide the most non-biased laboratory information and also will allow your team to create your own models as well. Done our initial research and development. Did this information allowed to be used and collect an invoice information other models within the system, but for security reasons and other aspects of this, we decided to disable dysfunction in our production side of our house however, in the research side on how it’s the best we still have it available but a production and more sensitive stuff because of security and the lack of guardrails and business ethics of those are outside the world of our development in ethics disabled this before it could be also used against us as well, but we get strong. Kudos to this platform and so thankful that it’s available.
It’s bonafide!
Watchman Reeves
As a recent convert the localized models and Ollama I must say I’m a believer. Frankly, I’m probably lucky to have started now because the selection of models that can run in my Mac mini is impressive—already finding quality that matches or surpasses my personal Gemini 2.5 pro and pro research — which has been the best so far. Now I’m running local for free and with Reins I can extend this newfound freedom and security to have a pretty seamless experience from clouds to home. Thank you. You are honored for making this and giving it away
Used to be decent, now charges a subscription (or $100) to use basic model features on local hosts
Olivine Cat
Thinking and tool use are basic features provided by the already running ollama server. I get gating some features, like search API features, or things that require your custom MCP server, but gating basic model features like thinking, image input support, or simple things like a server list, model parameters, and other things that the users are already getting just by running ollama is just kind of... clownishi could see paying $10 or so for some of those features. one time. the more premium lifetime access for the search feature and custom tools and whatever also makes sense, but a subscription fee for features that were already freely available before makes me sad. i could just pay $20 for for claude credits to deploy openwebui with a self update script or something and use that in safari at that point 😭
Exceptional
lubo2001
I’m not sure if i’ve ever written a review but I want to here. It’s a very professional looking app that gives me a solution i’ve been looking for for awhile - using my locally hosted AI in a package similar to the chatgpt app. My only complaint is that when it’s writing the response, the screen sort of jitters up and down if you try and scroll up to read while it’s still going. That is the only issue I’ve had. Bravo!
On-Device models are here. Run AI models directly on your iPhone and iPad. Share them across your local network if you want to.
• On-Device models: Powered by MLX for the best performance on Apple Silicon, with support for all of Reins' features.
• Local network sharing: Expose models on your iPhone or iPad over an Ollama- or OpenAI-compatible endpoint for other devices to use.
• Download any MLX model: Search for a specific model and download it even if it's not listed in the library.
• Tool calling is now available for MiniCPM5 on-device models.
• The LFM 2.5 2.6B on-device model now works. It previously failed to run.
• Bug fixes and improvements.
Version 2.3.2
The developer, Ibrahim Cetin, indicated that the app’s privacy practices may include handling of data as described below. For more information, see the developer’s privacy policy .
Data Not Collected
The developer does not collect any data from this app.
Privacy practices may vary, for example, based on the features you use or your age. Learn More
Accessibility
The developer has not yet indicated which accessibility features this app supports. Learn More