
Source: lifehacker.com
With the rise of freely available Large Language Models (LLMs), running them locally on a computer has become commonplace. This approach offers several benefits, including increased privacy, as you don’t have to send any data to the cloud, and the ability to work offline. While running LLMs on phones hasn’t been as straightforward, thanks to the limitations of smartphone hardware, we’re now at a point where most handsets are powerful enough to support these advanced AI models.

The benefits of running a local LLM on your phone are the same as those experienced on a computer. You get an LLM that’s always available and private to you, with no data sent back to Google, OpenAI, Anthropic, or anyone else. However, there are some downsides to consider, including slower and more limited performance from your AI, as you’re dealing with smaller and less powerful models, and a potential hit on battery life, as AI chats can be quite resource-intensive.

Despite these compromises, the local AI you’ll be able to get up and running on your phone is still capable of handling everyday tasks and chats, which means you can give Gemini or Siri (and perhaps your AI subscription) a rest. With a local LLM on your phone, you can enjoy the benefits of AI-powered conversations without relying on cloud-based services.

Any iPhone or Android phone launched in the last couple of years should be able to run a local LLM decently, thanks to the AI boom, which has led manufacturers to start building their devices with this kind of usage specifically in mind. However, older phones can still apply, but you might need to opt for smaller models and expect worse performance.
When it comes to running an AI model, RAM is actually a bigger consideration than chipset performance. You’ll need at least 6GB of RAM to run an LLM satisfactorily, and 8GB or more is recommended for the larger models. If you have less than 8GB of RAM, you’ll be better off sticking to 1-2B models, where ‘B’ refers to the number of parameters in a model, essentially determining how smart and versatile it’s going to be.
In terms of storage space, you don’t have too much to worry about. Even the largest models that use 7-8B will top out at around 5GB for how much room they take up on your phone, so you might even want to keep several AI models on your handset and switch between them as required.
There are several apps that can help you get AI models on your phone and running them. A couple of the most popular options are Atomic Chat (available on both Android and iOS) and PocketPal AI (also available on both Android and iOS). These apps offer different features and approaches, with Atomic Chat being easier to extend to desktop LLMs and PocketPal AI being a bit more lightweight.
When it comes to models themselves, you’ve got plenty of choice, too. For instance, Gemma is the umbrella name of the open-source models made available by Google for free, and some of these are specifically engineered for working in tighter spaces with fewer resources, like on your phone. Meta has its own open-source AI models given the Llama moniker, while Microsoft has its Phi-4 models, which are also highly rated for efficiency. You just need to look for the versions with the lowest number of parameters in front of the ‘B’ (or with ‘mini’ in the name) to find the packages that’ll work best with your phone.
At the moment, these SLMs are mostly text-only, although some of the newer, larger, and more advanced ones can analyze images and files. If you want to be able to generate images and videos, you’ll have to use the conventional cloud-based AI models, at least until the next leap forward in the technology.
To see how useful one of these local AI models might be on a phone, I installed PocketPal AI and one of the smaller Google Gemma models on my Pixel 9 Pro. I was pleased to find that the app’s selection wizard directed me straight to a suitable AI model for my phone, making it easy to get started. Once I’d downloaded and installed a couple of models, it was simple enough to load them up and start using them.
PocketPal AI also gives you access to optional ‘pals’ that can tailor AI models for your use. The default Pip option seems fine as a stand-in for what you might be used to with the standard Gemini or Siri AI apps. Prompting and following-up works as normal, with your chat history saved by default.
There is a noticeable (and expected) slowness to the responses when you run LLMs on your phone, and a noticeable difference between AI model sizes. Choosing a smaller model will get you answers significantly faster, even if they’re not quite as smart or complete. It’s worth experimenting with a few models just to find your own sweet spot between performance and speed.
With no web search or up-to-date knowledge available, this is best for brainstorming ideas, analyzing and refining existing text, composing new text, and getting fast facts or comparisons. As always, watch out for those hallucinations, and don’t take the word of any AI to be guaranteed as accurate.
Online Assistant