
Run language models on Android and serve an OpenAI-compatible API on your own Wi-Fi
AI & LLM · Android
Private offline AI chat on Android with GGUF models and llama.cpp
Llama.cpp Edge Gallery runs language models directly on an Android phone. You download a compatible GGUF model, load it once and chat locally with CPU inference. Prompts and responses stay in the app instead of being sent to a cloud service.
Running language models on a phone is mostly a question of memory and heat, and the app is built around that: RAM guidance helps choose a model size and quantization that fits the device, thread controls balance speed against battery and temperature, and a benchmark screen compares tokens per second across models and quantizations. Responses stream token by token, and models can be bookmarked and switched.
Yes, once a model is downloaded. Internet is only needed to download models.
The app shows which model sizes and quantization levels fit the RAM of your device. Speed also depends on the processor and the temperature of the device.

Run language models on Android and serve an OpenAI-compatible API on your own Wi-Fi

An AI router and LLM gateway that keeps your API keys on your phone

Charge alarm at 80, 85, 90 or 100 percent, plus battery temperature and charge cycles
Mobile apps
I build Android apps from the first idea to the release on Google Play. Tell me briefly what your app should do.