
An AI router and LLM gateway that keeps your API keys on your phone
AI & LLM · Android
Run language models on Android and serve an OpenAI-compatible API on your own Wi-Fi
Ollama Local AI turns an Android phone into a local AI server. It runs quantized GGUF models on the device through an embedded llama.cpp engine and exposes OpenAI-compatible endpoints (/v1/chat/completions, /v1/models, /health) to other machines on the same Wi-Fi. Desktop IDEs such as Cursor, VS Code or Windsurf can then use the phone as their model backend.
Under the hood the app is also a multi-provider gateway: requests can go to the on-device model, a self-hosted Ollama or llama.cpp server, or a cloud API such as OpenAI, Anthropic, NVIDIA NIM or Hugging Face. Credentials stay on the device, protected with hardware-backed AES-256 encryption, and cloud requests travel directly from the phone to the provider without a relay server.
Not for on-device models. Once a GGUF model is downloaded, inference runs locally on the phone. Internet is only needed when requests are routed to a cloud provider.
Any client that speaks the OpenAI API format and accepts a custom base URL, for example Cursor, VS Code extensions such as Continue, Cline or Roo Code, and Windsurf.

An AI router and LLM gateway that keeps your API keys on your phone

Private offline AI chat on Android with GGUF models and llama.cpp

Control other Android devices over Wi-Fi: shell, screen mirroring and screenshots
Mobile apps
I build Android apps from the first idea to the release on Google Play. Tell me briefly what your app should do.