About the app

Llama.cpp Edge Gallery runs language models directly on an Android phone. You download a compatible GGUF model, load it once and chat locally with CPU inference. Prompts and responses stay in the app instead of being sent to a cloud service.

Running language models on a phone is mostly a question of memory and heat, and the app is built around that: RAM guidance helps choose a model size and quantization that fits the device, thread controls balance speed against battery and temperature, and a benchmark screen compares tokens per second across models and quantizations. Responses stream token by token, and models can be bookmarked and switched.

Features

  • Offline AI chat with no account
  • GGUF model download and management
  • Real-time token streaming
  • CPU thread controls
  • Benchmark screen with tokens per second
  • RAM guidance for model size and quantization

Made for

  • Anyone who wants an AI chat that works without internet
  • Developers testing small models on real Android hardware
  • Privacy-minded users

Questions about the app

Does it work without internet?

Yes, once a model is downloaded. Internet is only needed to download models.

Which model size fits my phone?

The app shows which model sizes and quantization levels fit the RAM of your device. Speed also depends on the processor and the temperature of the device.

More apps

  • Ollama Local AI

    AI & LLM

    Run language models on Android and serve an OpenAI-compatible API on your own Wi-Fi

  • Ollama Router Pro

    AI & LLM

    An AI router and LLM gateway that keeps your API keys on your phone

  • Battery Guardian

    Tools

    Charge alarm at 80, 85, 90 or 100 percent, plus battery temperature and charge cycles

Mobile apps

Planning an app of your own?

I build Android apps from the first idea to the release on Google Play. Tell me briefly what your app should do.