TGViewer
Android - Reddit Android - Reddit @reddit_android · 1.33K subscribers
Post #34223 128
Has anyone successfully run a local LLM on Android without cloud? What's your experience with on-device AI inference performance?

I've been experimenting with running a small LLM fully on-device on Android (Snapdragon 8 Gen 2, 8GB RAM) using llama.cpp via NDK/JNI. No internet connection needed, everything processes locally.



So far I've tested Llama 3.2 3B and Phi-3.5 Mini with Q4_K_M quantization. Getting roughly 15–20 tokens/sec which is usable for a chat assistant.



Curious if others have done something similar:

\- What devices have you tested on and what performance did you get?

\- Have you tried GPU acceleration on Android (Vulkan/OpenCL)?

\- How do you handle the \~2GB model file distribution? Bundled in APK vs downloaded on first run?



Seems like on-device AI is becoming increasingly feasible on flagship Android hardware. Interested to hear others' experiences!

https://redd.it/1ufw07o
@reddit_android
Reddit From the Android community on Reddit Explore this post and more from the Android community
More from @reddit_android
  1. Oct 9, 2026First look: This Pro Max Android phone is everything the iPhone isn't https://www.androida…
  2. Oct 9, 2026Face Unlock vs. Fingerprint Scanner: What’s your main way to unlock your phone? Hey everyo…
  3. Oct 9, 2026Compact tablet with plenty of power – Lenovo Yoga Tab Gen 2 review https://www.notebookche…
  4. Oct 9, 2026Gemini as the Google Assistant is a disaster Gemini as the Google Assistant is a disaster.…
  5. Oct 9, 2026Google appears to be testing Phone Unlock for Googlebooks, with Android 17 QPR3 showing se…
  6. Oct 9, 2026gbos-vm - Run Googlebook OS in a GPU-accelerated VM on Apple Silicon Macs https://github.c…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →