Has anyone successfully run a local LLM on Android without cloud? What's your experience with on-device AI inference performance?
I've been experimenting with running a small LLM fully on-device on Android (Snapdragon 8 Gen 2, 8GB RAM) using llama.cpp via NDK/JNI. No internet connection needed, everything processes locally.
So far I've tested Llama 3.2 3B and Phi-3.5 Mini with Q4_K_M quantization. Getting roughly 15–20 tokens/sec which is usable for a chat assistant.
Curious if others have done something similar:
\- What devices have you tested on and what performance did you get?
\- Have you tried GPU acceleration on Android (Vulkan/OpenCL)?
\- How do you handle the \~2GB model file distribution? Bundled in APK vs downloaded on first run?
Seems like on-device AI is becoming increasingly feasible on flagship Android hardware. Interested to hear others' experiences!
https://redd.it/1ufw07o
@reddit_android
Post #34223
128