Advice for AI engineers
You can run production-level LLM inference on a laptop CPU or even on a phone.
No cloud accounts. No API keys. No internet.
LFM2.5-1.2B-Instruct from liquidai gives:
239 tokens/s on an AMD CPU
82 tokens/s on a mobile NPU
less than 1 GB of RAM
Get the link
••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1986
173
