Main point: local AI is no longer a weekend toy. The useful setup is not the biggest model, but the right model for the job and hardware.
🧩 Text: start with
Qwen3-4B/8B, Gemma-3-4B, or Llama-3.2-1B/3B. Qwen3 is neat because it has /think and /no_think: use slower reasoning only when needed. MiMo is worth watching too: Xiaomi's MiMo-7B-RL is on GitHub/HuggingFace, tuned for math, code and reasoning. The paper says the base model used 25T pretraining tokens, then RL on 130K verifiable math/code tasks.Video:
Lightricks/LTX-Video and LTXV-13B can run locally through Python/ComfyUI, but be honest with your laptop. The 13B line wants a serious GPU. For experiments, start with distilled/FP8 or the 2B branch. Lower quality, much faster iteration.Your docs: local RAG means Chroma or LanceDB, Ollama embeddings like
embeddinggemma or qwen3-embedding, then a small LLM. Important detail: use the same embedding model for indexing and search, or the answers will sound smart but miss the source.Jupyter AI also fits the stack: chat inside JupyterLab, attach files, ask about a notebook or cell, and connect it to local Ollama or vLLM.
⚠️ Hardware note: 16 GB RAM is fine for 1B to 4B quantized models. 32 GB RAM or a discrete GPU makes 7B to 8B much nicer. Long context eats memory fast: Ollama defaults to 4096 tokens, and raising
num_ctx hits RAM/VRAM.Best 2026 laptop stack: small LLM, local embeddings, RAG, Jupyter or IDE integration. You can build it without cloud calls and without a token bill.