Key features:
✅ 3.5B parameters - works even on 8GB VRAM (RTX 4060)
✅ Inside: Gemma-3-4B-it + Jina CLIP v2 for deep understanding of prompts
✅ Structured XML prompts: full control over characters without random clothing changes
✅ FLUX.1-dev 16-ch VAE - soft skin, fabric and metal textures
✅ Inference in ~20 steps, LoRA support, Apache-2.0 license + non-commercial use
✅ Trained on over 10M anime images with XML annotations - confidently handles multi-character scenesb
⚡ Up to 40 percent faster than models >8B and reliably handles prompts up to 500 characters in length.
🧠 Bonus: Noise → Context Refiner pipeline eliminates the classic DiT problem - "the image is beautiful, but the prompt is ignored".
Model
••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
