The AI Desk · Story 6
Stop asking what model to run. There are literally only two.
A popular post in r/LocalLLaMA argues that users should only consider running two models: Qwen and Gemma. The post and its 640 comments have received over 2,600 upvotes.
The discussion reflects real trade-offs the local LLM community faces — coding vs. creative tasks, hardware constraints, and finetuned model variants — as companies release ever more specialised open-weight models.
Key details
- The top comment says Gemma is better than Qwen for any creative task, at any quantisation level.
- A second commenter disagrees, claiming the two-model rule only holds for coding and that the best model for roleplaying is Gemma4 31B.
- One commenter with limited hardware asks for help getting more than one token per minute on a system with less than 16 GB of RAM and no GPU.
- A satirical comment lists a fictional, heavily modified model name to mock the rapid churn of specialised finetunes.
Discussion angles
- Gemma is considered superior to Qwen for creative and roleplaying tasks.
- The two-model simplification may apply only to coding, not to use cases like character chat.
- Users with low-RAM, no-GPU setups cannot run the recommended models and need alternatives.
- The pace of new community finetunes makes any simple recommendation quickly outdated.
Top comments
Commenters challenge the idea that only two models matter, pointing out that Gemma is preferred for creative and roleplaying use while Qwen may be better for coding. Users with weak hardware ask for realistic alternatives, and others satirise the flood of specialised finetunes.