🧠 Local LLMs · Ollama · Hardware Guide
Ollama Models List: Which Model Fits Your Computer
The Ollama catalog is massive, and most beginners make the exact same mistake: they pull the biggest model they've heard of, wait through a massive 40 GB download, and then watch their laptop completely freeze. The right question isn't "which model is the smartest?" but rather "which model actually fits my RAM and my specific job?"
This is a practical, no-fluff guide to the Ollama model library. We'll break down what each model family does best, exactly how much memory it eats, and the terminal commands you need to get running in under two minutes.
Where this lands: Model Sizes Explained
Model sizes (measured in parameters, like 7B or 8B) directly dictate your hardware requirements. Here is the general hierarchy:
- Small models (1B to 3B): Run on almost any modern laptop. They are incredibly fast and fine for short answers, simple summarization, and basic classification tasks.
- Mid-size models (7B to 9B): The sweet spot for most people with 16 GB of RAM. They offer a great balance of reasoning capability and speed for everyday chat and writing.
- Large models (13B to 14B): Require 16 GB to 24 GB of RAM. Noticeably smarter at complex logic and coding, but will slow down on a CPU.
- Massive models (27B to 70B+): Need a lot of memory (32 GB to 64 GB+) or a strong dedicated GPU with high VRAM. They are painfully slow without dedicated hardware.
⚠️ Note: Model names and versions change frequently as Meta, Mistral, and Google release updates. Always confirm what's current in the official catalog before downloading.
What these tools actually need from you
Ollama downloads models in a compressed quantized format (usually GGUF). A rough rule of thumb for RAM/VRAM requirements is:
- ~8 GB of RAM for 7B to 8B models
- ~16 GB of RAM for 13B to 14B models
- ~48 GB+ of RAM for 70B models
A GPU with enough VRAM speeds things up exponentially, but small models run perfectly fine on a standard CPU. If you haven't installed the software yet, start with our Ollama setup guide, or follow the specific steps to install Ollama on Windows if that's your OS.
Essential Terminal Commands
ollama list # see downloaded models
ollama pull llama3.2 # download a model
ollama run llama3.2 # chat with it
ollama rm llama3.2 # delete it to free up space
Models by job: Which one should you pull?
General Chat & Everyday Tasks
For drafting emails, brainstorming, and general Q&A:
- llama3.2 (1B, 3B) — Incredible for edge devices and low-RAM laptops.
- llama3.1 (8B) — The gold standard for 16 GB RAM setups.
- mistral (7B) — Fast, punchy, and great at following instructions.
- gemma2 (2B, 9B, 27B) — Google's open model, punches way above its weight class.
- qwen2.5 (many sizes) — Excellent multilingual support.
- phi3 (3.8B) — Microsoft's highly efficient small model.
Coding & Development
If you need help writing, debugging, or explaining code, general models will struggle. Use these instead:
- qwen2.5-coder — Currently dominating the open-source coding benchmarks.
- codellama — Meta's dedicated coding variant.
- deepseek-coder-v2 — Excellent for complex repository-level understanding.
For a deeper dive into parameter sizes and IDE integrations, check out our guide on the best Ollama models for coding.
Reasoning & Math
- deepseek-r1 — This model actually shows its "thinking" process (chain of thought) before it answers. It's slower, but significantly better at logic puzzles, math, and step-by-step problem solving.
Vision (Image Analysis)
Want to pass an image to your local AI and ask it questions?
- llava — The classic open-source vision model.
- llama3.2-vision — Meta's newer, highly capable multimodal model.
Embeddings (RAG & Search)
If you are building a local Retrieval-Augmented Generation (RAG) pipeline to chat with your own PDFs, you need embedding models to turn text into vectors:
- nomic-embed-text — Fast and highly accurate.
- mxbai-embed-large — Top-tier performance for complex semantic search.
The "too big for my computer" problem
Pull a model that's too large, and Ollama will still try to run it. It will use your hard drive as "swap" memory, which usually means your computer will completely lock up, generate one token every 10 seconds, and eventually crash.
🛑 How to avoid it: Always start one size smaller than you think you need. Test it on your real task. Only move up to a larger parameter count if the smaller model's answers aren't good enough. Our run Llama 3 with Ollama performance guide shows exactly what token speeds to expect on different hardware.
Comparing your options
| Your Computer | Start With | Good For | Speed on CPU |
|---|---|---|---|
| 8 GB RAM, no GPU | llama3.2 3B, phi3 | Short answers, drafts, simple chat | Fine (20-40 t/s) |
| 16 GB RAM | llama3.1 8B, mistral 7B | Everyday chat, writing, light coding | Usable (10-20 t/s) |
| Coding Work | qwen2.5-coder (7B or 14B) | Code completion, debugging | Depends on size |
| 32 GB+ or good GPU | gemma2 27B, llama3.1 70B | Complex reasoning, deep analysis | Very slow without GPU |
Before you commit
If you're unsure whether a terminal-based workflow is right for you, read our breakdown of Ollama vs LM Studio. LM Studio offers a beautiful graphical interface, while Ollama is better for developers and background API services.
Finally, remember that while the Ollama software is free, each model has its own license. Always check the license file on the model's library page before deploying it in a commercial product or customer-facing app.
Frequently asked questions
~/.ollama/models on Mac/Linux or C:\Users\YourUsername\.ollama\models on Windows). If your main drive is small, you can move this storage location to an external or secondary drive by setting the OLLAMA_MODELS environment variable to your preferred folder path.
Further reading & Official Resources
- Ollama model library — The always-current, official catalog of available models.
- ollama.com — Official downloads, API documentation, and release notes.