Model names pack a lot of information into a few words. Understanding that information makes it easier to choose a model that fits your hardware and the task you want to try.
This post looks at what model names actually mean, how quantization lets you squeeze a huge model into a regular GPU, and how to run open-source models on your own hardware. By the end you’ll understand why a model called Qwen3.5-Coder-9B-Instruct-Q4_K_M.gguf is named that way — and what every part of that name tells you.
What Are Parameters?
When you see a model called “9B” or “35B”, that number is the parameter count — billions of parameters.
Parameters are the numbers the model learned during training. Think of them as the model’s memory — every pattern it picked up from the training data gets encoded as a parameter. More parameters generally means the model can hold more knowledge and handle more complex reasoning.
Here’s how the sizes roughly break down:
| Size | What to expect |
|---|---|
| 1B–4B | Fast, lightweight. Good for simple tasks — autocomplete, classification, short answers. Won’t handle complex reasoning well. |
| 7B–14B | The sweet spot for local use. Can hold a real conversation, write decent code, follow multi-step instructions. |
| 30B–70B | Getting serious. Noticeably better at nuanced tasks, longer context, harder problems. Needs beefy hardware. |
| 100B+ | Cloud-scale. This is where Claude, GPT, and Gemini live. You’re not running these locally. |
The jump from 4B to 9B isn’t just “a bit better” — it’s a meaningful step up in what the model can do. But there’s a catch: bigger models need more memory to run, which is where quantization comes in.
Reading a Model Name
Model names look like alphabet soup until you know the pattern. Let’s break one down:
Qwen3.5-Coder-9B-Instruct-Q4_K_M.gguf
│ │ │ │ │ │
│ │ │ │ │ └─ File format (GGUF)
│ │ │ │ └─ Quantization level (4-bit, K_M variant)
│ │ │ └─ Variant (instruction-tuned)
│ │ └─ Size (9 billion parameters)
│ └─ Specialty (code-focused)
└─ Family and version (Qwen, generation 3.5)
Most models follow this pattern. Here’s what each part means:
Family and version — who made it and which generation. Qwen is from Alibaba, Gemma is from Google, Llama is from Meta. The version number (3.5, 4, etc.) tells you which generation — higher is usually newer and better.
Specialty — some models are fine-tuned for specific tasks. “Coder” means it was trained extra on code. Not all models have this — a model without a specialty tag is a generalist.
Size — the parameter count. 9B = 9 billion. This is the single biggest factor in both capability and hardware requirements.
Variant — how the model was prepared for use. This is important enough to get its own section.
Quantization and format — how compressed the model is and what file format it’s in. Also gets its own section.
Variants: Thinking vs Instruct vs Base
When you download a model, you’ll usually see a few variants available. The three main ones:
Base — the raw model straight from training. It’s trained to predict the next word, and that’s it. If you type “how do I fix a null pointer exception”, a base model might continue your sentence instead of answering your question. You almost never want this for practical use.
Instruct — the base model fine-tuned to follow instructions. This is what you want for chat, coding help, and question-answering. When you ask it something, it actually answers. Most of the time, instruct is what you’re looking for.
Thinking — a newer variant where the model is trained to reason step by step before giving an answer. It’s similar to Claude’s extended thinking mode — the model works through the problem internally before responding. The output is usually better for complex tasks (math, debugging, architecture decisions), but it’s slower because it’s generating more tokens behind the scenes.
Here’s how to choose:
| Use case | Pick this variant |
|---|---|
| Coding assistant, chat, general use | Instruct |
| Complex reasoning, math, tricky debugging | Thinking |
| Research, text generation experiments | Base (rare) |
If a model offers both thinking and instruct, try instruct first. Switch to thinking when you need the extra reasoning depth — same escalation pattern as Sonnet to Opus.
Quantization: Fitting Big Models Into Small Spaces
This is where it gets interesting. A 9-billion-parameter model at full precision needs about 18GB of memory just to load. That won’t fit on most consumer GPUs. Quantization solves this by reducing the precision of each parameter.
Think of it like image compression. A RAW photo might be 50MB. A high-quality JPEG of the same image is 5MB — you lose a tiny bit of detail, but you probably can’t tell the difference. Quantization does the same thing to model weights.
The bit levels
Each parameter in a model is stored as a number. The “bit” level determines how precisely that number is stored:
| Level | Bytes per parameter | 9B model size | Quality impact |
|---|---|---|---|
| FP16 (full) | 2.0 | ~18GB | None — this is the original |
| Q8 | ~1.0 | ~9GB | Virtually identical to full precision |
| Q6_K | ~0.75 | ~7GB | Very minor loss, hard to notice |
| Q4_K_M | ~0.5 | ~5GB | Slight quality loss, great trade-off |
| Q3_K | ~0.4 | ~3.5GB | Noticeable degradation on harder tasks |
| Q2_K | ~0.3 | ~2.5GB | Significant quality loss — last resort |
The sweet spot for most people is Q4_K_M — it’s roughly 4x smaller than full precision with surprisingly little quality loss. You’ll see this recommended everywhere, and for good reason.
What the suffixes mean:
- K = K-quant method (smarter quantization that keeps important weights at higher precision)
- M = medium. You’ll also see S (small/aggressive) and L (large/conservative)
- So Q4_K_M means: 4-bit quantization, using K-quant, medium aggressiveness
The rule of thumb
Here’s a quick way to estimate if a model will fit your GPU:
At Q4 quantization: ~0.5GB per billion parameters + 1-2GB overhead for context
So on a 10GB GPU:
- 4B model at Q4 → ~3-4GB ✓ fits easily, room for large context
- 9B model at Q4 → ~5-6GB ✓ fits well
- 14B model at Q4 → ~8-9GB ⚠️ tight, may need to limit context length
- 30B+ at Q4 → ~16GB+ ✗ won’t fit
Capabilities: Vision, Tool Use, and More
Some models can do more than just text. When you see these tags in a model name, here’s what they mean:
Vision — the model can look at images, not just text. You can paste a screenshot, a diagram, or a photo and ask questions about it. For example, show it a UI mockup and ask “what’s wrong with this layout?” or paste an error screenshot instead of typing it out. Claude has this built in. In local models, look for names containing “vision” or “VL” (vision-language).
Tool use / Function calling — the model can output structured calls to external tools instead of just free text. This is how an assistant calls MCP servers: it generates a structured tool call, such as reading a file or querying a sample database, rather than just suggesting that action. Locally, this is useful for building programs that use an LLM to make decisions and take actions.
Structured output — closely related to tool use. The model can output clean JSON or other formats that a program can parse. Instead of “the temperature is 72 degrees”, it returns {"temperature": 72, "unit": "fahrenheit"}. This is essential if you’re building anything that consumes LLM output programmatically.
Not all models support all capabilities. Check the model card (the description page) before downloading if you need a specific feature.
Running Models Locally with LM Studio
LM Studio is a desktop app that makes running local models simple. No command line, no Python environments, no dependency hell. You download it, search for a model, click download, and start chatting.
Why run models locally?
- Privacy — nothing leaves your machine. Useful for personal notes or private files.
- No cost per token — once the model is loaded, you can use it as much as you want.
- Experimentation — try different models, compare outputs, learn how they work.
- Offline access — works without internet once the model is downloaded.
- Build with it — LM Studio exposes an OpenAI-compatible API, so any tool that works with the OpenAI API works with your local model.
Setting it up
- Download LM Studio from lmstudio.ai
- Install and open it
- Search for a model (e.g., “qwen 3.5 coder 9B” or “gemma-4-e4b”)
- LM Studio will show you available quantizations — pick Q4_K_M for the best size/quality balance
- Click download and wait (a 9B Q4 model is about 5-6GB)
- Once downloaded, load it and start chatting
The model browser shows the model card, parameter count, available formats, and lets you pick your quantization before downloading.
That’s it. No API keys, no config files, no cloud accounts.
GGUF vs MLX — which format to pick
When you search for a model you’ll see it listed in different formats. The two you’ll encounter most:
GGUF — the standard format for Windows and Linux with NVIDIA GPUs. It’s what everything in this post uses. Built on llama.cpp, runs on your GPU (or CPU if needed).
MLX — Apple’s format, only for Macs with Apple Silicon (M1/M2/M3/M4). It uses the unified memory architecture where CPU and GPU share the same RAM pool — meaning you can run bigger models than your “GPU VRAM” spec suggests. On Apple Silicon, MLX will generally be faster than GGUF.
LM Studio detects your hardware and surfaces the right format automatically, so you usually don’t have to think about this. But if you’re ever unsure:
| Your hardware | Format |
|---|---|
| Windows / Linux + NVIDIA GPU | GGUF |
| Mac with Apple Silicon | MLX (or GGUF if MLX isn’t available) |
| CPU only | GGUF (slower, but it works) |
The local API server
This is the killer feature most people miss. LM Studio can run as a local API server — it exposes the same API format as OpenAI’s API on http://localhost:1234. That means:
- VS Code extensions that support OpenAI-compatible endpoints can use your local model
- Scripts and programs can call your local model the same way they’d call GPT or Claude
- You can test structured output and tool-use workflows without paying per token
Start the server from the “Developer” tab in LM Studio. Any tool that lets you configure a custom API endpoint can point at http://localhost:1234/v1.
The Developer tab shows the loaded model, all supported API endpoints, and live server logs. The endpoints are OpenAI-compatible — any tool that works with the OpenAI API works here.
Ollama — The Command-Line Alternative
If you prefer the terminal, Ollama does the same thing from the command line:
ollama run qwen3.5-coder:9b
That one command downloads and starts the model. Ollama also runs a local API server on http://localhost:11434.
The trade-off: Ollama is faster to get started and easier to script, but LM Studio gives you a visual interface for browsing models, comparing quantization options, and monitoring GPU usage. Use whichever fits your workflow — they both run the same models.
Real-World Setup: Running Models on an RTX 3080
Here’s what it actually looks like running local models day-to-day on a Windows gaming desktop with an RTX 3080 (10GB VRAM).
What I run
Qwen 3.5 Coder 9B (Q4_K_M) — my daily driver for local coding tasks. At Q4 quantization it uses about 5-6GB of VRAM, leaving room for context. Speed is a real advantage here — the RTX 3080 pushes around 100 tokens per second, which means responses feel instant. For context, that’s a full paragraph in about a second. Cloud APIs can’t match that on short tasks where you’re just waiting for a round trip. Good at writing functions, explaining code, and generating structured JSON output.
Gemma 4 E4B — Google’s efficient 4-billion parameter model. The “E” means it’s been architecturally optimised to punch above its weight for its size. I use this when the 9B feels like overkill — quick one-off questions, summarising a short doc, or when I want the model loaded fast without committing 5-6GB of VRAM. It also supports vision, so you can paste a screenshot and ask about it, which the Qwen coder model doesn’t do. At 4B it’s small enough to run at Q6 or Q8 and still fit comfortably, meaning you get better quality relative to size than you’d get running a 9B at the same quant.
Understanding the settings
A live chat session with Qwen in LM Studio. The settings popup and right panel expose everything you need to tune the model’s behaviour and hardware usage.
LM Studio exposes two categories of settings. The first group controls how the model is loaded — changing these requires a restart. The second group controls how it behaves at runtime and can be adjusted mid-conversation.
Model loading settings:
- Context Length — how many tokens the model can “see” at once. Higher means it can hold a longer conversation or read a larger file, but uses more VRAM. On a 10GB GPU with a 9B model you have room for around 8,000–16,000 tokens. The screenshot shows ~17,000, which is roughly 12,000 words of context.
- GPU Offload — how many of the model’s layers are loaded onto the GPU vs left on CPU RAM. Higher = faster but uses more VRAM. Set it as high as your VRAM allows. If you max it out and get out-of-memory errors, drop it a few layers.
- CPU Threads — how many CPU cores to use for the layers not on the GPU. Relevant if you’re partially offloading; otherwise leave it at the default.
- Flash Attention — a more memory-efficient way of computing attention that also runs faster. Leave this on unless you hit compatibility issues.
- Unified GPU Cache / CPU KV Cache / Offload KV Cache — these control where the KV cache lives. The KV cache is the memory the model uses to remember the conversation so far. Keeping it on GPU is fastest; offloading to CPU saves VRAM at the cost of some speed.
- Keep Model in Memory — stops the model from being evicted from RAM when idle. Useful if you switch between models and don’t want to wait for a full reload every time.
- Try Mmap — memory-mapped file loading, which lets the OS load only the parts of the model file actually needed. Faster initial load, slightly less predictable performance.
Runtime parameters:
- Prompt Format — the template used to wrap your messages before sending to the model. Models are trained expecting a specific format (e.g. Qwen uses a different wrapper than Llama). LM Studio picks this automatically for known models, but if responses seem off it’s worth checking.
- System Prompt — a set of instructions that’s silently prepended to every conversation. Use this to give the model a persona (“you are a senior Go developer”) or standing rules (“always return JSON”).
- Temperature — controls how random the output is. Lower (0.1–0.4) = more deterministic and focused, good for code and structured output. Higher (0.7–1.0) = more creative and varied, better for writing or brainstorming. For coding tasks, keep it low.
- Limit Response Length — caps how many tokens the model generates in a single reply. Useful if you want concise answers or are building something where response length matters.
- GPU Threads — fine-tuning how many GPU threads are used for inference. Generally leave this at the default unless you’re chasing performance.
- Enable Thinking — toggles the thinking/reasoning mode on models that support it (the blue toggle visible in the screenshot). When on, the model reasons through the problem before answering, similar to Claude’s extended thinking. Slower but better for complex tasks.
- Structured Output — forces the model to respond in a specific format (usually JSON schema). When you have a program consuming the output, turn this on and provide the schema — it prevents the model from adding prose around the JSON.
What works well
- Speed — 100 tokens per second means you’re not waiting. Ask it to write a function and the whole thing appears before you’ve even moved your cursor. For short tasks this is noticeably faster than a cloud round trip.
- Coding assistance — writing functions, refactoring, explaining code, generating boilerplate. It won’t trace a bug across 10 files, but for focused tasks within a single file or small context it’s genuinely solid.
- Structured output — ask it to return JSON and it does, consistently. Because it’s local you can hammer it with requests while iterating on a parser or schema without watching a token meter tick up.
- Quick chat — rubber-ducking an idea, thinking through a design, drafting something rough. Fast, free, and nothing leaves the machine.
Where it falls short
- Complex multi-file reasoning — this is where the gap between a 9B local model and Claude becomes obvious. Don’t expect a local model to trace a bug across 5 files the way Opus does.
- Long context — a 10GB GPU limits how much context you can fit alongside the model. Long conversations or large codebases will push up against that limit.
- Deep domain knowledge — local models have less training data and fewer parameters. They know less about niche topics.
The honest take
Local models are a complement to Claude, not a replacement. Think of them as your quick-and-dirty tool — private, instant, free to use. Claude is still the one you go to when the problem is genuinely hard or when you need deep codebase understanding. The value of running local models is understanding how all of this works, having a private option for sensitive work, and having a free playground to experiment with.
How This All Connects to Claude
Parameters, quantization, thinking modes, and context limits apply to Claude too, just at a much larger scale.
Claude Opus is estimated to be hundreds of billions of parameters (Anthropic doesn’t publish exact numbers). It runs on massive GPU clusters in the cloud, with far more VRAM than any consumer machine. That’s why it can hold huge conversations in context and reason through complex problems — it has the raw capacity to do it.
But the fundamentals are the same. When Claude “thinks” using extended thinking, it’s doing the same thing a local thinking model does — generating reasoning tokens before answering. When it processes a large file, it’s loading that into context the same way your local model does — and hitting the same trade-offs between context size and available memory.
Understanding these concepts makes you a better user of any LLM — whether it’s a 4B model on your GPU or Claude Opus in the cloud. You know why certain tasks are harder for smaller models, why long sessions cost more, and why sometimes a simpler model is the right choice.
Quick Reference
| I want to… | Model size | Quantization |
|---|---|---|
| Quick tasks, classification, autocomplete | 1B–4B | Q6 or Q8 (small enough to keep high quality) |
| Coding assistant, chat, structured output | 7B–9B | Q4_K_M (best trade-off) |
| Complex reasoning, detailed analysis | 14B+ | Q4_K_M (if it fits your VRAM) |
| Maximum quality, don’t care about size | Any | Q8 or FP16 (if you have the VRAM) |
GPU VRAM cheat sheet (at Q4_K_M)
| GPU | VRAM | Max comfortable model size |
|---|---|---|
| RTX 3060 | 12GB | ~14B |
| RTX 3080 | 10GB | ~12B |
| RTX 3090 / 4090 | 24GB | ~30B |
| Mac M1/M2/M3 (unified) | 16-128GB | Up to ~70B+ (using MLX) |