The most common sizing mistake in LLM inference is budgeting only for model weights. Real deployments also need memory for the KV cache, runtime workspace, fragmentation, batching, and the framework itself. That is why a model that "fits on paper" can still be unstable in production.
This guide gives you a practical VRAM planning framework for the search queries teams actually use:
- how much VRAM for 7B, 13B, 70B, and 405B models
- best GPU for LLM inference
- how much memory do I need for long-context inference
- when to choose RTX 4090, L40S, H100, or H200
If you want help mapping a specific model to hardware, talk to our team or browse inference-ready hardware.
TL;DR: Practical VRAM Targets by Model Class
| Model class | Practical VRAM target | Typical deployment shape | GPUs that commonly fit |
|---|---|---|---|
| 7B | 16GB to 24GB | Single GPU, low to moderate batch | RTX 4090, L40S |
| 13B | 24GB to 48GB | Single GPU, moderate context | RTX 4090, L40S, A100 40GB |
| 30B to 34B | 48GB to 80GB | Single GPU or 2-way split | L40S, A100 80GB, H100 |
| 70B | 96GB to 160GB | 2 to 4 GPUs or one very large-memory GPU depending on quantization | H100, H200, multi-GPU L40S |
| 405B class | 512GB+ | Cluster-class deployment | H200 clusters, multi-node training/inference fabric |
These are planning ranges, not immutable laws. Quantization, context window, concurrency target, and framework choice all move the number.
What Actually Consumes Memory During Inference
1. Model weights
This is the baseline footprint. A smaller quantized model can fit where full-precision weights cannot. That is why the same 70B model may be impossible on one GPU at BF16 but workable with aggressive quantization and tighter latency expectations.
2. KV cache
The KV cache is what catches people. Every additional request, longer prompt, or larger batch grows memory usage. If your application serves long documents, codebases, or agent traces, KV cache can dominate the deployment budget.
3. Runtime overhead
Inference engines need workspace memory for kernels, graph capture, communication buffers, and fragmentation. Leave headroom. A configuration that sits at 98% memory usage in a lab rarely survives real traffic.
4. Redundancy and replicas
If you need high availability, you are not buying memory for one model. You are buying memory for at least two serving paths, plus canary or staging capacity.
A Simple Sizing Formula
Use this as a first pass:
- Start with the expected weight footprint at your target precision.
- Add 20% to 35% for runtime overhead and fragmentation.
- Add KV cache budget based on context length and concurrency.
- Add extra headroom if you plan to batch aggressively or keep multiple models loaded.
If that total looks tight, move up a memory tier instead of assuming software optimization will save you later.
Practical Sizing by Model Family
7B Models
Examples include many entry-level instruction-tuned models and compact agent backends.
- Comfortable range: 16GB to 24GB
- Best fit: single GPU inference
- Good options: cloud RTX 4090 or L40S
This class is usually the easiest to deploy. If your workloads are short-context and latency-sensitive, a single 24GB GPU often works well. If you need more concurrency, step into 48GB territory rather than forcing extreme batching.
13B Models
This is where memory planning becomes less forgiving.
- Comfortable range: 24GB to 48GB
- Best fit: stronger single-GPU inference, moderate concurrency
- Good options: L40S, A100 40GB, A100 80GB
A 24GB card can sometimes run 13B models, but production headroom gets thin fast once context grows. For customer-facing applications, 48GB-class GPUs are materially easier to operate.
30B to 34B Models
This is the crossover zone between workstation-class and datacenter-class serving.
- Comfortable range: 48GB to 80GB
- Best fit: one larger GPU or two smaller GPUs
- Good options: L40S and H100-class deployments
If you need predictable p99 latency, staying on a single larger-memory GPU is usually cleaner than model-parallel splits.
70B Models
This is one of the most searched sizing questions because 70B is where cost, latency, and memory all become real tradeoffs.
- Comfortable range: 96GB to 160GB
- Best fit: H100/H200 or a multi-GPU deployment
- Good options: H100 and H200 systems
Teams often underestimate the operational difference between "it loads" and "it serves traffic." A quantized 70B can appear to fit on a smaller footprint, but long context or concurrent users quickly erase the margin.
405B and Other Very Large Models
At this point, you are no longer making a workstation decision. You are designing an inference system.
- Comfortable range: 512GB+
- Best fit: cluster-scale serving
- Good options: training and inference pods
For these deployments, network topology, memory bandwidth, and scheduling matter as much as per-GPU VRAM. Read our AI cluster networking guide before locking the architecture.
When H200 Beats H100 for Inference
H100 remains excellent when raw throughput is the priority and the model already fits comfortably. H200 becomes compelling when your workload is memory-bound:
- larger models
- longer context windows
- higher concurrency
- retrieval-augmented applications that keep larger prompts in flight
If the decision is driven by memory pressure rather than pure compute, H200 usually gives you a cleaner operating envelope. If you want a full comparison, read H100 vs H200.
Single GPU vs Multi-GPU: Which Is Better for Inference?
Prefer a single larger GPU when:
- latency matters more than absolute throughput
- you want a simpler software stack
- you are early in deployment and want fewer moving parts
- your model fits comfortably with headroom
Prefer multiple GPUs when:
- one GPU cannot hold the model and KV cache budget
- you need high throughput through batching
- you want to serve multiple replicas on shared infrastructure
- you are already operating a multi-node cluster
Multi-GPU inference works, but it adds communication overhead and operational complexity. Do not distribute the model unless memory or throughput forces you to.
Rent First or Buy First?
If you are still validating prompt format, concurrency assumptions, or framework choice, renting is usually the correct first step. gpu.fm Cloud lets you test H100 80GB and RTX 6000 Ada instances before committing to physical hardware.
Buying makes more sense when:
- your serving pattern is stable
- your utilization is high and predictable
- you need private deployment or data control
- you want fixed infrastructure under your own security boundary
If you are deciding between cloud and owned hardware, read GPU total cost of ownership.
A Better Buying Workflow
When teams ask "what GPU do I need for LLM inference," this is the information that matters most:
- Exact model name and precision target
- Average and worst-case context length
- Expected concurrent requests
- Latency target
- Whether the deployment is cloud, colo, or on-prem
With those five inputs, hardware selection becomes much more deterministic.
Request a recommendation if you want us to map your model and throughput target to a server or cloud configuration.
Frequently Asked Questions
Can I run a 70B model on a single GPU?
Sometimes, but only under specific conditions. A single large-memory GPU can be enough for some quantized 70B deployments, but production headroom is usually limited. For longer context or real concurrency, plan on more memory than the absolute minimum.
Is VRAM more important than raw TFLOPS for inference?
For many real deployments, yes. If the model or KV cache does not fit comfortably, extra compute does not help. Memory capacity and bandwidth often determine whether the system is usable.
Should I choose H100 or H200 for inference?
Choose H100 when the model already fits and you want strong throughput. Choose H200 when memory pressure, context length, or concurrency is the bottleneck.
Do quantized models eliminate the need for larger-memory GPUs?
No. Quantization reduces weight footprint, but you still need room for KV cache, runtime overhead, and batching. It helps, but it does not remove sizing discipline.



