Skip to main content
Guides—

GPU Memory Requirements for LLM Inference: How Much VRAM Do You Need?

A practical VRAM sizing guide for LLM inference: 7B, 13B, 70B, and 405B model classes, how quantization changes memory needs, and when to choose RTX 4090, L40S, H100, or H200.

GPU memory planning for LLM inference

The most common sizing mistake in LLM inference is budgeting only for model weights. Real deployments also need memory for the KV cache, runtime workspace, fragmentation, batching, and the framework itself. That is why a model that "fits on paper" can still be unstable in production.

This guide gives you a practical VRAM planning framework for the search queries teams actually use:

  • how much VRAM for 7B, 13B, 70B, and 405B models
  • best GPU for LLM inference
  • how much memory do I need for long-context inference
  • when to choose RTX 4090, L40S, H100, or H200

If you want help mapping a specific model to hardware, talk to our team or browse inference-ready hardware.


TL;DR: Practical VRAM Targets by Model Class

Model classPractical VRAM targetTypical deployment shapeGPUs that commonly fit
7B16GB to 24GBSingle GPU, low to moderate batchRTX 4090, L40S
13B24GB to 48GBSingle GPU, moderate contextRTX 4090, L40S, A100 40GB
30B to 34B48GB to 80GBSingle GPU or 2-way splitL40S, A100 80GB, H100
70B96GB to 160GB2 to 4 GPUs or one very large-memory GPU depending on quantizationH100, H200, multi-GPU L40S
405B class512GB+Cluster-class deploymentH200 clusters, multi-node training/inference fabric

These are planning ranges, not immutable laws. Quantization, context window, concurrency target, and framework choice all move the number.


What Actually Consumes Memory During Inference

1. Model weights

This is the baseline footprint. A smaller quantized model can fit where full-precision weights cannot. That is why the same 70B model may be impossible on one GPU at BF16 but workable with aggressive quantization and tighter latency expectations.

2. KV cache

The KV cache is what catches people. Every additional request, longer prompt, or larger batch grows memory usage. If your application serves long documents, codebases, or agent traces, KV cache can dominate the deployment budget.

3. Runtime overhead

Inference engines need workspace memory for kernels, graph capture, communication buffers, and fragmentation. Leave headroom. A configuration that sits at 98% memory usage in a lab rarely survives real traffic.

4. Redundancy and replicas

If you need high availability, you are not buying memory for one model. You are buying memory for at least two serving paths, plus canary or staging capacity.


A Simple Sizing Formula

Use this as a first pass:

  1. Start with the expected weight footprint at your target precision.
  2. Add 20% to 35% for runtime overhead and fragmentation.
  3. Add KV cache budget based on context length and concurrency.
  4. Add extra headroom if you plan to batch aggressively or keep multiple models loaded.

If that total looks tight, move up a memory tier instead of assuming software optimization will save you later.


Practical Sizing by Model Family

7B Models

Examples include many entry-level instruction-tuned models and compact agent backends.

This class is usually the easiest to deploy. If your workloads are short-context and latency-sensitive, a single 24GB GPU often works well. If you need more concurrency, step into 48GB territory rather than forcing extreme batching.

13B Models

This is where memory planning becomes less forgiving.

  • Comfortable range: 24GB to 48GB
  • Best fit: stronger single-GPU inference, moderate concurrency
  • Good options: L40S, A100 40GB, A100 80GB

A 24GB card can sometimes run 13B models, but production headroom gets thin fast once context grows. For customer-facing applications, 48GB-class GPUs are materially easier to operate.

30B to 34B Models

This is the crossover zone between workstation-class and datacenter-class serving.

If you need predictable p99 latency, staying on a single larger-memory GPU is usually cleaner than model-parallel splits.

70B Models

This is one of the most searched sizing questions because 70B is where cost, latency, and memory all become real tradeoffs.

  • Comfortable range: 96GB to 160GB
  • Best fit: H100/H200 or a multi-GPU deployment
  • Good options: H100 and H200 systems

Teams often underestimate the operational difference between "it loads" and "it serves traffic." A quantized 70B can appear to fit on a smaller footprint, but long context or concurrent users quickly erase the margin.

405B and Other Very Large Models

At this point, you are no longer making a workstation decision. You are designing an inference system.

For these deployments, network topology, memory bandwidth, and scheduling matter as much as per-GPU VRAM. Read our AI cluster networking guide before locking the architecture.


When H200 Beats H100 for Inference

H100 remains excellent when raw throughput is the priority and the model already fits comfortably. H200 becomes compelling when your workload is memory-bound:

  • larger models
  • longer context windows
  • higher concurrency
  • retrieval-augmented applications that keep larger prompts in flight

If the decision is driven by memory pressure rather than pure compute, H200 usually gives you a cleaner operating envelope. If you want a full comparison, read H100 vs H200.


Single GPU vs Multi-GPU: Which Is Better for Inference?

Prefer a single larger GPU when:

  • latency matters more than absolute throughput
  • you want a simpler software stack
  • you are early in deployment and want fewer moving parts
  • your model fits comfortably with headroom

Prefer multiple GPUs when:

  • one GPU cannot hold the model and KV cache budget
  • you need high throughput through batching
  • you want to serve multiple replicas on shared infrastructure
  • you are already operating a multi-node cluster

Multi-GPU inference works, but it adds communication overhead and operational complexity. Do not distribute the model unless memory or throughput forces you to.


Rent First or Buy First?

If you are still validating prompt format, concurrency assumptions, or framework choice, renting is usually the correct first step. gpu.fm Cloud lets you test H100 80GB and RTX 6000 Ada instances before committing to physical hardware.

Buying makes more sense when:

  • your serving pattern is stable
  • your utilization is high and predictable
  • you need private deployment or data control
  • you want fixed infrastructure under your own security boundary

If you are deciding between cloud and owned hardware, read GPU total cost of ownership.


A Better Buying Workflow

When teams ask "what GPU do I need for LLM inference," this is the information that matters most:

  1. Exact model name and precision target
  2. Average and worst-case context length
  3. Expected concurrent requests
  4. Latency target
  5. Whether the deployment is cloud, colo, or on-prem

With those five inputs, hardware selection becomes much more deterministic.

Request a recommendation if you want us to map your model and throughput target to a server or cloud configuration.


Frequently Asked Questions

Can I run a 70B model on a single GPU?

Sometimes, but only under specific conditions. A single large-memory GPU can be enough for some quantized 70B deployments, but production headroom is usually limited. For longer context or real concurrency, plan on more memory than the absolute minimum.

Is VRAM more important than raw TFLOPS for inference?

For many real deployments, yes. If the model or KV cache does not fit comfortably, extra compute does not help. Memory capacity and bandwidth often determine whether the system is usable.

Should I choose H100 or H200 for inference?

Choose H100 when the model already fits and you want strong throughput. Choose H200 when memory pressure, context length, or concurrency is the bottleneck.

Do quantized models eliminate the need for larger-memory GPUs?

No. Quantization reduces weight footprint, but you still need room for KV cache, runtime overhead, and batching. It helps, but it does not remove sizing discipline.