Skip to content
ZenoCloud Call
GPU comparison · India

H100 vs A100 vs H200: which GPU should you rent?

Pick A100 when the model fits in 80GB and you want the lowest monthly bill, and pick H100 when you need FP8 and Hopper throughput for training or busy serving. Pick H200 when 141GB is what lets a 70B-class model or a long context fit on one card.

A100 80GB
₹99,000per GPU card per month
H100 80GB
₹1,80,000per GPU card per month
H200 141GB
₹2,20,000per GPU card per month
See all seven GPUs
Side by side

Specs and monthly price

Dense Tensor Core figures from NVIDIA’s datasheets; the headline numbers NVIDIA prints are with sparsity and are about double. Prices are ZenoCloud’s public INR reference per GPU card per month.

Spec A100 80GBH100 80GBH200 141GB
Architecture AmpereHopperHopper
GPU memory 80GB HBM2e80GB HBM2e (PCIe) or 80GB HBM3 (SXM)141GB HBM3e
Memory bandwidth 1.94 TB/s PCIe, 2.04 TB/s SXM2.0 TB/s PCIe, 3.35 TB/s SXM4.8 TB/s
FP16 / BF16 Tensor Core (dense) 312 TFLOPS756 TFLOPS PCIe, 989 TFLOPS SXM835 TFLOPS NVL, 989 TFLOPS SXM
FP8 support No. Ampere has no FP8 Tensor Cores; use INT8 or 4-bit quantisation instead.Yes. FP8 Tensor Cores with Transformer Engine.Yes. FP8 Tensor Cores with Transformer Engine.
NVLink / form factor as offered PCIe or SXM. NVLink 600 GB/s on SXM. MIG up to 7 instances.PCIe or SXM. NVLink 900 GB/s on SXM.SXM or NVL. NVLink 900 GB/s.
Typical fit Fine-tuning and serving models that fit in 80GB on a mature CUDA stack, at the lowest monthly price of the three.Training and fine-tuning runs where time matters, and FP8 serving of models that still fit in 80GB with their KV cache.Serving 70B-class models on one card, long-context inference and any job where the H100 runs out of memory before it runs out of compute.
ZenoCloud monthly price ₹99,000per GPU card per month₹1,80,000per GPU card per month₹2,20,000per GPU card per month
Model page NVIDIA A100 80GB in IndiaNVIDIA H100 80GB in IndiaNVIDIA H200 141GB in India
Buyer scenarios

Which one should you pick?

Fine-tuning a 7B to 13B model

Pick: A100 80GB

A 13B model in BF16 is about 26GB of weights. LoRA or QLoRA fine-tuning adds a small adapter plus activations, so the whole job fits on one 80GB card. A100 does this at ₹99,000 per month. Move to H100 when you run many iterations and the wall-clock time of each run is what costs you money.

Serving a 70B model

Pick: H200 141GB

70B weights are about 140GB in FP16, 70GB in FP8 and 35GB at 4-bit. In FP8, an H100 just holds the weights, and once the serving engine reserves its working memory there are only a few GB left for KV cache, so it suits light traffic at best. One H200 (₹2,20,000) leaves about 70GB for cache. Two H100s (₹3,60,000 for the pair) also work, but you then split the model across cards.

Long-context inference or many concurrent users

Pick: H200 141GB

KV cache grows with context length times concurrent sequences. When prompts run to tens of thousands of tokens or you batch many users, memory and memory bandwidth decide throughput, not raw compute. H200 has 141GB at 4.8 TB/s against 80GB on the H100.

A budget-limited first project

Pick: A100 80GB

Start on A100 at ₹99,000 per month. The same PyTorch, CUDA and container images run on H100 and H200, so you move your checkpoints when a run or a traffic level proves you need more. Nothing about the A100 stops you from upgrading later.

Memory sizing

How much GPU memory do you need?

Rules of thumb, not hard limits. Your framework, context length and concurrency decide the final number; test on the card before you commit to a month.

WeightsCount about 2 bytes per parameter in FP16 or BF16, 1 byte in FP8 or INT8, and 0.5 bytes at 4-bit. A 70B model is therefore about 140GB, 70GB or 35GB before anything else loads.
KV cacheEvery token in every active sequence keeps keys and values in memory. For a 70B model in FP16 that is a few hundred MB per 1,000 tokens of context, per concurrent request. Long prompts and high concurrency can need more memory than the weights themselves.
Fine-tuningFull fine-tuning with mixed precision and an Adam-style optimiser needs roughly 16 to 18 bytes per parameter for weights, gradients and optimiser state, so even a 7B model wants more than one 80GB card. LoRA and QLoRA keep the base weights frozen (often at 4-bit) and fit 7B to 13B on one card.
HeadroomLeave 10 to 20 percent free for the framework, activations and fragmentation. Two cards do not pool their memory on their own; your framework has to split the model across them.
Weights only, by parameter count and precision
ModelFP16 / BF16FP8 / INT84-bit
7B14GB7GB3.5GB
13B26GB13GB6.5GB
70B140GB70GB35GB

Add KV cache and headroom on top. A100 and H100 give you 80GB per card; H200 gives you 141GB.

Comparison questions

What buyers ask before choosing.

Straight answers on price, upgrades and what each card can hold. Every price on this page is per GPU card per month, billed monthly.

Is H100 better than A100?

For training and for FP8 inference, yes: H100 adds FP8 Tensor Cores, the Transformer Engine and, in the SXM variant, HBM3 at 3.35 TB/s. If your model already fits in 80GB and meets its latency target on A100, the A100 at ₹99,000 per month is the better buy. Measure your own workload before paying ₹1,80,000 for H100.

H200 vs H100: what’s the difference?

Both are Hopper GPUs with the same compute class and FP8 support. H200 carries 141GB of HBM3e at 4.8 TB/s against 80GB on the H100, so it holds larger models, longer contexts and bigger batches on one card. If your job is memory-bound, H200 helps. If it fits comfortably on an H100, H200 mostly adds cost.

How much does an H100 cost per month in India?

₹1,80,000 per H100 80GB GPU card per month on ZenoCloud, billed monthly with a one-month minimum. A100 80GB is ₹99,000 and H200 141GB is ₹2,20,000. The written quote states the form factor (PCIe or SXM), card count, host CPU, RAM, storage and network.

Can I start on A100 and move to H100?

Yes. The same CUDA, PyTorch and container images run on both, and model checkpoints are portable. You will usually re-tune batch size and may switch to FP8 on H100. Tell us when you want to change and we quote the new configuration; the one-month minimum applies to each.

Does A100 support FP8?

No. FP8 Tensor Cores arrived with Hopper and Ada, so of these three cards FP8 needs H100 or H200. On A100 you get the same memory saving with INT8 or 4-bit quantisation, which most serving stacks support.

Which GPU do I need to run a 70B model?

At 4-bit (about 35GB of weights) a single A100 or H100 works with room for cache. In FP8 (about 70GB) a single H200 is the comfortable choice; a single H100 fits the weights but leaves little for KV cache. In FP16 (about 140GB) you need two H100s or two H200s, with the model split across cards.

Tell us the model and we’ll price the card.

Quote request

What should we quote?

Describe the configuration or problem, location and timing.