Skip to content

Llama 3.1 70B vs Qwen 2.5 72B

Comparing VRAM requirements, performance, and capabilities for running these models locally with Ollama.

Parameters

70B

Context

128K

VRAM Range

43.5–72 GB

Recommended

Q4_K_M (43.5 GB)

ByMeta·LicenseLlama 3.1 Community License
Parameters

72B

Context

128K

VRAM Range

44.7–74 GB

Recommended

Q4_K_M (44.7 GB)

ByAlibaba·LicenseQwen License

VRAM Requirements by Quantization

Side-by-side memory needs at each quality level.

QuantizationLlama 3.1 70BQwen 2.5 72BDifference
Q4_K_M43.5 GB44.7 GB-1.2 GB
Q8_072 GB74 GB-2.0 GB

Capabilities

Feature support comparison.

CapabilityLlama 3.1 70BQwen 2.5 72B
text generationYesYes
code generationYesYes
reasoningYesYes
multilingualYesYes
tool useYesYes
mathYesYes
creative writingYesYes
summarizationYesYes

Benchmark Scores

Higher is better. Scores from published evaluations.

BenchmarkLlama 3.1 70BQwen 2.5 72B
mmlu83.685.3

Hardware Compatibility

Can each model run at recommended quantization on common VRAM tiers?

VRAMLlama 3.1 70BQwen 2.5 72B
8 GBNoNo
12 GBNoNo
16 GBNoNo
24 GBNoNo
32 GBOffloadOffload
48 GBTightTight
64 GBRunsRuns
96 GBRunsRuns

Run Llama 3.1 70B

ollama run llama3.1:70b-instruct-q4_K_M

Run Qwen 2.5 72B

ollama run qwen2.5:72b-instruct-q4_K_M

Check your exact hardware

Use the compatibility checker to see how each model performs on your specific GPU or Mac.