Command-R 35B vs Qwen3 32B
Side-by-side VRAM requirements, benchmark scores, and GPU compatibility for local AI inference.
Quick verdict
Qwen3 32B is more hardware-efficient: it needs 23.9 GB at its Q4_K_M build vs 35.9 GB for Command-R 35B's Q4_K_M, fitting on 74 GPUs natively.
VRAM at each quantization (8k context)
FP32
Command-R 35B
168.8 GB
Qwen3 32B
148.4 GB
BF16
Command-R 35B
90.4 GB
Qwen3 32B
75.0 GB
FP16
Command-R 35B
90.4 GB
Qwen3 32B
75.0 GB
Q8_0
Command-R 35B
53.7 GB
Qwen3 32B
40.5 GB
Q6_K
Command-R 35B
44.2 GB
Qwen3 32B
31.7 GB
Q5_K_M
Command-R 35B
39.9 GB
Qwen3 32B
27.7 GB
Q4_K_M
Command-R 35B
35.9 GB
Qwen3 32B
23.9 GB
Q3_K_M
Command-R 35B
30.9 GB
Qwen3 32B
19.2 GB
Q2_K
Command-R 35B
27.0 GB
Qwen3 32B
15.5 GB
NVFP4
Command-R 35B
31.6 GB
Qwen3 32B
19.9 GB
| Quant | Command-R 35B | Qwen3 32B | Diff |
|---|---|---|---|
| FP32 | 168.8 GB | 148.4 GB | +14% |
| BF16 | 90.4 GB | 75.0 GB | +21% |
| FP16 | 90.4 GB | 75.0 GB | +21% |
| Q8_0 | 53.7 GB | 40.5 GB | +32% |
| Q6_K | 44.2 GB | 31.7 GB | +40% |
| Q5_K_M | 39.9 GB | 27.7 GB | +44% |
| Q4_K_M | 35.9 GB | 23.9 GB | +50% |
| Q3_K_M | 30.9 GB | 19.2 GB | +61% |
| Q2_K | 27.0 GB | 15.5 GB | +74% |
| NVFP4 | 31.6 GB | 19.9 GB | +59% |
Diff is Command-R 35B relative to Qwen3 32B. Green = lower VRAM (fits more GPUs).
Model specifications
| Spec | Command-R 35B | Qwen3 32B |
|---|---|---|
| Org | Cohere | Alibaba |
| Parameters | 35B | 32.8B |
| Architecture | Dense | Dense |
| Context | 125k tokens | 128k tokens |
| Modalities | text | text |
| License | CC-BY-NC 4.0 | Apache 2.0 |
| Commercial | No | Yes |
| Released | 2024-08-30 | 2025-04-29 |
| GPUs (native) | 53 / 119 | 74 / 119 |
Benchmark scores
Green = higher score (better). N/A = not yet available. ~ = inherited from a base model, not independently reported for that release itself.
GPUs that run only Command-R 35B(0)
Every GPU that runs Command-R 35B also runs Qwen3 32B.
GPUs that run only Qwen3 32B(21)
- NVIDIA RTX 409024 GB
- NVIDIA RTX 309024 GB
- NVIDIA RTX 3090 Ti24 GB
- NVIDIA RTX 4000 Ada20 GB
- NVIDIA RTX 4500 Ada24 GB
- AMD Radeon RX 7900 XTX24 GB
- AMD Radeon RX 7900 XT20 GB
- AMD Strix Halo (32GB)32 GB
- Apple M5 Pro (24GB)24 GB
- Apple M5 (32GB)32 GB
- +11 more
GPUs that run both natively(53)
- NVIDIA RTX 509032 GB
- NVIDIA B300 288GB288 GB
- NVIDIA B200 180GB180 GB
- NVIDIA H200 141GB141 GB
- NVIDIA H100 80GB80 GB
- NVIDIA A100 80GB80 GB
- NVIDIA A100 40GB40 GB
- NVIDIA L40S48 GB
- NVIDIA RTX A600048 GB
- NVIDIA RTX 5000 Ada32 GB
- NVIDIA RTX 6000 Ada48 GB
- NVIDIA RTX Pro 600096 GB
- +41 more GPUs run both
Which should you use?
Choose Command-R 35B if:
- • You want maximum capability and have a 36 GB+ GPU
Choose Qwen3 32B if:
- • You have limited VRAM: it's a smaller model needing 23.9 GB vs 35.9 GB
- • Long context matters: it supports 128k tokens vs 125k
- • You need commercial use rights
- • Benchmark quality matters: scores 65.5 vs 33.0 on MMLU-Pro
- • You need chain-of-thought reasoning
- • It's the newer release (2025-04-29 vs 2024-08-30); check the benchmark table above for what actually improved
Frequently asked questions
- Which is better, Command-R 35B or Qwen3 32B?
- Command-R 35B has 35B parameters vs 32.8B for Qwen3 32B, so Command-R 35B is the larger model. Qwen3 32B is more hardware-efficient, needing 23.9 GB at its Q4_K_M build vs 35.9 GB for Command-R 35B's Q4_K_M. Qwen3 32B runs on more GPUs natively (74 vs 53). On MMLU-Pro, Qwen3 32B scores higher (65.5 vs 33.0).
- How much VRAM does Command-R 35B need vs Qwen3 32B?
- At 8k context, Command-R 35B needs approximately 35.9 GB of VRAM at its Q4_K_M build, while Qwen3 32B needs 23.9 GB at its Q4_K_M build. At the largest build each ships, Command-R 35B requires 90.4 GB (FP16) vs 75.0 GB (FP16) for Qwen3 32B.
- Can you run Command-R 35B on the same GPUs as Qwen3 32B?
- Yes, 53 GPUs can run both natively in VRAM, including NVIDIA RTX 5090, NVIDIA B300 288GB, NVIDIA B200 180GB. However, no GPU can run Command-R 35B without also fitting Qwen3 32B, and 21 GPUs can run Qwen3 32B but not Command-R 35B.
- What is the difference between Command-R 35B and Qwen3 32B?
- Command-R 35B has 35B parameters (dense) with a 125k context window. Qwen3 32B has 32.8B parameters (dense) with a 128k context window. Licensing differs: Command-R 35B is CC-BY-NC 4.0 while Qwen3 32B is Apache 2.0.
- Which model fits in 24 GB of VRAM, Command-R 35B or Qwen3 32B?
- Neither fits in 24 GB: Command-R 35B needs 35.9 GB at Q4_K_M and Qwen3 32B needs 23.9 GB at Q4_K_M. Both require at least a 48 GB GPU.