Desktop GPU

NVIDIA RTX 5090

The Fastest Consumer GPU Ever Made

~$4,500 (street price, up from $1,999 MSRP due to GDDR7 shortage)
Market estimate
32 GB
GDDR7 VRAM
32 gigabytes of fast video memory. That's the workspace the GPU uses to hold model weights while generating tokens. Bigger = can run larger models without swapping to slow regular RAM. At full precision, this fits models up to about 25 billion parameters.
575 W
Total Power Draw
Thermal Design Power — how much electricity it draws under load. 575 W is roughly the same as running four gaming consoles at once, or about half an average home's continuous draw (which is ~1,200 W).
~$4,500 (street price, up from $1,999 MSRP due to GDDR7 shortage)
Street Price
MSRP at launch (January 2025) was $1,999, but a GDDR7 memory shortage drove street prices to $4,300–$4,900 by August 2026. You're paying a supply-chain premium — if that gap concerns you, consider waiting or stepping up to the RTX PRO 6000 Blackwell for more memory at a more predictable enterprise price.

⚡ Power in everyday terms

🏠 0.48× average homes running simultaneously (avg home = 1,200 W)
🔋 Drains a 90 kWh EV battery in 156.5 hours

✅ Good for

  • Inference on models up to 25 B params
  • Small teams experimenting with open-source LLMs
  • Single-user workstations

❌ Not the right fit for

  • The three models on our AI Model page (all require 340+ GB)
  • Any production workload needing high throughput
Form factor: PCIe Card (fits in a desktop)

Workstation GPU

NVIDIA RTX PRO 6000 Blackwell

Workstation Power, Pro-Grade Memory

$8,999
Market estimate
96 GB
GDDR7 VRAM
96 gigabytes — three times an RTX 5090. At full precision, this fits models up to about 80 billion parameters. For the models on our AI page you'd still need multiple cards or a server, but this card alone handles many useful open-source 70B models like Llama 3 70B.
600 W
Total Power Draw
600 W is about half an average home's continuous draw. Expect a dedicated 20 A circuit in your workspace.
$8,999
Street Price
Workstation GPUs carry a premium for ECC memory protection and pro drivers. Enterprise buyers often pay up to $13,000 for volume orders with support contracts.

⚡ Power in everyday terms

🏠 0.50× average homes running simultaneously (avg home = 1,200 W)
🔋 Drains a 90 kWh EV battery in 150.0 hours

✅ Good for

  • 70 B parameter models
  • Multi-GPU workstations (stack 4 for ~384 GB)
  • Serious researchers and ML teams

❌ Not the right fit for

  • The three models on our AI page as a single card
  • Data-center-scale throughput
Form factor: PCIe Card (requires a workstation, not a desktop)

Server System

NVIDIA H200 HGX 8-GPU Server

The Professional AI Server

$370,000
Market estimate
1,128 GB
HBM3e VRAM (8× GPUs)
HBM3e (High Bandwidth Memory) is a different type of memory than GDDR7 — it sits directly on the same package as the GPU chip, giving extreme bandwidth. Each H200 has 141 GB; eight of them connected via NVLink give you 1,128 GB of shared GPU memory. That's enough to run all three models on our AI page.
5,600 W
Total Power Draw
The whole server pulls 5.6 kilowatts — about 4.7 average homes running simultaneously. You'll need a dedicated power circuit, and data-center-grade cooling (CRAC units, raised floors, or liquid cooling). Not for a typical office.
$370,000
Street Price
This is professional infrastructure. At this level you're buying reliability, NVLink interconnects, enterprise support, and the ability to treat all 1,128 GB as one big shared pool of GPU memory. Compare to renting equivalent cloud compute, which runs $15–25/hr, meaning you break even in roughly 2–3 years.
NVLink
GPU Interconnect
NVLink lets all 8 GPUs share memory as if they were one giant GPU. This is real hardware-level clustering — not the software coordination trick of linking multiple desktop PCs. Data moves between GPUs at ~900 GB/s, vs ~64 GB/s for a PCIe desktop.

⚡ Power in everyday terms

🏠 4.67× average homes running simultaneously (avg home = 1,200 W)
🔋 Drains a 90 kWh EV battery in 16.1 hours
⚠️ Requires professional data-center power — dedicated circuits, industrial cooling, and proper infrastructure planning.

✅ Good for

  • All three models on our AI Model page
  • Production inference at scale
  • Organizations needing enterprise support

❌ Not the right fit for

  • Tight budgets
  • Offices without data-center power and cooling
Form factor: 8-GPU rack server (fits 4U rack space)

Datacenter Rack

NVIDIA GB300 NVL72 Rack

An Entire Data Center in One Cabinet

$3,700,000 – $6,500,000
Market estimate
20,480 GB
HBM3e VRAM (72× GPUs)
20 terabytes of GPU memory — all accessible as a single address space through NVLink. For reference, our largest listed model (GLM-5.2 at 744B params) only needs 893 GB. This rack can run that model with 22× headroom to spare, or run dozens of large models simultaneously.
137,000 W
Total Power Draw
137 kilowatts is the electricity draw of about 114 average homes running at the same time. At $0.10/kWh, that's roughly $120,000/year in electricity alone. This requires a dedicated data-center power delivery system (PDU, UPS, generator backup) and industrial-scale cooling — liquid cooling is effectively mandatory at this wattage.
$3,700,000 – $6,500,000
Street Price
Price varies by configuration, memory capacity, and negotiated enterprise discounts. At the high end this is equivalent to hiring 30+ senior engineers for a year. Most organizations at this scale are replacing $10M+/year cloud GPU bills, making the economics compelling within 12–18 months.
NVLink
GPU Interconnect
The GB300 NVL72's NVLink domain is categorically different from just buying multiple desktop PCs. All 72 GPUs share one memory address space — a model can use any part of the 20 TB freely. Desktop cards connected over a network must coordinate via slower message-passing, adding latency and reducing throughput by 10–100×.

⚡ Power in everyday terms

🏠 114.17× average homes running simultaneously (avg home = 1,200 W)
🔋 Drains a 90 kWh EV battery in 39 minutes
⚠️ Requires professional data-center power — dedicated circuits, industrial cooling, and proper infrastructure planning.

✅ Good for

  • The world's largest open-source models
  • Serving dozens of models simultaneously
  • Training, not just inference
  • Replacing significant cloud GPU spend

❌ Not the right fit for

  • Anywhere without industrial power and cooling infrastructure
  • Workloads under $1M/year cloud spend
Form factor: Full datacenter rack (~2 meters tall)