🚀 I'm a Small Startup Run one big model, lowest possible cost
or
🏢 I'm a Mid-Size Company Serious capacity, room to grow
🚀

The Small Startup Path

You want to self-host a frontier open-source model without paying cloud GPU bills month after month. You're cost-sensitive. You want something that actually works — not enterprise overhead.

1

What model do you want to run?

You're looking at the three models from our AI page — all between 284B and 753B parameters. All of them need 341–904 GB of GPU memory.

GLM-5.2 (Max) 753B params → 904 GB needed
Tencent Hy3 295B params → 354 GB needed
DeepSeek V4-Flash 284B params → 341 GB needed
2

Can any desktop or workstation card handle these?

No — not as a single card. The RTX 5090 has 32 GB. The RTX PRO 6000 has 96 GB. Neither fits even the smallest model (341 GB needed). You'd need multiple workstation cards with tensor parallelism, which adds latency and complexity — not practical for a startup's first deployment.
3

What's the minimum that actually works?

The 8× H200 HGX Server — $370,000. Eight H200 GPUs connected via NVLink give you 1,128 GB of shared memory. That's enough to run all three models, with headroom for the KV cache and concurrent requests.
STARTUP RECOMMENDATION

Small Startup: One Big Model, Lowest Cost

You want to self-host a frontier open-source model without paying cloud bills forever. The minimum hardware that can run any of the three models on our page.

Server System NVIDIA H200 HGX 8-GPU Server
× 1
$370,000
Total GPU Memory 1,128 GB
Total Power Draw 5,600 W
= 4.7× average homes | drains a 90 kWh EV in ~16 hours
Total Price $370,000

Why this build?

All three listed models need 341–893 GB of GPU memory. A single 8×H200 server has 1,128 GB — enough for any of them, with breathing room for inference overhead. There is no cheaper path to running these specific models on your own hardware.

When you need more

If you outgrow one server, add a second H200 system and connect them over InfiniBand for ~2,256 GB total.


🏢

The Mid-Size Company Path

You need serious capacity — multiple models, multiple teams, room to grow. Downtime costs you money. You're thinking in years, not months. Cloud GPU bills are becoming significant.

1

Why not just buy more H200 servers?

You could. Three H200 servers give you ~3,384 GB for $1.1M, connected over InfiniBand. But InfiniBand bandwidth (~400 GB/s) is lower than NVLink (~900 GB/s), and you're managing three separate machines, three power circuits, three failure domains. The GB300 NVL72 gives you more memory at a single point of management.
2

What does "room to grow" actually mean?

The GB300 NVL72's 20 TB of pooled memory can run all three listed models simultaneously — each getting 600–900 GB — while still having 17+ TB available for fine-tuning, new model evaluations, and future models that don't even exist yet. One rack. One management interface. One infrastructure contract.
3

Is the economics right for my company?

At $0.10/kWh, the GB300's 137 kW draw costs ~$120,000/year in electricity. Add staff, cooling, and facilities — call it $300–500K/year in operating costs. If your equivalent cloud GPU bill exceeds that, you're already ahead. Most mid-size companies crossing $2M/year in GPU cloud spend find on-premise payback under 3 years.
MID-SIZE RECOMMENDATION

Mid-Size Company: Serious Capacity and Room to Grow

You need to serve multiple models or multiple teams, can't afford downtime, and expect your workloads to grow. One rack handles it all.

Datacenter Rack NVIDIA GB300 NVL72 Rack
× 1
$3,700,000 – $6,500,000
Total GPU Memory 20,480 GB (20 TB)
Total Power Draw 137,000 W (~137 kW)
= 114× average homes | drains a 90 kWh EV in ~39 minutes
Total Price $3,700,000 – $6,500,000

Why this build?

20 TB of pooled GPU memory lets you run all three listed models simultaneously — each model gets 600–900 GB of headroom, still leaving 17+ TB for other workloads. NVLink at 72-GPU scale is true datacenter clustering: models can span all 72 GPUs with no software tricks, delivering maximum throughput per dollar at this scale.

When you need more

Most organizations configure two NVL72 racks: one production, one for training and evaluation. Your team can fine-tune on new data without touching the serving rack.

What makes this a "real" cluster?

The GB300 NVL72's 72 GPUs are connected by NVLink Switch — a dedicated high-speed fabric operating at ~900 GB/s between any two GPUs. All 72 GPUs share one memory address space. A 753B parameter model can be split across all 72 chips seamlessly, as if they were one massive GPU.

Compare this to buying 72 separate RTX 5090s and networking them: you'd get 72 isolated 32 GB pools, connected by Ethernet or InfiniBand at 10–100× lower bandwidth. Software (tensor parallelism) can coordinate them, but each inter-GPU data transfer crosses a slow link. For large-model inference, latency multiplies with each network hop.

MetricGB300 NVL72 (NVLink)72× RTX 5090 (networked)
Total memory20 TB (one pool)2,304 GB (72 isolated pools)
GPU-to-GPU bandwidth~900 GB/s25–400 GB/s (network)
Memory sharingHardware-level, seamlessSoftware coordination required
Inference latencyLow — no network hopHigher — crosses interconnect
Management overheadOne system72 nodes to manage
Approximate cost$3.7M–$6.5M~$173K (but can't run 744B+ models)