RTX PRO 6000 server rental: 8 GPUs and 768 GB of GPU memory.
Rent a dedicated bare-metal server with eight NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs. The whole machine is yours, with full root access, a fixed monthly price and a 99.5% SLA, hosted near Sofia, inside the EU.
FROM $1.70 PER GPU-HOURPOWER, COOLING, INTERNET AND RACK INCLUDED+359 882 490 689WHATSAPP, ANY TIME
One server, set up for AI work around the clock.
We deliver it with Ubuntu Server 24.04 LTS and NVIDIA drivers, or with your own operating system, after full hardware diagnostics.
- GPUs
- 8x NVIDIA RTX PRO 6000 Blackwell Server Edition
- GPU memory
- 96 GB GDDR7 per GPU, 768 GB in total
- GPU link
- PCIe 5.0 x16 per GPU (no NVLink)
- CPUs
- 2x AMD EPYC 9754, 256 cores and 512 threads in total
- System memory
- 1.5 TB DDR5-4800 ECC
- Storage
- 2x 4 TB Samsung 990 PRO NVMe
- Network
- 10 GbE interface on a dedicated 1 Gbps uplink, no traffic charges, static public IPv4
- Access
- Full root, plus BMC console, virtual media and power control
- Location
- Neterra data centre near Sofia, Bulgaria (EU)
- SLA
- 99.5% monthly availability, with service credits
- Support
- 24/7 for critical issues



More memory per card, and FP4 built in.
Large models fit on fewer cards. A 70B model in FP8 runs on a single GPU, with memory left for the KV cache.
Fifth-generation Tensor Cores run FP4 natively. It halves the memory per weight compared with FP8 and speeds up quantized models.
With Multi-Instance GPU (MIG), each card splits into up to four fully isolated instances, so one server can host up to 32 separate workloads.
The Server Edition has no fans of its own. The chassis cools it, and it runs at up to 600 W per card, around the clock.
How large models map onto eight GPUs.
Weights need about 2 bytes per parameter in FP16 or BF16, 1 byte in FP8 and half a byte in FP4. Leave another 10 to 30% for the KV cache, more for long contexts or large batches.
| Model | Precision | GPUs per copy | Copies per server |
|---|---|---|---|
| 70B dense model, such as Llama 3.3 70B | FP8 | 1 to 2 | 4 to 8 |
| gpt-oss-120b | MXFP4 | 1 | up to 8 |
| 235B mixture of experts, such as Qwen3-235B-A22B | FP8 | 4 | 2 |
| Very large mixture of experts, DeepSeek-V3 class | FP4 | 8 | 1 |
The GPUs talk to each other over PCIe 5.0, not NVLink. For inference, run each model on 1, 2 or 4 GPUs and scale with more copies. Splitting one model across all eight works, but each extra GPU adds less than on NVLink systems. We compare the numbers in RTX PRO 6000 vs H100 for LLM inference.
A whole server for the month, not GPUs by the hour.
| NEOXIS dedicated server | Hourly GPU marketplaces | |
|---|---|---|
| What you rent | The whole machine | Single GPUs or containers |
| Price | Fixed monthly, from $1.70 per GPU-hour, everything included | Per hour, changes with demand |
| Availability | Reserved for you for the whole term | Others can take it between your rentals |
| Access | Full root, your own OS, BMC console | Usually a container on someone else's host |
| SLA | 99.5% monthly, with service credits | Usually none |
| Best for | Steady inference, long jobs, private data | Short tests and bursts |
From $1.70 per GPU-hour, everything included.
That is about $9,930 a month for the whole server with eight GPUs, fixed for the term. There are no extra costs: we pay for power, cooling, internet (1 Gbps, unmetered) and rack space, and support is included. Longer commitments get better rates, and prices exclude VAT where it applies.
Tell us your workload and timeline, and we usually reply within one business day. For a quick answer, call +359 882 490 689 or message us on WhatsApp, at any time.