Server · RTX PRO 6000 Blackwell

RTX PRO 6000 server rental: 8 GPUs and 768 GB of GPU memory.

Rent a dedicated bare-metal server with eight NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs. The whole machine is yours, with full root access, a fixed monthly price and a 99.5% SLA, hosted near Sofia, inside the EU.

FROM $1.70 PER GPU-HOURPOWER, COOLING, INTERNET AND RACK INCLUDED+359 882 490 689WHATSAPP, ANY TIME

01 · Specifications

One server, set up for AI work around the clock.

We deliver it with Ubuntu Server 24.04 LTS and NVIDIA drivers, or with your own operating system, after full hardware diagnostics.

GPUs
8x NVIDIA RTX PRO 6000 Blackwell Server Edition
GPU memory
96 GB GDDR7 per GPU, 768 GB in total
GPU link
PCIe 5.0 x16 per GPU (no NVLink)
CPUs
2x AMD EPYC 9754, 256 cores and 512 threads in total
System memory
1.5 TB DDR5-4800 ECC
Storage
2x 4 TB Samsung 990 PRO NVMe
Network
10 GbE interface on a dedicated 1 Gbps uplink, no traffic charges, static public IPv4
Access
Full root, plus BMC console, virtual media and power control
Location
Neterra data centre near Sofia, Bulgaria (EU)
SLA
99.5% monthly availability, with service credits
Support
24/7 for critical issues
THE ACTUAL MACHINE
Open chassis of the NEOXIS server with eight NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs, three fan modules and two AMD EPYC CPUs
Before installation: eight RTX PRO 6000 GPUs, two EPYC 9754 CPUs and 1.5 TB of memory.
The NEOXIS 8-GPU server mounted in a rack at Neterra SDC Stolnik near Sofia
In the rack at Neterra SDC Stolnik, near Sofia.
Rear of the NEOXIS server with three hot-swap fan modules
Three hot-swap fan modules at the rear of the chassis.
02 · Why this GPU

More memory per card, and FP4 built in.

96 GB PER GPU

Large models fit on fewer cards. A 70B model in FP8 runs on a single GPU, with memory left for the KV cache.

FP4 ON BLACKWELL

Fifth-generation Tensor Cores run FP4 natively. It halves the memory per weight compared with FP8 and speeds up quantized models.

UP TO 32 ISOLATED INSTANCES

With Multi-Instance GPU (MIG), each card splits into up to four fully isolated instances, so one server can host up to 32 separate workloads.

BUILT FOR RACKS

The Server Edition has no fans of its own. The chassis cools it, and it runs at up to 600 W per card, around the clock.

03 · What runs on it

How large models map onto eight GPUs.

Weights need about 2 bytes per parameter in FP16 or BF16, 1 byte in FP8 and half a byte in FP4. Leave another 10 to 30% for the KV cache, more for long contexts or large batches.

ModelPrecisionGPUs per copyCopies per server
70B dense model, such as Llama 3.3 70BFP81 to 24 to 8
gpt-oss-120bMXFP41up to 8
235B mixture of experts, such as Qwen3-235B-A22BFP842
Very large mixture of experts, DeepSeek-V3 classFP481
A NOTE ON NVLINK

The GPUs talk to each other over PCIe 5.0, not NVLink. For inference, run each model on 1, 2 or 4 GPUs and scale with more copies. Splitting one model across all eight works, but each extra GPU adds less than on NVLink systems. We compare the numbers in RTX PRO 6000 vs H100 for LLM inference.

04 · Monthly or hourly

A whole server for the month, not GPUs by the hour.

NEOXIS dedicated serverHourly GPU marketplaces
What you rentThe whole machineSingle GPUs or containers
PriceFixed monthly, from $1.70 per GPU-hour, everything includedPer hour, changes with demand
AvailabilityReserved for you for the whole termOthers can take it between your rentals
AccessFull root, your own OS, BMC consoleUsually a container on someone else's host
SLA99.5% monthly, with service creditsUsually none
Best forSteady inference, long jobs, private dataShort tests and bursts
05 · Price

From $1.70 per GPU-hour, everything included.

That is about $9,930 a month for the whole server with eight GPUs, fixed for the term. There are no extra costs: we pay for power, cooling, internet (1 Gbps, unmetered) and rack space, and support is included. Longer commitments get better rates, and prices exclude VAT where it applies.

Tell us your workload and timeline, and we usually reply within one business day. For a quick answer, call +359 882 490 689 or message us on WhatsApp, at any time.

06 · Questions

About the RTX PRO 6000 server.

How soon can we start?
Typically within 5 business days after signing and the first payment, subject to availability. We use that time for a clean install and full hardware diagnostics.
Can we install our own operating system?
Yes. Mount your ISO through the BMC virtual media and install what you need. By default we deliver Ubuntu Server 24.04 LTS with NVIDIA drivers.
Does the server have NVLink?
No. The RTX PRO 6000 Server Edition connects over PCIe 5.0 x16. For inference we recommend running models on 1, 2 or 4 GPUs and adding copies. Splitting one model across all eight GPUs works, but scales less efficiently than on NVLink systems.
Can one server run several workloads?
Yes. Each GPU supports Multi-Instance GPU (MIG) with up to four isolated instances, and you are free to run containers or virtual machines on the host.
Can we resell the capacity to our own customers?
Yes. GPU clouds and resellers are welcome. We only ask that your terms with your customers include the same acceptable use and export control rules as ours.
Is there a minimum term?
Terms are flexible, and longer commitments get better rates. Tell us your plans and we will suggest the options that fit.