Private LLM hosting on your own GPU server.
Run open-weight models such as Llama, Qwen, Mistral or gpt-oss on a dedicated server with 768 GB of GPU memory, hosted in the EU. Your prompts and data stay on your hardware, and the monthly cost stays the same however many tokens you generate.
FROM $1.70 PER GPU-HOURPOWER, COOLING, INTERNET AND RACK INCLUDED+359 882 490 689WHATSAPP, ANY TIME
Why teams move from APIs to their own LLM server.
Prompts, documents and answers are processed on your server. Nothing is sent to a third-party API.
There is no per-token billing. Heavy use does not raise the bill, so you can plan the budget for the whole term.
The GPUs are yours. Capacity does not depend on someone else's traffic, and there is no queue at busy times.
Run open-weight models, your own fine-tuned versions or adapters, and switch models whenever you need to.
From request to your own API in four steps.
- STEP 01
Tell us the models and the load
Which models, how many users and what context length. We confirm the configuration and the start date.
- STEP 02
We hand over the server
A clean install of Ubuntu Server 24.04 LTS with NVIDIA drivers, or your own OS, after full hardware diagnostics.
- STEP 03
You deploy your stack
Run vLLM, SGLang or Ollama. vLLM and SGLang expose an OpenAI-compatible API, so existing clients only need a new base URL.
- STEP 04
Scale inside the server
Run several copies of a model, split GPUs with MIG, or move to larger models as your needs grow.
768 GB of GPU memory in one machine.
Eight NVIDIA RTX PRO 6000 Blackwell GPUs with 96 GB each. See the sizing guide for more models.
| Model | Precision | GPUs per copy | Copies per server |
|---|---|---|---|
| 70B dense model, such as Llama 3.3 70B | FP8 | 1 to 2 | 4 to 8 |
| gpt-oss-120b | MXFP4 | 1 | up to 8 |
| 235B mixture of experts, such as Qwen3-235B-A22B | FP8 | 4 | 2 |
When a private LLM pays off.
APIs charge per token, so the bill grows with usage. A dedicated server costs the same every month. Where the two lines cross depends on your volume, the model and the API you compare with:
- Monthly API cost is about the tokens per month divided by one million, times the price per million tokens, with input and output priced separately.
- Monthly private cost is the fixed server price, from $1.70 per GPU-hour or about $9,930 a month for all eight GPUs with power, cooling, internet and rack space included, plus the time your team spends running the stack.
With steady, high usage, private hosting usually costs less. For occasional or unpredictable use, an API is often cheaper. Send us your numbers and we will compare them with you. The full price breakdown is on the pricing page.
Single tenant, in the EU, under your control.
No other customers share the machine, its GPUs, its memory or its drives.
You get full root and manage your own firewall. We open ports at the network level only on your request.
The server runs at Neterra's data centre near Sofia, Bulgaria, so your data stays inside the European Union.
We do not log in to your operating system. At the end of the term, the drives are wiped before the server is used again.