Private LLM hosting

Private LLM hosting on your own GPU server.

Run open-weight models such as Llama, Qwen, Mistral or gpt-oss on a dedicated server with 768 GB of GPU memory, hosted in the EU. Your prompts and data stay on your hardware, and the monthly cost stays the same however many tokens you generate.

FROM $1.70 PER GPU-HOURPOWER, COOLING, INTERNET AND RACK INCLUDED+359 882 490 689WHATSAPP, ANY TIME

01 · Why private

Why teams move from APIs to their own LLM server.

YOUR DATA STAYS WITH YOU

Prompts, documents and answers are processed on your server. Nothing is sent to a third-party API.

A FIXED MONTHLY COST

There is no per-token billing. Heavy use does not raise the bill, so you can plan the budget for the whole term.

NO RATE LIMITS

The GPUs are yours. Capacity does not depend on someone else's traffic, and there is no queue at busy times.

YOUR CHOICE OF MODELS

Run open-weight models, your own fine-tuned versions or adapters, and switch models whenever you need to.

02 · How it works

From request to your own API in four steps.

  1. STEP 01

    Tell us the models and the load

    Which models, how many users and what context length. We confirm the configuration and the start date.

  2. STEP 02

    We hand over the server

    A clean install of Ubuntu Server 24.04 LTS with NVIDIA drivers, or your own OS, after full hardware diagnostics.

  3. STEP 03

    You deploy your stack

    Run vLLM, SGLang or Ollama. vLLM and SGLang expose an OpenAI-compatible API, so existing clients only need a new base URL.

  4. STEP 04

    Scale inside the server

    Run several copies of a model, split GPUs with MIG, or move to larger models as your needs grow.

03 · What fits

768 GB of GPU memory in one machine.

Eight NVIDIA RTX PRO 6000 Blackwell GPUs with 96 GB each. See the sizing guide for more models.

ModelPrecisionGPUs per copyCopies per server
70B dense model, such as Llama 3.3 70BFP81 to 24 to 8
gpt-oss-120bMXFP41up to 8
235B mixture of experts, such as Qwen3-235B-A22BFP842
04 · API or private

When a private LLM pays off.

APIs charge per token, so the bill grows with usage. A dedicated server costs the same every month. Where the two lines cross depends on your volume, the model and the API you compare with:

  • Monthly API cost is about the tokens per month divided by one million, times the price per million tokens, with input and output priced separately.
  • Monthly private cost is the fixed server price, from $1.70 per GPU-hour or about $9,930 a month for all eight GPUs with power, cooling, internet and rack space included, plus the time your team spends running the stack.

With steady, high usage, private hosting usually costs less. For occasional or unpredictable use, an API is often cheaper. Send us your numbers and we will compare them with you. The full price breakdown is on the pricing page.

05 · Security and privacy

Single tenant, in the EU, under your control.

SINGLE TENANT

No other customers share the machine, its GPUs, its memory or its drives.

YOUR SYSTEM, YOUR RULES

You get full root and manage your own firewall. We open ports at the network level only on your request.

IN THE EU

The server runs at Neterra's data centre near Sofia, Bulgaria, so your data stays inside the European Union.

CLEAN HANDOVER

We do not log in to your operating system. At the end of the term, the drives are wiped before the server is used again.

06 · Questions

About private LLM hosting.

Which models can we run?
Any open-weight model that fits in 768 GB of GPU memory, for example Llama, Qwen, Mistral, DeepSeek or gpt-oss, as well as your own fine-tuned versions.
Do you have access to our data?
No. You have full root access and manage the operating system yourself. We look after power and hardware, and at the end of the term the drives are wiped.
Can we offer an OpenAI-compatible API to our users?
Yes. vLLM and SGLang both provide OpenAI-compatible endpoints, so existing clients and SDKs work after a change of the base URL.
Does it help with GDPR?
The server is in Bulgaria, inside the EU, and you decide what data is processed and stored on it. That makes GDPR requirements easier to meet than sending data to a provider outside the EU.
What happens if a GPU fails?
Hardware faults are covered by our SLA, and critical issues are handled 24/7. We replace failed parts, and if availability falls below the SLA, service credits apply automatically.