Blog ·

Self-hosted AI agents: when your own GPU server beats the API bill

By the NEOXIS team · 7 min read

Agent frameworks such as OpenClaw and Hermes Agent made it easy to run an assistant that works all day: it reads mail, calls tools and writes back. The next question is the bill. Every step of an agent is a paid request, and the requests add up faster than in a chat.

This guide prices one real conversation three ways: on a frontier API, on an API for an open model, and on a dedicated GPU server that you rent by the month. The answer depends on volume, on the model you need and on where your data may go.

If Choose Why
You need the strongest model, or fewer than about 68,000 agent conversations a month A frontier API such as Claude No servers to run, the best models, you pay only for what you use
An open model is good enough and you want the lowest price per token An API for an open model, from $0.03 per million input tokens Hard to beat on price at almost any volume
An open model is good enough, the agents run around the clock, and data must stay in your hands Your own GPU server, such as our 8x RTX PRO 6000 from $9,930 a month A fixed bill, no rate limits, no data leaving your servers

Why agents burn tokens

A chat sends one request per message. An agent sends one request per step: it reads the conversation, decides to call a tool, gets the result back and asks the model again. Each request resends the conversation so far, so the input grows with every step.

Aura Digital, the team that builds this site, measured it on a production assistant for an online shop, which searches the product catalogue before it answers. One conversation of five messages from the customer made 10 requests to the model and used 60,507 input tokens and 2,593 output tokens. Two requests per message is a modest agent; frameworks that plan, browse and retry make more.

The ratio matters more than the total: the model read 23 tokens for every token it wrote. For agents, the price of input tokens decides the bill.

What one conversation costs on an API

The measured conversation, priced at the list prices on 9 October 2026:

Model and price per million tokens (input / output) One conversation
Claude Sonnet 5, $2 / $10, no caching $0.147
Claude Sonnet 5, with 80% of the input read from the prompt cache $0.066
Claude Haiku 5.5, $0.10 / $0.50 $0.007
gpt-oss-120b through an API, typical price $0.15 / $0.60 $0.011
gpt-oss-120b through an API, lowest price $0.03 / $0.17 $0.002

Prompt caching matters for agents, because most of each request repeats the previous one. On Claude, a cache read costs a tenth of the normal input price and writing to the cache costs a quarter more. The 80% share is our assumption for a long conversation; your share depends on how the agent builds its requests.

What a dedicated GPU server costs and what it runs

Our server has eight NVIDIA RTX PRO 6000 Blackwell Server Edition cards with 96 GB each, 768 GB in total. It costs from $1.70 per GPU-hour, about $9,930 a month for the whole machine, with power, cooling and a 1 Gbps connection included. The bill is the same whether the agents are quiet or busy.

An independent vLLM benchmark on one RTX PRO 6000 Server Edition shows what a single card does with 50 requests at once:

Model on one card Memory Output tokens per second First token
gpt-oss-120b, 4-bit 61 GB 1,525 342 ms
Qwen3-14B 28 GB 1,744 510 ms
Qwen3-8B 15 GB 2,658 319 ms

gpt-oss-120b fits on one card, so the server can run eight copies of it side by side, about 12,000 output tokens per second in total at that load. One caution: the benchmark used prompts of 100 tokens, while agents send tens of thousands. Reading long prompts takes extra GPU time. vLLM reduces it with automatic prefix caching, which keeps the processed conversation history and reuses it on the next request instead of reading it again.

Where the lines cross

Divide the monthly price of the server by the price of one conversation, and you get the number of conversations a month at which the server costs the same as the API:

Compared with Conversations a month to break even Per day
Claude Sonnet 5, no caching about 68,000 about 2,300
Claude Sonnet 5, 80% cached about 151,000 about 5,000
gpt-oss-120b API, typical price about 934,000 about 31,000
Claude Haiku 5.5 about 1,350,000 about 45,000

Against a frontier model, the crossover is within reach of a busy product or a team of agents that run all day: 2,300 to 5,000 conversations a day. At that volume the server writes on average 70 to 150 output tokens per second, about 1% of what eight cards can produce, so it has room to grow at no extra cost.

Against cheap APIs for open models, the crossover needs close to a million conversations a month. At that point the server must also read about 22,000 input tokens per second around the clock. That would have to be measured with your own prompts before you count on it.

The comparison with Claude is only fair if an open model can do the job. Test your agent on gpt-oss-120b or a Qwen model first; if it fails tasks that Claude handles, the price of the server does not matter.

When the API is still the better choice

  • Low or uneven volume. Below the numbers above, the API is cheaper, and you pay nothing in quiet months.
  • You need the strongest model. The best closed models are not available to run on your own hardware.
  • Nobody on the team runs servers. A server needs updates, monitoring and someone to restart the inference engine. We handle the hardware, power and network, but the software on the server is yours.
  • You only need an open model at a low price. APIs for open models start at a few cents per million tokens, and there is no idle time to pay for.

When the decision is not about price

  • Data that may not leave your servers. Mail, contracts, patient records or source code read by an agent stay on a machine you control, in the EU, with GDPR data processing terms in the contract. Our private LLM hosting page covers this in detail.
  • Agents that never stop. A personal agent that checks mail every few minutes has no quiet hours. A fixed price removes the worry about a runaway loop on a per-token bill.
  • No rate limits. The only limit is the hardware. Ten agents can run in parallel without waiting for a quota.
  • Your own models. Fine-tuned or merged models run on your server the same way as the public ones.

Running OpenClaw or Hermes Agent on your own model

Both are open source under the MIT licence and work with local models as well as cloud APIs. OpenClaw describes itself as an assistant that runs on your own computer and connects to more than 20 chat channels. Hermes Agent, from Nous Research, lists Ollama, vLLM and llama.cpp among its local back ends and can import an existing OpenClaw setup.

The usual set-up is short. Start vLLM on the server with the model you chose; it serves an OpenAI-compatible API. Then point the agent at that address instead of a cloud provider. On one 8-GPU server you can run one model per card for many agents, or one large model across several cards.

An agent can run commands and read files on the machine it controls, so give it its own user and only the access it needs. Our dedicated GPU server comes with full root access, and nothing on it is shared with other customers.

Questions

Can one RTX PRO 6000 run gpt-oss-120b? Yes. The 4-bit model takes 61 GB of the card's 96 GB, which leaves room for the context of many conversations at once.

Is an open model good enough for agents? It depends on the task. Simple tool calls and fixed workflows often work well; long planning and tricky reasoning are where the best closed models still lead. Test with your own tasks before you compare prices.

Does prompt caching change the result? Yes. It cuts the cost of our measured conversation on Claude by more than half and moves the break-even point from about 68,000 to about 151,000 conversations a month.

Sources