Skip to content
souveraen.ai
PRIVATE SPARK

A computer in your building, where your AI runs. Available now.

Private Spark is souveraen.ai on an NVIDIA DGX Spark that we put in your building. The models run on it, your knowledge sits on it, and your agents work on it. By default, it keeps a management connection to souveraen.ai.

  • 2.790 € net per month including Spark DGX hardware
  • 24-month term. Buy the DGX up front and the monthly price drops to 2.490 € per month.

Data sources

  • Contracts
  • Manuals
  • Emails

Local model

gpt-oss-120b

about 35 tokens/s

NVIDIA DGX Spark

128 GB unified memory

  • Data, index and models stay local
  • Managed via souveraen.ai

LOCAL SPEED

How fast a model answers on your appliance.

Choose a model. The answer appears exactly as fast as that model generates tokens on a DGX Spark – measured for a single request.

Model
souveraen.ai · Qwen 3.5 9B
You
souveraen.ai

The maintenance contract with Nordtech GmbH has been running since 1 March 2025 and renews for twelve months at a time. The key points: 1. Termination: three months before the end of the term, in writing. 2. Response time: four hours on working days, 24 hours at weekends. 3. Price: €1,850 net per month, adjusted annually by at most 3%. 4. Liability: capped at the annual fee, except in cases of intent. The next possible termination date is 28 February 2027; the notice period for it ends on 30 November 2026.

0 / 140

Tokens

36.8

Tokens per second

0.0

Seconds

For comparison: adults read silently at around four words per second.

Measured for a single request without a thinking phase. Qwen 3.5: llama.cpp with 4-bit quantisation (Q4_K_M), 512 generated tokens, published in the BAEM1N llm-bench benchmark (2026). gpt-oss-120b: our own measurement with vLLM and MXFP4 on a DGX Spark, without speculative decoding. Optimised runtimes reach more, for example around 52 tokens per second for Qwen 3.5 122B-A10B with speculative decoding. When several people work at the same time, they share the capacity. Measured on our reference DGX in Auerbach. The figures are for Qwen 3.5; the current generation in the catalogue is Qwen3.6.

IN BRIEF

What is a token?

Language models read and write text neither letter by letter nor word by word, but in tokens: short pieces of text from a fixed vocabulary. Common words are a single token; rare or compound words are split into several pieces.

Tokens per second therefore tell you how fast a model produces text. What matters most is how many parameters it has to read from memory for each token: mixture-of-experts models such as Qwen 3.5 35B-A3B use only part of their parameters per token, which makes them faster than smaller models that always read all of theirs.

This is how a Qwen model splits the word “unenforceable”:

4 tokens

German needs more tokens than English: the sample answer above has 91 words and 140 tokens, the German version 80 words and 180 tokens.

As a rule of thumb, from around 10 tokens per second a model writes faster than you can read along.

LOCAL OPERATION

What the price
includes.

  • Runs inside your infrastructure, external models optional
  • One DGX Spark is sized for a department, not a whole group
  • Onboarding workshop (4 h, remote) included

The DGX Spark is about the size of a thick book. It needs a power socket and a network cable, and it sits in your server room or under a desk. We set it up remotely; the 4-hour onboarding workshop is included. If it fails, you get a replacement within five working days. Your data is on the machine, so a backup is part of the quote. Talk to Klaus Bauer on +49 3744 365 2202. We will tell you whether one DGX Spark is enough for your team.

Start on the free plan

2.790 €

net per month

including Spark DGX hardware

24-month term

Buy the DGX up front and the monthly price drops to 2.490 € per month

NON-BINDING OFFER SUMMARY

Private Spark

This summary does not place an order. You can complete it in your account. The exact scope is set out in the quote. 2.790 € net per month, including Spark DGX hardware. Buy the DGX up front and the monthly price drops to 2.490 € per month. 24-month term. Billed monthly.