Local first. Your data stays in the building.
souveraen.ai runs on an NVIDIA DGX Spark in your server room. You use external models only if you want to, and you pay a fixed monthly price instead of tokens.
NVIDIA GB10 · 128 GB unified memory · data stays local
Everything the AI needs sits on one machine.
Your stores are read locally, the models compute locally, and the answers stay in the building. An external model only joins if you add it.
Models on the machine
Open models run on the appliance, with no interface to a vendor.
Knowledge on the machine
Passages, search indexes and the knowledge graph sit on the same appliance, with access rights on every passage.
Tools in a signed package
Reviewed MCP servers arrive as a signed package, nothing is fetched at runtime. Admins or, depending on permissions, users can add more servers.
Private Spark for a department, Air-Gap for the rest.
One DGX Spark is sized for a department. If you need more compute, redundancy or fully isolated operation, we size the hardware with you.
All prices net, excluding VAT. What is in your quote applies.
LOCAL SPEED
How fast a model answers on your appliance.
Choose a model. The answer appears exactly as fast as that model generates tokens on a DGX Spark – measured for a single request.
The maintenance contract with Nordtech GmbH has been running since 1 March 2025 and renews for twelve months at a time. The key points: 1. Termination: three months before the end of the term, in writing. 2. Response time: four hours on working days, 24 hours at weekends. 3. Price: €1,850 net per month, adjusted annually by at most 3%. 4. Liability: capped at the annual fee, except in cases of intent. The next possible termination date is 28 February 2027; the notice period for it ends on 30 November 2026.
0 / 140
Tokens
36.8
Tokens per second
0.0
Seconds
For comparison: adults read silently at around four words per second.
Measured for a single request without a thinking phase. Qwen 3.5: llama.cpp with 4-bit quantisation (Q4_K_M), 512 generated tokens, published in the BAEM1N llm-bench benchmark (2026). gpt-oss-120b: our own measurement with vLLM and MXFP4 on a DGX Spark, without speculative decoding. Optimised runtimes reach more, for example around 52 tokens per second for Qwen 3.5 122B-A10B with speculative decoding. When several people work at the same time, they share the capacity. Measured on our reference DGX in Auerbach. The figures are for Qwen 3.5; the current generation in the catalogue is Qwen3.6.
Included in the price
The size of a thick book. One socket, one network cable.
The DGX Spark sits in your server room or under a desk. We set it up remotely and train your team. Your data lives on the machine; how the backup works is set out in the quote.
Klaus Bauer can tell you whether one machine is enough for your team: +49 3744 365 2202.
0 h
onboarding workshop, remote
0 days
to a replacement unit, in working days
0 months
term, billed monthly
9 €
token costs for local models
What IT and management want to know first.
About hardware, network and operation.