Skip to content
souveraen.ai
On-premise · Private Spark

Local first. Your data stays in the building.

souveraen.ai runs on an NVIDIA DGX Spark in your server room. You use external models only if you want to, and you pay a fixed monthly price instead of tokens.

NVIDIA GB10 · 128 GB unified memory · data stays local

Local first
Models, knowledge and agents run on the appliance in your building. A management connection to souveraen.ai is kept.
External models optional
Larger models from external vendors can be added if you want them, with the vendor's token costs.
Privacy on the device
Your document collection never leaves the building, not even with an external model you have added.
No token costs
A fixed monthly price from €2,490. Local requests cost nothing extra, however much your team asks.
What stays in the building

Everything the AI needs sits on one machine.

Your stores are read locally, the models compute locally, and the answers stay in the building. An external model only joins if you add it.

External modelsOptional, with token costs
File server
SharePoint
DGX Spark · your building
Nextcloud
Confluence

Models on the machine

Open models run on the appliance, with no interface to a vendor.

Qwen3.6gpt-oss-120bMistral Small 4

Knowledge on the machine

Passages, search indexes and the knowledge graph sit on the same appliance, with access rights on every passage.

Keyword indexVector indexKnowledge graph

Tools in a signed package

Reviewed MCP servers arrive as a signed package, nothing is fetched at runtime. Admins or, depending on permissions, users can add more servers.

SignedReviewedNo fetching
Two ways into your building

Private Spark for a department, Air-Gap for the rest.

One DGX Spark is sized for a department. If you need more compute, redundancy or fully isolated operation, we size the hardware with you.

Private Spark

Most chosen

DGX Spark included, for one department.

2.790 €net / month

including Spark DGX hardware · 24-month term

Request a quote

Included

Runs in your infrastructure, external models optional

DGX Spark included in the price

Onboarding workshop (4 h, remote)

Replacement unit within five working days

Private Spark

DGX Spark bought up front, lower monthly rate.

2.490 €net / month

plus Spark DGX hardware · 24-month term

Request a quote

Included

Same software, same features

DGX Spark bought once, up front

Onboarding workshop (4 h, remote)

24-month term

Air-Gap Enterprise

Fully isolated, hardware sized to your needs.

On quote
Book a call

Included

No connection to the outside

Hardware sized for capacity and availability

Connected to the systems you already run

Acceptance and handover on your terms

All prices net, excluding VAT. What is in your quote applies.

LOCAL SPEED

How fast a model answers on your appliance.

Choose a model. The answer appears exactly as fast as that model generates tokens on a DGX Spark – measured for a single request.

Model
souveraen.ai · Qwen 3.5 9B
You
souveraen.ai

The maintenance contract with Nordtech GmbH has been running since 1 March 2025 and renews for twelve months at a time. The key points: 1. Termination: three months before the end of the term, in writing. 2. Response time: four hours on working days, 24 hours at weekends. 3. Price: €1,850 net per month, adjusted annually by at most 3%. 4. Liability: capped at the annual fee, except in cases of intent. The next possible termination date is 28 February 2027; the notice period for it ends on 30 November 2026.

0 / 140

Tokens

36.8

Tokens per second

0.0

Seconds

For comparison: adults read silently at around four words per second.

Measured for a single request without a thinking phase. Qwen 3.5: llama.cpp with 4-bit quantisation (Q4_K_M), 512 generated tokens, published in the BAEM1N llm-bench benchmark (2026). gpt-oss-120b: our own measurement with vLLM and MXFP4 on a DGX Spark, without speculative decoding. Optimised runtimes reach more, for example around 52 tokens per second for Qwen 3.5 122B-A10B with speculative decoding. When several people work at the same time, they share the capacity. Measured on our reference DGX in Auerbach. The figures are for Qwen 3.5; the current generation in the catalogue is Qwen3.6.

Included in the price

The size of a thick book. One socket, one network cable.

The DGX Spark sits in your server room or under a desk. We set it up remotely and train your team. Your data lives on the machine; how the backup works is set out in the quote.

Klaus Bauer can tell you whether one machine is enough for your team: +49 3744 365 2202.

0 h

onboarding workshop, remote

0 days

to a replacement unit, in working days

0 months

term, billed monthly

9 €

token costs for local models

Questions

What IT and management want to know first.

About hardware, network and operation.

A DGX Spark is sized for a department, not a whole group. When several people work at once, they share its capacity. We work out with you beforehand whether one machine is enough or Air-Gap Enterprise fits better.

In your own building

Talk to us about your server room.

Tell us how many people will work with it and which stores should be connected. You get a quote covering hardware, models and setup.

  1. Clarify what you need1
  2. Receive a quote2
  3. Set up the appliance3