Which AI models run for you.
Open models such as Qwen, Mistral and gpt-oss run on your appliance, and your data stays on it. Larger open models are added in hosting in Germany. External models are available in On Demand depending on the plan and are processed in the vendors' EU and zero-data-retention (ZDR) tenants.
- Local, or offline with Air-Gap
- Open models out of the box
- One interface for every task
Model catalogue
On the appliance, out of the box
22
Mistral Small 4
Mistral AI
- Context window
- 256,000 tokens
- Max output
- –
ChatAvailable locallyMinistral 3 8B
Mistral AI
- Context window
- 256,000 tokens
- Max output
- –
ChatAvailable locallyDevstral Small 2
Mistral AI
- Context window
- 256,000 tokens
- Max output
- –
CodeAvailable locallyMagistral Small 1.2
Mistral AI
- Context window
- 128,000 tokens
- Max output
- –
ChatAvailable locallyQwen3.6-35B-A3B
Qwen (Alibaba)
- Context window
- 256,000 tokens
- Max output
- –
ChatAvailable locallyQwen3.6-27B
Qwen (Alibaba)
- Context window
- 256,000 tokens
- Max output
- –
ChatAvailable locallyQwen3-Coder-30B-A3B
Qwen (Alibaba)
- Context window
- 262,144 tokens
- Max output
- –
CodeAvailable locallyQwen3-30B-A3B-2507
Qwen (Alibaba)
- Context window
- 262,144 tokens
- Max output
- –
ChatAvailable locallyQwen3-14B
Qwen (Alibaba)
- Context window
- 32,768 tokens
- Max output
- –
ChatAvailable locallyQwen3-8B
Qwen (Alibaba)
- Context window
- 32,768 tokens
- Max output
- –
ChatAvailable locallyQwen3 Embedding 8B
Qwen (Alibaba)
- Context window
- –
- Max output
- –
EmbeddingAvailable locallyGemma 4
Google
- Context window
- 128,000 tokens
- Max output
- –
ChatAvailable locallyGemma 3 27B
Google
- Context window
- 128,000 tokens
- Max output
- –
ChatAvailable locallygpt-oss-120b
OpenAI
- Context window
- 131,072 tokens
- Max output
- 131,072 tokens
ChatAvailable locallygpt-oss-20b
OpenAI
- Context window
- 131,072 tokens
- Max output
- 131,072 tokens
ChatAvailable locallyLlama 4 Scout
Meta
- Context window
- 10,000,000 tokens
- Max output
- –
ChatAvailable locallyLlama 3.3 70B
Meta
- Context window
- 128,000 tokens
- Max output
- –
ChatAvailable locallyLlama 3.1 8B
Meta
- Context window
- 128,000 tokens
- Max output
- –
ChatAvailable locallyPhi-4
Microsoft
- Context window
- 16,384 tokens
- Max output
- –
ChatAvailable locallyPhi-4-reasoning-plus
Microsoft
- Context window
- 32,768 tokens
- Max output
- –
ChatAvailable locallyLlama 3.3 Nemotron Super 49B
NVIDIA
- Context window
- 128,000 tokens
- Max output
- –
ChatAvailable locallyNemotron Nano 9B v2
NVIDIA
- Context window
- 128,000 tokens
- Max output
- –
ChatAvailable locally
In addition, hosted in Germany
4
Mistral Large 3
Mistral AI
- Context window
- 256,000 tokens
- Max output
- –
ChatEU hostedMistral Medium 3.5
Mistral AI
- Context window
- 256,000 tokens
- Max output
- –
ChatEU hostedCodestral
Mistral AI
- Context window
- 256,000 tokens
- Max output
- –
CodeEU hostedMistral Embed
Mistral AI
- Context window
- –
- Max output
- –
EmbeddingEU hosted
External models
25
Available in On Demand, through AI credits. Optional in on-premise installations, with the vendor's token costs. Not on Air-Gap Enterprise.
DeepSeek V4-Pro
DeepSeek
- Context window
- 1,000,000 tokens
- Max output
- –
ChatExternal · EU/ZDRDeepSeek V4-Flash
DeepSeek
- Context window
- 1,000,000 tokens
- Max output
- –
ChatExternal · EU/ZDRClaude Fable 5.1
Anthropic
- Context window
- 1,000,000 tokens
- Max output
- 128,000 tokens
ChatExternal · EU/ZDRClaude Fable 5
Anthropic
- Context window
- 1,000,000 tokens
- Max output
- 128,000 tokens
ChatExternal · EU/ZDRClaude Opus 5.5
Anthropic
- Context window
- 1,000,000 tokens
- Max output
- 128,000 tokens
ChatExternal · EU/ZDRClaude Opus 5
Anthropic
- Context window
- 1,000,000 tokens
- Max output
- 128,000 tokens
ChatExternal · EU/ZDRClaude Opus 4.8
Anthropic
- Context window
- 1,000,000 tokens
- Max output
- 128,000 tokens
ChatExternal · EU/ZDRClaude Sonnet 5
Anthropic
- Context window
- 1,000,000 tokens
- Max output
- 128,000 tokens
ChatExternal · EU/ZDRClaude Haiku 4.5
Anthropic
- Context window
- 200,000 tokens
- Max output
- 64,000 tokens
ChatExternal · EU/ZDRGemini 3.1 Pro
Google
- Context window
- 1,000,000 tokens
- Max output
- 64,000 tokens
ChatExternal · EU/ZDRGemini 3.5 Flash
Google
- Context window
- 1,000,000 tokens
- Max output
- 64,000 tokens
ChatExternal · EU/ZDRGemini 3.1 Flash-Lite
Google
- Context window
- 1,000,000 tokens
- Max output
- 64,000 tokens
ChatExternal · EU/ZDRGemini Embedding 2
Google
- Context window
- –
- Max output
- –
EmbeddingExternal · EU/ZDRGPT-5.6 Sol
OpenAI
- Context window
- 1,050,000 tokens
- Max output
- 128,000 tokens
ChatExternal · EU/ZDRGPT-5.6 Terra
OpenAI
- Context window
- 1,050,000 tokens
- Max output
- 128,000 tokens
ChatExternal · EU/ZDRGPT-5.6 Luna
OpenAI
- Context window
- 1,050,000 tokens
- Max output
- 128,000 tokens
ChatExternal · EU/ZDRGPT-5.5
OpenAI
- Context window
- 1,050,000 tokens
- Max output
- 128,000 tokens
ChatExternal · EU/ZDRGPT-5.4
OpenAI
- Context window
- 1,050,000 tokens
- Max output
- 128,000 tokens
ChatExternal · EU/ZDRGPT-5.4 mini
OpenAI
- Context window
- 400,000 tokens
- Max output
- 128,000 tokens
ChatExternal · EU/ZDRGPT-5.1
OpenAI
- Context window
- 400,000 tokens
- Max output
- 128,000 tokens
ChatExternal · EU/ZDRGPT-5
OpenAI
- Context window
- 400,000 tokens
- Max output
- 128,000 tokens
ChatExternal · EU/ZDRGPT Image 2
OpenAI
- Context window
- –
- Max output
- –
ImageExternal · EU/ZDRGLM-5.2
Z.ai (GLM)
- Context window
- 1,000,000 tokens
- Max output
- 128,000 tokens
ChatExternal · EU/ZDRKimi K2.6
Moonshot AI (Kimi)
- Context window
- 256,000 tokens
- Max output
- –
ChatExternal · EU/ZDRMiniMax-M3
MiniMax
- Context window
- 1,000,000 tokens
- Max output
- –
ChatExternal · EU/ZDR
51 of 51 models
Open-weight models run on the appliance and in On Demand out of the box. Larger open-weight models are available in addition in hosting in Germany. External models from OpenAI, Anthropic, Google, DeepSeek and other vendors are available in On Demand through AI credits and optional on-premise, there with the vendor's token costs. External models are processed in the vendors' EU and zero-data-retention (ZDR) tenants.
How models are run on souveraen.ai.
Fixed model versions, model and processing location shown with every answer, updates only after your approval.
Operating routes
First the data path, then the model.
A model name says nothing about where your request is processed. So you choose the operating route first; the models that fit follow from it.
What people ask most before they choose a model.
Was Geschäftsführung, IT und Datenschutz vor dem Start wissen wollen.
Lieber persönlich sprechen?
Wir klären Einsatzfall, Datenschutz und Betriebsart in einem kurzen Gespräch, ohne Verkaufsdruck. Telefon (03744) 365 2202.