Skip to content
souveraen.ai
Models

Which AI models run for you.

Open models such as Qwen, Mistral and gpt-oss run on your appliance, and your data stays on it. Larger open models are added in hosting in Germany. External models are available in On Demand depending on the plan and are processed in the vendors' EU and zero-data-retention (ZDR) tenants.

  • Local, or offline with Air-Gap
  • Open models out of the box
  • One interface for every task

Model catalogue

On the appliance, out of the box

22

  • Mistral Small 4

    Mistral AI

    Context window
    256,000 tokens
    Max output
    –
    ChatAvailable locally
  • Ministral 3 8B

    Mistral AI

    Context window
    256,000 tokens
    Max output
    –
    ChatAvailable locally
  • Devstral Small 2

    Mistral AI

    Context window
    256,000 tokens
    Max output
    –
    CodeAvailable locally
  • Magistral Small 1.2

    Mistral AI

    Context window
    128,000 tokens
    Max output
    –
    ChatAvailable locally
  • Qwen3.6-35B-A3B

    Qwen (Alibaba)

    Context window
    256,000 tokens
    Max output
    –
    ChatAvailable locally
  • Qwen3.6-27B

    Qwen (Alibaba)

    Context window
    256,000 tokens
    Max output
    –
    ChatAvailable locally
  • Qwen3-Coder-30B-A3B

    Qwen (Alibaba)

    Context window
    262,144 tokens
    Max output
    –
    CodeAvailable locally
  • Qwen3-30B-A3B-2507

    Qwen (Alibaba)

    Context window
    262,144 tokens
    Max output
    –
    ChatAvailable locally
  • Qwen3-14B

    Qwen (Alibaba)

    Context window
    32,768 tokens
    Max output
    –
    ChatAvailable locally
  • Qwen3-8B

    Qwen (Alibaba)

    Context window
    32,768 tokens
    Max output
    –
    ChatAvailable locally
  • Qwen3 Embedding 8B

    Qwen (Alibaba)

    Context window
    –
    Max output
    –
    EmbeddingAvailable locally
  • Gemma 4

    Google

    Context window
    128,000 tokens
    Max output
    –
    ChatAvailable locally
  • Gemma 3 27B

    Google

    Context window
    128,000 tokens
    Max output
    –
    ChatAvailable locally
  • gpt-oss-120b

    OpenAI

    Context window
    131,072 tokens
    Max output
    131,072 tokens
    ChatAvailable locally
  • gpt-oss-20b

    OpenAI

    Context window
    131,072 tokens
    Max output
    131,072 tokens
    ChatAvailable locally
  • Llama 4 Scout

    Meta

    Context window
    10,000,000 tokens
    Max output
    –
    ChatAvailable locally
  • Llama 3.3 70B

    Meta

    Context window
    128,000 tokens
    Max output
    –
    ChatAvailable locally
  • Llama 3.1 8B

    Meta

    Context window
    128,000 tokens
    Max output
    –
    ChatAvailable locally
  • Phi-4

    Microsoft

    Context window
    16,384 tokens
    Max output
    –
    ChatAvailable locally
  • Phi-4-reasoning-plus

    Microsoft

    Context window
    32,768 tokens
    Max output
    –
    ChatAvailable locally
  • Llama 3.3 Nemotron Super 49B

    NVIDIA

    Context window
    128,000 tokens
    Max output
    –
    ChatAvailable locally
  • Nemotron Nano 9B v2

    NVIDIA

    Context window
    128,000 tokens
    Max output
    –
    ChatAvailable locally

In addition, hosted in Germany

4

  • Mistral Large 3

    Mistral AI

    Context window
    256,000 tokens
    Max output
    –
    ChatEU hosted
  • Mistral Medium 3.5

    Mistral AI

    Context window
    256,000 tokens
    Max output
    –
    ChatEU hosted
  • Codestral

    Mistral AI

    Context window
    256,000 tokens
    Max output
    –
    CodeEU hosted
  • Mistral Embed

    Mistral AI

    Context window
    –
    Max output
    –
    EmbeddingEU hosted

External models

25

Available in On Demand, through AI credits. Optional in on-premise installations, with the vendor's token costs. Not on Air-Gap Enterprise.

  • DeepSeek V4-Pro

    DeepSeek

    Context window
    1,000,000 tokens
    Max output
    –
    ChatExternal · EU/ZDR
  • DeepSeek V4-Flash

    DeepSeek

    Context window
    1,000,000 tokens
    Max output
    –
    ChatExternal · EU/ZDR
  • Claude Fable 5.1

    Anthropic

    Context window
    1,000,000 tokens
    Max output
    128,000 tokens
    ChatExternal · EU/ZDR
  • Claude Fable 5

    Anthropic

    Context window
    1,000,000 tokens
    Max output
    128,000 tokens
    ChatExternal · EU/ZDR
  • Claude Opus 5.5

    Anthropic

    Context window
    1,000,000 tokens
    Max output
    128,000 tokens
    ChatExternal · EU/ZDR
  • Claude Opus 5

    Anthropic

    Context window
    1,000,000 tokens
    Max output
    128,000 tokens
    ChatExternal · EU/ZDR
  • Claude Opus 4.8

    Anthropic

    Context window
    1,000,000 tokens
    Max output
    128,000 tokens
    ChatExternal · EU/ZDR
  • Claude Sonnet 5

    Anthropic

    Context window
    1,000,000 tokens
    Max output
    128,000 tokens
    ChatExternal · EU/ZDR
  • Claude Haiku 4.5

    Anthropic

    Context window
    200,000 tokens
    Max output
    64,000 tokens
    ChatExternal · EU/ZDR
  • Gemini 3.1 Pro

    Google

    Context window
    1,000,000 tokens
    Max output
    64,000 tokens
    ChatExternal · EU/ZDR
  • Gemini 3.5 Flash

    Google

    Context window
    1,000,000 tokens
    Max output
    64,000 tokens
    ChatExternal · EU/ZDR
  • Gemini 3.1 Flash-Lite

    Google

    Context window
    1,000,000 tokens
    Max output
    64,000 tokens
    ChatExternal · EU/ZDR
  • Gemini Embedding 2

    Google

    Context window
    –
    Max output
    –
    EmbeddingExternal · EU/ZDR
  • GPT-5.6 Sol

    OpenAI

    Context window
    1,050,000 tokens
    Max output
    128,000 tokens
    ChatExternal · EU/ZDR
  • GPT-5.6 Terra

    OpenAI

    Context window
    1,050,000 tokens
    Max output
    128,000 tokens
    ChatExternal · EU/ZDR
  • GPT-5.6 Luna

    OpenAI

    Context window
    1,050,000 tokens
    Max output
    128,000 tokens
    ChatExternal · EU/ZDR
  • GPT-5.5

    OpenAI

    Context window
    1,050,000 tokens
    Max output
    128,000 tokens
    ChatExternal · EU/ZDR
  • GPT-5.4

    OpenAI

    Context window
    1,050,000 tokens
    Max output
    128,000 tokens
    ChatExternal · EU/ZDR
  • GPT-5.4 mini

    OpenAI

    Context window
    400,000 tokens
    Max output
    128,000 tokens
    ChatExternal · EU/ZDR
  • GPT-5.1

    OpenAI

    Context window
    400,000 tokens
    Max output
    128,000 tokens
    ChatExternal · EU/ZDR
  • GPT-5

    OpenAI

    Context window
    400,000 tokens
    Max output
    128,000 tokens
    ChatExternal · EU/ZDR
  • GPT Image 2

    OpenAI

    Context window
    –
    Max output
    –
    ImageExternal · EU/ZDR
  • GLM-5.2

    Z.ai (GLM)

    Context window
    1,000,000 tokens
    Max output
    128,000 tokens
    ChatExternal · EU/ZDR
  • Kimi K2.6

    Moonshot AI (Kimi)

    Context window
    256,000 tokens
    Max output
    –
    ChatExternal · EU/ZDR
  • MiniMax-M3

    MiniMax

    Context window
    1,000,000 tokens
    Max output
    –
    ChatExternal · EU/ZDR

51 of 51 models

Open-weight models run on the appliance and in On Demand out of the box. Larger open-weight models are available in addition in hosting in Germany. External models from OpenAI, Anthropic, Google, DeepSeek and other vendors are available in On Demand through AI credits and optional on-premise, there with the vendor's token costs. External models are processed in the vendors' EU and zero-data-retention (ZDR) tenants.

Missing a model?

We add further open models to your profile after a licence and quality review. Tell us which model you need.

Request a model
Model operation

How models are run on souveraen.ai.

Fixed model versions, model and processing location shown with every answer, updates only after your approval.

The right model per task

Chat, code, embedding and image: open models and, in On Demand, external models in EU and ZDR tenants.

You can see which model answered

Every answer names the model that produced it and where it ran.

Nothing changes overnight

A model is updated only after you approve the update.

Operating routes

First the data path, then the model.

A model name says nothing about where your request is processed. So you choose the operating route first; the models that fit follow from it.

Local on the appliance

Open models such as Qwen3.6, Mistral Small 4, Gemma 4, Llama and gpt-oss run on the DGX appliance in your building. External models are optional, with the vendor's token costs.

  • Data stays local
  • No token costs for local models
  • A fixed monthly price
See Private Spark

On Demand

The same software as on the appliance, in a data centre in Germany. We sign a GDPR data-processing agreement with you. German and European law governs the contract and the operation. External models are processed here in the vendors' EU and zero-data-retention (ZDR) tenants.

  • GDPR data processing agreement
  • A separate tenant
  • Start for free
Plans and pricing

External models

In On Demand, models from OpenAI, Anthropic, Google and other vendors are available depending on the plan. They are processed in the vendors' EU and zero-data-retention (ZDR) tenants. On Private Spark they are optional.

  • In the vendors' EU and ZDR tenants
  • On-premise optional, with token costs
  • Not on Air-Gap Enterprise
Discuss your use
Common questions

What people ask most before they choose a model.

Was Geschäftsführung, IT und Datenschutz vor dem Start wissen wollen.

Lieber persönlich sprechen?

Wir klären Einsatzfall, Datenschutz und Betriebsart in einem kurzen Gespräch, ohne Verkaufsdruck. Telefon (03744) 365 2202.

Beratung anfragen

You decide. On the appliance an open model such as Qwen3.6 or gpt-oss-120b does the work; which one depends on your tasks and how many people work at once, and is recorded in the offer. In On Demand, depending on the plan, external models in the vendors' EU and ZDR tenants are added.

The right profile

Which models belong on your appliance?

Tell us your tasks and how many people work at once. We propose the model profile and record it in the offer.

  1. Name tasks and user count1
  2. Receive a model profile2
  3. Fixed in the quote3