Source-led article
Soofi Consortium Launches Open Hybrid Mamba-Transformer MoE Model for German and English

A German research consortium has published the pretraining report for Soofi S 30B-A3B, an open base model designed for both German and English. This new foundation model is a Mixture-of-Experts (MoE) hybrid combining Mamba and Transformer architectures, totaling approximately 31.6 billion parameters and activating about 3.2 billion parameters per token. The model’s development was coordinated by the KI Bundesverband and funded by the German Federal Ministry for Economic Affairs and Energy.
The training of Soofi S 30B-A3B was conducted on Deutsche Telekom’s Industrial AI Cloud in Munich. Preview weights for the model are now available on Hugging Face. Notably, Soofi S has achieved high aggregate scores in both English and German among several fully open base models tested, demonstrating its proficiency across both languages.
Architectural Design and Efficiency
Soofi S adopts the Nemotron 3 Nano reference design without modifications, a choice made for deployability on existing stacks like vLLM, serving efficiency, and scientific control. The model’s network consists of 52 layers, including 23 Mamba-2 sequence-mixing layers, 23 granular MoE layers, and 6 Grouped-Query Attention (GQA) layers. Only the GQA layers maintain a KV cache. Each MoE layer incorporates 128 routed experts, activating six per token, alongside two shared experts.
The model also features a dimension of 2688, uses squared ReLU activation, RMSNorm, and does not employ positional embeddings. These design choices aim to optimize performance and efficiency in language processing tasks.
Training Methodology and Data
The training data recipe for Soofi S followed a Warmup–Stable–Decay (WSD) schedule. Phase 1 involved approximately 20 trillion tokens from a diverse, quality-tiered mixture. Phase 2 consumed about 6.58 trillion tokens of high-quality annealing data, with German language content deliberately increased from 7.2% in Phase 1 to 15.32%. This emphasis on German data, including sources like HPLT v3 and v4, German Commons, German FinePDFs, FineWiki, and commercially licensed Genios articles, differentiates Soofi S from models with lower non-English allocations. Phase 3 extended the usable context window up to 1 million tokens with 0.10 trillion tokens at a 1,048,576-token sequence length.
The infrastructure for training utilized up to 512 NVIDIA B200 GPUs, consuming approximately 253,000 B200 GPU-hours between March and May 2026. This highlights the significant computational resources invested in developing the model.
Performance and Deployment Implications
In evaluations against 16 other open base models using the lm-evaluation-harness pipeline, Soofi S showed notable improvements. It gained 1.8 points on the English aggregate and 4.2 points on German compared to its architecture-identical reference. While larger open-weight models like Qwen3.5 35B-A3B still lead in overall means, Soofi S registered competitive scores, achieving 70.1 in English (compared to Gemma 3 27B’s 70.3 and Ministral 3 14B’s 70.3) and leading in German with 79.1 (against 78.4 and 78.3 respectively).
Key facts:
| Feature | Detail |
|—|—|
| Model Name | Soofi S 30B-A3B |
| Type | Open Hybrid Mamba-Transformer MoE Foundation Model |
| Parameters | ~31.6 Billion total, ~3.2 Billion active per token |
| Languages | German and English |
| Consortium | KI Bundesverband, Fraunhofer IAIS, DFKI, TU Darmstadt, ellamind, Merantix Momentum |
The consortium suggests three primary deployment scenarios for Soofi S. First, for German document processing, its performance in GLP-DE (88.8) and INCLUDE-DE (61.2) makes it suitable for fine-tuning on policy PDFs within sectors like insurance. Second, for bilingual code assistance, its MBPP-DE score of 84.2 indicates utility for teams prompting in German for Python tasks. Third, for high-concurrency long-context serving, such as a support-ticket RAG system, it matches measured regimes at batch 32 and 40K context.
This release provides a new open-source option for developers and researchers, particularly those focused on bilingual applications involving German and English. The emphasis on an open architecture and robust German language capabilities could foster innovation in regions and industries that require strong multilingual AI solutions.
Source: MarkTechPost, https://www.marktechpost.com/2026/07/15/soofi-consortium-releases-soofi-s-30b-a3b-an-open-hybrid-mamba-transformer-moe-foundation-model-for-german-and-english/
Datos clave
| Punto | Detalle |
|---|---|
| Fuente | MarkTechPost |
| Fecha | 2026-07-15T21:02:48+00:00 |
| Tema | Soofi Consortium Releases Soofi S 30B-A3B: An Open Hybrid Mamba-Transformer MoE Foundation Model For German And English |