Cohere’s 218B MoE Translation Model Posts 83.6 WMT26 Under a Non-Commercial License

Cohere’s North Small Translate: 83.6 WMT26, 218B MoE, 2x H100 inference, and a CC BY-NC license blocking commercial use.
Introducing North Small Translate: A leading sovereign open-weight machine translation model
By Andres SEO Expert.

Key Takeaways

  • North Small Translate pairs 218B total parameters with 25B active in a mixture-of-experts design, scoring 83.6 on WMT26 across 50+ languages and running on 2x H100s at W4A4.
  • Throughput reaches 112 output tokens per second versus Gemma 4 31B’s 81, while long-context translation scores 48.9 against 21.3 for Google Translate.
  • Weights ship under CC BY-NC 4.0 for non-commercial research only, so enterprise deployment runs through the RWS Language Weaver partnership.

Sovereign Open-Weight Translation Arrives With Breakneck Benchmarks

On September 10, 2026, Cohere announced North Small Translate, a mixture-of-experts translation model that pairs a 218-billion-parameter total footprint with 25 billion active parameters. The release lands with an 83.6 WMT26 all-languages score and support for more than 50 languages.

It is the first dedicated translation model in the North family, building on the company’s earlier multilingual work with Aya and Command A Translate. The weights are available now on Hugging Face under a CC BY-NC 4.0 license for research and non-commercial use.

The commercial production path runs through RWS Language Weaver, a partnership that positions the model for enterprise translation workloads beyond the research preview. That split between open-weight access and commercial licensing is the central tension of this release.

Inside the Architectural Shift Behind North Small Translate

According to Cohere’s official announcement, North Small Translate is optimized for machine translation as a dedicated task, not as a general-purpose language model. Its mixture-of-experts design activates a 25-billion-parameter subset for each forward pass, which keeps inference costs lower than the 218-billion total parameter count might suggest.

  • Architecture: Mixture-of-experts
  • Model size: 218B total, 25B active
  • Context length: 16k input, 16k output
  • Input and output modalities: Text
  • Languages: 50+ supported, with full list available
  • Minimum hardware: 1x B200 at W4A4 or 2x H100s at W4A4

This hardware envelope matters. It means organizations can run a state-of-the-art translation model on a pair of H100s rather than a full accelerator rack.

Benchmark Positioning and Regional Consistency

The standard model posts an 83.6 overall WMT26 score, ahead of Qwen 3.5 397B A17B at 81.56, DeepL NextGen at 81.37, Gemma 4 31B at 79.46, GLM 5.2 FP8 at 76.50, and Google Translate at 68.20. An agentic variant raises the WMT26 mark to 84.36 by detecting and correcting translation errors within the same workflow.

Regionally, North Small Translate outperforms Gemma 4 31B outright in Europe and essentially matches it in South Asia. Both the standard and agentic versions exceed DeepL NextGen across every non-European region tested, with the widest advantages appearing in South Asia and MENA.

Throughput, Long-Context, and Unit Economics

Throughput testing shows up to 1.4 times higher output token generation than Gemma 4 31B TP1 under identical concurrency and hardware. The measured rates are 112 output tokens per second versus 81 at low concurrency, and 39 versus 30 at high concurrency.

The long-context evaluation, which measures translation quality across two book chapters in a single call, returns a 48.9 score. That compares with 21.3 for Google Translate and 19.4 for Gemma 4 31B.

In the pricing comparison, the model lands at $0.000676 per task at a 661-token average. Gemini 3.1 Pro Preview high, by contrast, lands at $0.038928 per task, a gap of 5,762 percent.

The Licensing Fault Line Beneath the Benchmark Triumph

The most consequential part of this launch may not be the WMT26 chart. It is the license.

According to a Reuters Practical Law journal analysis of the legal landscape around open-weight releases, the term is often conflated with open source. Most open-weight providers do not release training data, training code, evaluation data, or runtime software, and the Open Source Initiative’s current definition requires all of those elements.

A non-commercial license such as CC BY-NC 4.0 can be treated as a contract rather than a copyright condition. Enforcement can become murky if a user does not affirmatively accept the terms.

That does not erase North Small Translate’s benchmark lead, but it does mean the model is not a direct open-source drop-in for commercial production. The RWS Language Weaver route is where the commercial deployment conversation actually happens.

TechTarget’s enterprise reporting reinforces that point: customers that adopt open-weight systems take on accountability for security, model updates, and data governance that a managed API vendor would otherwise absorb. Benchmarks are only one variable in that evaluation.

Analysts in that reporting expect hybrid multi-model estates: open-weight translation for controlled or specialized internal workloads, and proprietary platforms where vendor accountability and advanced capabilities carry more weight.

Sovereign Translation Infrastructure Gets a New Production Path

North Small Translate gives AI teams a rare combination: state-of-the-art translation benchmarks, explicit sovereignty positioning, and a hardware profile that does not demand a full rack of accelerators. For teams evaluating open-weight translation models as part of multilingual AI pipelines, programmatic SEO and AI automation is how Andres SEO Expert approaches scaled deployment — contact us.

Frequently Asked Questions

What is Cohere North Small Translate?

Cohere North Small Translate is a dedicated open-weight machine translation model announced on September 10, 2026. It uses a mixture-of-experts architecture with 218 billion total parameters and 25 billion active parameters, supports more than 50 languages, and scored 83.6 on WMT26 all-languages.

How many parameters does North Small Translate have?

It has 218 billion total parameters but activates only 25 billion parameters per forward pass. That mixture-of-experts design keeps inference costs lower than the total parameter count might suggest.

What hardware is required to run North Small Translate?

The minimum hardware is 1x B200 at W4A4 or 2x H100s at W4A4. This means teams can run the model on a pair of H100s rather than a full accelerator rack.

Is North Small Translate open source?

No. It is open-weight, not open source. The weights are available on Hugging Face under a CC BY-NC 4.0 license, but Cohere does not release training data, training code, evaluation data, or runtime software as required by the Open Source Initiative definition.

Can North Small Translate be used commercially?

The open-weight release is limited to research and non-commercial use under CC BY-NC 4.0. For commercial production, the deployment path runs through RWS Language Weaver, Cohere’s partnership for enterprise translation workloads.

How does North Small Translate compare with Google Translate, DeepL, and Gemma?

North Small Translate scored 83.6 on WMT26, ahead of Qwen 3.5 397B A17B at 81.56, DeepL NextGen at 81.37, Gemma 4 31B at 79.46, GLM 5.2 FP8 at 76.50, and Google Translate at 68.20. Its agentic variant raises the score to 84.36.

What is the agentic variant of North Small Translate?

The agentic variant detects and corrects translation errors within the same workflow, raising the WMT26 score from 83.6 to 84.36. It is part of the same release and is positioned for higher-accuracy translation tasks.

Prev

Subscribe to My Newsletter

Subscribe to my email newsletter to get the latest posts delivered right to your email. Pure inspiration, zero spam.
You agree to the Terms of Use and Privacy Policy