Recent multilingual embeddings for finding relevant passages.
Apache-2.0 · Official model card ↗Beta tool · Open models · Private AI
Which open model can
actually help you?
Compare the model and its inference engine for your use case, GPU memory and number of users.
Recent multimodal model for long, varied document collections.
Apache-2.0 · Official model card ↗Multimodal edge model for summarising and answering from sources.
Apache-2.0 · Official model card ↗- Model class
- 12 to 14B · Q4
- Indicative workload
- 1 to 3 users
- Suggested engine
- LM Studio or Ollama for testing · vLLM or SGLang in production
Estimate to validate with your data · catalogue verified on 26 August 2026. Context, quantisation, KV cache and concurrency may change actual requirements.
More than a
file size.
VRAM alone is not enough to choose a model. The final tool will connect technical constraints to the actual work expected.
Choose the model
Use case, French-language quality, context, licence and tool-calling ability.
Qwen · Mistral · Gemma · Granite · specialist modelsSize the infrastructure
Weights, KV cache, runtime, quantisation and headroom for concurrency.
GPU · VRAM · RAM · storageChoose the engine
Local testing, a server on your premises or a private environment operated by Initial IA.
LM Studio · Ollama · vLLM · SGLangEvery recommendation
will need to show its source.
The catalogue is reviewed against each publisher's model cards. Recommendations, licences and French-language results are checked before publication.
- Official source and exact version
- Last verification date
- Estimate distinguished from actual measurement
- Explicit limits before any purchase
Official catalogues: Qwen ↗ · Mistral AI ↗ · Gemma 4 ↗ · IBM Granite ↗
The model is one component.
The system matters more.
We can size, host or install the complete environment around your use case.
Explore a configuration ↗Video call · 30 minutes · No obligation