Initial IADiscuss deployment ↗

Beta tool · Open models · Private AI

Which open model can
actually help you?

Compare the model and its inference engine for your use case, GPU memory and number of users.

Immediate resultVisible assumptionsNo documents sent
Test a configuration ↓
ENGINE PREVIEWIndicative estimate
01 · Your use case
02 · Available VRAM
MODELS TO EVALUATE · Documents & RAGIntermediate
Model class
12 to 14B · Q4
Indicative workload
1 to 3 users
Suggested engine
LM Studio or Ollama for testing · vLLM or SGLang in production

Estimate to validate with your data · catalogue verified on 26 August 2026. Context, quantisation, KV cache and concurrency may change actual requirements.

Receive the complete version

A model-by-model comparison, licences and deployment scenario.

By submitting this request, you agree to be contacted about this tool. Privacy.
The full version

More than a
file size.

VRAM alone is not enough to choose a model. The final tool will connect technical constraints to the actual work expected.

01

Choose the model

Use case, French-language quality, context, licence and tool-calling ability.

Qwen · Mistral · Gemma · Granite · specialist models
02

Size the infrastructure

Weights, KV cache, runtime, quantisation and headroom for concurrency.

GPU · VRAM · RAM · storage
03

Choose the engine

Local testing, a server on your premises or a private environment operated by Initial IA.

LM Studio · Ollama · vLLM · SGLang
INFORMATION THAT AGES QUICKLY

Every recommendation
will need to show its source.

The catalogue is reviewed against each publisher's model cards. Recommendations, licences and French-language results are checked before publication.

  • Official source and exact version
  • Last verification date
  • Estimate distinguished from actual measurement
  • Explicit limits before any purchase

Official catalogues: Qwen ↗ · Mistral AI ↗ · Gemma 4 ↗ · IBM Granite ↗

Already have a need in mind?

The model is one component.
The system matters more.

We can size, host or install the complete environment around your use case.

Explore a configuration ↗Video call · 30 minutes · No obligation