Resource · AI Models

Which AI model to choosefor an enterprise project?

Our dated and assumed ranking of major language models: what we actually deploy for our clients, and why. Not another benchmark—an agency perspective, focused on cost, latency, and compliance.

01 — The ranking

Our tier list of AI models.

One criterion: the value of a model in a production enterprise project. Reasoning ability, cost per million tokens, latency, API stability, and data hosting location.

Updated Septembre 202613 models ranked

S

What we deploy in production without hesitation

The models we choose when reasoning quality takes precedence over cost: tool-equipped agents, complex corpus analysis, code generation.

  • Claude Fable 5A step above for very long reasoning. The extra cost is only justified for a minority of tasks—rarely at scale.
  • GPT-5.6 SolOpenAI’s premium offering. The broadest ecosystem of tools and libraries, very strong on code.
A

The best quality-to-price ratio

What we deploy for most projects: good enough for almost everything, cheap enough to run continuously without breaking the bank.

  • Claude Opus 5Very strong on tool-using agents and long-running tasks, with stable behaviour from one version to the next. Fable now beats it when reasoning quality is the priority, but it remains an excellent default on most projects.
  • Gemini 3.7 FlashExtremely fast and very affordable. Ideal for high-volume document processing and batch jobs.
  • Mistral Medium 3.5Our go-to when data must stay in Europe. More than sufficient for document RAG, with explicit EU hosting.
B

Excellent for specific use cases

Excellent in their niche, less versatile. We use them when the use case justifies it—not by default.

  • Claude Sonnet 5The right choice for assistants and high-volume RAG, where its latency and price make the difference. On complex reasoning, the models in the tiers above take the lead.
  • GPT-5.6 LunaVery high value for money and good execution speed. An excellent candidate for targeted, high-volume use cases, especially if your stack already runs on OpenAI.
  • Grok 4.6High performance-to-cost ratio. Best for use cases where data sensitivity is not a concern.
  • Qwen3.8-MaxVery strong in multilingual tasks. Part of the range is self-hostable, making it a credible fallback option.
  • Mistral SmallOpen weights, runs on your own infrastructure. The choice when nothing must leave the network.
C

Reliable for self-hosting

Freely downloadable, to be self-hosted. Inference costs are controlled, but operational costs are real—GPUs, monitoring, updates.

  • DeepSeek-V4-ProExcellent capacity-to-cost ratio for reasoning. Sovereign hosting is a requirement here, not an option.
  • GLM-5.2 TurboFast and cost-effective. Ideal for high-volume structured tasks, less so for open-ended reasoning.
D

To watch, not yet validated by us

Promising models we haven’t yet deployed on a client project. We don’t recommend them until we have.

  • Kimi K3Reported strong performance on long contexts. Needs validation on real-world corpora before any recommendation.

This ranking is an agency opinion, not a benchmark. It evolves quickly: a model released last week could reshuffle the deck. We revise it with every significant update and revalidate it against our clients’ real-world use cases before making recommendations.

02 — Our criteria

Five criteria, no overall score.

No model is ever "the best" in absolute terms. Here’s what we look at before deciding, in this order.

01

Reasoning quality for YOUR use case

Public rankings measure generic tasks. We replay your own cases—your documents, your business rules, your output formats—across three or four models before deciding. It’s the only benchmark that commits someone.

02

Cost per million tokens

An internal assistant running all day costs ten to thirty times more on a top-tier model than on a mid-range one, with often no visible gain for the end user. The right approach: start high to validate feasibility, then step down until quality drops.

03

Perceived latency

In a chat, two extra seconds kill the experience. In overnight processing, no one notices. The right model depends first on whether someone is waiting in front of their screen.

04

Data hosting location

Healthcare, public sector, legal, defense: the question isn’t 'which is the best model' but 'which ones am I allowed to use.' This is decided upfront, not at the end—it’s what eliminates most candidates.

05

Long-term stability

A model deprecated six months after deployment means unbudgeted rework. We favor providers that announce their lifecycle, and we keep the application able to switch models without rewriting.

03 — Going further

The model is never the point.

Choosing the model takes one meeting; the rest takes the project. Here’s what actually determines the success of an AI use case.

Frame before choosing

Model selection comes at the end of scoping, not the beginning. First, define the use case, available data, and acceptable risk level—the model is just a fine-tuning variable.

AI scoping

An assistant connected to your data

The quality of a corporate assistant depends first on how it accesses your documents, and only then on the chosen model. That’s where the real difference lies.

AI assistants on your data

Stay sovereign over your data

Hosting in Europe, open-weight models, deployment on your own infrastructure: when data cannot leave, the selection narrows—but remains more than sufficient.

Sovereign applications

Put a model into your processes

A model only delivers value once integrated with your tools: CRM, ERP, email, DMS. This is where a small, fast model often beats a top-tier one.

Automations
04 — FAQ

Your questions, our answers.

The questions that come up whenever we discuss model selection with an IT department or business leadership.

On our hands-on experience, not public benchmarks. We rank models by their value in a production enterprise project: real-world reasoning quality, cost per million tokens, latency, API stability, and data hosting location. A model that tops academic leaderboards but costs thirty times more than a sufficient alternative will rank lower here. This is an agency opinion—explicitly so—and time-stamped, because it won’t stay true for six months.

Whenever the landscape shifts significantly, and at least every time we switch our default model on a project. The last update date is displayed at the top of the ranking: if it’s several months old, treat the table as indicative and ask us for the latest.

No, and it’s rarely the right move. Tier S is for validating whether a use case is feasible; then we step down until quality drops. On most of our projects, the model deployed is Tier A—far cheaper, with a difference users don’t notice. Paying top-tier for email sorting is waste, not a guarantee.

It depends on the data and the framework. For health records, legal files, or sensitive public data, the answer is often no—or only under strict conditions: EU hosting, contractual commitments, prior anonymization. It’s the first question we ask during scoping, because it eliminates the most candidates. When it blocks, European models and self-hosted open-weight models cover the vast majority of cases.

Nothing, if the application was built correctly. We isolate model access behind an abstraction layer: switching providers means updating a config and rerunning our test suite, not rewriting the app. It’s an architectural choice we make from day one, precisely because this ranking evolves.

By testing it on your data, not by reading a table. Our AI scoping includes a comparison phase: we replay your real cases on three or four models, measure quality, cost, and latency, and document the choice. This page tells you where to start; it doesn’t replace measurement.
Let’s discuss

An AI use case to decide, an assistant to connect to your data, a model choice to document? Let’s talk data, budget, and real constraints.

Contact details
20 Rue des Taillandiers
75011 Paris
Response within 24 business hours.