French Drift
Open source · Python · No dependencies

Is the French answer as good as the English one?

French Drift runs the same corpus, the same questions and the same metrics through an AI content pipeline in English and in Canadian French, side by side. Then it reports the gap. Two languages can each clear a quality bar while the French reader gets a thinner, less grounded answer, in a register that reads as foreign. Scoring each language on its own can't see that.

The report's opening panel: 69% service parity, and 6 queries where the French reader was not served.
The report leads with the worst thing a French reader experienced, not the mean.
69%
service parity

Queries where the French reader got a substantively equivalent answer.

80%
false corrections

Valid Quebec forms a naive “proofread this” pass rewrote.

+0.116
coverage gap

French answers carry less than their English counterparts.

58%
media accessible

Assets whose French descriptors are complete.

01 · The finding

The hypothesis it was built on didn't survive contact.

The harness started from a plausible argument: Canadian French vocabulary is weakly represented in retrieval models, so French retrieval diverges. The default backend models that gap from a hand-written lexicon table. Then we measured it with a real multilingual embedding model, bge-m3, run locally.

DimensionModelled gap (BM25) Measured gap (bge-m3)
Retrieval precision@1+0.039−0.010
Retrieval recall@3+0.017+0.003
Retrieval nDCG@3+0.022−0.005
Grounding in retrieved context+0.033+0.021
Content coverage+0.116+0.116
Canadian French register+0.059+0.059
Localization+0.015+0.015

On precision and nDCG, the sign flipped: French retrieval came out marginally ahead. The query the harness was built around, “where can people go to keep warm?”, retrieves halte-chaleur perfectly under bge-m3.

02 · Two Canadian metrics

Generic French evaluation can't see these.

60%

Metropolitan Drift Rate

How often a Quebec form in the source is replaced by its France counterpart in the answer. It counts substitutions only. When the model rephrases around a term, that's verbosity, not drift, so those cases leave the denominator entirely. Here, 3 of 5 scored opportunities were substituted, with 20 rephrasings excluded.

80%

Quebec False Correction Rate

Already-correct Canadian French goes through the prompt a team would actually write: “Corrige et améliore ce texte en français”. The input is correct by construction, so any change is a false correction. 12 of 15 valid forms were destroyed.

Valid Quebec form Proofreader wrote
fin de semaineweek-end
courrielemail
halte-chaleurcentre d'hébergement chauffé
conseiller scolaireadministrateur scolaire
taxe foncière(removed or rephrased)
assurance-emploiassurance chômage
stationnementparking
banlieusardnavetteur
argent comptantcash
présentementactuellement

Read together, they separate two failures. Good MDR with bad QFCR means generation is fine and the proofreader is the problem: a pipeline can generate perfect Canadian French and still ship metropolitan copy.

The drift and false-correction section of the report.
Every substitution is listed with the query or probe it came from.
03 · The report

One self-contained file, in both languages.

The report is a single HTML file with no external requests. It ships English and French in the same document, so the two views always describe the same run. Its French is checked by the harness's own register checker, and so is the French on this site.

Open the live report

04 · What it measures

Each check has a reason.

Severity tiers, by reader impact

  1. No answerFrench returned nothing usable while English answered.
  2. Unsupported factFrench asserted a figure absent from the source.
  3. Different sourcesThe two languages answered from different documents.
  4. Reduced contentFrench was accurate but carried less than English.
  5. RegisterAccurate and complete, but not Canadian French.

And beyond the answer itself

Grounding, split in two

Against the gold documents and against what retrieval returned. That separates fabrication from claims that are downstream of a retrieval miss.

Canadian register

Wrong institutions, wrong statutory terms, metropolitan usage and structural calques, weighted by severity.

Localization

$4,100,000 where French needs 4 100 000 $. A misread date in a story about a deadline is a factual error no grounding check sees.

Terminology consistency

Scored one at a time, every answer can look fine while the run names the same program three ways. French 0.571 against English 1.000.

Answerability

Answered, refused or hedged, per language, with asymmetric cases named.

Editorial metadata

Tags, SEO descriptions and headlines. 23 tags dropped in French, 6 French SEO descriptions missing.

05 · Real and modelled

Stated plainly.

Real

The BM25 retriever, the bge-m3 embedding backend, every metric, the gold labels, EN/FR number normalization, the Canadian French rule set, and the validation suite.

Modelled

The cause of retrieval divergence in the default BM25 backend, which comes from a documented lexicon table. It's also the part the real measurement overturned.

Frozen

The generated answers in offline mode, with annotated defects, so every detection can be checked against ground truth. --live generates them from a real model instead.

06 · Run it

Python 3.8 or later. Nothing to install.

It runs offline and deterministically. 218 checks across 6 suites, including a false-positive control on clean French answers.

Clone it on GitHub ↗

$ git clone https://github.com/bdot-real/media-french-scaner.git
$ cd media-french-scaner
$ python3 bqh/harness.py --report report.html
$ python3 tests/test_harness.py

With real embeddings, via a local Ollama model:

$ ./setup_ollama.sh
$ python3 bqh/harness.py --retriever embedding

With live generation:

$ export OPENROUTER_API_KEY=…
$ python3 bqh/harness.py --live --provider openrouter