Is the French answer as good as the English one?
French Drift runs the same corpus, the same questions and the same metrics through an AI content pipeline in English and in Canadian French, side by side. Then it reports the gap. Two languages can each clear a quality bar while the French reader gets a thinner, less grounded answer, in a register that reads as foreign. Scoring each language on its own can't see that.

Queries where the French reader got a substantively equivalent answer.
Valid Quebec forms a naive “proofread this” pass rewrote.
French answers carry less than their English counterparts.
Assets whose French descriptors are complete.
The hypothesis it was built on didn't survive contact.
The harness started from a plausible argument: Canadian French vocabulary is weakly represented in retrieval models, so French retrieval diverges. The default backend models that gap from a hand-written lexicon table. Then we measured it with a real multilingual embedding model, bge-m3, run locally.
| Dimension | Modelled gap (BM25) | Measured gap (bge-m3) |
|---|---|---|
| Retrieval precision@1 | +0.039 | −0.010 |
| Retrieval recall@3 | +0.017 | +0.003 |
| Retrieval nDCG@3 | +0.022 | −0.005 |
| Grounding in retrieved context | +0.033 | +0.021 |
| Content coverage | +0.116 | +0.116 |
| Canadian French register | +0.059 | +0.059 |
| Localization | +0.015 | +0.015 |
On precision and nDCG, the sign flipped: French retrieval came out marginally ahead. The query the harness was built around, “where can people go to keep warm?”, retrieves halte-chaleur perfectly under bge-m3.
Generic French evaluation can't see these.
Metropolitan Drift Rate
How often a Quebec form in the source is replaced by its France counterpart in the answer. It counts substitutions only. When the model rephrases around a term, that's verbosity, not drift, so those cases leave the denominator entirely. Here, 3 of 5 scored opportunities were substituted, with 20 rephrasings excluded.
Quebec False Correction Rate
Already-correct Canadian French goes through the prompt a team would actually write: “Corrige et améliore ce texte en français”. The input is correct by construction, so any change is a false correction. 12 of 15 valid forms were destroyed.
| Valid Quebec form | Proofreader wrote |
|---|---|
| fin de semaine | week-end |
| courriel | |
| halte-chaleur | centre d'hébergement chauffé |
| conseiller scolaire | administrateur scolaire |
| taxe foncière | (removed or rephrased) |
| assurance-emploi | assurance chômage |
| stationnement | parking |
| banlieusard | navetteur |
| argent comptant | cash |
| présentement | actuellement |
Read together, they separate two failures. Good MDR with bad QFCR means generation is fine and the proofreader is the problem: a pipeline can generate perfect Canadian French and still ship metropolitan copy.

One self-contained file, in both languages.
The report is a single HTML file with no external requests. It ships English and French in the same document, so the two views always describe the same run. Its French is checked by the harness's own register checker, and so is the French on this site.
Each check has a reason.
Severity tiers, by reader impact
- No answerFrench returned nothing usable while English answered.
- Unsupported factFrench asserted a figure absent from the source.
- Different sourcesThe two languages answered from different documents.
- Reduced contentFrench was accurate but carried less than English.
- RegisterAccurate and complete, but not Canadian French.
And beyond the answer itself
Grounding, split in two
Against the gold documents and against what retrieval returned. That separates fabrication from claims that are downstream of a retrieval miss.
Canadian register
Wrong institutions, wrong statutory terms, metropolitan usage and structural calques, weighted by severity.
Localization
$4,100,000 where French needs 4 100 000 $. A misread date in a story about a deadline is a factual error no grounding check sees.
Terminology consistency
Scored one at a time, every answer can look fine while the run names the same program three ways. French 0.571 against English 1.000.
Answerability
Answered, refused or hedged, per language, with asymmetric cases named.
Editorial metadata
Tags, SEO descriptions and headlines. 23 tags dropped in French, 6 French SEO descriptions missing.
Stated plainly.
Real
The BM25 retriever, the bge-m3 embedding backend, every metric, the gold labels, EN/FR number normalization, the Canadian French rule set, and the validation suite.
Modelled
The cause of retrieval divergence in the default BM25 backend, which comes from a documented lexicon table. It's also the part the real measurement overturned.
Frozen
The generated answers in offline mode, with annotated defects, so every detection can be checked against ground truth. --live generates them from a real model instead.
Python 3.8 or later. Nothing to install.
It runs offline and deterministically. 218 checks across 6 suites, including a false-positive control on clean French answers.
$ git clone https://github.com/bdot-real/media-french-scaner.git
$ cd media-french-scaner
$ python3 bqh/harness.py --report report.html
$ python3 tests/test_harness.py
With real embeddings, via a local Ollama model:
$ ./setup_ollama.sh
$ python3 bqh/harness.py --retriever embedding
With live generation:
$ export OPENROUTER_API_KEY=…
$ python3 bqh/harness.py --live --provider openrouter