Evaluation ·
What a refusal taught us
For this site's demonstration, we queried Knowledge AI on an extract of sixteen articles of the French Labour Code. One question was meant to be refused, two were meant to be answered. The result is more instructive than expected: it shows what a refusal proves, what it does not, and how to check it.
1. The protocol
The corpus consists of sixteen articles of the French Labour Code, published by Légifrance under the Open Licence, in three documents: outside companies and the prevention plan, chemical risk, risk assessment. Knowledge AI was queried from its main branch, on a dedicated database. The answers were produced by the Ministral 3 8B model, on CPU only.
Three questions, as a company director would ask them, in French. The first two have their answer in the corpus. The third has none: “What does the company risk if the prevention plan has not been drawn up?” None of the sixteen articles provides for a penalty for that absence.
2. Five refusals out of five
The third question passes the preliminary evidence check: all its terms appear in the passages retrieved. The decision therefore rests with the model alone. Over five completed trials, two through the programming interface and three in the product's interface, it refused five times.
The six passages it read deal with the prevention plan and risk assessment; none provides for a penalty when no prevention plan has been drawn up. A general-purpose model most likely knows labour law. It did not fill that silence with its general knowledge: it stayed silent.
3. What the refusal did not prove
The corpus contains a quantified penalty on a neighbouring subject: article R. 4741-1 punishes failing to record the risk assessment, not failing to draw up a prevention plan. We first presented the refusal as resisting that trap.
Before publishing it, we reconstructed the context actually passed to the model. The passage that mentions R. 4741-1 ends on the article's heading; its text, the penalty itself, sits in a separate chunk that was not passed. The model did not set the penalty aside: it never saw it. We withdrew that story.
The lesson applies to any evaluation: a refusal must be read against what the tool actually read, not against what the corpus contains.
4. Two wrongful refusals
The two questions whose answer was in the corpus were refused, before any generation, by the evidence check. This check measures the share of the question's terms found in the passages retrieved; it refuses below 25% and only answers from 50%.
| Question (translated from French) | Terms found | Decision |
|---|---|---|
| “A contractor is coming next week to service our degreasing tanks. What do we need to formalise with them?” | 10% | Refused |
| “One of our solvents is classified as carcinogenic. Are we required to replace it?” | 29% | Refused |
| “An outside company is going to carry out work in our workshops. What must we draw up with them?” | 63% | Answer allowed |
| “A carcinogenic chemical agent is used in our workshops. Must we substitute it?” | 50% | Answer allowed |
5. The defects recorded
These trials revealed five defects in the product. We recorded them rather than worked around them; no setting was changed for the demonstration:
- the evidence check compares words, not meaning: “contractor” instead of “outside company” is enough to have a legitimate question refused;
- the refusal does not say why the tool is silent;
- chunking can separate an article's heading from its text: 6 articles out of 16 are split;
- the audit trace names the model declared in the configuration, not the one that answered;
- a damaged PDF fails without a visible reason, whereas a protected PDF is now quarantined with an understandable reason.
A refusal can be checked
A refusal is only worth something if you know what the tool read. The scripts and results of these measurements are kept with this site's code and presented during a demonstration. What cannot be checked is not published.