We started like everyone else. Put the job ad and the CV into a language model and trust the answer. It did not work, and here is what we did instead.
A job ad is written to make people apply. It is sales copy. It contains sentences like “you are driven and enjoy a fast pace” and “here we are everything from bakers to snowplough drivers”. It is written to sound appealing, not to decide who fits.
Building an assessment on it is like valuing a home by the estate agent’s listing.
It shows the moment you measure. We ran four different ads for what was in practice the same role through the same model and got four different requirement lists. The difference was not that the jobs differed, but who happened to have written the ad.
Worse, requirements crept in that a CV cannot possibly answer. “Available evenings and weekends” is a reasonable requirement for the job, but it appears in no CV. So it was always unmet, for everyone, and blocked the top for everyone.
We had 45 real applications, assessed by hand as the answer key. Nine of them were top candidates. The system found one.
We assumed the fault was in how we had instructed the model, so we rewrote the instructions. Seven times. Nothing happened on any of them.
We did about thirty measurements before the assessment held. Sixteen of them were attempts to fix things by writing clearer instructions. Fifteen of the sixteen moved nothing at all.
That was the most expensive lesson and the most useful one. When something does not work in a system like this, it is rarely the wording. It is what you feed in, and who wrote it.
So we stopped patching the instructions and changed the input.
The role is now described once per position, by a person, in plain text. What the work consists of. What counts as the same kind of work. Where the lower limit goes. The text is written once and applies to every posting of the role, whoever writes the ad. The job ad is not used as the basis for assessment at all.
The requirements a CV cannot answer are gone from it. Availability we ask about in the application. Manner you hear in the conversation. A CV tells you what someone has done, not what they are like.
The assessor reads only the version of the CV the candidate has gone through and approved, never a model’s interpretation of a PDF the candidate has not seen.
Then it works in four steps.
First, every entry in the CV is rewritten as functions: “took payment and handled the till”, “served several customers at once during rush hour”. This is done without the role being taken into account, and we measure in code whether the model has borrowed words from the requirements profile that are not in the candidate’s own entry.
Then the same for the role. Then the match, requirement by requirement, as a table: covered, partly, or missing, with a line in the CV to point to for every answer.
And only then the level. The model does not set it. It is computed in code from the table, and the justification the recruiter reads is written from the same rows. The model can no longer argue for one outcome and deliver another, which is precisely what happened in version one.
The scale is deliberately coarse. A CV cannot separate the second best from the third best, and pretending otherwise is putting numbers on guesses.
The rules on top all point the same way. In doubt, the level goes up: a wrong “low” costs a candidate their chance, a wrong “high” costs the recruiter a moment. Old experience lowers one step, never to low. Missing dates never punish. A CV that does not exist is not the same thing as a weak CV.
The top group is clean now. Of the twelve candidates the assessor placed highest, none had a weak background by our own reading, and seven of our nine top candidates were there. The first version found one.
What it does not yet do well enough is separate twenty months from three weeks under the same title. That is the next thing we measure.
Every candidate gets two levels. Background, from the CV assessment. Profile, from the conversation, with the candidate’s own words under every level.
The profile weighs most. That is a choice, and it can be supported.
Frank Schmidt and John Hunter’s review of selection methods is the most cited in the field. In 2022, Paul Sackett and colleagues redid the calculation with a better method, and the list reshuffled. Structured interviews became the strongest single predictor of how someone performs on the job, stronger than logic tests and personality tests.
Years of experience came last. Van Iddekinge and colleagues went through 81 samples in 2019, with more than eleven thousand people, and found a relationship close to zero, even when the experience was directly relevant to the role.
That is uncomfortable, because experience is the first thing we look at.
The research decided how we built. The conversation carries the assessment, and a weak background is a reason to hesitate, not a cap. Only when the CV shows nothing resembling the work and the conversation also gives fewer than two strong answers does the candidate stop at Medium.
The difference is concrete. A candidate with four strong interview answers and a weak background can now end up at the top. For a long time she could not, however good she was in the conversation.
So we did not throw the CV away, we moved it down. Two measures that measure different things are together better than the best of them alone. A CV says what someone has done. A conversation says how.
We write about transparent recruitment, AI in hiring, and what we are building. Short, honest updates. No spam.