Roughly 99% of the Fortune 500 filter applicants through an automated score. The number looks like a measure of candidate quality. It isn't. It measures how much a résumé resembles a job description — and that is a different thing, with a much lower ceiling. Here is how well ATS scoring actually works, rated against what effectiveness really means.
The market says "ATS scoring" as if it were a single technology. It is two — and they fail in different ways.
Parse the résumé into fields, extract keywords, compute overlap with the job description. Output a match %. Below the cutoff, no human ever looks.
Whether the right words appear in the right place. A skill in a dated job bullet outscores the same skill in a flat list.
Map profile and role into vector space, score proximity. Recognises that "Software Engineer" ≈ "Computer Programmer" and infers unlisted skills.
Whether the résumé looks like the job — with real nuance. Still not whether the person can do the job.
Mechanics per Jobscan, Resume Optimizer Pro, CVCraft (commercial reverse-engineering) and vendor documentation.
Personnel-selection science ranks how well each method predicts job performance. Sackett et al. (2022) re-ran the canonical evidence and reordered the field. Watch where the résumé score can actually see.
The irony is the whole story. The methods that actually predict performance — structured interviews, work samples, job-knowledge tests — cannot be run on a PDF. What's left at the screening stage sits at the bottom of the hierarchy, and is asked to do the heaviest cutting in the funnel.
Bar lengths are relative and illustrative of the post-Sackett ordering, not exact coefficients.
Sackett, Zhang, Berry & Lievens (2022), revising Schmidt & Hunter (1998): validity estimates cut ≈.10–.20; structured interviews now top-ranked.
"How effective is the score?" has seven answers, not one. Neither generation earns the precision it displays.
| Dimension | Legacy keyword | AI talent intel. |
|---|---|---|
| Predictive validityDoes the score predict performance? | D− | C |
| Construct validityDoes it measure capability? | D | C+ |
| ReliabilitySame input → same score? | D+ | B− |
| RecallDoes it find the qualified? | F | C+ |
| Fairness / disparate impactEquitable across groups? | D | C− |
| TransparencyCan you see how it decided? | F | C |
| Provenance / defensibilityPer-decision record? | F | C− |
Screening is configured against a long list of disqualifiers, not for a short list of must-haves. The cost lands on qualified people.
Don't use this number. It traces to a 2012 sales pitch from Preptel — a company that folded in 2013 — with no published methodology. The 88% figure is the defensible one.
The honest counter-reading: the failure is human-configured criteria, not the software executing them. The category optimises for efficiency at the cost of recall.
University of Washington researchers varied 120 race- and gender-associated names across 550+ real résumés, ranked by three production-grade models against 500+ real jobs.
The human-in-the-loop defence is weakening. A 2025 UW follow-up found people mirror biased AI recommendations — the score doesn't just sit there, it propagates into the decisions downstream. Amazon scrapped its own recruiting AI in 2018 for the same failure mode.
Wilson & Caliskan, AAAI/ACM AIES 2024 · UW human-mirroring study, 2025 (N=528).
The first large-scale test of AI hiring tools in federal court — and it reached the vendor, not just the employer.
Parallel: Harper v. Sirius XM (2025, Title VII); Eightfold faces its own applicant suit; EEOC v. iTutorGroup already settled (AI age discrimination).
NYC Local Law 144 was the first regime to mandate public bias audits of hiring algorithms. In practice the mechanism barely bites.
Each square = one employer studied. Filled = posted a bias-audit report.
employers had posted an audit report (and just 13 a transparency notice). Cornell calls the gap "null compliance": employers hold enough discretion over scope that a missing audit can't even be read as non-compliance.
Cornell "Null Compliance" (ACM FAccT 2024) · NY State Comptroller audit of DCWP enforcement (Dec 2, 2025). Nearly all posted audits reported impact ratios above 0.8.
A résumé score asks one question: does this document resemble this job? Intelletto runs a different operation — it decomposes, grades, fuses and models the candidate across seven dimensions the applicant tracking system doesn't attempt.
Takes the JD as written — often bloated and brittle. The root cause of the recall failure.
Treats the JD as a first-class engineered artifact. Role-normative enrichment scores against what the role should require — JD alignment plus a market-expected skills delta.
One certification, one keyword. Matched or not matched — no structure beneath it.
Decomposes certifications and skills into their hard and soft dimensions, each scored on its own terms against the role.
Binary. Does the word appear on the page, and ideally in a dated bullet?
Grades proficiency by level — drawn from recency, duration and corroboration, not the candidate's self-assertion.
A flat list of past titles. Direction, velocity and shape of the career are invisible.
Models progression, velocity and direction as a predictive signal in its own right.
One self-reported document plus an application form. That is the entire input — and the entire ceiling.
Fuses professional, code and social signal from multiple corroborating sources, every record provenance-stamped. This is what lifts the score off the résumé-only ceiling.
Most vendors explicitly exclude it — scoring only literal, job-relevant fields.
Models fit to the specific company's context — held inside the audit and compliance layer so it never becomes a proxy for a protected trait.
A once-a-year bias audit at best — under-enforced, gameable, no per-decision trail.
Answerable by decision, built for NYC LL144, the EU AI Act and EEOC from the substrate up.
Data fusion, proficiency grading and trajectory modelling lift Intelletto off the résumé-only input that bounds the ATS category — so the ceiling on the panel above is the ATS's ceiling, not Intelletto's. One discipline remains, and it is a position of strength: criterion validity against real outcomes. Run the study at production scale — score against actual performance and retention, publish the methodology, including the shadow-compare deltas. That turns the category's universal weakness (asserted, retention-keyed validity claims) into the one place Intelletto proves what the rest only state.