If you have ever seen a chart ranking hiring methods by how well they predict job performance, it almost certainly came from a 1998 paper. The numbers on that chart were revised in 2022, and the order changed.

Key Takeaway

Sackett, Zhang, Berry and Lievens found that range restriction corrections used across decades of personnel selection meta-analyses systematically overcorrected, meaning the validity of many selection procedures has been substantially overestimated[1]. In the revised estimates, structured interviews emerged as the strongest predictor at r = .42, ahead of job knowledge tests (.40), empirically keyed biodata (.38), work samples (.33) and cognitive ability (.31)[2]. Sackett described it as the most important paper of my career and a course correction for the field[2].

Our Grade For This Claim

Applying the scheme from the first article in this series.

Grade A for the direction. The finding that prior estimates were inflated by systematic overcorrection is published in the Journal of Applied Psychology[1], has generated a follow-up applied paper by the same authors in a Cambridge journal[3], and is reported by the field's own professional society as a course correction[2].

Grade B for the specific coefficients. They come to us through a professional society article rather than from the paper's own tables, though that article was reviewed by the lead author before publication[2], which is unusually good provenance for a secondary source.

One thing distinguishes this from the disputes in the previous articles. The revision was not produced by opponents attacking a finding. It came from within the same research tradition, correcting its own methodology, and the authors of the corrected work published guidance on how to do it properly.

That is what a field looks like when it is functioning, and it is worth saying so after four articles that mostly showed the opposite.

A Note On Method

Everything here is verified to August 2026.

We did not obtain the full text of Sackett and colleagues (2022) or of Schmidt and Hunter (1998). We rely on the published abstract of the former[1] and on a detailed article in the professional society's publication[2].

That society article is written by researchers at a personnel research organisation and states that Paul Sackett reviewed previous drafts[2]. We treat it as high-quality secondary reporting for that reason and say so wherever we rely on it.

We do not report Schmidt and Hunter's original coefficient for cognitive ability, because we did not obtain it. Only one before-and-after pair is available to us, for interests, and we give it.

We describe the statistical mechanism as the society article explains it, and we are not in a position to evaluate the underlying argument independently.

This article reviews selection research. It is not employment law, human resources or hiring advice. Selection practice is subject to human rights and employment legislation that varies by jurisdiction and that this article does not address.

A Source We Rejected

Something that happened while researching this article, reported because it bears directly on how anyone should use material of this kind.

The specific validity coefficients above first reached us through a webpage that presented them cleanly, cited the paper correctly, and had the numbers right.

That page also carried a byline stating it was written by agents, meaning generated by an AI system, with a note that pages personally verified by the site's owner say so and this one did not.

We did not use it. We went and found the figures in the professional society's own publication instead, which is where they now come from in this article.

Three observations, ours.

The AI-generated page was, as far as we can tell, accurate. This is not a story about a machine inventing numbers.

But we had no way to know that in advance, and the cost of checking was one additional retrieval. Where a figure is going to be published and relied on, that is not a meaningful cost.

And the general rule we would draw is not never use AI-generated sources. It is that a number's provenance is part of the number. A coefficient that cannot be traced to a document produced by an accountable author is not yet a fact you can publish, however plausible it looks.

We include this because the same discipline is available to any reader, and because a series about research quality that quietly used an unverified secondary source would be worth very little.

Fifty Years Of Orthodoxy

What was believed, and how firmly.

The society's own account puts it plainly. Within industrial and organisational psychology there had been at least one fundamental truth: cognitive ability is the best predictor of work performance. It was rooted in numerous meta-analyses, including Schmidt and Hunter (1998), and had been confidently proclaimed for over half a century, with decades of research and hiring and promotion methods built on this proclamation[2].

The article's framing of how unusual that certainty was is worth noting: researchers ordinarily take great pains to emphasise the tentative nature of their conclusions, and this was one of the few things the field asserted without hedging[2].

The source paper is Schmidt and Hunter (1998), The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings, Psychological Bulletin, 124, 262 to 274[2].

Two observations, ours.

Eighty-five years of research findings is an enormous evidentiary base, and the paper's authority was earned rather than assumed.

And the practical reach was total. If you have used a validity ranking chart in a hiring discussion, attended a recruitment training course, or read almost any article on selection published before 2022, the numbers came from this lineage.

The Statistical Error

What was wrong, stated as the paper's own abstract states it.

The abstract records that corrections for range restriction in meta-analyses of predictor-criterion relationships in personnel selection typically involve the use of an artifact distribution; that after outlining and critiquing five approaches commonly used to create and apply such distributions, the authors conclude that each has significant issues that often result in substantial overcorrection; and that therefore the validity of many selection procedures for predicting job performance has been substantially overestimated[1], a conclusion reproduced identically in the publisher aggregation record[4].

The society article summarises the target of the critique: the authors argue that commonly used corrections systematically inflate relations among personnel selection assessments and job performance, and are particularly critical of one widespread practice[2].

That practice is the subject of the next two sections, because it is comprehensible without statistical training and is the crux of the whole matter.

Predictive Versus Concurrent

The distinction that does all the work.

The society article explains it directly. Predictive validation designs include actual job applicants who are hired on the basis of the assessment. Concurrent validation designs involve administering the same assessment to current employees[2].

The problem: the widespread practice being criticised uses range restriction estimates generated from predictive validation studies to correct the full set of studies included in a meta-analysis, which also includes many adopting concurrent designs[2].

And the reason that fails: because current employees were not selected based on the assessment administered in concurrent studies, applying across-the-board corrections overinflates validity estimates, sometimes to a substantial degree[2].

Our own restatement, for a reader who does not work with this material.

Range restriction correction exists to fix a real problem. If you only hire the people who scored well on a test, the range of test scores among your employees is narrower than among applicants, which artificially lowers the observed correlation between test score and performance. Correcting for that is legitimate.

But the size of the correction depends on how much restriction actually occurred. And in a concurrent study, where existing staff take a test that was never used to select them, the restriction on that test is much smaller or absent. Applying a correction calibrated for genuine selection to a study where no selection on that measure occurred inflates the result.

Why That Inflates Everything

The consequence, and why it affected the whole table rather than one entry. This section is our own analysis.

Three features made this a systematic rather than a random problem.

The correction was applied across the board, to predictive and concurrent studies alike, so it did not average out.

It ran in one direction. Overcorrection always inflates, never deflates, so errors accumulated rather than cancelling.

And it was invisible in the outputs. A corrected coefficient looks exactly like an uncorrected one. Nothing in a published validity table signals how much of the number came from the data and how much from the adjustment.

Two consequences worth drawing out.

The effect on the ordering was not uniform, which is the interesting part. If every estimate had been inflated by the same proportion, the ranking would have survived and only the absolute numbers would have fallen. The ranking changed, which means different predictors had different proportions of concurrent studies behind them.

And this connects directly to the first article in this series. It is another instance of the general finding that published effect sizes in this domain run substantially above what better methods produce. Here the mechanism is not publication bias but a technical correction, and the direction is the same.

The Revised Ordering

The result, with figures as the society article reports them.

Structured interviews emerged as the strongest predictors of job performance[2], with the highest mean operational validity at r = .42.

The article reports that several other job-specific assessments appeared among the top five: job knowledge tests at .40, empirically keyed biodata at .38, and work sample tests at .33. And that cognitive ability rounded out this list with a validity estimate of .31[2].

Two observations, ours.

The top four are all job-specific. Structured interviews, job knowledge, biodata keyed to the role and work samples all involve assessing the candidate against the particular work, rather than measuring a general capacity.

The society article notes this pattern explicitly, describing the group as job-specific assessments that fared quite well[2].

That is a coherent story rather than a reshuffled list, and it points at a practical principle: the closer an assessment sits to the actual job, the better it appears to predict performance in it.

The Reframing In The Authors' Words

What the authors say the result means, which is more than a change of ranking.

The society article quotes them: the finding suggests a reframing. Although Schmidt and Hunter (1998) positioned cognitive ability as the focal predictor, with others evaluated in terms of their incremental validity over cognitive ability, one might propose structured interviews as the focal predictor against which others are evaluated[2].

Three consequences, ours, and this is the part with the most practical bite.

Under the old frame, the question about any hiring method was what does it add beyond a cognitive test. That framing built the test into the baseline of every comparison.

Under the new frame, the baseline is a well-run structured interview, and every other instrument has to earn its place against that.

Which reverses the burden in a practical way. A vendor selling an assessment product is no longer answering is this better than nothing. They are answering is this better than interviewing properly, which is a considerably harder question and one most organisations can test themselves.

Sackett's own characterisation of the paper is worth recording: I view this as the most important paper of my career, offering a course correction to the field's cumulative knowledge about the validity of personnel selection assessments[2].

What These Numbers Actually Mean

Translating correlations into something usable, with our own arithmetic.

A correlation squared gives the proportion of variance associated with the predictor. On the revised figures, our own calculations give:

Structured interviews at .42: about 17.6 percent of variance in job performance.

Job knowledge tests at .40: about 16.0 percent.

Empirically keyed biodata at .38: about 14.4 percent.

Work samples at .33: about 10.9 percent.

Cognitive ability at .31: about 9.6 percent.

The gap between the top and bottom of that list is 8.0 percentage points of variance, meaning structured interviews are associated with roughly 1.8 times as much explained variance as cognitive ability tests.

Two cautions, ours. These are our own calculations from reported correlations and appear in no source. And squaring a correlation is a conventional but crude translation, which understates practical usefulness in selection contexts where you are choosing among candidates rather than predicting an individual's score.

The Humbling Part

The observation we think matters most for anyone running a hiring process. This section is our own.

Take the best number in the revised table. A structured interview at r = .42 is associated with about 17.6 percent of the variance in job performance.

Which leaves roughly 82 percent unexplained by the single best selection tool the field has identified after ninety years of research.

Three consequences.

Hiring is irreducibly uncertain, and a good process is one that improves the odds rather than one that identifies the right person. Any vendor or consultant implying otherwise is selling something the research does not support.

Which means what happens after the hire matters at least as much as the selection. Onboarding, early feedback, the willingness of a new person to ask questions, and a manageable exit if it is wrong are all operating on the 82 percent.

And it argues for probation periods used seriously rather than as a formality, since direct observation of actual work is the only instrument that gets at what selection cannot.

We would put the practical version plainly. Improving your interview from unstructured to structured is among the highest-value changes available. It will still be wrong often, and a process designed on the assumption that it will not is a process that has misread the evidence.

The Caveat On Structured Interviews

The qualification the authors attach to their own headline finding, which is rarely repeated.

The society article records that while structured interviews had the highest mean operational validity, they also showed a relatively high degree of spread around that mean[2].

It attributes this partly to the wide range of constructs targeted by structured interviews, and notes the finding is a compelling call for researchers to identify the factors responsible for this variation and the approaches to developing, administering and scoring structured interviews that foster strong validities[2]. It flags digital interviewing and AI-based interview scoring as adding to the urgency[2].

Three consequences, ours.

"Structured interview" is a category, not a method. The .42 is an average across implementations that differ substantially, and yours will not be the average.

Which means the headline number is not a promise. Adopting the label does not confer the validity; the design, administration and scoring do, and the research has not yet specified which features matter most.

And it is an honest thing for the authors to have foregrounded, since it complicates the clean message their own paper produced. We note it here for the same reason.

The One That Went Up

The single before-and-after pair we can report, and it moved in the unexpected direction.

The society article records that compared to Schmidt and Hunter's work, the operational validity of interests increased from .10 to .24, and attributes the increase to Sackett and colleagues defining interests in a fit-based way, being between personal interests and unique job demands, rather than in a general way, being the relation between a general type of interest such as artistic or investigative and overall job performance[2].

Our own arithmetic: that moves explained variance from about 1.0 percent to about 5.8 percent, roughly a sixfold increase.

Three observations, ours.

Nothing new was measured. The increase came from redefining the construct, from a general trait to a fit between a person and a specific role.

That is the same pattern as the rest of the revised table. Job-specific beats general, and here the identical underlying idea moves from near-useless to modestly useful purely by being made specific.

And it is a caution about reading any single coefficient. A validity figure is a property of a construct as operationalised, not of the concept in the abstract. "Do interests predict performance" has no answer until you say what you mean by interests.

Adding Two Words To Each Question

The most immediately usable finding in the whole revision.

The society article reports that tailoring personality items to the job context increases their predictive validity, describing the operation as adding "at work" to each item, or asking applicants to respond in terms of how they behave at work. And that the validities were so much stronger for contextualised personality assessments that Sackett and colleagues suggest viewing them as essentially a different type of assessment relative to more general personality inventories[2].

Two observations, ours.

The intervention is close to free. It is a wording change to an existing instrument, not a new instrument.

And the effect was large enough that the authors propose treating the contextualised version as a different assessment altogether, which is a strong claim about a change of two words.

We did not obtain the figures behind that claim and report none. What we take from it is the same principle that runs through everything above: specificity to the job is doing the work, repeatedly, across unrelated instrument types.

For a reader who wants one transferable idea from this article, that is it. Whatever you are assessing, ask about it in the context of the actual job rather than in general.

Three Principles For Anyone Reading Research

The authors' methodological guidance, which generalises well beyond hiring.

The society article sets out three principles the authors advocate for future meta-analytic work[2].

Critically evaluate your assumptions. The authors advocate an end to the practice of simply assuming a degree of restriction with no empirical basis.

Be conservative, expressed as when in doubt, don't correct. If you conclude you do not have a credible estimate of range restriction or unreliability for a given study, it is better to be conservative and not correct than to apply an inaccurate correction.

Think locally. Rather than base corrections on general rules of thumb, ask whether more local sources of information would provide a more accurate estimate.

Two observations, ours.

The second principle is the striking one, because it runs against the instinct of anyone doing quantitative work. The natural impulse on discovering a known bias is to correct it. These authors argue that an uncorrected number with a known direction of error is safer than a corrected number with an unknown one.

And all three amount to a single discipline: an adjustment is a claim, and it needs evidence like any other. That applies as much to a business adjusting its own figures as to a meta-analyst.

What Structured Actually Means

Definitional groundwork, and we mark clearly what is ours.

The 2022 paper's domain included structured and unstructured interviews as separate categories[2], and the 2023 follow-up notes that Schmidt and Hunter's earlier review had likewise separated interviews into structured versus unstructured[3].

We did not obtain either paper's operational definition of structure, and we are therefore not going to state one as though it came from the research.

What we can say, as our own account of what the term conventionally denotes, is that structure refers to constraining the interview so that candidates face comparable conditions: the same questions in the same order, questions derived from an analysis of the job, and answers scored against defined criteria rather than assessed holistically.

Two consequences, ours.

The society article notes that practitioners' experience is that well-crafted structured interviews, grounded in detailed job analytic data and conducted by well-trained interviewers, are among the best selection tools available[2]. Those three qualifiers are the substance.

One practical qualification we did locate, from a consultancy summary rather than from the research: the effectiveness of empirical biodata and structured interviews varies, and while both can be adapted for new and experienced applicants, the content must be adjusted to be job-non-specific for inexperienced candidates[5]. We report that as a commercial publication’s observation and did not verify it against the papers. If it holds, it matters for any firm hiring junior staff, because the job-specific content that carries the validity is the very thing an inexperienced candidate cannot yet speak to.

And the honest position is that a business wanting the benefit should obtain the papers or professional guidance rather than working from our description, because the spread noted above means implementation details determine whether you get the validity at all.

The Small Business Problem

Why this matters more at small scale, in our assessment.

Three features of an owner-managed firm change the calculation.

Each hire is a larger fraction of the organisation. A bad hire in a ten-person business is ten percent of the workforce, and frequently a much larger share of a specific capability.

There is no assessment budget. Cognitive tests, biodata instruments and work-sample platforms carry licensing costs and administration overhead that small firms rarely absorb.

And the interview is happening anyway. Every small business interviews. The question is only whether it does so in a way that predicts anything.

Two consequences.

The revised ordering is unusually favourable to small firms. The instrument that now ranks highest is the one they already use and can improve at close to zero marginal cost.

And the old ordering was unusually unfavourable, because it pointed at cognitive testing, which is precisely the tool a small business is least able to buy, validate and defend.

We offer this as reasoning from the revised figures rather than as a finding, and note that we did not locate research on selection validity specifically in small organisations.

The Cheapest Available Upgrade

Our own view of where the effort should go.

The comparison that matters for most readers is not between a structured interview and a cognitive test. It is between a structured interview and the unstructured conversation they are currently having.

We did not obtain a validity figure for unstructured interviews from any source in this article, and we are not going to quote one from memory. Both the 1998 and 2022 analyses treated the two as distinct categories[2][3], which is itself informative: the field would not maintain the distinction if it did not matter.

What we would say, as reasoning rather than as evidence.

The costs of moving from unstructured to structured are a job analysis, a fixed question set, and a scoring sheet. That is preparation time, not expenditure.

The change is testable in your own organisation, which the first article in this series argued is stronger evidence for you than any published coefficient. Record your ratings, wait a year, and compare them against how people actually performed.

And almost nobody does that. Most organisations have no idea whether their hiring process predicts anything, because they never record a prediction in a form that could later be checked. A scoring sheet is worth having for that reason alone, independent of any validity claim.

A boundary this article does not cross, stated explicitly.

Selection procedures are subject to human rights and employment legislation, and in Canada those obligations arise at both federal and provincial levels depending on the employer.

Three things follow, and they are cautions rather than advice.

A method's predictive validity is not the same as its legal defensibility. Those are separate enquiries with separate standards.

The 2023 follow-up paper's own reference table includes subgroup difference statistics alongside validity figures[3], which indicates the research community treats differential impact as inseparable from the validity question. We did not obtain those figures and report none.

And structure itself has an obvious relationship to defensibility, since asking every candidate the same job-related questions and scoring against defined criteria produces a record. We are not qualified to say what weight that carries.

Anyone changing a selection process should take employment law advice in their own jurisdiction. This publication is not providing it.

What To Do

Stop using pre-2022 validity charts. The numbers were revised, the ordering changed, and a chart lifted from the 1998 lineage without the correction is out of date.

Treat the structured interview as the baseline. The authors themselves propose it as the focal predictor against which others should be evaluated.

Ask vendors the harder question. Not whether an instrument beats nothing, but whether it beats a well-run structured interview, and on what evidence.

Make everything job-specific. The top four predictors are all job-specific, interests improved sixfold by being defined as fit to a role, and contextualising personality items changed them enough that the authors call them a different assessment.

Add "at work" to your questions. The single cheapest intervention identified in the whole revision is a wording change.

Write the scoring sheet even if you doubt the research. It converts your hiring judgment into a recorded prediction, which is the only way you will ever learn whether your process works.

Do not expect certainty. The best available tool is associated with under a fifth of the variance in performance, so design the process after the hire on the assumption that selection will often be wrong.

Check where a number came from. The coefficients in this article first reached us via an AI-generated page. They were right, and we replaced the source anyway.

Take employment law advice separately. Validity and legal defensibility are different questions.

The Limits Of This Analysis

Several caveats matter. This article reviews selection research and is not employment law, human resources or hiring advice; selection practice is governed by human rights and employment legislation that varies by jurisdiction and is not addressed here. Everything is verified to August 2026. We did not obtain the full text of Sackett and colleagues (2022) or of Schmidt and Hunter (1998), and rely on the former's published abstract and on an article in the professional society's publication. All specific validity coefficients in this article come from that society article rather than from the paper's own tables, though that article states the lead author reviewed previous drafts. We do not report Schmidt and Hunter's original coefficient for cognitive ability because we did not obtain it, and the only before-and-after pair we can give is for interests. We did not obtain a validity figure for unstructured interviews and quote none. We did not obtain either paper's operational definition of interview structure, and our description of what structure conventionally means is our own, not drawn from the research. We did not obtain the figures behind the contextualised personality finding and state none. We did not obtain the subgroup difference statistics that accompany the validity estimates in the 2023 follow-up, and report none. Our variance calculations are our own, derived by squaring reported correlations, appear in no source, and use a conventional but crude translation that understates practical usefulness in selection contexts. We are not in a position to evaluate the underlying statistical argument independently. The small business analysis, the observation about the 82 percent unexplained, the reframing consequences and the source-provenance discussion are our own reasoning.

Frequently Asked Questions

What actually changed in 2022?
Sackett and colleagues found that range restriction corrections used across decades of selection meta-analyses systematically overcorrected, meaning validity had been substantially overestimated. In the revised estimates structured interviews rank highest at .42, ahead of cognitive ability at .31, reversing the previous ordering.
Why did the correction inflate the numbers?
Corrections calibrated from predictive studies, where applicants were actually hired on the assessment, were applied across the board to meta-analyses containing many concurrent studies, where current employees took an assessment that had never been used to select them. Little restriction occurred in those, so the correction inflated the result.
Is cognitive ability testing now useless?
No. It appears in the top five predictors at .31. What changed is its position: it is no longer the focal predictor others are measured against. The authors suggest structured interviews might take that role instead.
Does adopting a structured interview guarantee the .42?
No, and the authors flag this themselves. Structured interviews showed a relatively high degree of spread around that mean, partly because of the wide range of constructs they target. The label does not confer the validity; the job analysis, question design, interviewer training and scoring do.
What is the single cheapest improvement?
Probably contextualising your questions. The revision found that tailoring personality items to the job context, by adding "at work" or asking how the person behaves at work, raised validity enough that the authors suggest treating it as a different assessment. That is a wording change to an existing instrument.
How much of hiring can any of this actually explain?
Squaring the best coefficient gives about 17.6 percent of variance in job performance, leaving roughly 82 percent unexplained by the strongest tool the field has found in ninety years. Design your onboarding and probation on the assumption that selection will frequently be wrong.
IB

About The Insight Bureau Research Desk

The Insight Bureau is GSH Financial's research publication, written for Canadian business owners and the students who will eventually advise them. This article documents a source it rejected during research, and states which figures it could not obtain rather than filling the gaps from memory.

References

  1. Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068. DOI 10.1037/apl0000994, on the paper systematically revisiting prior meta-analytic conclusions about the criterion-related validity of personnel selection procedures and particularly the effect of range restriction corrections; on such corrections typically involving the use of an artifact distribution; and on the authors' conclusion, after outlining and critiquing five commonly used approaches, that each has significant issues that often result in substantial overcorrection and that therefore the validity of many selection procedures for predicting job performance has been substantially overestimated. Note: we obtained the published abstract only and no figures from the paper's own tables. pubmed.ncbi.nlm.nih.gov
  2. O'Shea, P. G., & Luscombe, A. F., Human Resources Research Organization. Is Cognitive Ability the Best Predictor of Job Performance? New Research Says It's Time to Think Again. TIP, Society for Industrial and Organizational Psychology, volume 60, number 3, issue 603, on cognitive ability having been treated within I-O psychology as a fundamental truth rooted in numerous meta-analyses including Schmidt and Hunter (1998) and confidently proclaimed for over half a century; on Sackett's statement that he views this as the most important paper of his career, offering a course correction; on the critique of using range restriction estimates generated from predictive validation studies to correct meta-analyses also containing concurrent designs, and on the distinction between predictive designs including actual applicants hired on the basis of the assessment and concurrent designs administering the assessment to current employees who were not selected on it; on structured interviews emerging as the strongest predictors with the highest mean operational validity of r = .42 but a relatively high degree of spread around that mean; on job knowledge tests, empirically keyed biodata and work sample tests appearing among the top five with validities of .40, .38 and .33 respectively and cognitive ability rounding out the list at .31; on the operational validity of interests increasing from .10 to .24 through a fit-based rather than general definition; on contextualised personality assessments showing validities so much stronger that the authors suggest viewing them as essentially a different type of assessment; on the three principles of critically evaluating assumptions, being conservative and thinking locally; on the authors' proposed reframing of structured interviews as the focal predictor; and on the note that Paul Sackett reviewed previous drafts of the article. Note: a professional society publication written by personnel research practitioners; all specific coefficients in this article come from here rather than from the paper's tables. siop.org
  3. Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors. Industrial and Organizational Psychology, Cambridge University Press, received 14 September 2022, revised 12 January 2023, accepted 13 January 2023, first published online 9 May 2023, on the 2022 paper having identified previously unnoticed flaws in the way range restriction corrections had been applied and having offered revised estimates often quite different from prior estimates; on the present paper attempting to draw out the applied implications; on Schmidt and Hunter (1998) having separated interviews into structured versus unstructured; on subsequent meta-analyses having led the authors to differentiate empirically keyed from rationally keyed biodata, knowledge-based from behavioural tendency-based situational judgment tests, ability-based from personality-based emotional intelligence measures, and contextualised from decontextualised personality measures; on correction factors generally coming from predictive studies and overcorrecting substantially if applied to concurrent studies; and on the reference table including subgroup difference statistics alongside validity estimates. Note: we obtained portions only, and did not obtain the subgroup difference figures. cambridge.org
  4. Semantic Scholar record for Sackett, Zhang, Berry and Lievens (2022), reproducing the paper's stated conclusion that after outlining and critiquing five approaches commonly used to create and apply range restriction artifact distributions, each has significant issues often resulting in substantial overcorrection, and that the validity of many selection procedures for predicting job performance has therefore been substantially overestimated. Note: a publisher aggregation record, used to corroborate the abstract at reference 1. semanticscholar.org
  5. Master HR. Insights from Sackett et al. (2022, 2023): Towards a Fuller Understanding of the Validity of Selection Tools, giving the full citations for Sackett, Zhang, Berry and Lievens (2022), Journal of Applied Psychology, 107(11), 2040–2068, and Schmidt, F. L., & Hunter, J. E. (1998), The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings; and noting that the effectiveness of empirical biodata and structured interviews varies, and that content must be adjusted to be job-non-specific for inexperienced candidates. Note: a commercial consultancy publication, used only for citation corroboration and for the observation about inexperienced candidates. master-hr.com

This article reviews selection research and is not employment law, human resources or hiring advice. Neither primary paper was obtained in full. All validity coefficients come from a professional society article rather than from the papers' own tables, and that article states the lead author reviewed it before publication. Figures not obtained, including the original cognitive ability estimate, any unstructured interview estimate, and the subgroup difference statistics, are identified as such and not supplied from any other source. Variance calculations are the authors' own and appear in no source.