Verifiable Screening Narrows Early-Career Callback Gaps

A 36,880-application résumé audit links discrimination to evaluative discretion, with gaps shrinking where jobs use more structured criteria.

Editorial Desk·July 28, 2026·5 min readmoderate

Underlying Paper

Hiring Discrimination and the Task Content of Jobs: Evidence from a Large-Scale Resume Audit

We conducted a large-scale resume audit of 36,880 applications to 9,220 job advertisements for new college graduates across the United States. Firms express task preferences through job-advertisement text, which we link to occupation-level task measures from O*NET and the American Community Survey. We develop a model in which discrimination increases with evaluative discretion, defined as the share of hiring decisions driven by subjective rather than verifiable assessment. Callback gaps vary systematically with the task content of jobs. In management occupations, callbacks are 28 to 43 percent lower for Black men, Black women, White women, and Hispanic men than for otherwise identical White men. Broad occupation categories conceal important variation in task demands. When jobs are grouped by task intensity, discrimination concentrates in positions combining high analytical and interpersonal demands with low routine content. Decomposing task content into subjective-evaluation and objective-precision components, we find that subjective evaluation widens callback gaps while objective precision compresses them. Customer contact amplifies this divergence, widening gaps in non-routine jobs but not in routine jobs. Randomly assigned resume credentials that increase callbacks on average reduce gaps in low-discretion jobs but not in high-discretion jobs. Early-career exclusion from high-return task bundles may entrench long-run demographic gaps in employment outcomes.

arXiv:2604.01933Submitted: Jul 24, 2026v2

Hiring discrimination is usually measured as an average callback gap, but the economically important question is often where that gap appears. This paper studies that margin in entry-level jobs for college graduates: whether otherwise identical applicants face larger screening penalties in jobs whose task demands leave more room for subjective judgment. The authors combine a national résumé audit with occupation-level task measures and advertisement text, then test whether verifiable résumé credentials compress gaps in lower-discretion settings.

Core Contribution

The paper’s main contribution is to treat job tasks as part of the screening technology. Analytical and interpersonal work is modeled as harder to evaluate from a résumé because employers must infer fit, communication, and judgment from thin signals. Routine cognitive work and explicit screening instruments, by contrast, give employers more verifiable criteria. The theoretical object is an evaluative-discretion index, Ej=B^j/(B^j+P^j)E_j^* = \hat{B}_j/(\hat{B}_j + \hat{P}_j), where B^j\hat{B}_j proxies subjective evaluation difficulty and P^j\hat{P}_j proxies objective precision.

The paper does not claim that task content is randomly assigned. It is careful on that point. Applicant race, ethnicity, gender, and credentials are randomized within job advertisements, so callback gaps within an ad are experimentally identified. The relationship between those gaps and a job’s task content is a cross-occupation pattern, supported by several checks rather than by direct manipulation of tasks.

Technical Approach

The audit sent four randomly generated résumés to each of 9,220 job advertisements in 2016 and 2017, yielding 36,880 applications to 4,969 firms. Names signaled White men, Black men, Hispanic men, White women, Black women, and Hispanic women. Résumé controls included university, major, GPA, internships, computer skills, study abroad, language ability, and other credentials.

Job advertisements were mapped to 175 detailed occupations using the ONET-SOC Autocoder and linked to ONET and ACS task measures. The task taxonomy covers analytical, interpersonal, routine cognitive, routine manual, physical, and contact intensity. The sample is concentrated in four broad occupation groups: management, business and financial operations, sales, and office and administrative support, which together cover about 94 percent of audited advertisements.

Figure 2 shows why sample composition matters: the audit jobs are not a random slice of the labor market. They are less routine-manual and less physical than all ACS occupations, while somewhat more analytical, interpersonal, routine-cognitive, and contact-intensive. That makes the study most informative about early-career, white-collar job search rather than the full occupational distribution.

Figure 2. Distributional Comparisons of Tasks, Occupations in the Audit and All Occupations in the ACS

Results and Analysis

The headline result is that callback gaps are concentrated in high-discretion task bundles. In the full sample, Black men face a 2.1 percentage-point callback gap and Black women a 1.4-point gap relative to White men. The broad occupation split is sharper: in management jobs, four non-reference groups have significant gaps from 3.3 to 5.1 points. The Black male gap in management is 5.1 points, or 28 percent relative to the White-male callback rate; in the Poisson specification, significant management callback ratios range from 0.72 for Black men to 0.82 for White women.

The K-means task clusters give a more task-specific view. The largest gaps appear in the cluster with high analytical and interpersonal intensity and low routine cognitive content. Four of five non-reference groups have significant gaps there, from 2.3 to 3.6 percentage points; the Black male gap of 3.6 points is 42 percent of the White-male callback rate in that cluster. A second cluster has similarly high analytical intensity but higher routine cognitive content, and no group shows a significant gap there. That contrast is the paper’s cleanest descriptive evidence that objective structure, not just job complexity, matters.

The credential experiment is the strongest mechanism test. Social internships, programming-plus-data skills, and study abroad raise callbacks on average by 1.1, 1.0, and 0.8 percentage points. In low-discretion jobs, the two larger-return credentials reduce non-White-male gaps by 4.1 and 3.7 points. In high-discretion jobs, the same credentials provide no comparable differential benefit, and the joint test for the positive-return credential triples rejects at p = 0.048. The interpretation is direct: verifiable signals help when the screening environment can use them, but they do little in jobs where employers still rely on subjective résumé evaluation.

Limits and Significance

The evidence is strongest for callback-stage discrimination and for randomized credential effects within advertisements. It is weaker for pinning down the causal role of task content, because jobs differ in wages, firms, applicant pools, and screening practices along with tasks. The authors address this with job-ad fixed effects, alternative task groupings, advertisement-text measures, Poisson estimates, screening-instrument analyses, and occupation-level bootstrap checks, but task content itself remains observational.

The practical implication is narrow but important. Early-career exclusion is not evenly distributed across jobs; it appears most in the analytical and interpersonal roles that are often better paid and more career-shaping. The paper also warns that standardized screening will not automatically remove gaps if it leaves the key judgment about fit untouched.

Evidence Box

moderate

Key Claims

  • Callback discrimination concentrates in high-discretion task bundles
  • Objective screening criteria compress non-White-male callback gaps
  • Customer contact amplifies gaps mainly in non-routine jobs
  • Verifiable credentials reduce gaps only where evaluation is structured

Key Results

  • 36,880 applications sent to 9,220 advertisements at 4,969 firms
  • Black men face a 2.1 percentage-point full-sample callback gap and Black women a 1.4-point gap relative to White men
  • Management gaps range from 3.3 to 5.1 percentage points, with the Black male gap equal to 28 percent of the White-male callback rate
  • Social internships and programming-plus-data credentials reduce low-discretion non-White-male gaps by 4.1 and 3.7 percentage points

Limitations & Caveats

  • Task content is not experimentally assigned across job advertisements
  • Callbacks measure only the initial screening stage, not offers or wages
  • Sample excludes many licensed and specialized fields such as medicine, law, engineering, and accounting
  • Occupation-level task measures may miss within-occupation variation in individual postings

Related Articles

Readers are encouraged to consult the original arXiv paper for complete details. SOTA Papers does not make claims beyond what is supported by the authors' reported evidence.