The Earnings Test: How safe is your program?
Simulations, specific numbers, and a dataset to check your own program
The Department of Education estimates that only 1% of programs/majors at the Top 100 liberal arts colleges will fail the Earnings Test, a.k.a. “Do No Harm” metric, next year. However, that estimate was based on a single snapshot with a calculation method that does not match what the Earnings Test will actually do. The simulations I show here suggest that the real risk of failure over the next 5 years is much higher, even for those with true earnings well above the benchmark, and it depends a lot on the size of your cohorts.
As one of my PhD advisors is fond of saying: an academic paper is not a mystery novel - just tell us the punchline and then go back to fill in the details. Ok, Jeff, will do!

Quick recap: What is the Earnings Test?
Starting in 2027, college programs/majors will have to pass the Earnings Test to retain eligibility for federal student loans. The Department of Education will match graduates who used Title IV funding to support their education (Pell grants and federal student loans) to their earnings 4 years after graduation.1 If the median graduate of the college program earns more than the median high school graduate aged 25-34, the program passes. If not, the program has to warn its majors that soon they may not be allowed to use federal student loans. If the program fails twice over three years, the program loses federal student loan eligibility. Important detail: if a program has fewer than 30 graduates who can be matched to their earnings data in a single cohort, cohorts are aggregated together until the sample size reaches 30, first bundling up to 5 cohorts of the same program (e.g. Econ alums of ‘21 + Econ alums of ‘20 + … Econ alums of ‘17), then bundling across programs too (see my last post for more details). If no amount of reasonable aggregation can get to 30 matched graduates, the program is exempt from the test.
The figure above shows that a college programs’ grads need to earn several thousand dollars more than high school grads for the program to be safe in the Earnings Test
The figure above shows that the likelihood a program passes the Earnings Test depends on the true median earnings of those graduates AND the size of the program’s cohort. The y-axis shows the true median earnings for graduates in that test bucket. The x-axis shows the number of graduates matched to their earnings in a cohort. The dark dashed line shows the official bar that programs need to clear: $34,808 in 2024 dollars. The shading in the graph shows the risk that the program will fail the Earnings Test at least once in five years and will have to send a warning letter to their majors that soon they may not be able to use federal student loans. If you want to be 95% sure that your program won’t have to send scary letters, your median graduate needs to earn more than the solid orange line. For a program with 29 matched graduates per year, that’s $41,232, or $6,424 more than the official benchmark. For a program with 30 matched graduates per year, that’s $44,214, or $9,406 more than the official benchmark.
But those grads do make more than high school grads. Where does the risk of failing come from?
In short: noise. Random variation. Some grads will make more, some will make less, and in small samples, measured median earnings are going to bounce around a LOT.

The figure above shows what measured median earnings might look like for three different sized programs if we didn’t aggregate across cohorts. Let’s look at the top row first. The solid navy line shows the distribution of earnings for graduates of this program:2 some make more, some make less, for a wide variety of reasons. The height of the line shows how likely it is that an individual graduate will earn the amount of money shown on the x-axis. The distribution line is highest around $30k, meaning that’s the most likely number for any individual graduate. The true median or middle of this distribution is $40k (shown as the vertical dashed navy line), well above the median earnings of high school graduates ($34,808, the vertical dashed orange line) so this program should pass the Earnings test. But this program only has 6 matched graduates per cohort. In a sample that small, the measured median will often be pretty far away from the true median. In this example, you can see 6 little dots representing the 6 graduates in one cohort. The little line shows the median for this particular cohort. In this case, it’s below the high school grad benchmark, and if this program were judged only on these six graduates, it would fail the Earnings Test. The light blue bar shows where the measured median earnings will fall 95% of the time for cohorts of 6 students each; you can see that most of the time this program would be evaluated as “passing” but it would be declared “failing” quite often too.
Bigger samples mean less noise in the measured median. The second row of the above figure shows what happens with a cohort of 14 students. You can see that the light blue bar is narrower than in the first row of the figure. A program with 14 matched graduates per cohort is more likely to pass the Earnings Test, but there’s still quite a bit of the light blue bar that falls below the benchmark. The bottom row of the figure shows what happens for a program with 30 matched graduates per cohort. The light blue bar is narrower still, and a program with true median earnings of $40k and 30 matched graduates per cohort is less likely to fail the Earnings Test due to simple bad luck.
Aggregating across cohorts will increase sample sizes and decreases noise, huzzah! Bad news though: it adds a different and more consequential problem: correlated outcomes across tests.
A program with 6 matched graduates per cohort needs 5 cohorts aggregated together to get to 30 grads in the test bucket, which means that the same graduates show up in multiple consecutive tests. In small programs, ‘21 grads and their 2025 earnings will show up in the 2027, 2028, 2029, 2030, and 2031 Earnings Tests. I hope they make good money! Because if not, one year of bad luck will contaminate five tests in a row.
Small programs that have to aggregate several cohorts together are more likely to have failures that come in streaks that result in losing loan eligibility, not just one-offs that require embarrassing warning letters. The figure below shows the likelihood of losing loan eligibility by cohort size and true median earnings, with darker shading indicating a higher likelihood of losing access to federal student loans.

A program with 6 matched graduates per cohort needs their median earnings to be at least $41,670 - that’s $6,862 above the benchmark - to have less than a 5% chance of losing loan eligibility over the next 5 years. This is shown in the dashed orange line on the figure above. A program with 50 matched completers per year only needs $3,909 of clearance, or true median earnings of $38,717.
Punchline 1: The likelihood of bad outcomes in the Earnings Test is much higher than 1% at the Top 100 LACs
According to the Department of Education’s prediction, only 1.1% of programs at the Top 100 liberal arts colleges are predicted to fail the Earnings Test, shown as red dots in the figure above. However, these simulations show that an additional 3.8% have a greater than 5% chance of losing loan eligibility over the next 5 years. An additional 2.5% have a greater than 5% chance of being declared failing at least once in the next 5 years and must notify their majors that they may lose federal student loan eligibility. The programs that were not previously predicted to fail, but I think should be nervous, are shown as orange dots in the figures above. In total, given the information we have now, I think 7.4% of programs at liberal arts colleges should expect to fail the Earnings Test at least once in the next five years.
Punchline 2: These estimates are a floor; the true likelihood of failing the Earnings Test may be higher
These simulations assume that the type of graduates included in each test bucket is stable across years. Meaning: it’s always 5 cohorts of Fine Art majors with the same earnings distribution in a test bucket, or it’s always 3 cohorts of Dance, Theater, Film, Music, Fine Art and Art History majors in a test bucket, so Earnings Tests across years are comparable and we can say something about how measured median earnings might bounce around and therefore how verdicts of the tests might play out.
That’s not how this is actually going to work.
In reality, in some years the Fine Arts major will have been large enough to be evaluated on its own with 5 cohorts of graduates in a test bucket. The very next year, Fine Arts may have fewer matched graduates over the preceding 5 cohorts and will have to be aggregated together with Art History, maybe over 5 cohorts, or maybe over fewer. And the following year, Fine Arts and Art History together may have been smaller still, so in the third round of the Earnings Test, they may be aggregated together with graduates across All of the Arts and may only need a single cohort to fill the bucket. Maybe graduates from each of those separate majors have similar earnings distributions, or maybe not. Maybe the earnings distributions will be stable across years, or maybe not. (Think: recessions. Also: AI.)
So how are the actual aggregation methods going to play out with respect to variability in the likelihood of Earnings Test failures and loan eligibility sanctions? Inconsistently. Erratically. With risk and ambiguity and genuine Knightian Uncertainty. With outcomes that are difficult to predict and even more difficult to try to improve.
Big Flashing Reminder: We don’t have any idea what might happen for 78% of programs at the Top 100 LACs
At the Top 100 LACs, only 22% of programs have any prediction at all about whether they will pass or fail, so 78% have no idea what might happen in the first Earnings Test next year. The Department of Education’s predictions of failing and passing are based on Scorecard data, which suppresses information for programs with fewer than 16 matched graduates over 2 cohorts at the 4-digit cipcode level due to privacy concerns.
What do we know about your program?
I’ve put together a dataset here where you can look up the numbers and make your own prediction for your program. Use the AHEAD/Scorecard Earnings data tab to see if the Department of Education has been able to estimate median earnings for your program. Use the filters on the top row (the little upside down triangles) to find your school and program.
It’s very likely that the Department of Education has not published an estimate of median earnings for your program (see Big Flashing Reminder above). If that’s the case for you, use the filters to see what the median earnings for similar programs at similar schools are like. Then use the IPEDS Completer data tab (which does include information for all programs) to find your school and program to guess how many matched graduates your program might have in recent cohorts. Remember that match rates vary with the fraction of students who use federal funding, so multiply the total number of graduates in recent cohorts by .25-.6 to get an estimate of the number of matched graduates per cohort.3 Remember also that programs will be aggregated first across years and then together with similar programs in their cipcode to try to get to 30 matched graduates to put in a test bucket. My last post explains all the weeds-y details.
Why not all graduates? Turns out the Education Department is banned by law from trying to measure earnings for everyone. I know! I was surprised too! See my previous post for details.
For my super-nerds: This earnings distribution is log-normal, median $40k, sigma=0.4585; sigma was calibrated to match Scorecard national earnings quartiles.
Other simulation details: A program "fails" in a given year if its measured median graduate earnings is at or below the national high-school-grad median (the benchmark); "Any failure" = fail at least 1 of the 5 modeled tests (failing even once triggers required notification of current and prospective students); Lose loan eligibility ("sanction") = fail 2 of any 3 consecutive annual tests. Standardized-distance Monte Carlo, 100,000 runs per cohort size, deterministic (fixed seeds). The initial coding of the simulation was done with Claude Fable 5. Later tune-ups were done with Opus 4.8. Let me know if you want other details! I’m not a simulation expert, but I did co-author a paper with simulations like these a few years ago, so I’m not a total newb. Also, I’m going to brag a bit: our article is one of AJAE’s top 10 cited papers published in 2024! I really like that paper.
Match rates vary both by school and by program, so if you really want to get fancy, you could use the match rates shown in the Earnings data tab by program and school to better estimate the match rate for your specific program at your specific school.



Interesting post! Do you think this will lead to universities combining more programs to create bigger alumni pools or breaking apart programs to make the sample size smaller? Like in English, instead of having Creative Writing and Poetry and English as separate majors - would they combine them? Or would it be in their best interests to separate them out more?
This is so fascinating! Something that I noticed in the programs with verdict = fail data is that so many of them are at religious universities. Two of the five failing programs with the most students are at BYU campuses, in the "Human Development, Family Studies, and Related Services" category. I wonder if this is because these are majors dominated by women in a religion where young marriage + having the husband as the primary earner is idealized, so many of the graduates are not working (and not looking for work) a few years after graduation. Along a similar theme, there are several programs in Rabbinical colleges with large numbers of students (like possibly more than half of their undergraduate population--seems like a crisis!) who may lose funding; presumably, that is because after graduation, they are employed in religious vocations that do not formally pay.
It indicates to me that there is a contingent of American society that sees the purpose of a college education as being orthogonal to individually improving one's earning or career prospects, and uses federal student loans to fund a completely different end. Interesting!