We could prove the h-index was unfair. That was its one virtue.
UK Research and Innovation tells assessors they will never be asked to score the individual modules of a narrative CV. That is a deliberate and defensible design choice, and it also removes the audit trail that made the case against the h-index possible in the first place. Research assessment has moved from a bad number to a judgment nobody has yet measured at scale, and the first careful measurements, presented in late 2025, are ambiguous rather than reassuring.
UK Research and Innovation's guidance for the Resume for Research and Innovation contains a sentence that is easy to read past. Assessors use the document to inform the assessment of the application overall, it says, and they will not view it in isolation. Then: "They will never be asked to score individual modules."
That is a deliberate choice and, on its own terms, a good one. Scoring the four modules separately would turn them into a scoreboard within a year, and a scoreboard is the thing the narrative CV was invented to escape. But it has a second consequence, less discussed. No separable number is produced at any point in the process. There is nothing to correlate, nothing to disaggregate, nothing to hold up against an applicant's gender or first language or institution and ask whether the ranking tracked what it claimed to track.
The case against the metrics we are replacing was made with metrics. It is worth sitting with that for a moment before deciding it does not matter.
The recantation, and what made it possible
Jorge Hirsch proposed the h-index in 2005 as a single number summarising an individual's research output. Fifteen years later, in an essay in Physics and Society, the American Physical Society forum newsletter that states plainly that its articles are not peer reviewed, he wrote:
I proposed the H-index hoping it would be an objective measure of scientific achievement. By and large, I think this is believed to be the case. But I have now come to believe that it can also fail spectacularly and have severe unintended negative consequences.
Hirsch's own complaint was about intellectual conformity, that a high h-index confers an authority which discourages the naive question. The wider objection was different and more measurable. The h-index tracks field size, citation culture, career length and self-citation habits at least as strongly as it tracks contribution. Everyone who wanted to argue that could argue it, because the number sat there in public and could be recomputed by anyone, for anyone, and cross-tabulated against whatever you suspected it was really measuring. Its legibility is what made its unfairness demonstrable.
The reform that followed moved quickly. The Royal Society's Resume for Researchers gave the format its shape. UKRI committed to a version of it in April 2021 and mandated the resulting R4RI across most of its funding opportunities over the following three years. The Swiss National Science Foundation, the Dutch Research Council, Science Foundation Ireland and, from an October 2024 announcement, Canada's three federal granting agencies have each moved in the same direction.
A scoping review published in F1000Research in December 2025 by Albert and colleagues at the University of Ottawa Heart Institute went looking for every narrative CV template a research funder provides. It found that the number of funders actually requiring one remains low, that some provide no formal completion guidance at all, and, in the line that matters most here, that among the templates identified "there is little insight into the evaluation of NCVs". The review carries F1000Research's open peer review status of one approved and two approved with reservations, which is worth stating rather than rounding up.
So: the format has been mandated across a national funding system for four years, and the question of how it is actually evaluated is still open. That ordering is the argument of this piece.
The judgment we already had, measured
It helps to know what the baseline is, and for grant review we do know, because someone built the apparatus to find out.
In 2018 Elizabeth Pier and colleagues at Wisconsin reconstructed the NIH peer review process outside the NIH. They recruited 43 oncology researchers into four simulated study sections and gave them 25 real R01 applications, sixteen of which the NIH had funded first time and nine of which had been funded only after resubmission. Every application was rated by between two and four reviewers, producing 83 primary reviews under the NIH's own instructions and its nine-point scale.
The intraclass correlation for the ratings was 0. Krippendorff's alpha was 0.024. Reviewers did not agree on the ratings, did not agree on the number of strengths and weaknesses, and did not even agree on how a given count of strengths and weaknesses should translate into a score. As the authors put it, multiple ratings of the same application were about as similar as ratings of different applications. Which reviewer you drew mattered more than which proposal you wrote.
That finding is not an argument for metrics. It is an argument that expert human judgment of research merit, applied under structured criteria by people with relevant expertise, was already producing something close to noise, and that we only know this because a research team spent years building a replica in order to measure it. No equivalent measurement exists yet for narrative CVs at national scale.
The first measurements, and they are not reassuring
They are, however, starting. In November 2025 DORA convened four research groups working on narrative CV implementation, and the summary it published is the most useful single snapshot currently available.
Hannah Frith presented the Breaking Barriers project, which ran from August 2024 to August 2025 and combined interviews with twenty researchers from under-represented groups, a linguistic analysis of 27 UKRI-templated narrative CVs using BERT and LIWC22, and evaluation of those CVs by external reviewers. Two findings sit together uncomfortably. Some minority groups were less likely to use agentic language and language conveying certainty. Reviewers rated applications with less agentic and less certain language less positively on expertise, on success and on leadership.
If that holds, the format has a sorting rule, and the rule is how confidently you can write about yourself in English. Applicants reported this directly, describing the burden of adopting an unfamiliar register and the anxiety that came with it, most acutely those for whom English is not a first language and those who are disabled or neurodivergent.
Luisa Ciampi presented preliminary results from a randomised controlled trial at Cambridge comparing standard and narrative CVs in postdoctoral recruitment. The quantitative signal so far is that the move to narrative CVs benefited men, applicants from the Global South, and applicants self-identifying as white. That is a genuinely mixed result and should be reported as one. It does not support a simple story in either direction, and the trial is not finished.
Karen Stroobants summarised a review funded by the UK's Department for Science, Innovation and Technology and UKRI's metascience fund, and produced the sharpest formulation of the problem. Introducing a narrative CV format achieves the objective of recognising broader contributions. As a standalone intervention it does not guarantee that those contributions are valued. The primary barrier she identified was the lack of support and training for the reviewers, who struggled particularly when publication lists were discouraged.
All of this is preliminary, presented at a meeting, and in most cases not yet through peer review. That is precisely the point. It is the earliest serious attempt to measure a change that has already been made.
The strongest case against this argument
Reliability is not validity, and demanding the first can be a way of protecting a measure that lacks the second.
The h-index is highly reproducible. Two people computing it for the same researcher will agree. It is also, on the strongest reading, a reliable measure of something other than contribution, and the reason we can criticise it so precisely is not a virtue of the metric but an accident of arithmetic. Meanwhile the things narrative CVs were built to surface, mentorship, data stewardship, teaching, building infrastructure other people use, sustaining research in an under-resourced setting, are not badly measured by citation counts. They are structurally invisible to them. A noisy measure of the right thing may well beat a precise measure of the wrong one.
Nor is inter-rater reliability obviously the right standard. We do not demand it of job interviews, doctoral vivas, or editorial decisions, all of which we accept as judgment rather than measurement. UKRI's refusal to score individual modules is not evasion. It is a defence against exactly the process that ruined publication counts, where a proxy adopted for convenience became the target and then became the goal.
I think that case is largely right, and it is why the conclusion here is not that funders should go back to counting papers.
Where it lands
A claim to validity is still a claim. Stroobants' distinction between recognising a contribution and valuing it is the crux: recognition is a property of the form, and the form has changed. Valuing is a property of the panel, and the reform mostly did not fund the reviewer training, the guidance, or the measurement that would change that. The Breaking Barriers linguistic analysis found what it found because someone deliberately built an instrument to look. Nothing about the narrative CV surfaces that pattern on its own, and nothing in the process produces a number that would make it visible by accident.
There is a further wrinkle, and I will label it as inference rather than finding, because no study I am aware of has yet tested it. If reviewers reward agentic, confident prose, and a language model will produce agentic, confident prose on request, then the trait being rewarded has just become cheap to manufacture. Whether that equalises anything depends entirely on who feels licensed to use the tool and who has been advised not to, which is not a distribution anyone is currently tracking. This is the same shape of failure that overtook publication counts, running on a much shorter clock.
The lesson is not that research assessment reform was a mistake. It is that reform of this kind is a move from one form of judgment to another, not from judgment to its absence and not from bias to its absence. What made the old form criticisable was that it was written down in a shape anyone could interrogate. If the new form is to be better rather than merely newer, the instrumentation has to be part of the reform rather than something a handful of research groups start assembling five years after the mandate.
Bizarus has argued something adjacent before, in a piece on how peer review learned to detect AI-written reviews with real statistical rigour while never measuring review quality at all. The pattern repeats because it is structural. It is far easier to change what a process asks for than to find out what the process is doing.