The Office for Students (OfS) has begun setting out the future of the Teaching Excellence Framework (TEF), the system used to assess teaching quality in higher education in England. Some of the most controversial elements of its original proposals have been revised following consultation. That is good news. But the bigger question remains: how do we measure teaching quality accurately?
Quality regulation does much more than produce ratings. It shapes where students choose to study, how universities invest in teaching and support, and which institutions face regulatory intervention. If the data behind these decisions are flawed, the incentives created by regulation can push institutions in the wrong direction rather than encouraging genuine improvement.
What do existing measures of quality tell us?
One of the most debated proposals is the use of graduate earnings data from the Longitudinal Education Outcomes (LEO) dataset. At first sight, higher earnings look like a sensible indicator of educational success. Students and policy-makers understandably care about employment and pay after graduation.
The problem is that earnings do not measure teaching quality directly. Graduate salaries reflect many other factors, including prior attainment, family background, subject studied, gender and the state of the labour market.
Research from the Institute for Fiscal Studies (IFS) shows that these factors continue to influence earnings long after graduation (Britton et al, 2020). In other words, universities can appear to be highly successful because of whom they recruit rather than how well they teach.
This is why benchmarking matters. More recent IFS work argues that any earnings measure used for regulation should adjust carefully for students’ characteristics and prior attainment (Britton et al, 2025). Without that adjustment, institutions serving more disadvantaged students could be unfairly penalised, while those with more advantaged intakes could receive credit for outcomes that are not a consequence of their teaching.
There is also a timing issue. Earnings data arrive years after students graduate. A university that improves its teaching today may not see the results reflected in earnings measures until the next decade. That makes earnings useful as one source of evidence, but less useful as a tool for encouraging immediate improvement.
The OfS has recognised many of these challenges and committed to careful benchmarking and contextual assessment. That represents a significant step forward.
A second challenge concerns coverage. The LEO dataset does not fully capture international students or many graduates who attended school outside England. For institutions with large international student populations, this means that earnings measures may reflect only part of their graduate community. The resulting picture of institutional performance may therefore be incomplete.
Should we be cautious about student satisfaction scores?
The second major source of evidence in TEF comes from the National Student Survey (NSS). Students’ feedback is valuable and provides important information about their experience of higher education. But research increasingly shows that satisfaction scores can be influenced by factors unrelated to teaching quality (Bell and Brooks, 2018).
Studies from several countries show that women academics often receive lower teaching evaluations than male colleagues even when students achieve similar outcomes (Boring, 2017; Mengel et al, 2019; Ayllón et al, 2026). Other research finds evidence that ethnicity and other demographic characteristics can influence ratings (Heffernan, 2021). These differences appear even when teaching quality is held constant.
This matters because universities employ different academic workforces. If satisfaction scores are influenced by staff characteristics, institutions with more diverse teaching staff could be disadvantaged through no fault of their own. Benchmarking student characteristics alone may not fully address this issue.
There is also evidence that some subjects receive systematically lower satisfaction scores (Ayllón et al, 2026). Disciplines with a strong analytical component, including economics, law and engineering, can be more demanding for students. Lower satisfaction may therefore reflect challenge and academic rigour rather than weaker teaching.
The key lesson is not that student surveys should be abandoned. Rather, satisfaction scores should be treated as measures of perceptions rather than neutral indicators of quality. They provide useful evidence, but they cannot be assumed to tell the whole story.
What should happen next?
The good news is that the OfS seems to be listening to the experts. One example is the decision to abandon the proposal that a provider’s overall TEF rating should automatically be determined by its lowest score in any category.
Critics argued that this would exaggerate small measurement errors and obscure important differences between institutions. The OfS went further and removed the overall rating altogether, instead reporting separate assessments for student experience and student outcomes.
This approach gives students more information rather than less. Universities often have different strengths and weaknesses, and separate ratings provide a more accurate picture than a single headline score. At the same time, the concern that the lowest of the ratings will be used to penalise universities remains.
There is also an example elsewhere in the regulatory system that points towards a more evidence-based model. Access and participation plans increasingly require universities to evaluate their interventions using robust methods and to build evidence about what works, though this raises questions about capacity-building (Moores et al, 2023).
The wider lesson is that quality indicators are not the same as quality itself. Earnings data, student surveys and continuation rates are all proxies. They can be informative when used carefully, but they have limitations that policy-makers must recognise.
A stronger system would continue to refine earnings benchmarking, explore whether workforce characteristics should play a role in NSS adjustments, place greater emphasis on contextual evidence and remain realistic about what available data can and cannot measure.
Higher education regulation may sound technical, but it is ultimately about protecting students. Decisions based on imperfect measures affect millions of pounds of investment, the reputation of institutions and the opportunities available to future graduates.
The OfS has shown a willingness to listen to evidence during the first stage of reform. The next stage should continue that approach by drawing on the growing body of research about what quality indicators reveal, what they miss and how they can be improved.




