A better-information benchmark for AI

People are increasingly turning to AI chatbots, based on large language models (LLMs), for information that shapes their lives. For example, Google says its AI Overviews feature, now the default in most Google Searches, has 2.5 billion users a month. As these tools become embedded in everyday life, the quality and reliability of their outputs become a matter of public interest on par with the accuracy of broadcast media or government statistics.

Earlier this year we conducted a trial where these models made dozens of major errors when assessing false claims.

More needs to be done to understand and address the potential harms caused by AI-generated misinformation so that we can make informed decisions about when to rely on AI, and public trust can be grounded on evidence rather than marketing claims.

What are we doing?

Full Fact is building a publicly available, auditable benchmark to independently evaluate the output of leading AI models.

We are asking these language models the same set of questions every day and recording their responses. We will annotate and analyse these responses and benchmark their performance against these five key dimensions:

  • Factuality. Do responses contain verifiably accurate claims, avoid hallucination and correctly represent the evidence base?
  • Transparency. Does the language model communicate uncertainty, cite or attribute high quality sources, acknowledge limitations, and distinguish fact from opinion?
  • Timeliness. Do responses reflect current information rather than outdated data, and does the language model recognise when its knowledge may be stale?
  • Consistency. Does the language model give materially the same answer to the same question over time and when asked in different ways?
  • Civic responsibility. Is information about democratically important questions balanced, does it not amplify misinformation, and does it support informed participation?

The benchmark will apply the same rigorous editorial standards that Full Fact uses when assessing claims made by public figures and reported in the media. By systematically testing these models over time, we will expose any inconsistencies in their responses, and provide the general public with an independent assessment on which AI tools they can trust.

Transparency is at the heart of the project. The results will be published in full, to be used by both individuals who use AI to source information in their everyday lives, and institutions and professionals who hold tech companies accountable for producing good information.

The first report on this project will be released in mid-October.

Watch this space.