GoodwillyGiving, lending, passing on

How Charities Work

Why Charity Rating Sites Disagree With Each Other

Two evaluators can look at the same organisation and reach opposite conclusions. The disagreement is usually about what they decided to measure.

A man and woman at a boutique counter discussing a clothing purchase.
Photograph by MART PRODUCTION via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

The points below about charity evaluation and ratings are ordered by how much difference they make, not by how often they get repeated.

What matters most

  • Financial ratings and impact evaluation measure different things.
  • Small organisations are systematically disadvantaged by rating criteria.
  • Coverage is partial, so absence from a list means little.

Two different questions

One family of evaluators assesses financial health, governance and transparency, all of which can be read from published documents. Another assesses whether the work achieves anything, which requires evidence about outcomes that most organisations do not generate. These produce different answers because a well-governed organisation can do ineffective work and a chaotic one can do valuable work.

Confusing the two is the single most common misreading of any rating, and the sites themselves generally state the distinction clearly. Knowing which question a rating answers tells you what it can and cannot be used for.

Why financial ratings favour large organisations

Criteria involving audited accounts, published policies, formal websites and detailed reporting all require administrative capacity. A small organisation doing excellent work with three staff will score poorly against those criteria while being perfectly sound. Ratings that penalise high overhead ratios compound the effect, since fixed costs are a larger share of a small budget.

On the shop floor, the result is a systematic tilt towards established organisations, which the more careful evaluators acknowledge openly. For local giving, published ratings are usually irrelevant and direct enquiry is more informative.

Why impact evaluation covers so little

Rigorous outcome evidence exists for a narrow set of interventions, mostly in health and economic development where trials are feasible. Evaluators focusing on measurable impact therefore recommend a short list, and that shortness is a feature of the method rather than a claim about everything else.

Whole categories, including advocacy, arts, heritage and rights work, resist this kind of measurement almost entirely. Absence from such a list is not evidence of ineffectiveness; it usually means the work is not amenable to the evaluation approach used. Reading the methodology page is the fastest way to understand what a list is actually claiming.

Data quality and lag

Most ratings are built on filed accounts, which are historical by the time they are published and can be a year or more out of date. Organisations change quickly, particularly small ones, so a rating can describe a situation that no longer exists.

Self-reported information introduces a different problem, since organisations with capacity to fill in surveys are not a random sample. Coverage varies enormously by country, and many jurisdictions have no rating infrastructure at all. A missing rating usually means nobody looked rather than that somebody looked and disapproved.

Where ratings are genuinely useful

They are good at flagging organisations that fail basic transparency, file late, or have unexplained governance gaps. They are good at surfacing organisations you had not heard of, which is a real service given how much giving follows advertising. For the specific domains where outcome evidence exists, the recommendations rest on substantially better evidence than most donors could assemble.

Bought used, they provide a common vocabulary for comparison, which improves donor conversations even when the ratings themselves are contested. Used as a filter rather than a verdict, they save time without deciding anything.

Assembling your own view

Start with the regulator's register for legal status and filing history, which is free and jurisdiction-specific. Read the accounts and the trustee report yourself, since fifteen pages tell you more than any summary score. For local organisations, ask people who use the service, which is available evidence nobody else is gathering.

Treat conflicting ratings as a prompt to read the methodologies rather than as a puzzle to be resolved by averaging. Accept that no available process will tell you with confidence whether a given organisation is effective.

Everything above, in order of what to do first

  1. Two different questions. One family of evaluators assesses financial health, governance and transparency, all of which can be read from published documents.
  2. Why financial ratings favour large organisations. Criteria involving audited accounts, published policies, formal websites and detailed reporting all require administrative capacity.
  3. Why impact evaluation covers so little. Rigorous outcome evidence exists for a narrow set of interventions, mostly in health and economic development where trials are feasible.
  4. Data quality and lag. Most ratings are built on filed accounts, which are historical by the time they are published and can be a year or more out of date.
  5. Where ratings are genuinely useful. They are good at flagging organisations that fail basic transparency, file late, or have unexplained governance gaps.
  6. Assembling your own view. Start with the regulator's register for legal status and filing history, which is free and jurisdiction-specific.

The takeaway

Find out which question the rating answers before you let it answer yours.

Unrestricted money is the most useful gift and the least satisfying to make.

Questions readers ask

Should I only give to top-rated charities?

Only if the rating measures something you care about, and most measure financial transparency rather than results. Coverage is also partial, so many good organisations are simply absent.

Are cost-per-outcome estimates reliable?

They are useful for comparing orders of magnitude and depend on contested assumptions at finer resolution. The better evaluators publish their uncertainty, and that is worth reading.

How Charities Workratingsevaluationmethodology
More in How Charities Work
Nirmala Saxena
Editor, Goodwilly

Nirmala edits Goodwilly and asks of every project whether it would survive without the grant.

Also by Nirmala Saxena