Back to Knowledge hub

25 Aug 2026

Supplier scorecards can change behaviour. But only when the score matters.

Almost 9,000 audits of Rome's electricity grid show what happens when a supplier score moves from reporting into decision-making: compliance jumped from around 25% to more than 80%, blackouts got shorter, and prices did not rise. The lesson for procurement teams is not a better scorecard.

Rome illuminated at dusk: St Peter's Basilica and a lit bridge reflected in the Tiber

By Matthew Spencer, Co-founder and CEO of FlockScore

What almost 9,000 audits of Rome's electricity grid tell us about supplier performance.

Most of the criticism aimed at supplier scorecards is fair.

They tell you what has already gone wrong. By the time a delivery miss reaches a quarterly review, the line may already have stopped. Some of the criteria are subjective, the data often comes from the same organisation that selected and manages the supplier, and the finished deck can easily become an expensive PDF that nobody uses.

An article in Logistics Viewpoints made a similar case earlier this year: a lagging measurement tool is not the same thing as active supplier management. I agree with that. But it leaves a more useful question unanswered. Under what conditions does a supplier score actually change behaviour?

An unusual case from Rome gives us one of the clearest answers I have seen.

In short

  • Supplier compliance rose from around 25% to more than 80% before the new scoring system was even used to award a contract. Suppliers changed because they knew future business would depend partly on their performance.
  • The improvement was visible in the electricity network, not only in audit scores, and there was no clear increase in the prices Acea paid.
  • The system worked well for suppliers Acea already knew. For new suppliers, Acea still had the same problem procurement teams face today: no performance history of its own.

What Acea changed

Acea operated Rome's electricity network, serving roughly 1.6 million power customers. In 2015 it spent around $200m on grid works covering cables, transformers, poles and substations.

These contracts had traditionally been awarded through price-only auctions. Acea did inspect work sites, but the old process relied on written reports and contractual penalties were rarely enforced.

In October 2007, Acea introduced random, structured work-site audits. Inspectors used a fixed list of 136 quality and safety checks across 12 categories. Two months later, Acea told contractors that these results would form a reputation index and eventually count for 25% of the score in future auctions. Price would count for the remaining 75%.

The price-only auctions continued for another 37 months while suppliers built up a performance history. Acea explained the system, showed suppliers how a stronger rating would affect the result and gave each firm visibility of its own score.

During that period, compliance rose from around 25% to more than 80%. It later reached roughly 90% and remained high. Suppliers improved most on the measures carrying the greatest weight, which is a good indication that they were responding to the incentive rather than simply getting better at being audited.

The study covers 8,973 audits, 634 contracts and 84 contractors between 2007 and 2017. That is a far stronger evidence base than the usual supplier-management case study.

Did the work actually improve?

A higher audit score is not enough. Suppliers can learn how to pass an audit without the underlying result getting much better.

The researchers therefore checked electricity service data reported to Italy's regulator. Compared with other large Italian distributors, Acea's unplanned interruptions became shorter and less frequent after the reform. The estimated annual duration of longer interruptions fell by about 43 minutes per customer. Acea's water network, which was not part of the reform, showed no similar improvement.

Acea also improved its audit process at the same time, so we cannot put all of the improvement down to the scoring mechanism. Even so, the overall pattern is convincing: suppliers responded to the measures that mattered most, and the improvement appeared in an operating outcome outside the scorecard itself.

That distinction is important to how we think about FlockScore too. A score should not be treated as the answer on its own. The value comes from the evidence behind it and whether it helps someone make a better sourcing, allocation or supplier-management decision.

Did better performance cost more?

This is where many internal business cases become difficult. If quality carries more weight, surely the buyer has to pay for it.

That did not happen here. Winning discounts increased when the scoring auctions began, and the researchers found no visible increase in the prices paid by Acea.

They estimated the reduction in blackouts was worth around €6.6m a year, with a further €3.5m to €5.3m in estimated safety benefits. I would not present that as a clean 5% procurement ROI: part of the benefit is modelled and the calculation does not subtract the cost of running the system. But the more important conclusion holds. Acea achieved a major improvement in supplier performance without evidence that it had to pay higher contract prices for it.

The scorecard was not the point

The obvious response would be to copy Acea's methodology or build a more detailed scorecard. I think that misses the lesson.

Acea already inspected its suppliers. What changed was that suppliers could see how performance was being measured and knew it would affect future business. The score moved from reporting into decision-making.

That consequence does not need to be a fixed 25% in every sourcing event. It could affect share of business, preferred-supplier status, access to new opportunities, development support or escalation. The important part is that good and poor performance lead to something different.

There is a warning in the case too. The scoring auctions lasted only a few months before supplier complaints and legal objections caused Acea to shelve them. Performance nevertheless remained high under a later system that returned to price-based awards but kept rigorous audits.

If a score affects revenue, the supplier needs to understand the method, see the evidence and have a reasonable route to challenge mistakes. That is not extra governance around the edges. It is part of making the score usable.

What procurement teams should take from it

  1. Put the score into a real decision. If it has no effect on allocation, status, escalation or future sourcing, suppliers will quite reasonably treat it as reporting.
  2. Prioritise evidence over false precision. Hard quality and delivery measures transfer more easily than a decimal score for "collaboration". Softer information can still be useful, but it needs context.
  3. Check whether outcomes improved. A cleaner scorecard is not the objective. Better delivery, quality, safety and continuity are.
  4. Design for suppliers you do not know yet. Past performance is most valuable before an award, which is also when an individual buyer is least likely to have it.

The gap an internal scorecard cannot close

Acea could use past performance because it already had experience of the contractors doing the work. For new or lightly audited firms, it assigned the average bidder score until it had completed enough audits. That avoided unfairly penalising a new entrant, but it did not tell Acea how the supplier was likely to perform.

Industrial sourcing still works in much the same way. A manufacturer may have very good data on its current suppliers. When it considers a supplier for the first time, however, the most relevant performance evidence usually sits inside other customers' systems. The buyer falls back on references, certificates, supplier documents and its own assessments.

This is the gap FlockScore is built around: bringing anonymised, real-world supplier performance evidence from across companies into sourcing decisions earlier. It does not replace the buyer's own assessment or remove the need for context. It gives the buyer relevant evidence before it has years of direct experience of its own.

The Rome study does not test shared supplier intelligence. What it does show is that past performance can be valuable enough to change supplier behaviour and improve operational outcomes. It also shows exactly why that information is missing at the moment it could be most useful.

A better internal scorecard cannot solve that on its own.


Sources