Rating method

The HVAC Bench Score

A rating for an equipment family is only worth reading if you know what went into it. This page describes exactly that, including the conditions under which we publish no number at all.

8 families tracked0 currently scored5 weighting profiles

Three kinds of evidence, kept apart

Most equipment ratings online blend two things that should never be mixed: what a manufacturer says a machine does, and what the people who bought it found. The first is documentation and the second is testimony, and treating them as one number is how a specification sheet ends up looking like a reliability record.

  1. Manufacturer facts. Model scope, features, and documented behaviour, taken from service and operation manuals. These establish what a system is built to do. They are never used to score how it was found in service, because a manufacturer is not a witness to that.
  2. Owner experience. First-hand accounts from identifiable owners, documented field experience, and independent testing. Each observation is stored with where it was read, when it was collected, which model family it concerns, and what it actually said. This is the only class that can move a score.
  3. The calculated score. Derived from the observations above and the published weights for that equipment class. There is no field anywhere in the system that lets anyone type a score in by hand.

When a family gets a number

A score publishes only when the evidence behind it clears every one of these. A near miss is still a miss, and the card says so instead of showing a hedged figure.

  • At least 12 distinct owner observations about that family, after duplicates are removed.
  • Observations from at least 3 independent sources, so one forum thread or one retailer page cannot carry a rating.
  • At least 5 observations from 2 sources before any single dimension is scored. Dimensions nobody spoke to stay blank rather than being filled in.
  • Enough scored dimensions to cover at least 50% of what matters for that equipment class.

Rules the system enforces on itself

  • An observation about one family is not evidence about another, even within a brand. A brand-wide impression never substitutes for a specific family rating.
  • Syndicated reviews are collapsed to one observation. The same retailer text appears across several sites, and counting it repeatedly would manufacture a sample.
  • An observation votes only on the dimensions it actually addresses. Somebody complaining about noise does not get a reliability vote.
  • Sample size and confidence are shown next to the number. Confidence describes how much evidence stands behind a score, not whether the score is good.

Weighting by equipment class

What matters depends on the machine. Serviceability weighs heavily on light commercial equipment because downtime costs money; heating output weighs more on a heat pump than on a cooling-only split, which is not offered a heating score at all. The weights below are configuration rather than code, and revising them changes every card of that class without touching a single page.

Ductless mini-split (heating and cooling)
Reliability
25%
Cooling performance
15%
Heating performance
12%
Noise
10%
Efficiency and value
10%
Build quality
10%
Serviceability and parts
7%
Controls and usability
6%
Owner satisfaction
5%

Reliability carries the most weight because a ductless failure removes heating or cooling from a single room with no backup, and because owner reports are more consistent about failures than about capacity.

Cooling-only split system
Reliability
28%
Cooling performance
20%
Noise
12%
Efficiency and value
12%
Build quality
10%
Controls and usability
8%
Serviceability and parts
7%
Owner satisfaction
3%

Heating performance is not offered at all on this class. A cooling-only system cannot be scored on something it does not do, and an absent dimension is clearer than a dimension scored at zero.

Air source heat pump
Reliability
24%
Heating performance
20%
Efficiency and value
14%
Cooling performance
10%
Noise
10%
Build quality
8%
Serviceability and parts
8%
Controls and usability
6%

Heating carries more weight than cooling because a heat pump is usually the primary heating system, and cold weather output is where owner experience varies most. Owner satisfaction is omitted because it duplicates the dimensions above rather than adding to them.

Light commercial ducted and cassette
Reliability
30%
Serviceability and parts
18%
Cooling performance
14%
Heating performance
12%
Efficiency and value
12%
Controls and usability
8%
Noise
6%

Serviceability weighs far more heavily here than on residential equipment, because downtime has a commercial cost and parts availability drives it. Build quality and owner satisfaction are dropped: the people reporting are rarely the people who bought it.

Thermostats and controllers
Controls and usability
40%
Reliability
30%
Build quality
15%
Owner satisfaction
15%

A controller has no airflow, no capacity, and no efficiency of its own, so those dimensions are not offered. Usability leads because that is what a controller is for.

Families being tracked

Each of these has its model scope established from manufacturer documentation, which is what a score would be attached to. Where a family shows no score, the owner evidence has not reached the threshold above.

No family currently carries a score. The owner evidence record holds 0 verified observations, which is short of what any family needs. That is a statement about our evidence, not about the equipment.