WattPair Appliance Load & Surge Provenance Dataset
38 appliance classes. Running watts, start-up surge, duty cycle and annual energy — and attached to every single figure, the document it was read from, the date it was read, and how much we trust it. Free to reuse under CC BY 4.0.
Read this before you use a number
254 of the
524 figures in this file are
null — 48%. That is not an oversight and it is
not an unfinished import. It is the design: where we opened the documentation and found no figure, the
field stays empty and keeps a status of unknown, because an
estimate dressed as a specification is the failure mode this dataset was built against.
Anyone can produce a wattage chart. Very few say which document each row came from, or admits which rows do not exist. That is the part worth downloading.
No sign-up, no email, no rate limit. Both are static files on this domain. The CSV is the same data one row per figure — appliance, field, value and the whole provenance of that one number on a single line — for reading in a spreadsheet. The JSON is the canonical form.
What is in the file
The last row is not a defect list, it is data. Those records carry a
ready_for_page_generation: false and a count of the reasons; they
stay in the file so a reuser can see which classes this project does not consider settled
rather than discovering the gap later. Two engine constant tables ship alongside the records —
surge duration bands and per-load-type power-factor defaults — because a duration in
milliseconds cannot be used correctly without knowing which band it falls into.
Where each start-up figure comes from
Start-up surge is the number that decides whether a machine can be run at all, and it is the number the rest of the web quotes without saying where it got it. Each of the 38 records is labelled with one of four bases.
nameplate_lra 7 Locked-rotor amperage from a compressor or motor data sheet. The strongest evidence in the file: a measured stall current, published by the part’s own manufacturer.
manufacturer_stated 5 A start-up figure the appliance brand published itself, in watts or as a multiple.
category_rule_of_thumb 3 A category envelope this project applied because we have found no published figure for the class. Labelled as such on every page that uses it, and it is not a manufacturer number.
unknown 23 No start-up figure, and the field is null. We have found no published start-up figure for these classes. The field stays null rather than being filled with a category rule of thumb.
Every field, and what it licenses you to say
Each figure in the file is an object, not a bare number. These are its keys.
value The figure. Null when no source we read carried it — never an estimate standing in for one.
unit W, VA, ms, ratio, h, kWh, Wh. Stated per field rather than assumed from the field name.
value_min Lower end of the range, where the class genuinely spans one (model to model).
value_max Upper end of the same range.
display_precision Decimal places this figure supports. A figure read to 1 dp is not a figure known to 3.
test_conditions What was measured, under what load, at what voltage — and where the source states it, the statement quoted word for word. This is the field that decides whether a figure is usable: 7 W and 7 W with the humidifier off are different claims.
source_tier official · measured · user. Everything in this release is official: this project has no laboratory and says so.
source_type What kind of document it was — specification sheet, manual PDF, government standard, support article, product page.
source_url The document itself. This is the field that makes every other one checkable.
source_label What was read out of that document for this figure. Usually the document’s name; narrower where one manual covers several model numbers.
verified_date When a human last opened that document and read this figure out of it.
confidence high · medium · low. Low is common and is not an apology — a derived figure or a borrowed coefficient says so here.
status verified · unknown · modelled · conflict. Unknown means no reliable source was found, and such a field is always empty — an estimate is never written into one. Modelled means the opposite kind of gap: the figure is a bound this project constructed rather than anything a document states, so it carries a value but never a source URL or a source tier, and its confidence is capped at low. A status of conflict means two sources disagree and the disagreement is unresolved, so the figure is not used.
absence_verified The negative record: on this date, these documents were opened and this figure was not in any of them. Distinguishes “we did not look” from “we looked, it is not there”.
is_derived True when the figure was computed from other published figures rather than read directly (e.g. annual kWh ÷ hours ÷ duty cycle).
Each record also carries a sources array — the documents behind
it, with source_url, source_name, source_tier, source_type, covers_fields, verified_date, fetch_status.
Which fields in this record that document backs — so one source cannot silently be credited for the whole record.
What a source tier means
Three tiers exist in this project's schema. Only one of them appears in this release, and that is worth saying plainly rather than leaving to be inferred from the data.
official Read out of a document the manufacturer, a government programme or a standards body published. Every figure in this release is this tier.
measured Measured on a bench by this project. There are none. WattPair has no laboratory, and rather than blur the distinction the tier stays empty until that changes.
user Submitted by a reader who measured it. Held in a review pool; a pairing is only labelled user-verified once three independent logs agree.
official sources
203 Licence and how to credit it
CC BY 4.0. Use it commercially, modify it, build a product on it. The one condition is attribution, and the licence lets us specify the wording — so here it is, ready to copy:
WattPair Appliance Load & Surge Provenance Dataset
© WattPair — CC BY 4.0
Source: https://wattpair.com/data/appliance-load/
The same three lines are embedded inside the JSON, under
dataset.attribution, because people who copy data copy
fields and not documentation. One URL, and it is this page.
A note on what a licence can and cannot do, stated because pretending otherwise would be the dishonest move: under US law facts are not copyrightable, and every figure here is a fact read from a public document. Anyone who retypes these numbers owes us nothing, licence or no licence. Attribution is a social norm, not a lever — we are asking, and most people who write for a living say yes. This is not legal advice.
How it was compiled
Compiled with AI assistance from manufacturer documents, government efficiency filings and standards bodies, then checked against the cited source. Every non-null figure carries the URL it was read from, so any claim here can be re-derived from public documents without trusting this file.
This dataset is deliberately incomplete and says where. For several appliance classes we have found no published start-up surge figure; those fields are null with a status of unknown rather than estimated, and the site makes no claim that the figure exists nowhere. Where a null field carries an absence_verified block, that block names the documents that were read and the date they were read on. Read stats.null_fields, stats.absence_verified_fields and stats.surge_basis before using any figure.
What is deliberately not in it
Not published. That library is what this site sells against, and opening it would be giving away the commercial core rather than the research.
Not published — though its thresholds are, in full, on the methodology page. The constants were never the secret; the constraint ordering is.
Where two documents disagree, the field is marked conflict and
the figure is withheld — but the internal queue of unresolved disagreements, with its
quoted excerpts, stays internal.
About half the figures carry a working note recording how the figure was corrected and what was rejected on the way. Those notes stay internal and stay in the language they were written in: they are a record of this project's process, not a fact about the appliance.
The research behind this dataset was written in Chinese. Version 0.1.0 shipped the machine-readable provenance only and named the free-text fields it was holding back rather than omitting them quietly; v0.2.0 adds them, translated in place — the per-field test conditions, the source label on every figure, the name of every source document, the description of every appliance class, and the reason each held-back record is held back.