A reproducible pipeline that assembles a registry of commercial data centers
— colocation, wholesale and hyperscale — from public sources, and attaches
power estimates that carry an explicit evidence tier and confidence
interval. Originally built for California; now generalised to any US state
via --state.
Read this first. No figure in this dataset is a metered electricity reading. Per-facility consumption is confidential utility data and is not publicly available for any US data center. Every power number here is either a cited third-party claim or a model output, and each is labelled as such. See LIMITATIONS.md.
uv run data-center-dataset build --state TX # or GA, AZ, IL, VA, CA
uv run data-center-dataset build-all # every supported state, one after anotherEach state writes to its own data/processed/<STATE>/ directory. Two
enrichments are California-only because no free nationwide equivalent exists
(see states.py for the reasoning): CEC electric-utility territories, and the
CEQAnet attested-MW crawl. Every other source (OSM, PeeringDB, EPA NEI) is
queried per state.
| State | Facilities | With power est. | Sum IT MW | Bottom-up TWh/yr | vs. anchor |
|---|---|---|---|---|---|
| California (CA) | 221 | 118 | ~1,473 | ~13.3 | within [12, 26] |
| Texas (TX) | 271 | 185 | ~5,298 | ~46.8 | outside [12, 34] |
| Georgia (GA) | 70 | 34 | ~3,009 | ~26.2 | outside [4, 14] |
| Arizona (AZ) | 67 | 38 | ~1,303 | ~11.6 | within [4, 16] |
| Illinois (IL) | 100 | 52 | ~1,333 | ~12.0 | within [4, 15] |
| Virginia (VA) | 390 | 335 | ~9,225 | ~81.4 | outside [18, 48] |
The "outside anchor" cases are not treated as bugs to fix by tuning: the
anchors are crude population/market-share heuristics (see
states.STATES[...].annual_twh_anchor, documented as order-of-magnitude only),
and Texas, Georgia and Virginia specifically host some of the largest
hyperscale campuses in the country (Meta Newton County GA, Microsoft/Google/Meta
TX megasites, AWS's ~111 mapped Northern Virginia buildings). A bottom-up total
several times the anchor is at least as likely to mean the anchor is too low
as the model is too high — which is exactly why calibration is off by default
and, when enabled, never touches attested figures. See
LIMITATIONS.md for the full discussion.
The rest of this document describes the California build in depth, which received the most validation; the same pipeline and caveats apply to every other state.
| Metric | Value |
|---|---|
| Facilities | 221 |
| With building footprint geometry | 87 |
| With a power estimate | 118 |
| Corroborated by 2+ sources | 53 |
| Bottom-up IT load | ~1,473 MW |
| Bottom-up annual energy | ~13.3 TWh/yr |
Power evidence tiers: 45 facilities Tier B (generator permits), 73 Tier C (floor-area model), 0 Tier A (no curated citations yet — see Adding attested figures).
Utility attribution is one of the more analytically useful outputs, and it confirms a known feature of California's data center geography: a single municipal utility carries the largest concentration in the state.
| Utility | Sites | Sum IT MW |
|---|---|---|
| SVP (Silicon Valley Power, Santa Clara) | 72 | 777 |
| PG&E | 68 | 468 |
| SMUD | 8 | 132 |
| LADWP | 36 | 59 |
| SCE | 17 | 18 |
| SDG&E | 13 | 18 |
uv sync --extra dev
uv run data-center-dataset build # fetch, transform, export
uv run data-center-dataset validate # check schema contracts
uv run data-center-dataset report # summary tables
uv run pytest # 43 offline testsThe first build downloads a 157 MB EPA archive once; later runs reuse it. All
HTTP traffic is disk-cached and identifies itself as
data-center-dataset/0.1 (aerith.netzer@northwestern.edu).
Optional flags:
uv run data-center-dataset build --with-ceqanet # crawl CEQA filings (slow, low recall)
uv run data-center-dataset build --apply-calibration # rescale modelled tiers to the state anchor
uv run data-center-dataset build --refresh # bypass caches| Source | Contributes | Access | Licence |
|---|---|---|---|
| OpenStreetMap via Overpass | Location, building footprint polygons, operator, year built | Free API | ODbL-1.0 |
PeeringDB /api/fac |
Colocation registry, interconnection counts, CLLI | Free API | CC-BY-4.0 |
| EPA NEI 2020 Region 9 | Facility recall + backup generator inventory (Tier B) | Free bulk | US Public Domain |
| CEC GIS | Electric utility service territories | Free API | US Public Domain |
| CEQAnet | Attested project MW (opt-in) | Free scrape | US Public Domain |
Deliberately not used. Baxtel publishes per-facility MW but paywalls it
(isAccessibleForFree: false; values render as ░░░). datacenters.com and
datacentermap.com sit behind bot-protection walls. Neither was circumvented.
Reconnaissance found NEI to be the strongest free per-facility source, because every large data center holds an air permit for its backup diesel generators. It supplies sites that neither OSM nor PeeringDB records — Google Mountain View, Vantage Santa Clara, RagingWire Sacramento, AWS Santa Clara and Hayward, Microsoft Santa Clara — each with coordinates.
Written to data/processed/:
| File | Contents |
|---|---|
facilities.{csv,parquet} |
One row per resolved site, with the preferred power figure |
facilities.geojson |
Same, as points, for mapping |
power_estimates.{csv,parquet} |
Every estimate from every method, uncollapsed |
facility_sources.csv |
facility_id → contributing source records |
exclusions.csv |
Records filtered out, each with the rule that fired |
dedupe_review.csv |
Uncertain match pairs awaiting human judgement |
tier_agreement.csv |
How far apart independent methods land |
reconciliation.json |
Bottom-up total vs top-down anchor |
datapackage.json |
Frictionless descriptor with field documentation |
facilities.best_power_mw is a convenience column. For anything analytical,
join power_estimates and filter on method so you control which evidence you
are willing to rely on.
Three tiers, in precedence order. power_tier records which one produced
best_power_mw.
A figure stated by a source, with a URL, retrieval date and verbatim quote. The schema contract rejects any Tier A row without a citation, so an uncited number cannot acquire the authority of attested evidence.
Backup generation exists to carry critical load through an outage, so installed capacity physically bounds what a site can serve.
nameplate_mw = parsed ratings + (unrated units × per-unit prior)
critical_mw = nameplate_mw / redundancy_factor
it_load_mw = critical_mw × 0.85
The per-unit prior is measured, not guessed. Of 817 California data-center generator units in NEI, 60 state a nameplate in free text; their median is 2,116 kW with an IQR of 885–2,190 kW. Those values set the prior.
The dominant uncertainty is the redundancy divisor: NEI does not record whether a site is N+1 or 2N, and the two imply very different loads for the same fleet. The interval spans 1.10–2.00 accordingly.
white_space = footprint_sqft × min(storeys, 3) × 0.60
it_load_mw = white_space × W_per_sqft / 1e6
Storeys are capped at three. Purpose-built data centers are one to three storeys;
above that the facility is a tenant in an office or carrier-hotel tower, and
crediting every floor measures the building rather than the facility. Such records
carry partial_occupancy = true and a lower bound assuming a single storey.
Density and PUE priors live in data/reference/power_density_priors.csv, keyed on
facility class and vintage, and are editable without touching code.
A facility with no footprint measurement gets no Tier C estimate. Treating a missing polygon as zero area would fabricate a zero-power data center.
annual_gwh = it_load_mw × PUE × utilization × 8760 / 1000, with PUE from the
vintage-keyed priors and utilization at 0.70 (range 0.55–0.85).
Tier A is currently empty because no figure was added that could not be
independently verified at build time. To contribute one, append a row to
data/reference/manual_overrides.csv:
match_name,match_operator,basis,value_mw,source_url,retrieved_at,quote
Colovore,Colovore,critical_load,9.0,https://example.gov/doc,2026-08-19,"...verbatim sentence..."basis is it_load, critical_load or total_facility; the pipeline converts
between them. Rows lacking a value or URL are skipped with a warning.
src/data_center_dataset/
cli.py fetch | build | validate | report
config.py paths, EPSG:3310, every model prior
http.py cached client: rate limit, retry, identifying UA
pipeline.py ingest -> classify -> resolve -> enrich -> power -> export
sources/ osm, peeringdb, epa_nei, ceqanet, cec_gis
normalize/ schema (pandera), classify, dedupe, geometry
power/ evidence (A), generators (B), model (C), reconcile
export.py parquet/csv/geojson + datapackage
data/
raw/<source>/<date>/ immutable dated snapshots (committed, except the NEI zip)
reference/ operator aliases, power priors, curated overrides
processed/ published tables
All areas and distances are computed in EPSG:3310 (California Albers, equal-area, metres). Computing areas in WGS84 degrees is a common and badly wrong shortcut.
data/raw/<source>/<date>/ snapshots are committed, so a build is reproducible
without re-querying upstream. The one exception is the 157 MB NEI archive, which
is gitignored; the pipeline re-downloads it from the documented URL and the
committed artefact is the filtered California extract it produces.
The dataset incorporates OpenStreetMap geometry and is therefore a derived database under ODbL-1.0. Redistribution carries attribution and share-alike obligations. If you need a permissively licensed product, rebuild while excluding the OSM source — you will lose all footprint geometry, and with it Tier C.
PeeringDB is CC-BY-4.0; EPA and CEC material is US public domain.