How We Clean the Data

Our approach to data quality, in detail.

Air quality data from real-world sensors is messy. Instruments malfunction. Readings spike for no apparent reason. A sensor freezes and reports the same number for hours. A government station switches units without warning.

If you publish this data as-is, people will draw wrong conclusions. If you silently delete the bad parts, people cannot verify what you did. Neither option is acceptable to us.

Our approach is to be thorough and transparent: check everything we can, flag every problem openly, and let people see exactly what we did and why. Here is how.

Architecture

Three layers, one principle

The original data is sacred. We never modify it. Instead, we build clean data on top of it — and keep the link between the two, so anyone can trace any published value back to its raw source.

Layer 0

Raw storage

The exact responses we receive from each API, stored as-is. One table per source. Nothing is changed, nothing is lost. This is our ground truth.

Layer 1

Clean & harmonized

All sources converted to common units, all measurements in a single format. Every row carries quality flags and a link to its raw source. This is where cleaning and validation happen.

Layer 2

Published open data

Only measurements that passed every quality check. This is what you download — compressed CSV files, a CSV station registry, GeoJSON for mapping. Clean, reliable, ready to use.

Quality control

Four layers of checks

Each layer catches different kinds of problems. A measurement must pass all four before it reaches the published dataset.

1

At the door

Before data even enters the system, the database itself enforces basic rules — required fields cannot be empty, duplicates are rejected, coordinates must be valid, only known parameter codes are accepted. Broken or incomplete API responses are caught here.

2

During cleaning

As we transform raw data into the harmonized format, every measurement is checked individually: is this value physically possible? Is PM2.5 higher than PM10 (which usually indicates a sensor problem)? Is this concentration negative? Each issue gets a specific flag.

3

Automated cleaning rules

Every hourly value goes through the same rules before publication. Values that fail are kept in our database with a flag and left out of the published files:

  • Hard cap — values at or above a physical limit (1,000 µg/m³ for PM2.5) are invalid
  • Constant station — a station-month in which a single value makes up 70% or more of the readings
  • Stuck sensor — the same value repeated for 6 or more consecutive hours
  • Dead dust channel — a PM2.5, PM10 or TSP monitor whose daily median stays below 2 µg/m³ on ten or more days within a month (KazHydroMet channels fail this way: they keep reporting about 1 µg/m³)
  • Duplicated dust channel — a PM2.5 reading that repeats the PM10 value of the same monitor (within 2%) for at least half of a day's hours. PM2.5 is a part of PM10 and cannot equal it for long, so the PM2.5 values of such days, and of the period in which they keep occurring, are not published
  • Raised floor — a dust channel that never comes down: 30 days in which even the cleanest hours of every day stay at 30 µg/m³ or above. Real air always has clean hours within a month; a monitor that does not show them is adding an offset or is stuck high
  • Noisy channel — a dust channel that jumps from hour to hour: on twelve or more days within a month the typical hourly change is 40% of the daily level or more. Real dust levels rise and fall over hours, not at random from one reading to the next
  • Cluster outlier — a station whose daily average is more than 3 robust standard deviations away from the median of its neighbourhood
  • Duplicate sources — when two sources report the same station and hour, one value is kept
  • Outside Almaty, Astana and Karaganda most towns have a single monitor, so the rules compare a station only with itself: hard caps, the analyser's saturation value, the same value for 24 hours or longer, zero dust for 6 hours or longer, and the dead, duplicated, raised-floor and noisy dust channel rules above. Short runs at an analyser's detection limit and one-hour peaks are kept — near industry such peaks are real
4

Final validation

Before anything is published, we run two independent validation systems across the entire dataset. If either one finds a problem — even a single failing check — publication is blocked until the issue is resolved.

The details

Validation checks

These are the specific checks that run before every publication. Errors block export entirely. Warnings are logged and reviewed.

Check Severity What it catches
out_of_range error Values outside physical bounds
negative_values error Negative concentrations
invalid_floats error NaN, Infinity values
orphan_measurements error Measurements without a registered station
duplicate_measurements error Should-not-exist duplicates
pm25_exceeds_pm10 warning Known AirKaz sensor issue
stuck_sensors warning Frozen sensor readings
spikes warning Sudden value jumps
stale_stations warning Active stations not reporting 48+ hours

We know this level of detail is not for everyone. Most people just want to download the data and trust that it is clean. We hope they can.

But for those who want to look under the hood — researchers, engineers, anyone who has been burned by bad data before — we want everything to be visible. The methodology, the flags, the raw originals. That is the kind of project we want to be.