Missing data doesn't delete itself — dropna() decides which rows deserve to stay.
You loaded a CSV, ran the analysis, got nonsense — and the culprit was eight blank cells you never looked at.
dropna() removes rows or columns that contain missing values. On a DataFrame (axis=0 default) it inspects every row and drops it if at least one watched cell is missing — NaN, None and NaT all count. The three knobs control strictness: how='all' drops a row only when EVERY watched cell is missing, subset=['a','b'] restricts the check to chosen columns, and thresh=N keeps a row if at least N real values survive. On a Series a single call just discards the missing entries. pandas 3.0 ships three missing markers — np.nan in float columns, pd.NaT in datetime64, pd.NA in arrow-backed dtypes — and dropna() treats all three uniformly.
Real tables arrive with holes: unlogged telemetry, skipped form fields, failed API calls. Downstream code handles those holes inconsistently — mean() skips them, groupby() may silently drop whole groups, an ML pipeline refuses to train. dropna() is the deliberate statement of which rows are complete enough to analyze, with readable arguments instead of a boolean-mask chain. It sits at the center of a cleaning family: fillna() repairs instead of discarding, notna() tests before dropping, drop_duplicates() removes repeats rather than gaps, and interpolate() bridges time-series gaps.
df.dropna(thresh=2)
city day pm25 no2 0 Berlin 2026-10-01 24.0 41.0 1 Berlin 2026-10-02 NaN 38.0 2 Berlin 2026-10-03 22.0 NaN 3 Vienna 2026-10-01 15.0 29.0 5 Rome 2026-10-01 31.0 55.0 6 Rome 2026-10-02 NaN 52.0 7 Rome 2026-10-03 19.0 NaN
thresh=2: keep rows carrying at least 2 real values. Row 4 (Vienna) misses BOTH sensors — junk, gone. Every partially-logged row survives.
df.dropna(subset=['pm25', 'no2'], how='all')
city day pm25 no2 0 Berlin 2026-10-01 24.0 41.0 1 Berlin 2026-10-02 NaN 38.0 2 Berlin 2026-10-03 22.0 NaN 3 Vienna 2026-10-01 15.0 29.0 5 Rome 2026-10-01 31.0 55.0 6 Rome 2026-10-02 NaN 52.0 7 Rome 2026-10-03 19.0 NaN
subset locks the check to the sensor columns; how='all' drops a row only when BOTH are missing. city/day are protected — a row with real identity but dead sensors still counts as data.
ship.dropna(ignore_index=True).to_string(index=False)
carrier days delivered
DHL 2.0 2026-10-01
UPS 1.0 2026-09-28
ignore_index=True (pandas ≥ 2.0) returns a fresh 0..n-1 index — no trailing .reset_index(drop=True). The zoo of missing markers (np.nan, None, NaT) drops in one call.
| Flag | Meaning |
|---|---|
axis=0 / 1 | drop rows (0, default) or columns (1) that contain missing values |
how='any' / 'all' | any = drop if ONE watched value is missing (default); all = drop only if ALL watched are missing |
thresh=N | keep a row only if at least N real values survive; wins over how when both are given |
subset=['col', ...] | restrict the missing-check to these columns (or rows, with axis=1) instead of the whole frame |
ignore_index=True | reset the result to a fresh 0..n-1 index (pandas ≥ 2.0, April 2023); default False keeps original labels |
inplace=True | modify in place instead of returning a copy; pandas team discourages it — assign the result |
df.isna().sum() | the decision-helper idiom: missing-per-column counts BEFORE you drop anything |
DataFrame.dropna(axis, how, thresh, subset) appeared with its full modern signature in 0.5.0. Series.valid() — 'copy with NaN entries dropped' — is visible in the 0.3.0 tag (February 2011) and was renamed to Series.dropna() before v0.4.1; the old valid() name was deprecated in 0.23.0 (2018) and removed in 1.0.0 (2020).
0.23.0 (May 2018) deprecated passing a list of axes to dropna (removed in 1.0.0); 2.0 (April 2023) added ignore_index, closing a decade in which every dropna() notebook call trailed a .reset_index(drop=True).
dropna() builds the notna() boolean mask once, reduces it along the chosen axis into a keep-vector, then gathers surviving rows in C. It is not per-cell surgery: cost is one reduction plus one take(), which is why five million rows drop in well under a second. The missing markers differ per memory layout — float columns carry the IEEE-754 NaN payload, datetime64 carries the NaT bit pattern, arrow-backed columns use a validity bitmap — and isna() normalizes all of them into that single boolean mask before the drop.