kmail.at
← learning

pandas · difficulty ◆◆

pandas fillna() — repair missing values instead of deleting the rows

dropna() throws rows away; fillna() decides what the hole should have said.

For six years, calling df.fillna() with no arguments quietly forward-filled your data. Then pandas deleted the keyword that made it possible.

2026-10-05 · 5 min read

$ df.fillna()

What it does

fillna() replaces missing values — NaN, None, NaT, pd.NA — with something you choose. The value can be a scalar, a dict mapping column names to fill values, a Series aligned on columns, or (new in 3.0) None, which fills each non-object column with that dtype's own NA marker. limit=N caps how many holes get filled per column, counting down each column independently. axis=1 flips the mapping so a dict or Series is read per row instead of per column. What fillna() no longer has is method: forward- and backward-filling moved out into ffill() and bfill(), and passing method= now dies with a TypeError.

Why it matters

Almost every real table arrives with holes, and the useful question is never 'is there a missing value' but 'what does this particular hole mean'. A missing sensor reading in a float column is plausibly the column mean. A missing category after a left join is a genuinely new label — UNASSIGNED, not zero. A missing count is zero. dropna() can only make that call globally; fillna() lets you write a different policy per column in one call, which is why it sits next to dropna() as the repair half of the cleaning family. Get it wrong and you either delete 80% of a table or invent 80% of it.

Example

$ import pandas as pd

df = pd.DataFrame({
    "city": ["Berlin", "Berlin", "Berlin", "Vienna", "Rome"],
    "day":  ["2026-10-01", "2026-10-02", "2026-10-03", "2026-10-01", "2026-10-02"],
    "pm25": [24.0, None, 22.0, 15.0, 31.0],
    "no2":  [41.0, 38.0, None, 29.0, None],
})
print(df.fillna({"pm25": df["pm25"].mean(), "no2": 0.0}).to_string())
     city         day  pm25   no2
0  Berlin  2026-10-01  24.0  41.0
1  Berlin  2026-10-02  23.0  38.0
2  Berlin  2026-10-03  22.0   0.0
3  Vienna  2026-10-01  15.0  29.0
4    Rome  2026-10-02  31.0   0.0

The dict is the whole point: pm25 gets the column mean (23.0), no2 gets a hard zero, and the identity columns are never touched. Unrecognised keys are skipped in silence, so a typo in a column name fails quietly rather than loudly.

$ orders = pd.DataFrame({"order_id": [1001, 1002, 1003, 1004],
                       "sku": ["A-1", "B-2", "C-3", "A-1"]})
stock = pd.DataFrame({"sku": ["A-1", "B-2"],
                      "warehouse": ["Berlin", "Vienna"]})

merged = orders.merge(stock, on="sku", how="left")
merged["warehouse"] = merged["warehouse"].fillna("UNASSIGNED")
print(merged.to_string(index=False))
 order_id sku  warehouse
     1001 A-1     Berlin
     1002 B-2     Vienna
     1003 C-3 UNASSIGNED
     1004 A-1     Berlin

Order 1003 has a real SKU but no matching stock row, so the left join leaves NaN — that NaN means 'no warehouse mapped', and fillna() turns it into a label the business can filter on. The column is already str dtype in pandas 3.0, so no astype() is needed.

$ log = pd.DataFrame({"route": ["Wien", "Wien", "Graz"],
                    "delay_min": [None, None, 5.0],
                    "load": [None, 812.0, None]})
print(log.fillna({"delay_min": 0.0, "load": 900.0}, limit=1).to_string())
  route  delay_min   load
0  Wien        0.0  900.0
1  Wien        NaN  812.0
2  Graz        5.0    NaN

limit=1 counts per column, not per frame: row 0 is filled in both columns, and row 1's load is a real reading that needs nothing. The unfilled delay_min at row 1 stays NaN — that is the cap doing its job instead of letting one rule rewrite a whole column's history.

$ shifts = pd.DataFrame({"store": ["A", "A", "A", "B", "B", "B"],
                       "sales": [100.0, None, 300.0, 40.0, None, 60.0]})

shifts["sales_filled"] = shifts.groupby("store")["sales"].transform(
    lambda s: s.fillna(s.median())
)
print(shifts.to_string(index=False))
store  sales  sales_filled
    A  100.0         100.0
    A    NaN         200.0
    A  300.0         300.0
    B   40.0          40.0
    B    NaN          50.0
    B   60.0          60.0

The groupby idiom that replaced the deleted groupby.fillna(): store A's median (200) and store B's (50) land only on the missing slots, and index alignment puts each row back where it belongs. Filling both with the frame-wide mean of 125.0 would have been wrong for both stores.

$ mixed = pd.DataFrame({
    "pm25": [24.0, None, 31.0],
    "zone": pd.array(["centre", None, "north"], dtype="string"),
})
print(mixed.fillna(value=None).to_string(index=False))
 pm25   zone
 24.0 centre
  NaN   <NA>
 31.0  north

value=None is new in 3.0 and looks like a no-op, but it is not: it re-fills each non-object column with the NA that dtype expects — nan for float64, <NA> for the string dtype — which is what you want before writing to Parquet or an Arrow-backed store. On an object column it genuinely does nothing.

Common flags

value=0
the common case — one scalar for every hole in every column. Filling a float64 column with a string silently upcasts that column to object dtype.
value={"col": v}
per-column policy. Keys that are not column names are ignored without a warning, so check the spelling yourself.
value=pd.Series({...})
same as the dict, aligned on the column index; with axis=1 it is aligned on the row index instead.
value=None
new in 3.0: fills each non-object column with its own NA marker — nan for float64, NaT for datetime64, <NA> for Int64 and string. A no-op on object columns.
axis=1
new in 3.0.0 for dict and Series arguments: read the mapping by row label instead of column name. Needs every column to share a dtype, or it raises ValueError.
limit=N
fill at most N holes per column, counted from the top of each column, independently.
inplace=True
fills in place and, since 3.0, returns the modified frame (self) instead of None. Chained as df['a'].fillna(0, inplace=True) it raises ChainedAssignmentError under Copy-on-Write.
df.ffill() / df.bfill()
the replacements for the removed method= keyword, with their own limit_area='inside' | 'outside' argument for edge gaps.
groupby().transform()
group-wise imputation now that groupby.fillna was removed in 3.0: df.groupby(k)[col].transform(lambda s: s.fillna(...)).

History

0.4.3, October 2011 — and a default nobody remembers

DataFrame.fillna shipped in pandas 0.4.3, published 2011-10-25, with the signature fillna(value=None, method='pad'). That default mattered for years: bare df.fillna() silently forward-filled, because 'pad' was the method and None meant 'use the method'. Series.fillna is older still — it is already present in the 0.3.0 source archive from February 2011, before the repo's tags existed, so tracing it means reading a release tarball rather than a commit. The default finally flipped to None in 0.21.0 (October 2017), the release where fillna() stopped being ffill() in disguise.

2023 to 2026 — the keyword diet

pandas 2.1.0 (2023-08-30) deprecated the method and limit keywords on fillna and pointed users at ffill()/bfill() (GH 53394, opened 2023-05-25, closed two weeks later). 2.2.0 (2024-01-20) deprecated the silent-downcasting behaviour, so a fill that used to quietly turn floats into ints stopped doing that. 3.0.0 (2026-01-21) finished the job: method is gone (GH 57760), GroupBy.fillna and Resample.fillna are gone (GH 55719), value=None became legal (GH 57723), and inplace=True started returning self (GH 63207). The signature went from seven parameters in 1.0.0 to five today.

Fun facts

Pros & cons

pros

  • + One call encodes a whole cleaning policy: a dict gives every column its own rule, with fill values computed from the same frame you are fixing
  • + No silent dtype rewrites — 3.0 enforced the downcasting deprecation, so filling an Int64 column with np.nan leaves it Int64 instead of quietly making it float64
  • + Pairs naturally with merge(how='left') to convert unmatched rows into a legible label instead of a NaN nobody can group by

cons

  • − Every fillna tutorial written before 2023 is now wrong: method= raises TypeError, groupby.fillna() and Resample.fillna() are gone, and none of the fixes is a find-and-replace
  • − Silent failure modes: unknown dict keys are ignored, and a string fill on a numeric column upcasts it to object without a warning
  • − axis=1 with a dict rejects an ordinary all-float64 frame with ValueError 'not a suitable type to fill into float64'; passing a Series instead works, which is not discoverable

Takeaways

  1. 1Use a dict for per-column policy — df.fillna({'pm25': df['pm25'].mean(), 'no2': 0.0}) — and remember that unknown keys are skipped in silence.
  2. 2Decide before you fill: a hole after a left join is a label ('UNASSIGNED'), a hole in a sensor reading is a statistic, a hole in a count is zero. One call, three different values.
  3. 3limit=N counts per column from the top, so it caps an imputation instead of letting one rule rewrite a whole column's history.
  4. 4Group-wise imputation is now df.groupby(k)[col].transform(lambda s: s.fillna(s.median())) — groupby.fillna was removed in 3.0.
  5. 5Never chain inplace: df[col].fillna(v, inplace=True) raises ChainedAssignmentError and changes nothing; assign the result instead.

Related commands

← all learning