kmail.at
← learning

pandas · difficulty ◆◆

pandas rename() — relabel columns and index labels without retyping the header

rename() relabels one axis at a time — say columns= out loud and it never surprises you.

I ran df.rename({'revenue': 'REVENUE'}) and got the same column name back. No error, no warning. pandas did exactly what it promised: it renamed the index — the axis I never mentioned.

2026-10-08 · 5 min read

$ df.rename()

What it does

rename() swaps labels on one axis: column names, row labels, or both. You hand it a mapper and it hands you back a new frame with new labels — the values are shared, not copied. A dict maps name to name and leaves everything else alone, a callable is applied to every label, a Series works like a dict, and a scalar means something only a Series understands (it sets .name). Since 0.21 you say which axis with columns= or index=; the older df.rename(mapper, axis=1) still works, and the really old positional df.rename({0:1}, {0:2}) form does not. It is additive by design: a label you never mention survives untouched, and a label you mention that is not there is ignored unless you ask for errors='raise'.

Why it matters

Header hygiene is the cheapest correctness work in a pipeline. 'Order ID', 'CUST name ', ' Revenue (SAR)' turn into df['Order ID'] quoting fights, break df.query() outright, and get silently mangled the moment anyone retypes them in a script. rename() fixes them once, by rule instead of by hand, and leaves the raw frame intact for lineage. It is also the standard repair step after an operation that invents labels: merge() appends _x/_y or your own suffixes, groupby().agg() with a list builds MultiIndex columns, pivot() stacks (measure, year) tuples on the header. All three get cleaned with the same call, and because the default is 'ignore' it is safe to run twice.

Example

$ raw = pd.DataFrame({
    "Order ID":        [1001, 1002, 1003, 1004, 1005],
    "CUST name ":      ["Yanbu Steel", "Jubail Pumps", "Riyadh Cement", "Yanbu Steel", "Jubail Pumps"],
    "  Revenue (SAR)": [12400, 8300, 21950, 6150, 9820],
    "Cost (SAR)":      [9100, 5200, 16800, 4400, 7300],
})
report = raw.rename(columns={
    "Order ID": "order_id", "CUST name ": "customer",
    "  Revenue (SAR)": "revenue", "Cost (SAR)": "cost"})
print(report)
   order_id       customer  revenue   cost
0      1001    Yanbu Steel    12400   9100
1      1002   Jubail Pumps     8300   5200
2      1003  Riyadh Cement    21950  16800
3      1004    Yanbu Steel     6150   4400
4      1005   Jubail Pumps     9820   7300

The workhorse: four ugly source headers in, four usable names out, and `raw` still holds the originals — rename returns a new object, so you keep the raw frame and hand the clean one downstream. Spell the mapping as a literal dict; six months from now that dict is the only record of what the ERP called things.

$ def slug(c):
    return re.sub(r"[^a-z0-9]+", "_", c.strip().lower()).strip("_")

print(raw.rename(columns=slug))
   order_id      cust_name  revenue_sar  cost_sar
0      1001    Yanbu Steel        12400      9100
1      1002   Jubail Pumps         8300      5200
2      1003  Riyadh Cement        21950     16800
3      1004    Yanbu Steel         6150      4400
4      1005   Jubail Pumps         9820      7300

The callable form: every label goes through the function, so a column that appears in next month's export is cleaned too — no dict to keep in sync. ' Revenue (SAR)' became revenue_sar because the regex eats the padding and the parentheses. This is the version that belongs inside a load_clean(path) helper every script imports.

$ sales = pd.DataFrame({"city": ["Yanbu", "Jubail", "Riyadh", "Yanbu", "Jubail"],
                      "revenue": [12400, 8300, 21950, 6150, 9820]})
per_city = sales.groupby("city")["revenue"].sum()
print(per_city.rename(index={"Yanbu": "Yanbu plant 1",
                             "Jubail": "Jubail plant 2",
                             "Riyadh": "Riyadh plant 3"}))
city
Jubail plant 2    18120
Riyadh plant 3    21950
Yanbu plant 1     18550
Name: revenue, dtype: int64

index= on a Series, right after an aggregation. The group keys became the index, so the report would ship with codes instead of site names; rename(index=...) is the fix and it keeps the object a Series — the index name `city` and the series `.name` revenue both survive the rename. pandas sorts the result alphabetically, which is why Jubail is first.

$ print(raw.rename(columns={"Order ID": "order_id", "Ordr ID": "oops"}).columns.tolist())
try:
    raw.rename(columns={"Order ID": "order_id", "Ordr ID": "oops"}, errors="raise")
except KeyError as e:
    print("KeyError:", e)
['order_id', 'CUST name ', '  Revenue (SAR)', 'Cost (SAR)']
KeyError: "['Ordr ID'] not found in axis"

The default is errors='ignore' — a typo'd key is not an error, it is a no-op. The first call renamed Order ID and quietly dropped the miss. Flip it to errors='raise' (0.25.0, July 2019) and the same code names the culprit: 'Ordr ID' not found in axis. My rule: 'ignore' while poking in a notebook, 'raise' in anything that runs on a schedule.

$ orders = pd.DataFrame({"order_id": [1001, 1002, 1003],
                       "customer": ["Yanbu Steel", "Jubail Pumps", "Riyadh Cement"],
                       "revenue": [12400, 8300, 21950]})
plan = pd.DataFrame({"customer": ["Yanbu Steel", "Jubail Pumps", "Riyadh Cement"],
                     "revenue": [13000, 8900, 21950]})
comparison = (orders.merge(plan, on="customer", suffixes=("_actual", "_target"))
                     .rename(columns={"revenue_actual": "revenue_ytd",
                                      "revenue_target": "revenue_plan"}))
print(comparison)
   order_id       customer  revenue_ytd  revenue_plan
0      1001    Yanbu Steel        12400         13000
1      1002   Jubail Pumps         8300          8900
2      1003  Riyadh Cement        21950         21950

The merge-cleanup pattern, and the rename I write most often. merge() appended the suffixes on its own; rename() takes them back off in the same expression, so no intermediate frame named tmp ever exists. Keep it inline in the chain — the day someone reorders the merge, the rename is right beside it.

$ wide = pd.DataFrame([[12400, 9800]],
                    columns=pd.MultiIndex.from_tuples([("revenue", "sum"), ("revenue", "mean")]))
print("before     :", wide.columns.tolist())
print("tuple key  :", wide.rename(columns={("revenue", "mean"): "avg"}).columns.tolist())
print("level=1    :", wide.rename(columns={"mean": "avg"}, level=1).columns.tolist())
print("level=0 fn :", wide.rename(columns=lambda x: x.upper(), level=0).columns.tolist())
before     : [('revenue', 'sum'), ('revenue', 'mean')]
tuple key  : [('revenue', 'sum'), ('revenue', 'mean')]
level=1    : [('revenue', 'sum'), ('revenue', 'avg')]
level=0 fn : [('REVENUE', 'sum'), ('REVENUE', 'mean')]

The trap that costs people an afternoon. On MultiIndex columns rename walks each level separately, so the tuple key ('revenue','mean') matches nothing and quietly returns the frame unchanged, while the bare 'mean' hits every level that carries that value. When you want one level, name it with level=: level=1 on 'mean' changed exactly one header and left ('revenue','sum') alone.

Common flags

columns={old: new}
Rename by column name. The spelling to type by reflex — a dict leaves unlisted columns alone, and since 1.0.0 (January 2020) it cannot be confused with the positional mapper any more.
index={old: new}
Rename row labels, keeping the index type. Works on a Series too — the natural cleanup after groupby() turns your key column into the index.
mapper (positional)
A dict-like or callable applied to the axis named by axis=. Since 1.0.0 only this first argument may be positional; df.rename({'a':'b'}, {'c':'d'}) is now a TypeError.
errors='ignore'|'raise'
What to do about a key that is not in the axis. 'ignore' is the default (keyword added in 0.25.0, July 2019) and makes a cleaning function safe to re-run; 'raise' turns a typo into KeyError instead of a silent no-op.
level=N
For MultiIndex axes, relabel only that level (added in 0.20.0, issue #4160). Without it a dict key is matched against the full tuple, and a tuple key never matches.
copy=True
Deprecated in 3.0.0 and ignored: Copy-on-Write makes rename always lazy, so the payload is never duplicated. Leave it out of new code; it is slated for removal in pandas 4.0.
inplace=True
Mutates and returns None. Unlike fillna()/replace() in 3.0, rename still hands back None — assign the result instead so the next reader can see the rename happened.

History

pandas 0.1 (December 2009) had no rename at all

The 0.1 sdist on PyPI — uploaded 25 December 2009, 139 files — has no rename method on DataFrame or Series. It arrived in the commit titled 'rename methods' by Wes McKinney on 21 April 2010 (c834263d, touching core/frame.py, matrix.py, panel.py and series.py), and shipped in 0.2 on 18 May 2010 with the signature it kept for a decade: DataFrame.rename(self, index=None, columns=None). Read that first implementation and today's makes sense — it did result = self.copy(), then self.index = [mapper(x) for x in self.index], then rebuilt the whole column dict as self._series = dict((mapper(k), v) for k, v in self._series.iteritems()). A full data copy plus a rebuilt block dictionary, on every single rename.

The copy-free rewrite, and fourteen years of keywords

Commit edd9f194, 'ENH: rename DataFrame axes without copying data', 21 September 2011, gave the method copy=True and pushed the real work into BlockManager.rename_items(mapper, copydata=False) — the frame keeps its data and only the item index is rebuilt. It shipped in 0.4.1 that September, and that release's notes say it plainly: 'DataFrame.rename has a new copy parameter to rename a DataFrame in place.' The rest arrived one keyword at a time: level= in 0.20.0 (May 2017, #4160), axis= in 0.21.0 (October 2017, #12392), errors= in 0.25.0 (July 2019, #13473), then 1.0.0 (January 2020, #29136) locked everything but the mapper behind keyword-only after the ambiguous df.rename({0:1}, {0:2}) form. 3.0.0 (January 2026, #57347) deprecated copy= itself.

Fun facts

Pros & cons

pros

  • + Additive by construction: name only what changes, everything else survives — no need to re-list every column just to fix one header
  • + Effectively free on large frames: it swaps Index objects over a shared BlockManager, 0.291 ms on an 80 MB frame versus 20.124 ms for a deep copy in my measurement
  • + Idempotent with the errors='ignore' default, so it drops straight into a cleaning function or a re-run notebook without a guard clause

cons

  • − The default axis is 0, so the shortest call df.rename({'a': 'b'}) renames rows and misses your column — with no error and no warning
  • − MultiIndex columns are genuinely confusing: a tuple key is a silent no-op and a bare key relabels every level carrying that value
  • − The keywords now overlap (mapper+axis, index=, columns=, level=, errors=) and rename_axis sits right next to it for names, so two codebases rarely agree on one spelling

Takeaways

  1. 1Always spell the axis: df.rename(columns={...}) or df.rename(index={...}). A bare dict is read as axis=0 and quietly touches row labels, not your header.
  2. 2Use errors='raise' in scheduled jobs and keep the 'ignore' default for notebook work — otherwise a typo'd key is a silent no-op instead of a crash.
  3. 3df.rename(columns=str.strip) is the whole whitespace fixup for an Excel-derived CSV, and a callable mapper keeps working when next month's export adds a column.
  4. 4On MultiIndex columns, name the level (level=1) or flatten with a list comprehension — a tuple key matches nothing and hands the frame back unchanged.
  5. 5rename() relabels, rename_axis() names the axis, and in 3.0 rename_axis rejects a label dict outright — chain the two when a downstream Excel reader needs clean names and a named axis.

Related commands

← all learning