pandas · difficulty ◆◆
pandas rename() — relabel columns and index labels without retyping the header
rename() relabels one axis at a time — say columns= out loud and it never surprises you.
I ran df.rename({'revenue': 'REVENUE'}) and got the same column name back. No error, no warning. pandas did exactly what it promised: it renamed the index — the axis I never mentioned.
$ df.rename()What it does
rename() swaps labels on one axis: column names, row labels, or both. You hand it a mapper and it hands you back a new frame with new labels — the values are shared, not copied. A dict maps name to name and leaves everything else alone, a callable is applied to every label, a Series works like a dict, and a scalar means something only a Series understands (it sets .name). Since 0.21 you say which axis with columns= or index=; the older df.rename(mapper, axis=1) still works, and the really old positional df.rename({0:1}, {0:2}) form does not. It is additive by design: a label you never mention survives untouched, and a label you mention that is not there is ignored unless you ask for errors='raise'.
Why it matters
Header hygiene is the cheapest correctness work in a pipeline. 'Order ID', 'CUST name ', ' Revenue (SAR)' turn into df['Order ID'] quoting fights, break df.query() outright, and get silently mangled the moment anyone retypes them in a script. rename() fixes them once, by rule instead of by hand, and leaves the raw frame intact for lineage. It is also the standard repair step after an operation that invents labels: merge() appends _x/_y or your own suffixes, groupby().agg() with a list builds MultiIndex columns, pivot() stacks (measure, year) tuples on the header. All three get cleaned with the same call, and because the default is 'ignore' it is safe to run twice.
Example
$ raw = pd.DataFrame({
"Order ID": [1001, 1002, 1003, 1004, 1005],
"CUST name ": ["Yanbu Steel", "Jubail Pumps", "Riyadh Cement", "Yanbu Steel", "Jubail Pumps"],
" Revenue (SAR)": [12400, 8300, 21950, 6150, 9820],
"Cost (SAR)": [9100, 5200, 16800, 4400, 7300],
})
report = raw.rename(columns={
"Order ID": "order_id", "CUST name ": "customer",
" Revenue (SAR)": "revenue", "Cost (SAR)": "cost"})
print(report) order_id customer revenue cost
0 1001 Yanbu Steel 12400 9100
1 1002 Jubail Pumps 8300 5200
2 1003 Riyadh Cement 21950 16800
3 1004 Yanbu Steel 6150 4400
4 1005 Jubail Pumps 9820 7300The workhorse: four ugly source headers in, four usable names out, and `raw` still holds the originals — rename returns a new object, so you keep the raw frame and hand the clean one downstream. Spell the mapping as a literal dict; six months from now that dict is the only record of what the ERP called things.
$ def slug(c):
return re.sub(r"[^a-z0-9]+", "_", c.strip().lower()).strip("_")
print(raw.rename(columns=slug)) order_id cust_name revenue_sar cost_sar
0 1001 Yanbu Steel 12400 9100
1 1002 Jubail Pumps 8300 5200
2 1003 Riyadh Cement 21950 16800
3 1004 Yanbu Steel 6150 4400
4 1005 Jubail Pumps 9820 7300The callable form: every label goes through the function, so a column that appears in next month's export is cleaned too — no dict to keep in sync. ' Revenue (SAR)' became revenue_sar because the regex eats the padding and the parentheses. This is the version that belongs inside a load_clean(path) helper every script imports.
$ sales = pd.DataFrame({"city": ["Yanbu", "Jubail", "Riyadh", "Yanbu", "Jubail"],
"revenue": [12400, 8300, 21950, 6150, 9820]})
per_city = sales.groupby("city")["revenue"].sum()
print(per_city.rename(index={"Yanbu": "Yanbu plant 1",
"Jubail": "Jubail plant 2",
"Riyadh": "Riyadh plant 3"}))city
Jubail plant 2 18120
Riyadh plant 3 21950
Yanbu plant 1 18550
Name: revenue, dtype: int64index= on a Series, right after an aggregation. The group keys became the index, so the report would ship with codes instead of site names; rename(index=...) is the fix and it keeps the object a Series — the index name `city` and the series `.name` revenue both survive the rename. pandas sorts the result alphabetically, which is why Jubail is first.
$ print(raw.rename(columns={"Order ID": "order_id", "Ordr ID": "oops"}).columns.tolist())
try:
raw.rename(columns={"Order ID": "order_id", "Ordr ID": "oops"}, errors="raise")
except KeyError as e:
print("KeyError:", e)['order_id', 'CUST name ', ' Revenue (SAR)', 'Cost (SAR)']
KeyError: "['Ordr ID'] not found in axis"The default is errors='ignore' — a typo'd key is not an error, it is a no-op. The first call renamed Order ID and quietly dropped the miss. Flip it to errors='raise' (0.25.0, July 2019) and the same code names the culprit: 'Ordr ID' not found in axis. My rule: 'ignore' while poking in a notebook, 'raise' in anything that runs on a schedule.
$ orders = pd.DataFrame({"order_id": [1001, 1002, 1003],
"customer": ["Yanbu Steel", "Jubail Pumps", "Riyadh Cement"],
"revenue": [12400, 8300, 21950]})
plan = pd.DataFrame({"customer": ["Yanbu Steel", "Jubail Pumps", "Riyadh Cement"],
"revenue": [13000, 8900, 21950]})
comparison = (orders.merge(plan, on="customer", suffixes=("_actual", "_target"))
.rename(columns={"revenue_actual": "revenue_ytd",
"revenue_target": "revenue_plan"}))
print(comparison) order_id customer revenue_ytd revenue_plan
0 1001 Yanbu Steel 12400 13000
1 1002 Jubail Pumps 8300 8900
2 1003 Riyadh Cement 21950 21950The merge-cleanup pattern, and the rename I write most often. merge() appended the suffixes on its own; rename() takes them back off in the same expression, so no intermediate frame named tmp ever exists. Keep it inline in the chain — the day someone reorders the merge, the rename is right beside it.
$ wide = pd.DataFrame([[12400, 9800]],
columns=pd.MultiIndex.from_tuples([("revenue", "sum"), ("revenue", "mean")]))
print("before :", wide.columns.tolist())
print("tuple key :", wide.rename(columns={("revenue", "mean"): "avg"}).columns.tolist())
print("level=1 :", wide.rename(columns={"mean": "avg"}, level=1).columns.tolist())
print("level=0 fn :", wide.rename(columns=lambda x: x.upper(), level=0).columns.tolist())before : [('revenue', 'sum'), ('revenue', 'mean')]
tuple key : [('revenue', 'sum'), ('revenue', 'mean')]
level=1 : [('revenue', 'sum'), ('revenue', 'avg')]
level=0 fn : [('REVENUE', 'sum'), ('REVENUE', 'mean')]The trap that costs people an afternoon. On MultiIndex columns rename walks each level separately, so the tuple key ('revenue','mean') matches nothing and quietly returns the frame unchanged, while the bare 'mean' hits every level that carries that value. When you want one level, name it with level=: level=1 on 'mean' changed exactly one header and left ('revenue','sum') alone.
Common flags
- columns={old: new}
- Rename by column name. The spelling to type by reflex — a dict leaves unlisted columns alone, and since 1.0.0 (January 2020) it cannot be confused with the positional mapper any more.
- index={old: new}
- Rename row labels, keeping the index type. Works on a Series too — the natural cleanup after groupby() turns your key column into the index.
- mapper (positional)
- A dict-like or callable applied to the axis named by axis=. Since 1.0.0 only this first argument may be positional; df.rename({'a':'b'}, {'c':'d'}) is now a TypeError.
- errors='ignore'|'raise'
- What to do about a key that is not in the axis. 'ignore' is the default (keyword added in 0.25.0, July 2019) and makes a cleaning function safe to re-run; 'raise' turns a typo into KeyError instead of a silent no-op.
- level=N
- For MultiIndex axes, relabel only that level (added in 0.20.0, issue #4160). Without it a dict key is matched against the full tuple, and a tuple key never matches.
- copy=True
- Deprecated in 3.0.0 and ignored: Copy-on-Write makes rename always lazy, so the payload is never duplicated. Leave it out of new code; it is slated for removal in pandas 4.0.
- inplace=True
- Mutates and returns None. Unlike fillna()/replace() in 3.0, rename still hands back None — assign the result instead so the next reader can see the rename happened.
History
pandas 0.1 (December 2009) had no rename at all
The 0.1 sdist on PyPI — uploaded 25 December 2009, 139 files — has no rename method on DataFrame or Series. It arrived in the commit titled 'rename methods' by Wes McKinney on 21 April 2010 (c834263d, touching core/frame.py, matrix.py, panel.py and series.py), and shipped in 0.2 on 18 May 2010 with the signature it kept for a decade: DataFrame.rename(self, index=None, columns=None). Read that first implementation and today's makes sense — it did result = self.copy(), then self.index = [mapper(x) for x in self.index], then rebuilt the whole column dict as self._series = dict((mapper(k), v) for k, v in self._series.iteritems()). A full data copy plus a rebuilt block dictionary, on every single rename.
The copy-free rewrite, and fourteen years of keywords
Commit edd9f194, 'ENH: rename DataFrame axes without copying data', 21 September 2011, gave the method copy=True and pushed the real work into BlockManager.rename_items(mapper, copydata=False) — the frame keeps its data and only the item index is rebuilt. It shipped in 0.4.1 that September, and that release's notes say it plainly: 'DataFrame.rename has a new copy parameter to rename a DataFrame in place.' The rest arrived one keyword at a time: level= in 0.20.0 (May 2017, #4160), axis= in 0.21.0 (October 2017, #12392), errors= in 0.25.0 (July 2019, #13473), then 1.0.0 (January 2020, #29136) locked everything but the mapper behind keyword-only after the ambiguous df.rename({0:1}, {0:2}) form. 3.0.0 (January 2026, #57347) deprecated copy= itself.
Fun facts
Pros & cons
pros
- + Additive by construction: name only what changes, everything else survives — no need to re-list every column just to fix one header
- + Effectively free on large frames: it swaps Index objects over a shared BlockManager, 0.291 ms on an 80 MB frame versus 20.124 ms for a deep copy in my measurement
- + Idempotent with the errors='ignore' default, so it drops straight into a cleaning function or a re-run notebook without a guard clause
cons
- − The default axis is 0, so the shortest call df.rename({'a': 'b'}) renames rows and misses your column — with no error and no warning
- − MultiIndex columns are genuinely confusing: a tuple key is a silent no-op and a bare key relabels every level carrying that value
- − The keywords now overlap (mapper+axis, index=, columns=, level=, errors=) and rename_axis sits right next to it for names, so two codebases rarely agree on one spelling
Takeaways
- 1Always spell the axis: df.rename(columns={...}) or df.rename(index={...}). A bare dict is read as axis=0 and quietly touches row labels, not your header.
- 2Use errors='raise' in scheduled jobs and keep the 'ignore' default for notebook work — otherwise a typo'd key is a silent no-op instead of a crash.
- 3df.rename(columns=str.strip) is the whole whitespace fixup for an Excel-derived CSV, and a callable mapper keeps working when next month's export adds a column.
- 4On MultiIndex columns, name the level (level=1) or flatten with a list comprehension — a tuple key matches nothing and hands the frame back unchanged.
- 5rename() relabels, rename_axis() names the axis, and in 3.0 rename_axis rejects a label dict outright — chain the two when a downstream Excel reader needs clean names and a named axis.