kmail.at
← learning

pandas · difficulty ◆◆

pandas drop() — delete rows and columns by label (without the axis trap)

The delete key for rows and columns: name what you are losing, and say which axis out loud.

`df.drop('city')` looks like it deletes a column. It does not — pandas walks down the rows, finds no row labelled 'city', and hands back a KeyError. The method that really deletes columns is the same one; the difference is the word `columns=`.

2026-10-07 · 5 min read

$ df.drop()

What it does

drop() removes labels from one axis of a frame. Rows are the default, columns are the thing you actually want nine times out of ten. Modern pandas gives you two self-documenting spellings — df.drop(columns=['a','b']) and df.drop(index=[1002,1003]) — and keeps the older df.drop(['a','b'], axis=1) form working for the code you inherited. Everything else is a modifier on that idea: level= picks which level of a MultiIndex to peel, errors='ignore' turns a missing label from a crash into a shrug, inplace=True mutates and returns None. It returns a new object; the original is left alone.

Why it matters

Half of real-world data work is subtraction: strip the PII columns before an export, drop the audit and debug fields your ERP pads onto every report, throw out the stale SKUs a validator rejected, collapse a two-row MultiIndex header after a pivot so the CSV opens cleanly. If you select instead of dropping you have to name every surviving column, which breaks the day someone adds one. drop() lets you name only what should go — the smaller, stabler list. The two ways to get it wrong are equally common and equally silent: the default axis is 0, so a bare label is read as a row, and a duplicated index label takes every matching row with it.

Example

$ sales.drop(columns=["internal_notes", "cost"])
   order_id    city  revenue
0      1001   Yanbu    12400
1      1002  Jubail     8300
2      1003  Riyadh    21950
3      1004  Jubail     6150
4      1005   Yanbu     9820

The setup for every example here: a five-row orders frame with two columns nobody outside the team should see. columns= takes a list, drops both, and leaves the original frame untouched — assign it: `report = sales.drop(columns=[...])`.

$ orders = sales.set_index("order_id")
orders.drop(index=[1002, 1003])
            city  revenue
order_id                 
1001       Yanbu    12400
1004      Jubail     6150
1005       Yanbu     9820

A hard-coded cancellation list. index= removes those row labels by name — no boolean mask, no .isin(). With a named index the order_id stays a column heading in the output, which is what you want when the next step is to_csv().

$ sales.drop(sales[sales.revenue < 10000].index)
   order_id    city  revenue   cost internal_notes
0      1001   Yanbu    12400   9100            vip
2      1003  Riyadh    21950  16800  late delivery

The reject-list pattern: build a frame of the rows you do not trust and drop its index. Equivalent to sales[sales.revenue >= 10000] — I checked with .equals() and it came back True — so drop the losers when the 'bad' rule is the natural way to say it, keep the winners when it is not.

$ sales.drop(columns=["cost", "internal_notes", "audit_flag"], errors="ignore")
['order_id', 'city', 'revenue']

errors='ignore' (0.16.1, May 2015) drops what exists and forgets what does not. This is what makes a shared cleaning function safe: a caller can hand it audit_flag from a system that has that column and one that does not. The printed list proves cost and internal_notes went and the missing audit_flag raised nothing.

$ wide.drop(columns="qty", level=0)
     revenue       
       Yanbu Jubail
Sept   12400   8300
Oct    21950   6150

After a pivot or a groupby on two columns you get MultiIndex columns. level=0 names a level rather than a full label, so one call deletes every qty column — both cities — while keeping the revenue half. Flatten to single-level names first and this becomes a list comprehension.

$ empty = raw.columns[raw.isna().all()]
raw.drop(columns=empty)
['legacy_ref', 'audit_flag']

   order_id    city  revenue
0      1001   Yanbu    12400
1      1002  Jubail     8300
2      1003  Riyadh    21950
3      1004   Yanbu     9820

The ingest-hygiene idiom I use most: find the columns that are completely empty and drop them in the same breath. The two dead columns from the source system are gone, and the report keeps its revenue. The printed list is the audit trail — log it and you can answer 'why did that column disappear?' six months later.

Common flags

labels (positional)
The label or list of labels to remove. Since 2.0 it is the ONLY argument you may still pass positionally — df.drop('a', 1) now raises TypeError.
columns=[...]
Drop by column name. The form you actually want; added in 0.21.0 (October 2017) as a readable alternative to axis=1.
index=[...]
Drop by row label. Works on any index type — dates, ids, MultiIndex tuples.
axis=0|1
0 = rows (the default), 1 = columns. The legacy sibling of the two keywords above — and you cannot pass axis=1 and columns= in the same call.
level=N
For a MultiIndex axis, drop from that level only. level=0 on MultiIndex columns removes the whole first header row.
errors='raise'|'ignore'
Whether a label that is not there is an error. 'ignore' is what makes cleaning functions idempotent.
inplace=False
Mutate the frame and return None instead of a copy. Works, discouraged — assign the result so the next reader can see the change.

History

A generic method, one commit, July 2011

pandas 0.3.0 had no drop at all — DataFrame offered only dropEmptyRows and dropIncompleteRows. The generic version arrived in commit 4f99020 by Wes McKinney on 24 July 2011, titled 'ENH: generic drop method for pandas objects', touching a single file (pandas/core/generic.py). It shipped in 0.4.0, tagged 12 September 2011, and the signature was already recognition-proof: drop(self, labels, axis=0) — rows first, same default as today.

Fourteen years of keywords bolted on

level= came in 0.7.2 (March 2012, issue #159), inplace in 0.13.0 (January 2014, alongside the great 'all methods take an inplace kwarg — but return None' cleanup), errors= in 0.16.1 (May 2015, issue #6736, added to suppress a ValueError), and the index=/columns= keywords in 0.21.0 (October 2017, issue #12392). Two later changes sharpened the edges: 0.23.0 (May 2018, issue #19186) made KeyError replace ValueError for a label missing from a duplicated axis, and 2.0 (April 2023, issue #41486) locked everything except labels behind keyword-only.

Fun facts

Pros & cons

pros

  • + One verb for both axes — columns=, index=, or labels plus axis= — instead of a separate 'remove column' function to remember
  • + errors='ignore' makes a cleaning function idempotent: callers can pass columns that may or may not exist and the pipeline keeps going
  • + Under Copy-on-Write a column drop is a zero-copy slice, so trimming a wide frame before an export is about as cheap as it can be

cons

  • − axis=0 is the default, so df.drop('city') — the call everyone types first — is a KeyError instead of a column removal
  • − A duplicated index label takes every matching row with it; there is no 'drop only the first match' mode, which bites on non-unique ids after a merge
  • − Four overlapping spellings (labels+axis, index=, columns=, and del/pop) mean two codebases rarely agree on one, and reviewers have to read twice

Takeaways

  1. 1Never write df.drop('col'). axis=0 is the default, so that is read as a row label and raises KeyError — spell it df.drop(columns='col').
  2. 2columns=/index= (0.21.0, October 2017) are the readable spelling; reach for the older labels + axis=1 only when you are matching surrounding code.
  3. 3Add errors='ignore' to any reusable cleaning helper so a caller passing a column that no longer exists cannot break the pipeline.
  4. 4df.drop(df[mask].index) and df[~mask] give identical frames — pick drop when the bad set comes from a validator or cancellation list, boolean-index when the keep rule is the natural expression.
  5. 5Check df.index.is_unique before dropping by label: on a duplicated index, drop(index=label) removes every matching row, not one.

Related commands

← all learning