pandas memory_usage() — see how much RAM each column actually costs

Without deep=True you're auditing a building but skipping every room: pass deep=True and read the real bytes.

Your 'wide report CSV' didn't get slow overnight — one string column of repeated ZIP codes is quietly paying rent on 60 bytes per row. memory_usage() names the tenant.

What it does

df.memory_usage() returns a Series of byte counts, one row per column plus the index; sum() it for a frame total. Without arguments, each numeric column reports items*itemsize — int64 or float64 means 8 bytes a row. Strings and object columns are the blind spot: by default they're counted at a placeholder 8 bytes per pointer, so pass deep=True to make pandas walk the objects with sys.getsizeof and count the real characters. Two knobs: index=False drops the index row, and df.info(memory_usage='deep') prints the same deep total inside the familiar info() layout.

Why it matters

The naive number lies to you exactly when you care most. A float64 revenue column genuinely costs 8 bytes a row, but a warehouse-code string costs 50+ — and the default report shows that string as another tidy 8. One deep=True and the column ordering flips: the innocent-looking text column becomes the whale, category dtype converts it to a dictionary code, and a 60 MB frame drops to 5. When a read_csv of a 2 GB CSV crashes your laptop, memory_usage tells you which columns to downcast before you re-read with dtype= and usecols.

Examples

import pandas as pd

cities = pd.DataFrame({
    "city":       ["Vienna", "Graz", "Linz", "Salzburg", "Innsbruck"],
    "district":   ["Innere Stadt", "Lend", "Lindgra", "Pond", "Wiltra"],
    "population": [1_931_452, 331_745, 206_809, 156_852, 132_493],
    "sales_2025": [8_250_300.75, 1_410_220.10, 903_441.50, 702_115.00, 611_980.25],
})

cities.memory_usage()
Index         132
city           40
district       40
population     40
sales_2025     40
dtype: int64

Five rows x 8 bytes = 40 for every numeric/text column, because object columns are counted as bare pointers at this level. The index row is charged too — a RangeIndex reports its overhead as a constant (132 bytes here); pass index=False to drop that row.

raw = pd.DataFrame({
    "warehouse": ["W-101"] * 300_000 + ["W-102"] * 300_000,
    "pallets":   [12.0, 7.5, 8.25] * 200_000,
    "shipped":   [1, 0] * 300_000,
})
raw.memory_usage(deep=True) / 1e6   # MB, per column
Index         0.000132
warehouse    32.400000
pallets       4.800000
shipped       4.800000
dtype: float64

deep=True exposes the trap: the warehouse string column burns 32.4 MB — 108 bytes per row for a 5-character code — while both numeric columns sit at 4.8. One downcast wave later, the same data fits in a quarter of the room.

opt = pd.DataFrame({
    "warehouse": raw["warehouse"].astype("category"),
    "pallets":   raw["pallets"].astype("float32"),
    "shipped":   raw["shipped"].astype("int8"),
})
opt.memory_usage(deep=True) / 1e6
Index        0.000132
warehouse    0.600108
pallets      2.400000
shipped      0.600000
dtype: float64

54.6x smaller overall, and the string whale is now a 0.6 MB dictionary-encoded column, exactly what a Parquet file does on disk. Memory_usage after astype is your proof the conversion actually paid.

Flags

FlagMeaning
deep=Truewalk string/object cells with sys.getsizeof for true byte counts — the default undercounts object columns as 8 bytes per row.
index=Falsedrop the Index row from the Series; handy with sum() so the constant block doesn't dilute per-row math.
info(memory_usage='deep')same deep total, printed inside the df.info() report — the quickest one-look audit.
.astype('category')dictionary-encode low-cardinality strings; the single biggest saving deep counts point at.
.astype('float32') / 'int32'halve numeric columns when the value range tolerates it; pair with deep sums before/after.
dtype={...} in read_csvspecify dtypes at load time so the shrunken frame never allocates the fat version.
Series.memory_usage(index=False)one column at a time; Series has the same method, minus the index row by default.

Shipped in pandas 0.15.0 (October 2014)

GitHub issue #6852 added df.memory_usage() plus the memory line in df.info(); the 0.15.0 release notes list both under 'Memory Usage'. The deep=True flag for object columns followed in 0.18.0 (March 2016), right after sys.getsizeof() was made pandas-aware in issue #11597.

Why the default is the 'shallow' number

Walking every Python object with sys.getsizeof costs O(rows) time, so pandas kept shallow counting as the default: numeric blocks report items*itemsize in C, no Python objects touched. The design echoes R's object.size(); R counts string bytes natively, while pandas must reach into each Python str — that's the exact gap deep=True closes.

Under the hood: numutils and sys.getsizeof

With deep=False the count is pure C: each block reports items*itemsize, no Python objects touched — constant time. String blocks skip the +8 pointer term entirely and route through memory_usage_of_objects, a Cython loop in pandas/_libs/lib.pyx that calls __sizeof__() (sys.getsizeof) on every cell, plus the block's own pointer array. That's the audit's cost: one Python call per row. Categoricals sidestep the walk — codes are int8/16/32, categories counted once. On this box (CPython 3.12) sys.getsizeof('') == 41, so short strings cost far more than their characters. Caveat for 2.x readers: pandas 2.x's default arrow-backed strings don't allocate a Python str per cell, so deep numbers differ there; 3.0 switched the default str dtype back to Python objects.

Fun facts

Pros

Cons

Takeaways