One .agg() call replaces the five stacked df[col].sum() lines — declare the summary you want, get it back as one table.
Run df.agg({'orders': 'sum', 'amount': 'mean'}) and the last line of output reads dtype: float64 — a summary table built of nothing but NaN.
agg() applies one or more aggregation functions and returns the collapsed result: one value per function. Pass a string ('sum', 'mean', 'median', 'count', 'nunique'), a callable, a list mixing both, or a dict mapping column names to any of these. A list gives a function-by-function summary across every column; a dict gives exactly what you asked per column, nothing else. And the same spec language runs identically on groupby, resample, and rolling objects — learn one syntax, use it in every aggregation context pandas has.
describe() gives you every stat including the ones you don't want and no custom ones. Stacking manual df[col].sum() lines gives you one stat per line and no coherent table. agg() is the middle path: declare exactly what you want and return it as a clean object that drops into a report, a plot, or a groupby. The groupby variant is where it earns its keep: three lines of Python replacing the GROUP BY clause of the SQL report you were about to write.
sales = pd.DataFrame({
"region": ["west", "west", "east", "east", "west", "north", "east", "west"],
"city": ["Rotterdam", "Utrecht", "Berlin", "Hamburg", "Antwerp", "Bremen", "Dresden", "Gent"],
"orders": [412, 87, 190, 45, 333, 128, 76, 210],
"amount": [89.50, 12.00, 130.25, 8.75, 64.10, 24.80, 11.30, 45.60],
})
sales.agg({"orders": "sum", "amount": ["mean", "max"]})orders amount sum 1481.0 NaN mean NaN 48.2875 max NaN 130.2500
Dict layout, function-by-function: 1481.0 sits in the sum row under orders, 48.2875 in the mean row under amount. Every square a function+column pair didn't ask for is NaN — read dict output row-wise, not column-wise.
sales.groupby("region").agg({"orders": "sum", "amount": "mean"})orders amount region east 311 50.1 north 128 24.8 west 1042 52.8
The real workhorse: same dict spec, per group. Three lines of Python replacing the GROUP BY every SQL report starts with.
sales.groupby("region").agg(
total_orders=("orders", "sum"),
avg_amount=("amount", "mean"),
big_orders=("orders", lambda s: (s > 100).sum()),
)total_orders avg_amount big_orders region east 311 50.1 1 north 128 24.8 1 west 1042 52.8 3
Named aggregation, pandas 0.25 (2019): keyword = output name, value = (column, function). Clean column names, and the lambda counts orders over 100 — a custom stat no built-in string can give you.
sales["orders"].agg(["sum", "median", "max"])
sum 1481.0 median 159.0 max 412.0 Name: orders, dtype: float64
Series variant: a list of strings builds a summary Series in one call, no loop, no stacking df[col].sum() lines.
| Flag | Meaning |
|---|---|
df.agg('sum') | single string — one aggregation applied down every column |
df.agg(['sum', 'median']) | list: a summary across columns, one block row per function |
df.agg({'orders': 'sum'}) | dict: exactly the columns you name — but read it row-wise, per function |
df.agg({'orders': ['sum', 'median']}) | dict of lists: several stats per chosen column |
df.agg('sum', axis=1) | axis=1 aggregates across each row — row totals in one call |
df.groupby('region').agg(...) | the workhorse: same spec, one summary row per group |
df.agg(total=('col', 'mean')) | named aggregation: (column, function) tuple renamed to the keyword |
0.20.0 shipped the ".agg() API" for Series and DataFrame — promoting the aggregation protocol groupby, rolling, and resample already spoke to top-level objects (issue 1623). Since then, one spec style covers whole-table, per-group, per-window, and per-resample summaries: pass the same dict to DataFrame.agg and groupby.agg and both understand it. Aggregate is a strict alias — same function, longer name.
Pandas 0.20 deprecated the dict-of-dicts renaming trick, and pandas 1.0.0 (January 2020) removed it (issue 29608). If an old post shows {'amount': {'avg': 'mean'}}, that door is sealed — named aggregation, added in 0.25 (2019), is the documented replacement.
The pandas implementation normalizes the spec and dispatches: strings resolve through the internal dispatch table ('sum' routes to the C-optimized paths, 'nunique' to factorize+hashtable), and callables run through the same machinery as apply — per column for a frame, per group for a groupby (issue 1623's design: one protocol, many hosts). No numexpr, no fused kernel: a dict spec is looped key by key. Heavy multi-stat summaries on wide frames are the case to profile, since every function is an independent pass.