pandas agg() — build any summary table in one call

One .agg() call replaces the five stacked df[col].sum() lines — declare the summary you want, get it back as one table.

Run df.agg({'orders': 'sum', 'amount': 'mean'}) and the last line of output reads dtype: float64 — a summary table built of nothing but NaN.

What it does

agg() applies one or more aggregation functions and returns the collapsed result: one value per function. Pass a string ('sum', 'mean', 'median', 'count', 'nunique'), a callable, a list mixing both, or a dict mapping column names to any of these. A list gives a function-by-function summary across every column; a dict gives exactly what you asked per column, nothing else. And the same spec language runs identically on groupby, resample, and rolling objects — learn one syntax, use it in every aggregation context pandas has.

Why it matters

describe() gives you every stat including the ones you don't want and no custom ones. Stacking manual df[col].sum() lines gives you one stat per line and no coherent table. agg() is the middle path: declare exactly what you want and return it as a clean object that drops into a report, a plot, or a groupby. The groupby variant is where it earns its keep: three lines of Python replacing the GROUP BY clause of the SQL report you were about to write.

Examples

sales = pd.DataFrame({
    "region": ["west", "west", "east", "east", "west", "north", "east", "west"],
    "city": ["Rotterdam", "Utrecht", "Berlin", "Hamburg", "Antwerp", "Bremen", "Dresden", "Gent"],
    "orders": [412, 87, 190, 45, 333, 128, 76, 210],
    "amount": [89.50, 12.00, 130.25, 8.75, 64.10, 24.80, 11.30, 45.60],
})
sales.agg({"orders": "sum", "amount": ["mean", "max"]})
      orders    amount
sum   1481.0       NaN
mean     NaN   48.2875
max      NaN   130.2500

Dict layout, function-by-function: 1481.0 sits in the sum row under orders, 48.2875 in the mean row under amount. Every square a function+column pair didn't ask for is NaN — read dict output row-wise, not column-wise.

sales.groupby("region").agg({"orders": "sum", "amount": "mean"})
        orders  amount
region                
east       311    50.1
north      128    24.8
west      1042    52.8

The real workhorse: same dict spec, per group. Three lines of Python replacing the GROUP BY every SQL report starts with.

sales.groupby("region").agg(
    total_orders=("orders", "sum"),
    avg_amount=("amount", "mean"),
    big_orders=("orders", lambda s: (s > 100).sum()),
)
        total_orders  avg_amount  big_orders
region                                        
east             311        50.1           1
north            128        24.8           1
west            1042        52.8           3

Named aggregation, pandas 0.25 (2019): keyword = output name, value = (column, function). Clean column names, and the lambda counts orders over 100 — a custom stat no built-in string can give you.

sales["orders"].agg(["sum", "median", "max"])
sum      1481.0
median    159.0
max       412.0
Name: orders, dtype: float64

Series variant: a list of strings builds a summary Series in one call, no loop, no stacking df[col].sum() lines.

Flags

FlagMeaning
df.agg('sum')single string — one aggregation applied down every column
df.agg(['sum', 'median'])list: a summary across columns, one block row per function
df.agg({'orders': 'sum'})dict: exactly the columns you name — but read it row-wise, per function
df.agg({'orders': ['sum', 'median']})dict of lists: several stats per chosen column
df.agg('sum', axis=1)axis=1 aggregates across each row — row totals in one call
df.groupby('region').agg(...)the workhorse: same spec, one summary row per group
df.agg(total=('col', 'mean'))named aggregation: (column, function) tuple renamed to the keyword

agg() arrived in pandas 0.20.0 (May 2017)

0.20.0 shipped the ".agg() API" for Series and DataFrame — promoting the aggregation protocol groupby, rolling, and resample already spoke to top-level objects (issue 1623). Since then, one spec style covers whole-table, per-group, per-window, and per-resample summaries: pass the same dict to DataFrame.agg and groupby.agg and both understand it. Aggregate is a strict alias — same function, longer name.

Design note: nested renaming is dead, named aggregation took over

Pandas 0.20 deprecated the dict-of-dicts renaming trick, and pandas 1.0.0 (January 2020) removed it (issue 29608). If an old post shows {'amount': {'avg': 'mean'}}, that door is sealed — named aggregation, added in 0.25 (2019), is the documented replacement.

Under the hood: spec normalization into apply/gb.agg

The pandas implementation normalizes the spec and dispatches: strings resolve through the internal dispatch table ('sum' routes to the C-optimized paths, 'nunique' to factorize+hashtable), and callables run through the same machinery as apply — per column for a frame, per group for a groupby (issue 1623's design: one protocol, many hosts). No numexpr, no fused kernel: a dict spec is looped key by key. Heavy multi-stat summaries on wide frames are the case to profile, since every function is an independent pass.

Fun facts

Pros

Cons

Takeaways