pandas · difficulty ◆◆
pandas transform() — per-group results in the original row shape
Aggregation collapses your rows; transform keeps every one of them and hands each row its group's number.
Aggregation throws your rows away and makes you merge them back. transform() is the line that skips the merge.
$ df.groupby().transform()What it does
transform() runs a function per group and returns a result with the same index and same length as the input — so it drops straight into a new column. Give it a string kernel like 'mean', 'size' or 'cumsum' and pandas computes the group value (or running value) once and broadcasts it across the group; give it a lambda and it calls your function per group, then reassembles by index. On a DataFrame, transform('mean') applies column by column and you get a frame back, not a scalar; df.transform(f, axis=1) is the row-wise variant. The contract is strict: your function must return something the same shape as the group, or a scalar it can broadcast.
Why it matters
Most real reporting needs the group's number attached to each row, not a summary table: revenue as a share of the region's total, this reading versus its sensor's baseline, this ticket flagged because its team has a breach. The old way was groupby().mean() then merge() back on the key — more code, more chances to get the join wrong, and a full copy of the frame in memory. transform() does it in one assignment, preserves row order, and works with non-unique indexes. Once this clicks, groupby stops being 'how I summarise' and becomes 'how I enrich'.
Example
$ import pandas as pd
import numpy as np
sales = pd.DataFrame({
"region": ["Yanbu", "Yanbu", "Jubail", "Jubail", "Yanbu", "Jubail"],
"month": ["Jan", "Feb", "Jan", "Feb", "Mar", "Mar"],
"revenue": [120.0, np.nan, 90.0, np.nan, 150.0, 110.0],
})
sales["revenue"] = sales.groupby("region")["revenue"].transform(
lambda s: s.fillna(s.mean())
)
print(sales.to_string()) region month revenue
0 Yanbu Jan 120.0
1 Yanbu Feb 135.0
2 Jubail Jan 90.0
3 Jubail Feb 100.0
4 Yanbu Mar 150.0
5 Jubail Mar 110.0The everyday cleaning case: fill a missing reading with that group's own average, not the global one. (120+150)/2 = 135.0 for Yanbu, (90+110)/2 = 100.0 for Jubail. A plain df.revenue.fillna(df.revenue.mean()) would have written the same number into both rows and quietly misstated one of them.
$ df = pd.DataFrame({
"team": ["Alpha", "Alpha", "Beta", "Beta", "Beta"],
"rep": ["A1", "A2", "B1", "B2", "B3"],
"score": [80, 100, 60, 70, 90],
})
grp = df.groupby("team")["score"]
df["z"] = (df["score"] - grp.transform("mean")) / grp.transform("std")
print(df[["team", "rep", "score", "z"]].round(3).to_string()) team rep score z
0 Alpha A1 80 -0.707
1 Alpha A2 100 0.707
2 Beta B1 60 -0.873
3 Beta B2 70 -0.218
4 Beta B3 90 1.091Within-group z-score in two calls, no groupby object to juggle. Reusing grp keeps both transforms tied to the same grouping, and the result lines up with df row for row, so the subtraction just works. transform('std') is sample std (ddof=1) — for Beta, (90-73.333)/19.438 = 1.091.
$ orders = pd.DataFrame({
"team": ["Alpha", "Alpha", "Beta", "Beta"],
"order": ["O1", "O2", "O3", "O4"],
"late": [False, True, False, False],
})
orders["team_has_late"] = orders.groupby("team")["late"].transform("any")
print(orders.to_string()) team order late team_has_late
0 Alpha O1 False True
1 Alpha O2 True True
2 Beta O3 False False
3 Beta O4 False FalseBroadcast a group-level boolean onto every member row — the shape of 'every order this team touched is under review'. This is where agg() cannot help: agg('any') gives two rows (Alpha, Beta) and you are back to merging. Any reduction kernel works as a string kernel here: 'sum', 'max', 'first', 'nunique'.
$ budget = pd.DataFrame({
"plant": ["Yanbu", "Jubail"],
"maint": [30.0, 10.0],
"ops": [20.0, 30.0],
"capex": [50.0, 60.0],
})
spend = ["maint", "ops", "capex"]
budget[spend] = budget[spend].transform(lambda r: r / r.sum(), axis=1)
print(budget.round(3).to_string()) plant maint ops capex
0 Yanbu 0.3 0.2 0.5
1 Jubail 0.1 0.3 0.6DataFrame.transform() with no groupby — axis=1 runs the function per row, giving each plant's spend mix as a share of its own total. Yanbu: 30/100, 20/100, 50/100. This is the axis that catches people: the grouped form takes no axis at all, so groupby().transform(f, axis=1) raises TypeError.
$ g = pd.DataFrame({"team": ["A", "A", "B"], "v": [1, 3, 10]})
print(g.groupby("team")["v"].transform("size").tolist())
print(g.groupby("team")["v"].apply("size").tolist())
print(g.groupby("team")["v"].transform(lambda s: pd.Series([1, 2])).tolist())[2, 2, 1]
[2, 1]
[1.0, 2.0, nan]Three answers, one lesson. transform('size') broadcasts the group count onto each row — 3 rows in, 3 values out. apply('size') gives one row per group. And that last line is the trap: a callable returning the wrong length does not raise, it index-aligns and pads with NaN. If a column of your output is mysteriously NaN, count what your lambda returns.
Common flags
- func
- a kernel name ('mean', 'sum', 'cumsum', 'ffill', 'rank' …) or a callable. Strings hit pandas' fast paths; callables take the general per-group path
- engine='numba'
- JIT-compile your callable instead of using the Cython path — the function then receives (values, index) as its first two arguments
- axis
- DataFrame.transform only: 0 = each column (default), 1 = each row. The grouped version accepts no axis at all
- 'size' vs 'count'
- size counts rows including NaN; count counts only non-null. transform('size') is the cheapest 'how big was this group' column
- Same-shape contract
- the function must return the group's shape or a scalar to broadcast. Wrong length = silent NaN padding, not an error
- Assignment idiom
- df['new'] = df.groupby(key)['col'].transform(...) — index-aligned, so it lands in row order with no merge and no sort
- Non-unique index
- transform works on frames with duplicated index labels — it reassembles by position, unlike a join on the key
History
0.4.0, September 2011 — the method arrives
transform() first appears in pandas/core/groupby.py at the 0.4.0 tag (released 12 September 2011), the same release that let GroupBy aggregations broadcast instead of collapse. Its docstring then was a single paragraph — 'Call function producing a like-indexed Series on each group and return a Series with the transformed values' — and the lone example was the z-score one. That example still runs unchanged today.
1.1.0 and 1.3.0 — engines and dtype honesty
1.1.0 added engine and engine_kwargs to the groupby transform, letting a Numba-jitted function replace the Cython path (GH 32854, GH 33388). 1.3.0 fixed a quieter problem: transform used to cast the result back to the input dtype whenever the values still matched under np.allclose, so a transform that legitimately produced floats from integer input could come back as ints. Since 1.3.0 no such casting happens (GH 21240) — the dtype you compute is the dtype you get.
Fun facts
Pros & cons
pros
- + Enrichment without a merge: one assignment, original index and row order preserved
- + String kernels ride the Cython paths, so transform('mean') pays no per-row Python overhead
- + Expresses ratios, deviations, rankings and flags as ordinary column assignments — readable to anyone who knows pandas
cons
- − Callables run a per-group Python loop, which is the slow path when you have millions of groups
- − Silent NaN padding on a wrong-length return hides the real bug until you go hunting for missing data
- − The same method name means three different things across Series/DataFrame and grouped/ungrouped objects, and only the DataFrame form takes axis
Takeaways
- 1Need the group's number on every row? Assign it: df['x'] = df.groupby(key)['v'].transform(...) — never groupby + merge.
- 2Prefer a string kernel when one exists; reach for a lambda only when the logic genuinely is not one of the 34 allowlisted names.
- 3transform preserves length and index — if your result has fewer rows than you started with, you wanted agg().
- 4Select numeric columns first on mixed dtypes, or the string column gets the same kernel and raises TypeError.
- 5For row-wise maths on an ungrouped frame use df.transform(f, axis=1); the grouped form has no axis.