kmail.at
← learning

pandas · difficulty ◆◆

pandas transform() — per-group results in the original row shape

Aggregation collapses your rows; transform keeps every one of them and hands each row its group's number.

Aggregation throws your rows away and makes you merge them back. transform() is the line that skips the merge.

2026-10-04 · 5 min read

$ df.groupby().transform()

What it does

transform() runs a function per group and returns a result with the same index and same length as the input — so it drops straight into a new column. Give it a string kernel like 'mean', 'size' or 'cumsum' and pandas computes the group value (or running value) once and broadcasts it across the group; give it a lambda and it calls your function per group, then reassembles by index. On a DataFrame, transform('mean') applies column by column and you get a frame back, not a scalar; df.transform(f, axis=1) is the row-wise variant. The contract is strict: your function must return something the same shape as the group, or a scalar it can broadcast.

Why it matters

Most real reporting needs the group's number attached to each row, not a summary table: revenue as a share of the region's total, this reading versus its sensor's baseline, this ticket flagged because its team has a breach. The old way was groupby().mean() then merge() back on the key — more code, more chances to get the join wrong, and a full copy of the frame in memory. transform() does it in one assignment, preserves row order, and works with non-unique indexes. Once this clicks, groupby stops being 'how I summarise' and becomes 'how I enrich'.

Example

$ import pandas as pd
import numpy as np

sales = pd.DataFrame({
    "region": ["Yanbu", "Yanbu", "Jubail", "Jubail", "Yanbu", "Jubail"],
    "month":  ["Jan", "Feb", "Jan", "Feb", "Mar", "Mar"],
    "revenue": [120.0, np.nan, 90.0, np.nan, 150.0, 110.0],
})

sales["revenue"] = sales.groupby("region")["revenue"].transform(
    lambda s: s.fillna(s.mean())
)
print(sales.to_string())
   region month  revenue
0   Yanbu   Jan    120.0
1   Yanbu   Feb    135.0
2  Jubail   Jan     90.0
3  Jubail   Feb    100.0
4   Yanbu   Mar    150.0
5  Jubail   Mar    110.0

The everyday cleaning case: fill a missing reading with that group's own average, not the global one. (120+150)/2 = 135.0 for Yanbu, (90+110)/2 = 100.0 for Jubail. A plain df.revenue.fillna(df.revenue.mean()) would have written the same number into both rows and quietly misstated one of them.

$ df = pd.DataFrame({
    "team":  ["Alpha", "Alpha", "Beta", "Beta", "Beta"],
    "rep":   ["A1", "A2", "B1", "B2", "B3"],
    "score": [80, 100, 60, 70, 90],
})

grp = df.groupby("team")["score"]
df["z"] = (df["score"] - grp.transform("mean")) / grp.transform("std")
print(df[["team", "rep", "score", "z"]].round(3).to_string())
    team rep  score      z
0  Alpha  A1     80 -0.707
1  Alpha  A2    100  0.707
2   Beta  B1     60 -0.873
3   Beta  B2     70 -0.218
4   Beta  B3     90  1.091

Within-group z-score in two calls, no groupby object to juggle. Reusing grp keeps both transforms tied to the same grouping, and the result lines up with df row for row, so the subtraction just works. transform('std') is sample std (ddof=1) — for Beta, (90-73.333)/19.438 = 1.091.

$ orders = pd.DataFrame({
    "team":  ["Alpha", "Alpha", "Beta", "Beta"],
    "order": ["O1", "O2", "O3", "O4"],
    "late":  [False, True, False, False],
})

orders["team_has_late"] = orders.groupby("team")["late"].transform("any")
print(orders.to_string())
    team order   late  team_has_late
0  Alpha    O1  False           True
1  Alpha    O2   True           True
2   Beta    O3  False          False
3   Beta    O4  False          False

Broadcast a group-level boolean onto every member row — the shape of 'every order this team touched is under review'. This is where agg() cannot help: agg('any') gives two rows (Alpha, Beta) and you are back to merging. Any reduction kernel works as a string kernel here: 'sum', 'max', 'first', 'nunique'.

$ budget = pd.DataFrame({
    "plant": ["Yanbu", "Jubail"],
    "maint": [30.0, 10.0],
    "ops":   [20.0, 30.0],
    "capex": [50.0, 60.0],
})

spend = ["maint", "ops", "capex"]
budget[spend] = budget[spend].transform(lambda r: r / r.sum(), axis=1)
print(budget.round(3).to_string())
    plant  maint  ops  capex
0   Yanbu    0.3  0.2    0.5
1  Jubail    0.1  0.3    0.6

DataFrame.transform() with no groupby — axis=1 runs the function per row, giving each plant's spend mix as a share of its own total. Yanbu: 30/100, 20/100, 50/100. This is the axis that catches people: the grouped form takes no axis at all, so groupby().transform(f, axis=1) raises TypeError.

$ g = pd.DataFrame({"team": ["A", "A", "B"], "v": [1, 3, 10]})

print(g.groupby("team")["v"].transform("size").tolist())
print(g.groupby("team")["v"].apply("size").tolist())
print(g.groupby("team")["v"].transform(lambda s: pd.Series([1, 2])).tolist())
[2, 2, 1]
[2, 1]
[1.0, 2.0, nan]

Three answers, one lesson. transform('size') broadcasts the group count onto each row — 3 rows in, 3 values out. apply('size') gives one row per group. And that last line is the trap: a callable returning the wrong length does not raise, it index-aligns and pads with NaN. If a column of your output is mysteriously NaN, count what your lambda returns.

Common flags

func
a kernel name ('mean', 'sum', 'cumsum', 'ffill', 'rank' …) or a callable. Strings hit pandas' fast paths; callables take the general per-group path
engine='numba'
JIT-compile your callable instead of using the Cython path — the function then receives (values, index) as its first two arguments
axis
DataFrame.transform only: 0 = each column (default), 1 = each row. The grouped version accepts no axis at all
'size' vs 'count'
size counts rows including NaN; count counts only non-null. transform('size') is the cheapest 'how big was this group' column
Same-shape contract
the function must return the group's shape or a scalar to broadcast. Wrong length = silent NaN padding, not an error
Assignment idiom
df['new'] = df.groupby(key)['col'].transform(...) — index-aligned, so it lands in row order with no merge and no sort
Non-unique index
transform works on frames with duplicated index labels — it reassembles by position, unlike a join on the key

History

0.4.0, September 2011 — the method arrives

transform() first appears in pandas/core/groupby.py at the 0.4.0 tag (released 12 September 2011), the same release that let GroupBy aggregations broadcast instead of collapse. Its docstring then was a single paragraph — 'Call function producing a like-indexed Series on each group and return a Series with the transformed values' — and the lone example was the z-score one. That example still runs unchanged today.

1.1.0 and 1.3.0 — engines and dtype honesty

1.1.0 added engine and engine_kwargs to the groupby transform, letting a Numba-jitted function replace the Cython path (GH 32854, GH 33388). 1.3.0 fixed a quieter problem: transform used to cast the result back to the input dtype whenever the values still matched under np.allclose, so a transform that legitimately produced floats from integer input could come back as ints. Since 1.3.0 no such casting happens (GH 21240) — the dtype you compute is the dtype you get.

Fun facts

Pros & cons

pros

  • + Enrichment without a merge: one assignment, original index and row order preserved
  • + String kernels ride the Cython paths, so transform('mean') pays no per-row Python overhead
  • + Expresses ratios, deviations, rankings and flags as ordinary column assignments — readable to anyone who knows pandas

cons

  • − Callables run a per-group Python loop, which is the slow path when you have millions of groups
  • − Silent NaN padding on a wrong-length return hides the real bug until you go hunting for missing data
  • − The same method name means three different things across Series/DataFrame and grouped/ungrouped objects, and only the DataFrame form takes axis

Takeaways

  1. 1Need the group's number on every row? Assign it: df['x'] = df.groupby(key)['v'].transform(...) — never groupby + merge.
  2. 2Prefer a string kernel when one exists; reach for a lambda only when the logic genuinely is not one of the 34 allowlisted names.
  3. 3transform preserves length and index — if your result has fewer rows than you started with, you wanted agg().
  4. 4Select numeric columns first on mixed dtypes, or the string column gets the same kernel and raises TypeError.
  5. 5For row-wise maths on an ungrouped frame use df.transform(f, axis=1); the grouped form has no axis.

Related commands

← all learning