pandas nunique() — count distinct values per column

value_counts() shows you the breakdown; nunique() counts the distinct values — usually the only thing the report actually asked for.

A million-row Series with a few thousand distinct values answers in milliseconds — where does the distinct-count number come from that fast?

What it does

nunique() counts distinct values: the DataFrame version gives one number per column, the Series version one number. It hashes every element into int64 codes and deduplicates them with the same factorize machinery that powers groupby internals, then returns the code count — NaN is skipped unless dropna=False. Because it's a hash-based pass, cost tracks distinct values, not rows, so big low-cardinality columns answer almost instantly.

Why it matters

Every profiling pass starts with the same questions — how many distinct regions, how many customers, is this join key really unique — and nunique() is the cheapest honest answer. If nunique() == len(df), your key has no duplicates; if it isn't, you found the bug before the join multiplied rows. It also separates a genuinely single-valued column from one that merely looks empty. With dropna=False the null itself becomes a countable category, so missingness is part of the same pass. And pandas 0.20 shipped nunique on DataFrame and groupby together, so one method name covers whole-table, per-column, and per-group — no set() loops over rows, no value_counts() size gymnastics.

Examples

import pandas as pd

visits = pd.DataFrame({
    "page": ["/", "/pricing", "/", "/blog", "/pricing", "/", "/blog", "/blog"],
    "city": ["Gent", "Antwerp", "Gent", "Berlin", "Antwerp", "Gent", None, "Vienna"],
    "seconds": [12, 95, 8, 210, 77, 5, 143, 61],
})
visits.nunique()
page       3
city       4
seconds    8
dtype: int64

Per-column distinct counts in one call. city counts 4 because NaN is skipped by default. This is the check you run before any groupby or join.

visits.nunique(dropna=False)
page       3
city       5
seconds    8
dtype: int64

dropna=False counts None as its own category — city goes from 4 to 5. The gap between the two runs is exactly your missing-data count in this example: one NaN.

sites = pd.DataFrame({
    "site": ["A", "B", "A", "A", "B", "A", "A", "B"],
    "reading": [19.9, 20.3, 20.1, 19.9, 20.4, 20.1, 20.6, 20.3],
})
sites.groupby("site").nunique()
reading
site
A          3
B          2

Distinct readings grouped by site — the sensor-health check. Internally this is SeriesGroupBy.nunique, the groupby sibling added in the same release as the DataFrame method.

sites.groupby("site")["reading"].transform("nunique")
0    3
1    2
2    3
3    3
4    2
5    3
6    3
7    2
Name: reading, dtype: int64

transform broadcasts each group's distinct count back to every row — a per-row group trait you can feed into queries or merges without changing the frame's shape.

Flags

FlagMeaning
df.nunique()distinct count per column, NaN skipped — the quick profile pass
df.nunique(dropna=False)count NaN as its own category — missingness joins the profile
df['col'].nunique()Series form: one number for one column
df.groupby('key').nunique()distinct counts per group, one row per group
df.groupby('key')['col'].transform('nunique')group distinct count broadcast back to every row
df['col'].agg('nunique')same machinery through the agg() spec language

Series got it first (0.13), DataFrame and groupby followed together in 0.20.0 (May 5, 2017)

Series.nunique existed since pandas 0.13 (January 2014), but the DataFrame method and its groupby sibling landed in the same release: 0.20.0, tagged May 5, 2017, announced in one sentence — "DataFrame and DataFrame.groupby() have gained a nunique() method to count the distinct values over an axis" (GH 14336, GH 15197). Even the 1.0.0 migration guide still points people to Series.nunique() as the replacement for the removed unique().nunique() workaround. So the version you cite depends on what you're calling it on: Series 2014, DataFrame and groupby 2017.

Design note: describe() has a 'unique' row, and it is not nunique

describe() on object columns returns a 'unique' entry, which reads like nunique but is part of describe's own counting layout — don't assume the two agree when NaNs are involved, because nunique skips NaN by default and describe's object summary is built by a different path. The dedicated distinct-count method is nunique, with explicit dropna control. Same word in both outputs, two different questions — trust the one whose signature shows a dropna flag.

Under the hood: factorize + hash tables, not a Python set

pandas runs nunique through factorize-style hashing: elements go into int64 codes via the internal hash-table kernels (separate object, integer, and float tables), and the distinct count is derived from the code table with the NaN mask handled explicitly. That's why a million-row, low-cardinality Series costs about the same as a small one — the hot loop is C hashing over rows, and the answer falls out of the codes, not a Python set. Related trivia from the changelog: 0.16.0 (2015) already notes an nunique speedup from 'calling unique instead of value_counts', and pandas 3.0 (2026) adds a dedicated fast SeriesGroupBy.transform('nunique') kernel.

Fun facts

Pros

Cons

Takeaways