pandas unique() — distinct values without sorting them

distinct values in order of appearance, NA included, nothing sorted — the category inventory, not the sorted report.

np.unique hands you a sorted alphabet; pandas unique hands you the alphabet in the order your file listed it — and that difference is the whole feature.

What it does

s.unique() hashes a Series and returns each distinct value once, in order of first appearance — no sorting, and NA counts as a plain category. The output type tracks the input: a modern str column hands back the StringArray block a notebook displays, a datetime Series returns an ndarray, Categorical in means Categorical out. The free function pd.unique(values) is the same machine that also takes a raw ndarray or Index. When all you need is the distinct set, it deliberately does less work than value_counts(), which hashes the values AND counts them.

Why it matters

I run it in three loops: inventory of categories (which regions does this file even contain?), validity checks (does the column hold exactly what the spec allows?), and report loops (distinct values straight into a for loop for per-category plots or one CSV per region). It is the read side of the categorical story — see what exists, then astype('category') to store it compactly once. And because it is hash-based and appearance-ordered, it is the only “distinct” that answers both what exists and the order it first showed up in — which happens to be the order a stakeholder wants the legend printed in.

Examples

import pandas as pd

orders = pd.DataFrame({
    'order_id': [5001, 5002, 5003, 5004, 5005],
    'region':   ['EMEA', 'APAC', 'EMEA', 'LATAM', 'APAC'],
    'amount':   [1200.0, 340.5, 780.0, 99.9, 340.5],
})

orders['region'].unique()
<StringArray>
['EMEA', 'APAC', 'LATAM ']
Length: 3, dtype: str

Verified on pandas 3.0.3. Two traps live in one block: uniques come back in appearance order (EMEA first — np.unique would print ['APAC', 'EMEA', 'LATAM']), and the StringArray display pads the shortest entry to the widest one, so LATAM prints trailing-spaced: it is display cosmetics, not data.

# "Web", "WEB " and "web" are three categories until you normalize.
tickets = pd.Series(['web', 'Web', 'WEB ', 'web', 'phone'], name='channel')

tickets.unique()
<StringArray>
['web', 'Web', 'WEB ', 'phone']
Length: 4, dtype: str

Verified on 3.0.3: a capital letter and a trailing space make genuinely distinct values. This is the raw material for the cleanup idiom — tickets.str.strip().str.lower().unique() then returns just ['web', 'phone'].

# Does the region column hold exactly what the spec allows?
observed = orders['region'].unique()
allowed  = {'EMEA', 'APAC', 'LATAM'}

set(observed) == allowed
True

The load-gate pattern: wrap in a set and compare — appearance order stops mattering. unique() plus a set comparison is my one-line schema check before any file gets ingested; one misspelled 'EMEA ' shows up instantly.

responses = pd.Series(['yes', 'no', None, 'no'], name='answer')

responses.unique()
<StringArray>
['yes', 'no', <NA>]
Length: 3, dtype: str

Verified on 3.0.3: NA is included, so unique() reports missing as its own category — np.unique would silently drop it. For “distinct non-null”, count with s.nunique() (dropna=True by default) or .dropna().unique().

c = pd.Series(pd.Categorical(['m', 'l', 'm', 'l', 's'], categories=['s', 'm', 'l']))

c.unique()
['m', 'l', 's']
Categories (3, str): ['s', 'm', 'l']

Categorical stays Categorical: values print in appearance order, while the Categories line keeps the (s, m, l) order — the sorted category system is untouched. Use .unique() to see what occurs, .cat.categories to see what is allowed.

Flags

FlagMeaning
order of appearancevalues come back first-seen-first; the docstring says so explicitly: “This does NOT sort.”
NA is a categoryNone / np.nan / pd.NA is a distinct entry; np.unique() sorts it away, and float nan still lands at the end.
output dtype tracks inputstr in → StringArray block on 3.0.3, numeric → plain ndarray, datetime → datetime64[us] ndarray, Categorical in → Categorical out.
pd.unique(values)the top-level twin also takes a raw ndarray or Index; Index in → Index out (Series.unique returns ndarray/ExtensionArray).
works on Categoricalreturns a Categorical holding the seen categories in appearance order — the category system itself is preserved.
s.nunique()the count-only twin: returns just the number of distinct values, NA excluded by default (dropna=True).
df.drop_duplicates()row-level dedupe on one or more columns; for distinct values per group, groupby('g')['c'].unique().

Shipped in pandas 0.7.0 (February 2012), hash tables from day one

I diffed the old source trees: v0.6.1's series.py has no def unique at all; v0.7.0 ships one whose entire body is return nanops.unique1d(self.values), with the docstring already bragging “Significantly faster than numpy.unique”. Just like value_counts(), it was a pure-Python dispatch (nanops.unique1d) that immediately handed the work to the Cython hashtables — lib.Float64HashTable, Int64HashTable, PyObjectHashTable — the khash machinery it still uses today. nunique() (then literally len(value_counts())) arrived in the very same release.

The free function arrived in 0.8.0, and the return type became law

pandas 0.8.0 (June 2012) added the top-level def unique(values) in pandas/core/algorithms.py — “Compute unique values (not necessarily sorted) efficiently” — so raw ndarrays and Index objects got the same hash-table treatment. Its return-type contract is quirky and ancient: Index in gives Index out, Categorical in gives Categorical out, a Series or ndarray gives an ndarray/ExtensionArray. On pandas 3.0.3 the Series method is a one-liner over the shared NDFrame base (return super().unique()) and the engine underneath is unique_with_mask(values) — same name, same promise, same tables.

Under the hood: one hash table, no sort

pd.unique dispatches over the dtype — khash for numeric/object, a different code path for categorical/object — and the whole job is a single hash-table walk that remembers which values it has already seen. Nothing sorts anything; that is the entire difference to np.unique. The same khash machinery powers value_counts(), factorize(), and Index.unique(), which is why they all agree on order-of-appearance semantics. The docstring's “Significantly faster than numpy.unique” line exists because hash tables are O(n) while numpy pays O(n log n) for the sort plus a pass to collapse runs.

Fun facts

Pros

Cons

Takeaways