distinct values in order of appearance, NA included, nothing sorted — the category inventory, not the sorted report.
np.unique hands you a sorted alphabet; pandas unique hands you the alphabet in the order your file listed it — and that difference is the whole feature.
s.unique() hashes a Series and returns each distinct value once, in order of first appearance — no sorting, and NA counts as a plain category. The output type tracks the input: a modern str column hands back the StringArray block a notebook displays, a datetime Series returns an ndarray, Categorical in means Categorical out. The free function pd.unique(values) is the same machine that also takes a raw ndarray or Index. When all you need is the distinct set, it deliberately does less work than value_counts(), which hashes the values AND counts them.
I run it in three loops: inventory of categories (which regions does this file even contain?), validity checks (does the column hold exactly what the spec allows?), and report loops (distinct values straight into a for loop for per-category plots or one CSV per region). It is the read side of the categorical story — see what exists, then astype('category') to store it compactly once. And because it is hash-based and appearance-ordered, it is the only “distinct” that answers both what exists and the order it first showed up in — which happens to be the order a stakeholder wants the legend printed in.
import pandas as pd
orders = pd.DataFrame({
'order_id': [5001, 5002, 5003, 5004, 5005],
'region': ['EMEA', 'APAC', 'EMEA', 'LATAM', 'APAC'],
'amount': [1200.0, 340.5, 780.0, 99.9, 340.5],
})
orders['region'].unique()<StringArray> ['EMEA', 'APAC', 'LATAM '] Length: 3, dtype: str
Verified on pandas 3.0.3. Two traps live in one block: uniques come back in appearance order (EMEA first — np.unique would print ['APAC', 'EMEA', 'LATAM']), and the StringArray display pads the shortest entry to the widest one, so LATAM prints trailing-spaced: it is display cosmetics, not data.
# "Web", "WEB " and "web" are three categories until you normalize. tickets = pd.Series(['web', 'Web', 'WEB ', 'web', 'phone'], name='channel') tickets.unique()
<StringArray> ['web', 'Web', 'WEB ', 'phone'] Length: 4, dtype: str
Verified on 3.0.3: a capital letter and a trailing space make genuinely distinct values. This is the raw material for the cleanup idiom — tickets.str.strip().str.lower().unique() then returns just ['web', 'phone'].
# Does the region column hold exactly what the spec allows?
observed = orders['region'].unique()
allowed = {'EMEA', 'APAC', 'LATAM'}
set(observed) == allowedTrue
The load-gate pattern: wrap in a set and compare — appearance order stops mattering. unique() plus a set comparison is my one-line schema check before any file gets ingested; one misspelled 'EMEA ' shows up instantly.
responses = pd.Series(['yes', 'no', None, 'no'], name='answer') responses.unique()
<StringArray> ['yes', 'no', <NA>] Length: 3, dtype: str
Verified on 3.0.3: NA is included, so unique() reports missing as its own category — np.unique would silently drop it. For “distinct non-null”, count with s.nunique() (dropna=True by default) or .dropna().unique().
c = pd.Series(pd.Categorical(['m', 'l', 'm', 'l', 's'], categories=['s', 'm', 'l'])) c.unique()
['m', 'l', 's'] Categories (3, str): ['s', 'm', 'l']
Categorical stays Categorical: values print in appearance order, while the Categories line keeps the (s, m, l) order — the sorted category system is untouched. Use .unique() to see what occurs, .cat.categories to see what is allowed.
| Flag | Meaning |
|---|---|
order of appearance | values come back first-seen-first; the docstring says so explicitly: “This does NOT sort.” |
NA is a category | None / np.nan / pd.NA is a distinct entry; np.unique() sorts it away, and float nan still lands at the end. |
output dtype tracks input | str in → StringArray block on 3.0.3, numeric → plain ndarray, datetime → datetime64[us] ndarray, Categorical in → Categorical out. |
pd.unique(values) | the top-level twin also takes a raw ndarray or Index; Index in → Index out (Series.unique returns ndarray/ExtensionArray). |
works on Categorical | returns a Categorical holding the seen categories in appearance order — the category system itself is preserved. |
s.nunique() | the count-only twin: returns just the number of distinct values, NA excluded by default (dropna=True). |
df.drop_duplicates() | row-level dedupe on one or more columns; for distinct values per group, groupby('g')['c'].unique(). |
I diffed the old source trees: v0.6.1's series.py has no def unique at all; v0.7.0 ships one whose entire body is return nanops.unique1d(self.values), with the docstring already bragging “Significantly faster than numpy.unique”. Just like value_counts(), it was a pure-Python dispatch (nanops.unique1d) that immediately handed the work to the Cython hashtables — lib.Float64HashTable, Int64HashTable, PyObjectHashTable — the khash machinery it still uses today. nunique() (then literally len(value_counts())) arrived in the very same release.
pandas 0.8.0 (June 2012) added the top-level def unique(values) in pandas/core/algorithms.py — “Compute unique values (not necessarily sorted) efficiently” — so raw ndarrays and Index objects got the same hash-table treatment. Its return-type contract is quirky and ancient: Index in gives Index out, Categorical in gives Categorical out, a Series or ndarray gives an ndarray/ExtensionArray. On pandas 3.0.3 the Series method is a one-liner over the shared NDFrame base (return super().unique()) and the engine underneath is unique_with_mask(values) — same name, same promise, same tables.
pd.unique dispatches over the dtype — khash for numeric/object, a different code path for categorical/object — and the whole job is a single hash-table walk that remembers which values it has already seen. Nothing sorts anything; that is the entire difference to np.unique. The same khash machinery powers value_counts(), factorize(), and Index.unique(), which is why they all agree on order-of-appearance semantics. The docstring's “Significantly faster than numpy.unique” line exists because hash tables are O(n) while numpy pays O(n log n) for the sort plus a pass to collapse runs.