pandas · difficulty ◆◆
pandas astype() — cast a whole report into real dtypes in one dict
astype() is where a column stops being text and starts being a number.
Every column pandas reads from a CSV is lying to you about what it is — until one dict of column-to-dtype tells the whole table the truth at once.
$ df.astype()What it does
astype() converts a Series or DataFrame to a different dtype. Pass one dtype to convert everything, or a dict of {column: dtype} to give each column its own target — the dict form is the one you actually use, because it turns a freshly parsed CSV into a typed table in a single call. Targets can be numpy names ('int64', 'float32', 'bool'), datetime units ('datetime64[s]'), or pandas types ('Int64', 'category', a CategoricalDtype with an explicit order). Missing values are the sharp edge: a plain int64 cannot hold NaN, so astype('int64') on a column with holes raises IntCastingNaNError instead of inventing a value. errors= takes two values only — 'raise' (default) and 'ignore' (return the column unchanged) — and notably not 'coerce', which lives on pd.to_numeric instead. Since 3.0 the copy= keyword is deprecated and Copy-on-Write, not astype, decides whether anything gets copied.
Why it matters
dtype is not cosmetic: it decides whether '19.90' * 40 equals 796.0 or '19.9019.90...'. A CSV hands you str columns (object before 3.0), so the first honest thing a session does is declare what each column means. Fixing dtypes at the boundary buys three things at once — arithmetic that works, memory you did not know you were burning (10,000 rows of a repeated city name cost 546 kB as str and 10 kB as category), and operations that only exist for certain types: the .dt accessor, .cat ordering, groupby on a categorical that keeps unused categories. Do it once, right after load, and every later line gets simpler.
Example
$ import pandas as pd
raw = pd.DataFrame({
"order_id": ["1001", "1002", "1003"],
"qty": ["12", "3", "40"],
"unit_price": ["19.90", "5.00", "2.25"],
"ordered_at": ["2026-09-01", "2026-09-02", "2026-09-03"],
"shipped": ["True", "False", "True"],
})
typed = raw.astype({
"order_id": "int64",
"qty": "int64",
"unit_price": "float64",
"ordered_at": "datetime64[ns]",
"shipped": "bool",
})
print(typed.dtypes)
print((typed["qty"] * typed["unit_price"]).round(2).sum())order_id int64
qty int64
unit_price float64
ordered_at datetime64[ns]
shipped bool
dtype: object
343.8The whole reason the method exists: one dict, five declared types, and the revenue line multiplies instead of concatenating. Drop the round(2) and the same sum prints 343.79999999999995 — float64 does not contain 19.90, only something very close to it. Columns you leave out of the dict keep their dtype, so the mapping doubles as a schema note.
$ sales = pd.Series([1200, None, 980], name="sales")
print(sales.astype("Int64").to_string())
try:
sales.astype("int64")
except Exception as exc:
print(f"{type(exc).__name__}: {str(exc).split('.')[0]}")
print(sales.astype("Int64").sum())0 1200
1 <NA>
2 980
IntCastingNaNError: Cannot convert non-finite values (NA or inf) to integer
2180Capital I: Int64 is the nullable pandas integer, and it keeps the hole as <NA> while sum() still adds the real values. Lowercase int64 refuses to hold a missing value and says so — that error is the feature, because the old pandas quietly promoted your integers to floats instead. Watch the truncation rule too: pd.Series([2.9, -2.9]).astype("int64") gives [2, -2], so call .round() first when 2.9 should not become 2.
$ audit = pd.DataFrame({"venue": ["Berlin", "Vienna", "Rome", "Berlin", "Vienna"] * 2000})
print("rows:", len(audit), "| dtype:", audit["venue"].dtype)
print("bytes as str: ", audit["venue"].memory_usage(deep=True))
cat = audit["venue"].astype("category")
print("bytes as category:", cat.memory_usage(deep=True))
print("categories:", cat.cat.categories.tolist())rows: 10000 | dtype: str
bytes as str: 546132
bytes as category: 10295
categories: ['Berlin', 'Rome', 'Vienna']A ~53x memory drop from a one-word astype, on 10,000 rows that hold three distinct values. That dtype line also dates the session: it says str, not object, which only happens on pandas 3.0+. Categories sort alphabetically here — pass pd.CategoricalDtype(["low", "med", "high"], ordered=True) when the real order is a business rule, and .min()/.max() will then return low and high instead of Berlin and Vienna.
$ log = pd.DataFrame({
"ts": ["2026-10-01 06:00", "2026-10-01 06:15", "2026-10-01 06:30"],
"pm25": ["24.5", "22.1", "23.8"],
})
log["ts"] = log["ts"].astype("datetime64[s]")
log["pm25"] = log["pm25"].astype("float32")
print(log.dtypes)
print(log.to_string(index=False))
print("pm25 bytes float32:", log["pm25"].memory_usage(deep=True),
"| float64:", log["pm25"].astype("float64").memory_usage(deep=True))
print("gap between samples:", log["ts"].iloc[2] - log["ts"].iloc[1])ts datetime64[s]
pm25 float32
dtype: object
ts pm25
2026-10-01 06:00:00 24.500000
2026-10-01 06:15:00 22.100000
2026-10-01 06:30:00 23.799999
pm25 bytes float32: 144 | float64: 156
gap between samples: 0 days 00:15:00Sensor and log pipelines are where astype earns its keep: on a million readings float32 is 4,000,132 bytes against float64's 8,000,132, and datetime64[s] states the resolution you actually have. The 23.799999 is float32 precision, not a bug — declare the dtype and you own its arithmetic. Note what astype is not doing here: parsing. astype("datetime64[ns]") on a column with one junk string raises DateParseError and casts nothing; pd.to_datetime(..., errors="coerce") is the parser.
$ raw = pd.DataFrame({"sku": ["A-1", "B-2"], "qty": ["4", "bad"]})
ignored = raw.astype({"qty": "Int64"}, errors="ignore")
print(ignored.dtypes)
print("still strings:", ignored["qty"].tolist())
try:
raw.astype({"qty": "Int64"}, errors="coerce")
except Exception as exc:
print(f"{type(exc).__name__}: {str(exc).split(': Error')[0]}")
print(pd.to_numeric(raw["qty"], errors="coerce").to_string())sku str
qty str
dtype: object
still strings: ['4', 'bad']
ValueError: Expected value of kwarg 'errors' to be one of ['raise', 'ignore']. Supplied value is 'coerce'
0 4.0
1 NaNerrors='ignore' is per column and all-or-nothing: one bad cell means the perfectly good '4' stays a string as well. And 'coerce' does not exist here — the error message literally lists the two legal values. When a dirty numeric column needs salvaging, astype is the wrong tool; pd.to_numeric(..., errors='coerce') is the one that turns the bad cell into NaN and keeps the rest.
Common flags
- dtype="int64"
- everything passed goes through pandas_dtype() first, so numpy names, Python types and pandas strings all work — but the whole object ends up as that one type.
- dtype={"qty": "int64"}
- the dict form: per-column targets, unlisted columns untouched. A key that is not a column raises KeyError — where fillna() ignores unknown keys, astype fails loudly.
- errors="raise" | "ignore"
- the only two values. 'ignore' returns the original column when the cast fails; there is no 'coerce' and no partial cast.
- dtype="Int64" / "Float64" / "boolean"
- capitalised nullable dtypes keep <NA> next to real values; their lowercase numpy twins raise IntCastingNaNError on a hole.
- dtype="category"
- the cheapest big win on repeated labels. Hand it pd.CategoricalDtype([...], ordered=True) when sort order is a business rule rather than alphabetical.
- dtype="datetime64[s]"
- explicit resolution instead of the default ns; the 3.0 notes recommend exactly this to stay unit-agnostic. It parses a narrow, uniform set of strings and nothing more.
- copy=True/False
- deprecated in 3.0 (GH 57347) and effectively a no-op under Copy-on-Write. Use .copy() when you want an eager copy; the keyword is slated for removal in 4.0.
History
0.4.0, September 2011 — a one-line method with no options
There is no astype anywhere in the pandas 0.3.0 sdist from February 2011 — grep the tarball and 'def astype' returns nothing. It arrived in 0.4.0, published 2011-09-12, as NDFrame.astype with a single argument, and the whole implementation in pandas/core/generic.py was one line: return self._constructor(self._data, dtype=dtype). The docstring says only 'Cast object to input numpy.dtype'. The dict-of-columns form that this tutorial leads with did not exist until 0.19.0 (2016-10-02), after issue 12086 sat open from January to June 2016; before that, per-column casting meant a loop and a concat. Series itself was a numpy subclass until 0.13.0 (2014-01-16), when it was refactored onto NDFrame — which is what let one implementation serve both containers.
2017 to 2026 — errors, an exception with a name, and the dying copy keyword
The errors keyword landed in 0.20.0 (2017-05-05, GH 14878), replacing raise_on_error in the same release and offering only raise and ignore from day one. The NaN-to-integer failure got its own identity in 1.3.0 (2021-07-02): IntCastingNaNError, defined in pandas/errors/__init__.py. The v1.2.0 source raises a bare ValueError('Cannot convert non-finite values (NA or inf) to integer') and never mentions the class, while v1.3.0's cast.py raises it twice — and deliberately as a ValueError subclass, so a decade of existing except ValueError clauses kept working. Then 3.0.0 (2026-01-21) deprecated copy= across astype and a dozen sibling methods (GH 57347, opened and closed inside February 2024) in favour of Copy-on-Write, and made astype(str) return the new dedicated str dtype instead of object.
Fun facts
Pros & cons
pros
- + One call expresses the whole schema decision — a dict per column, applied consistently, instead of five assignments and a temp variable
- + Reaches past numpy to extension dtypes and a full CategoricalDtype, so 'make it smaller' and 'make it sortable in business order' are the same method
- + Fails where it should: unknown column gives KeyError, NaN-to-int gives IntCastingNaNError, and an illegal errors= value prints the two legal ones
cons
- − No 'coerce' — a single bad cell forces the whole column to stay untyped under errors='ignore', and the real salvaging tool is a different function (pd.to_numeric)
- − Numeric surprises are built in: float-to-int truncates toward zero, out-of-range values wrap silently (300 into int8 gives 44), and float32 visibly loses precision
- − The copy= keyword is deprecated in 3.0, so the copy semantics described in every pre-2024 tutorial no longer match what the method does
Takeaways
- 1Cast right after load, with a dict: df.astype({"qty": "int64", "price": "float64", "ts": "datetime64[ns]"}) documents the schema inside the code that reads the file.
- 2A column with holes cannot be int64. Use "Int64" (capital I) for nullable integers, or accept IntCastingNaNError from lowercase int64 — that error is the feature, not a nuisance.
- 3astype() truncates floats toward zero. Call .round() first if 2.9 should become 3, and never assume it rounds for you.
- 4errors= offers only 'raise' and 'ignore'; when the column is genuinely dirty, switch to pd.to_numeric(..., errors="coerce") instead of fighting astype.
- 5astype("category") on a repeated-label column is the cheapest memory win in pandas — 546 kB of str became 10 kB of codes across 10,000 rows.