kmail.at
← learning

pandas · difficulty ◆◆

pandas astype() — cast a whole report into real dtypes in one dict

astype() is where a column stops being text and starts being a number.

Every column pandas reads from a CSV is lying to you about what it is — until one dict of column-to-dtype tells the whole table the truth at once.

2026-10-09 · 5 min read

$ df.astype()

What it does

astype() converts a Series or DataFrame to a different dtype. Pass one dtype to convert everything, or a dict of {column: dtype} to give each column its own target — the dict form is the one you actually use, because it turns a freshly parsed CSV into a typed table in a single call. Targets can be numpy names ('int64', 'float32', 'bool'), datetime units ('datetime64[s]'), or pandas types ('Int64', 'category', a CategoricalDtype with an explicit order). Missing values are the sharp edge: a plain int64 cannot hold NaN, so astype('int64') on a column with holes raises IntCastingNaNError instead of inventing a value. errors= takes two values only — 'raise' (default) and 'ignore' (return the column unchanged) — and notably not 'coerce', which lives on pd.to_numeric instead. Since 3.0 the copy= keyword is deprecated and Copy-on-Write, not astype, decides whether anything gets copied.

Why it matters

dtype is not cosmetic: it decides whether '19.90' * 40 equals 796.0 or '19.9019.90...'. A CSV hands you str columns (object before 3.0), so the first honest thing a session does is declare what each column means. Fixing dtypes at the boundary buys three things at once — arithmetic that works, memory you did not know you were burning (10,000 rows of a repeated city name cost 546 kB as str and 10 kB as category), and operations that only exist for certain types: the .dt accessor, .cat ordering, groupby on a categorical that keeps unused categories. Do it once, right after load, and every later line gets simpler.

Example

$ import pandas as pd

raw = pd.DataFrame({
    "order_id":   ["1001", "1002", "1003"],
    "qty":        ["12", "3", "40"],
    "unit_price": ["19.90", "5.00", "2.25"],
    "ordered_at": ["2026-09-01", "2026-09-02", "2026-09-03"],
    "shipped":    ["True", "False", "True"],
})
typed = raw.astype({
    "order_id":   "int64",
    "qty":        "int64",
    "unit_price": "float64",
    "ordered_at": "datetime64[ns]",
    "shipped":    "bool",
})
print(typed.dtypes)
print((typed["qty"] * typed["unit_price"]).round(2).sum())
order_id               int64
qty                    int64
unit_price           float64
ordered_at    datetime64[ns]
shipped                 bool
dtype: object
343.8

The whole reason the method exists: one dict, five declared types, and the revenue line multiplies instead of concatenating. Drop the round(2) and the same sum prints 343.79999999999995 — float64 does not contain 19.90, only something very close to it. Columns you leave out of the dict keep their dtype, so the mapping doubles as a schema note.

$ sales = pd.Series([1200, None, 980], name="sales")
print(sales.astype("Int64").to_string())
try:
    sales.astype("int64")
except Exception as exc:
    print(f"{type(exc).__name__}: {str(exc).split('.')[0]}")
print(sales.astype("Int64").sum())
0    1200
1    <NA>
2     980
IntCastingNaNError: Cannot convert non-finite values (NA or inf) to integer
2180

Capital I: Int64 is the nullable pandas integer, and it keeps the hole as <NA> while sum() still adds the real values. Lowercase int64 refuses to hold a missing value and says so — that error is the feature, because the old pandas quietly promoted your integers to floats instead. Watch the truncation rule too: pd.Series([2.9, -2.9]).astype("int64") gives [2, -2], so call .round() first when 2.9 should not become 2.

$ audit = pd.DataFrame({"venue": ["Berlin", "Vienna", "Rome", "Berlin", "Vienna"] * 2000})
print("rows:", len(audit), "| dtype:", audit["venue"].dtype)
print("bytes as str:     ", audit["venue"].memory_usage(deep=True))
cat = audit["venue"].astype("category")
print("bytes as category:", cat.memory_usage(deep=True))
print("categories:", cat.cat.categories.tolist())
rows: 10000 | dtype: str
bytes as str:      546132
bytes as category: 10295
categories: ['Berlin', 'Rome', 'Vienna']

A ~53x memory drop from a one-word astype, on 10,000 rows that hold three distinct values. That dtype line also dates the session: it says str, not object, which only happens on pandas 3.0+. Categories sort alphabetically here — pass pd.CategoricalDtype(["low", "med", "high"], ordered=True) when the real order is a business rule, and .min()/.max() will then return low and high instead of Berlin and Vienna.

$ log = pd.DataFrame({
    "ts":   ["2026-10-01 06:00", "2026-10-01 06:15", "2026-10-01 06:30"],
    "pm25": ["24.5", "22.1", "23.8"],
})
log["ts"] = log["ts"].astype("datetime64[s]")
log["pm25"] = log["pm25"].astype("float32")
print(log.dtypes)
print(log.to_string(index=False))
print("pm25 bytes float32:", log["pm25"].memory_usage(deep=True),
      "| float64:", log["pm25"].astype("float64").memory_usage(deep=True))
print("gap between samples:", log["ts"].iloc[2] - log["ts"].iloc[1])
ts      datetime64[s]
pm25          float32
dtype: object
                 ts      pm25
2026-10-01 06:00:00 24.500000
2026-10-01 06:15:00 22.100000
2026-10-01 06:30:00 23.799999
pm25 bytes float32: 144 | float64: 156
gap between samples: 0 days 00:15:00

Sensor and log pipelines are where astype earns its keep: on a million readings float32 is 4,000,132 bytes against float64's 8,000,132, and datetime64[s] states the resolution you actually have. The 23.799999 is float32 precision, not a bug — declare the dtype and you own its arithmetic. Note what astype is not doing here: parsing. astype("datetime64[ns]") on a column with one junk string raises DateParseError and casts nothing; pd.to_datetime(..., errors="coerce") is the parser.

$ raw = pd.DataFrame({"sku": ["A-1", "B-2"], "qty": ["4", "bad"]})
ignored = raw.astype({"qty": "Int64"}, errors="ignore")
print(ignored.dtypes)
print("still strings:", ignored["qty"].tolist())
try:
    raw.astype({"qty": "Int64"}, errors="coerce")
except Exception as exc:
    print(f"{type(exc).__name__}: {str(exc).split(': Error')[0]}")
print(pd.to_numeric(raw["qty"], errors="coerce").to_string())
sku    str
qty    str
dtype: object
still strings: ['4', 'bad']
ValueError: Expected value of kwarg 'errors' to be one of ['raise', 'ignore']. Supplied value is 'coerce'
0    4.0
1    NaN

errors='ignore' is per column and all-or-nothing: one bad cell means the perfectly good '4' stays a string as well. And 'coerce' does not exist here — the error message literally lists the two legal values. When a dirty numeric column needs salvaging, astype is the wrong tool; pd.to_numeric(..., errors='coerce') is the one that turns the bad cell into NaN and keeps the rest.

Common flags

dtype="int64"
everything passed goes through pandas_dtype() first, so numpy names, Python types and pandas strings all work — but the whole object ends up as that one type.
dtype={"qty": "int64"}
the dict form: per-column targets, unlisted columns untouched. A key that is not a column raises KeyError — where fillna() ignores unknown keys, astype fails loudly.
errors="raise" | "ignore"
the only two values. 'ignore' returns the original column when the cast fails; there is no 'coerce' and no partial cast.
dtype="Int64" / "Float64" / "boolean"
capitalised nullable dtypes keep <NA> next to real values; their lowercase numpy twins raise IntCastingNaNError on a hole.
dtype="category"
the cheapest big win on repeated labels. Hand it pd.CategoricalDtype([...], ordered=True) when sort order is a business rule rather than alphabetical.
dtype="datetime64[s]"
explicit resolution instead of the default ns; the 3.0 notes recommend exactly this to stay unit-agnostic. It parses a narrow, uniform set of strings and nothing more.
copy=True/False
deprecated in 3.0 (GH 57347) and effectively a no-op under Copy-on-Write. Use .copy() when you want an eager copy; the keyword is slated for removal in 4.0.

History

0.4.0, September 2011 — a one-line method with no options

There is no astype anywhere in the pandas 0.3.0 sdist from February 2011 — grep the tarball and 'def astype' returns nothing. It arrived in 0.4.0, published 2011-09-12, as NDFrame.astype with a single argument, and the whole implementation in pandas/core/generic.py was one line: return self._constructor(self._data, dtype=dtype). The docstring says only 'Cast object to input numpy.dtype'. The dict-of-columns form that this tutorial leads with did not exist until 0.19.0 (2016-10-02), after issue 12086 sat open from January to June 2016; before that, per-column casting meant a loop and a concat. Series itself was a numpy subclass until 0.13.0 (2014-01-16), when it was refactored onto NDFrame — which is what let one implementation serve both containers.

2017 to 2026 — errors, an exception with a name, and the dying copy keyword

The errors keyword landed in 0.20.0 (2017-05-05, GH 14878), replacing raise_on_error in the same release and offering only raise and ignore from day one. The NaN-to-integer failure got its own identity in 1.3.0 (2021-07-02): IntCastingNaNError, defined in pandas/errors/__init__.py. The v1.2.0 source raises a bare ValueError('Cannot convert non-finite values (NA or inf) to integer') and never mentions the class, while v1.3.0's cast.py raises it twice — and deliberately as a ValueError subclass, so a decade of existing except ValueError clauses kept working. Then 3.0.0 (2026-01-21) deprecated copy= across astype and a dozen sibling methods (GH 57347, opened and closed inside February 2024) in favour of Copy-on-Write, and made astype(str) return the new dedicated str dtype instead of object.

Fun facts

Pros & cons

pros

  • + One call expresses the whole schema decision — a dict per column, applied consistently, instead of five assignments and a temp variable
  • + Reaches past numpy to extension dtypes and a full CategoricalDtype, so 'make it smaller' and 'make it sortable in business order' are the same method
  • + Fails where it should: unknown column gives KeyError, NaN-to-int gives IntCastingNaNError, and an illegal errors= value prints the two legal ones

cons

  • − No 'coerce' — a single bad cell forces the whole column to stay untyped under errors='ignore', and the real salvaging tool is a different function (pd.to_numeric)
  • − Numeric surprises are built in: float-to-int truncates toward zero, out-of-range values wrap silently (300 into int8 gives 44), and float32 visibly loses precision
  • − The copy= keyword is deprecated in 3.0, so the copy semantics described in every pre-2024 tutorial no longer match what the method does

Takeaways

  1. 1Cast right after load, with a dict: df.astype({"qty": "int64", "price": "float64", "ts": "datetime64[ns]"}) documents the schema inside the code that reads the file.
  2. 2A column with holes cannot be int64. Use "Int64" (capital I) for nullable integers, or accept IntCastingNaNError from lowercase int64 — that error is the feature, not a nuisance.
  3. 3astype() truncates floats toward zero. Call .round() first if 2.9 should become 3, and never assume it rounds for you.
  4. 4errors= offers only 'raise' and 'ignore'; when the column is genuinely dirty, switch to pd.to_numeric(..., errors="coerce") instead of fighting astype.
  5. 5astype("category") on a repeated-label column is the cheapest memory win in pandas — 546 kB of str became 10 kB of codes across 10,000 rows.

Related commands

← all learning