kmail.at
← learning

shell · difficulty ◆◆

tr — translate or delete characters

One stream in, one stream out, character maps only — the smallest filter that still earns its slot in every pipeline.

`tr -d '\r' < data.csv > data.csv` empties the file to 0 bytes and prints nothing at all. No error, no warning — the shell opens the output for writing before tr ever reads a character.

2026-10-08 · 5 min read

$ tr

What it does

tr reads standard input, writes standard output, and does exactly one of three things between them. Give it two arrays and it substitutes character by character: the first character of ARRAY1 becomes the first of ARRAY2, and so on. Add -d and it deletes instead of substituting. Add -s and it squeezes every run of a repeated character down to one. -c (or -C) inverts ARRAY1, so you can say 'everything except the digits' — useful with -d, fatal-looking without it, because the complement of a small set is every other byte in the character set and you get a flood of substitutions. That is the whole language. There is no regex engine hiding in there, no context, no line buffer: POSIX says it plainly in its own notes, the operands 'are not regular expressions'. tr ships in coreutils, BSD, busybox, toybox — it is the one filter you can count on existing on a switch, a NAS and a container image with nothing else in it.

Why it matters, and why it eats a day when it goes wrong

The reason tr earns its place is whitespace. Almost every classic Unix tool pads fields for human eyes: `df -h` right-aligns its Size/Used/Avail columns, syslog pads a single-digit day to two spaces ('Oct 8' versus 'Oct 08'), ps and docker ps line up their columns with runs of spaces. Feed any of that to `cut -d' ' -f5` and you get an empty string, because field 5 is the gap between two runs of padding. `tr -s ' '` fixes it in front of the cut and costs nothing. The second job is the invisible bytes: a UTF-8 BOM at the start of a CSV header, a carriage return at the end of every line from a Windows export, a NUL byte in a dump that makes a parser stop early. Each is one tr invocation and each is a silent disaster otherwise — and the failure mode is the nasty kind, where the file parses and the numbers are subtly wrong. The third job is that tr is a byte filter, which is a feature: on a binary blob or a never-ending stream, tr -d is as predictable as it gets.

The two traps: no filename operand, and the redirect that truncates

First trap, small but constant: tr takes no file arguments. `tr 'a-z' 'A-Z' page.txt` does not read page.txt, it fails with `tr: extra operand 'page.txt'` — every other filter in the toolbox accepts files, this one does not, deliberately. Use `<` or a pipe. Second trap, expensive: `tr -d '\r' < f > f` reads nothing, because the shell has already opened f for writing and truncated it to zero bytes before tr starts; on a copy here it went from 17 bytes to 0 with no diagnostic. There is no -i flag and no in-place mode, so the fix is always the same two-liner — write to a temp file, then mv it over the original, with mv giving you a rename that is atomic on the same filesystem. Same care applies to the arrays: -d and -s may not share an array (`tr -ds 'abc' 'abc'` fails, `tr -ds 'a' 'b'` works), and a reversed range such as 'z-a' fails immediately with `tr: range-endpoints of 'z-a' are in reverse collating sequence order`.

Example

$ df -h | cut -d' ' -f5 | head -4
df -h | tr -s ' ' | cut -d' ' -f5 | head -4



421G
Use%
1%
36%
97%

Both halves are real, from this box. df pads its Size/Used/Avail columns to a fixed width, so with a plain single-space delimiter the fifth field is blank and the column you wanted only surfaces when the padding happens to run out — here on one line, printing 421G. Put `tr -s ' '` in front and field 5 is Use% down the whole file: 1% for /run, 36% for the EFI vars, 97% for /. This one idiom is worth more than half the flags.

$ tr -s ' ' < sys.log | cut -d' ' -f4
cut -d' ' -f4 sys.log
web1
web1
web2

08:12:03
08:12:05
web2

Three lines of syslog, and the third line is the tell: syslog pads a one-digit day to 'Oct 8' with a double space, so without -s the fourth field is the timestamp on two lines and lands on the hostname only on the one line that happens to use 'Oct 08'. The pipeline does not crash, it just quietly reports '08:12:03' as a host name. `tr -s ' ' < sys.log | cut -d' ' -f5-` then gives you the message text of every line.

$ docker logs kmail-app --tail 4000 | grep -E '^[0-9]+\.[0-9]+\.[0-9]+\.[0-9]+ ' > acc.log
tr -s ' ' < acc.log | cut -d' ' -f9 | sort | uniq -c | sort -rn
    410 404
    140 200

550 real request lines from the kmail-app container on this host. Nginx separates its fields with single spaces, but tr -s is the habit that makes the same pipeline work on an Apache log or a line with a doubled space after the timestamp. Swap -f9 for -f7 and you get the URL tally: 74 hits each on /api/trpc, /api/auth/signin, /api/auth/session and /api/auth/callback — bot traffic fishing for admin endpoints, which is why 410 of those 550 requests were 404s.

$ tr -cs '[:alpha:]' '[\n*]' < page.txt | tr '[:upper:]' '[:lower:]' | sort | uniq -c | sort -rn | head -5
      4 the
      3 is
      1 what
      1 turns
      1 tr

Word frequency in one pipe, over a two-sentence paragraph — 166 bytes, 4 occurrences of 'the'. `-cs` is the pairing to remember: complement plus delete throws away everything that is not a letter, and squeeze collapses each run of the leftover newlines into one. '[\n*]' is the ARRAY2 filler that repeats newline as many times as ARRAY1 is long, so you never have to count characters; POSIX uses this exact line as its official example. Drop -s and you get one blank line per deleted character, hundreds of them.

$ file orders.csv
cut -d, -f1 orders.csv | head -2 | xxd
tr -d '\357\273\277' < orders.csv > orders.clean.csv && file orders.clean.csv
orders.csv: CSV Unicode text, UTF-8 (with BOM) text
00000000: efbb bf6f 7264 6572 5f69 640a 3130 3031  ...order_id.1001
00000010: 0a                                       .
orders.clean.csv: CSV ASCII text

The BOM: three bytes ef bb bf sitting in front of 'order_id', so the header column is really named '\ufefforder_id' and every `if col == 'order_id'` check fails while the file looks perfect in a spreadsheet. Octal is the argument form tr wants — '\357' is 0xEF, '\273' is 0xBB, '\277' is 0xBF. file(1) confirms the cure: the copy is plain ASCII now.

$ od -c nul.txt
tr -d '\000' < nul.txt | cat -A
tr -cd '[:print:]\n' < ctrl.log | cat -v
0000000   a  \0   b  \0   c  \n
0000006
abc$
ok^Abad^Bline^G
okbadline

Two invisible-byte cases with the evidence shown. NUL bytes in a Postgres COPY or a captured dump make C string parsers stop at the first one; `tr -d '\000'` removes them and leaves the text. The second command strips every control character from a captured log — cat -v is how you see them at all (^A, ^B, ^G). If your locale is misbehaving, the byte-exact form is `tr -cd '\11\12\15\40-\176'`, which keeps tab, newline, CR and printable ASCII and nothing else.

$ find crlf2 -name '*.csv' -type f -exec sh -c 'tr -d "\r" < "$1" > "$1.tmp" && mv "$1.tmp" "$1"' _ {} \;
find crlf2 -name '*.csv' -type f -exec wc -c {} \;
printf 'before: %s bytes\n' "$(wc -c < clobber.csv)"
tr -d '\r' < clobber.csv > clobber.csv
printf 'after:  %s bytes\n' "$(wc -c < clobber.csv)"
15 crlf2/a.csv
14 crlf2/nested/b.csv
before: 17 bytes
after:  0 bytes

Top half: a small Windows-exported tree, one carriage return removed per line, 17→15 and 16→14 bytes, written to a .tmp and moved back over the original so the rename is atomic. Bottom half: the same conversion written the way it feels natural, which empties the file. `tr -d '\r' < f > f` is the single most expensive five seconds in shell: the redirection truncates before tr reads, and tr exits 0 because tr never had an error.

$ printf 'café ÜBER naïve\n' | tr 'a-z' 'A-Z'
printf 'Grüße\n' | tr '[:lower:]' '[:upper:]'
CAFé ÜBER NAïVE
GRüßE

Real output from coreutils 9.4 here, and the lesson is the shape of what did not change: the accented characters stayed exactly as they were. tr works on bytes, so it can see é as the two bytes c3 a9 but has no idea they are one letter — and case conversion, class deletion and -cs word splitting all inherit that blind spot. `LC_ALL=C tr` gives the same answer, just deliberately. If you need real Unicode case folding, tr is the wrong tool; iconv, sed or the Heirloom implementation are the ones that can see characters.

$ tr 'A-Za-z' 'N-ZA-Mn-za-m' <<< 'Pack My Box With Five Dozen Liquor Jugs'
tr 'A-Za-z' 'N-ZA-Mn-za-m' <<< 'Cnpx Zl Obk Jvgu Svir Qbmra Yvdhbe Whtf'
Cnpx Zl Obk Jvgu Svir Qbmra Yvdhbe Whtf
Pack My Box With Five Dozen Liquor Jugs

ROT13, and the only one of these you can do on any box without installing anything. Note how ARRAY2 is built: N-Z wraps the upper half of the alphabet back round to A-M, lowercase after it, which is why tr needs no wrapping logic — the rotation is baked into the two ranges. The same trick gives you ROT-N by shifting both endpoints, as long as the range stays ascending.

Common flags

-d, --delete
Delete the characters in ARRAY1 instead of translating. Only one array is used: `tr -d '\000'`, `tr -d '[:punct:]'`, `tr -d '\r'`. Combine with -c to keep only a set instead of removing one.
-s, --squeeze-repeats
Replace each run of a repeated character with a single instance. Squeezing happens after translation or deletion, and only for the characters in the last array given — so `tr -s '0-9' '#'` also turns every run of digits into one #.
-c, -C, --complement
Use everything except ARRAY1. `tr -cd '[:alpha:]'` keeps letters and deletes the rest. GNU accepts -c and -C as synonyms; the historical split was -c by byte value and -C by collating element, so they only differ in the corners of an exotic locale.
-t, --truncate-set1
Truncate ARRAY1 to the length of ARRAY2 first. Without it, `tr 'abcd' 'x'` pads ARRAY2 by repeating its last character and gives you xxxx; with -t you get xbcd. Only meaningful when translating.
\NNN
Octal escape, one to three digits — the reliable way to name bytes you cannot type: \000 NUL, \011 tab, \012 newline, \015 CR, \357\273\277 the UTF-8 BOM. Use all three digits whenever a real digit follows, or tr reads \0001 as one number.
[:class:]
POSIX character classes: [:alpha:] [:digit:] [:space:] [:blank:] [:print:] [:punct:] [:cntrl:] [:xdigit:]. In ARRAY2 only [:lower:] and [:upper:] are legal, and only opposite each other for case conversion — anything else fails with 'when translating, the only character classes that may appear in string2 are upper and lower'.
[x*n]
ARRAY2 only: repeat x as many times as needed to match ARRAY1's length — '[\n*]' for the famous newline filler, or '[x*]' to blank out a field. Give a count (or a leading-zero octal count) to pin the repetition instead of auto-extending.

History

Version 4 Unix, 1973 — written by the man who proposed the pipe

tr was written at Bell Labs by M. Douglas McIlroy and first shipped in Version 4 Unix in November 1973. McIlroy is the one who suggested wiring programs together with pipes, and he went on to write diff, sort, join, echo, speak and tr itself. Look at that list and the design falls into place: tr is a pipe citizen before it is a program, which is exactly why it reads stdin and refuses filenames — it was designed for the slot between two other commands, not for the desktop. The GNU version was rewritten by Jim Meyering and lives in coreutils; the original single-byte design has never been replaced, only patched.

Fifty years of one-line patches

The coreutils NEWS file shows how long a two-array byte map keeps getting fixed. 6.6 (2006-11-22) stopped tr mishandling a second operand that began with a dash. 6.9 (2007-03-22) stopped tr -c aborting when ARRAY2 was longer than the complement of ARRAY1 — tagged '[present in the original version, in 1992]' — and stopped it rejecting an unmatched [:lower:] or [:upper:] in ARRAY1, same vintage. 6.9.90 (2007-12-01) taught it to warn about a trailing unescaped backslash. 6.9.92 (2008-01-12) fixed case conversion failing in a locale where the numbers of upper and lower case characters differ. 8.6 (2010-10-15) made case-conversion classes consistent — and then 9.0 (2021-09-24) had to fix a crash using --complement with certain invalid combinations of those classes, a bug introduced by the 8.6 fix. Ten years between a patch and the regression it caused, in the smallest program on the shelf.

Fun facts

Pros & cons

pros

  • + In every base install and every busybox image, and small enough to reason about completely — three modes, no regex engine, no state
  • + The cheapest correct fix for the two things that break text pipelines in practice: runs of padding whitespace (tr -s ' ') and invisible bytes (BOM, CR, NUL, control characters)
  • + Streams and bytes only, so it works unchanged on a tail -f, a multi-gigabyte log or a binary blob where sed and awk want to think about lines and encodings

cons

  • − Byte-oriented, so on UTF-8 it cannot case-fold, count or split accented letters — `tr 'a-z' 'A-Z'` turned café into CAFé here, leaving the é untouched, and GNU tr still has no real multibyte mode
  • − Character maps only: no strings, no regexes, no lookahead, and its operands are explicitly not regular expressions, so substituting a word or a pattern needs sed or perl
  • − Two silent failure modes that cost real data: it accepts no filename operand at all, and redirecting onto its own input truncates the file to zero bytes with a success exit code

Takeaways

  1. 1Put `tr -s ' '` in front of every `cut -d' '` on ps, df, docker ps or syslog output — without it you are cutting on the padding, not the field
  2. 2`tr -cs '[:alpha:]' '[\n*]'` explodes text into one word per line; pipe that into sort | uniq -c | sort -rn for word frequency
  3. 3The three invisible bytes: `tr -d '\r'` for CRLF, `tr -d '\357\273\277'` for a UTF-8 BOM, `tr -d '\000'` for NUL in a dump
  4. 4Never write `tr ... < f > f` — it empties the file silently; always go through f.tmp and mv it back
  5. 5tr maps bytes, not characters: check any accented or CJK input before trusting case folding or class deletion, and keep a Unicode-aware tool around for real text work

Related commands

← all learning