linux · difficulty ◆◆
dd — copy and convert, block by block
The block-level copier that images disks, patches sectors, and measures real disk speed.
dd will happily read a raw disk, write over a running filesystem, and report success — because it never asks what you meant. It is the least forgiving command in coreutils, and the one you want in your hands at 3 a.m.
$ ddWhat dd is, and why it looks like nothing else
dd copies input to output with a block size you choose, optionally converting the data on the way through. That is the whole tool. cp cares about files; dd does not know what a file is — hand it a block device, a pipe, a byte range inside a 500 GB image, or /dev/zero, and it copies blocks. The operands are all `key=value` with no dashes at all, which is the first thing that confuses people coming from ls and grep, and it is a direct inheritance from mainframe job control: the coreutils manual says the operands' syntax 'was inspired by the DD (data definition) statement of OS/360 JCL'. The second confusing thing is that dd prints to stderr and only ever talks about blocks, never about files. Once you internalise those two facts, dd stops being scary and becomes the tool you reach for when nothing else can touch the data.
`N+M records` is the whole story
Every dd run ends with a counter in the form `full+partial records in` / `full+partial records out`. `1+0` means one complete block. `0+1` means a single partial block — and that is dd telling you the copy was short. This is where the classic footgun lives: `count=1 bs=8` against a pipe that delivers 4 bytes stops after 4 bytes, prints `0+1 records in` and exits 0. No error. Your data is half-copied and the exit code says everything is fine. The fix is `iflag=fullblock`, which makes dd loop on read() until the block is genuinely full. The same counter explains the byte-suffix rule too: `count=1` counts one *block*, so `count=1 bs=512` is 512 bytes, while `count=1B` (coreutils 9.1 and later) means one byte. Reading that counter correctly is most of using dd well.
dd is the honest speedometer — if you make it one
Ask any ops person how they test a disk and they will write a dd line. The catch is that a plain write lands in the page cache, so `bs=4M count=64` reports gigabytes per second and measures your RAM. Add `conv=fsync` (or `fdatasync`) and the kernel waits for the device before returning; add `oflag=direct` and you bypass the cache from the first write. On this box the two numbers are more than 4x apart for the exact same 256 MiB — 3.5 GB/s with the cache and 768 MB/s once the data actually has to land. If a storage benchmark looks fantastic, it is almost always measuring the wrong thing, and dd is the one tool that makes the difference visible in one line.
Example
$ sudo dd if=/dev/nvme0n1 of=/root/nvme-mbr.bak bs=512 count=11+0 records in
1+0 records out
512 bytes copied, 2.0261e-05 s, 25.3 MB/scount=1 with bs=512 collects exactly one sector. Offset 446 of that sector holds the four-entry partition table, offset 510 the 0x55 0xAA boot signature. This is the cheapest insurance you can buy before repartitioning: 512 bytes, twenty microseconds, and the map of the disk you are about to rearrange. Inspect what you saved with `dd if=nvme-mbr.bak bs=1 skip=446 count=66 | hexdump -C`.
$ dd if=/dev/zero of=prog.img bs=4M count=200 status=progress200+0 records in
200+0 records out
838860800 bytes (839 MB, 800 MiB) copied, 0.488756 s, 1.7 GB/sstatus=progress rewrites a single line on stderr about once a second, so a copy that runs for an hour is no longer a black box — this is the shape of every disk clone you will write. Swap the input for a real device (`sudo dd if=/dev/nvme0n1 of=/mnt/backup/nvme.img bs=64M status=progress conv=fsync`) and expect hundreds of MB/s and a duration in the thousands of seconds. Add conv=fsync before trusting the finish line, otherwise dd exits while the last writes are still in RAM.
$ dd if=/dev/zero of=test.img bs=1M count=100 conv=sparse
ls -lh test.img && du -h test.img100+0 records in
100+0 records out
104857600 bytes (105 MB, 100 MiB) copied, 0.00793915 s, 13.2 GB/s
-rw-rw-r-- 1 kmail kmail 100M Oct 9 09:01 test.img
0 test.imgconv=sparse tells dd to seek over runs of NUL input instead of writing them, so the file *reports* 100 MiB and *costs* zero bytes on ext4. That is the trick behind instant multi-gigabyte test files, and behind container images that unpack from 2 GB to nearly nothing. Without conv=sparse the identical command really does write 100 MiB of zeros and takes real time to do it.
$ dd if=/dev/zero of=pagecache.img bs=4M count=64
dd if=/dev/zero of=flushed.img bs=4M count=64 conv=fdatasync64+0 records in
64+0 records out
268435456 bytes (268 MB, 256 MiB) copied, 0.0773873 s, 3.5 GB/s
64+0 records in
64+0 records out
268435456 bytes (268 MB, 256 MiB) copied, 0.349528 s, 768 MB/sSame 256 MiB, same block size, two lines apart: 3.5 GB/s into the page cache, 768 MB/s once fdatasync forces the wait for the device. If a storage number makes you proud, reproduce it with conv=fsync or oflag=direct before you believe it — the first line here is RAM speed, not disk speed.
$ ( printf 'AAAA'; sleep 0.2; printf 'BBBB' ) | dd of=seg.bin bs=8 count=1
( printf 'AAAA'; sleep 0.2; printf 'BBBB' ) | dd of=full.bin bs=8 count=1 iflag=fullblock0+1 records in
0+1 records out
4 bytes copied, 2.0867e-05 s, 192 kB/s
1+0 records in
1+0 records out
8 bytes copied, 0.201286 s, 0.0 kB/sThe footgun, reproduced exactly: `0+1 records in` is one partial block, count=1 stopped dd right there, and BBBB never made it out of the pipe — exit code 0, no warning. With iflag=fullblock dd keeps calling read() until the 8-byte block is genuinely full: `1+0 records in`, 8 bytes, both halves present. This fires for real whenever dd reads from a pipe, a socket, or a disk that is starting to fail.
$ dd if=/dev/zero of=disk.img bs=4M count=64
head -c 4096 /dev/urandom > blk.bin
dd if=blk.bin of=disk.img bs=512 count=4 seek=8 conv=notrunc64+0 records in
64+0 records out
268435456 bytes (268 MB, 256 MiB) copied, 0.289822 s, 926 MB/s
4+0 records in
4+0 records out
2048 bytes (2.0 kB, 2.0 KiB) copied, 0.000116232 s, 17.6 MB/sseek=8 skips eight 512-byte output blocks, writes four, and stops — everything after byte 6144 is untouched because conv=notrunc stopped dd truncating the file at open. This is how you patch a header, a UUID, or a superblock inside a 500 GB image in a tenth of a millisecond instead of rewriting the whole file. Drop conv=notrunc and the same command leaves you with a 6 KiB file and no error message.
Common flags
- if=FILE / of=FILE
- Input and output. Defaults are stdin and stdout, which is why `dd` with no operands just sits there. of= truncates its target on open unless conv=notrunc is given.
- bs=BYTES
- Sets ibs and obs at once and overrides both. The 512-byte default is a 1970s tape block; bs=1M or bs=64M is where throughput lives on NVMe. Suffixes: K=1024, M=1024^2, KB=1000, MB=1000^2.
- count=N
- Copy N input blocks — blocks, not bytes, so the byte count depends on bs. Since coreutils 9.1 a trailing `B` switches N to a byte count: `count=512B` is half a kilobyte whatever bs says.
- skip=N / seek=N
- Skip N input blocks (skip, alias iseek) or N output blocks (seek, alias oseek) before copying. The aliases and the `B` byte suffix are GNU extensions over POSIX — the manual's own tape example uses iseek=512B to leave the first sector alone.
- conv=notrunc
- Do not truncate the output file before writing. Mandatory for in-place patches, and mandatory for any of= pointing at a device whose contents you intend to keep.
- iflag=fullblock
- Accumulate full input blocks, re-reading after a short read (coreutils 7.0 and later). Turns the silent `0+1 records in` partial copy into a complete one. Since 8.11 dd also warns you when a partial read might bite.
- status=progress
- Print transfer statistics about once a second (8.24 and later) — the only progress meter dd ever had on Linux, apart from signalling it. status=none (8.20 and later) silences everything but errors.
History
The operands are 1960s mainframe syntax, and the name is a joke that stuck
dd is not an abbreviation of anything sensible in Unix. Its operand style is modelled on the DD statement in IBM OS/360 Job Control Language, where DD stood for Data Definition — and you can still read the ancestry in the installed manual: the operands' 'syntax was inspired by the DD (data definition) statement of OS/360 JCL'. That is why there are no dashes, why `if=` and `of=` read like card images, and why the tool has survived fifty years of shell evolution looking exactly the same. The GNU implementation is credited in `dd --version` to Paul Rubin, David MacKenzie, and Stuart Kemp — three names from the era when coreutils was still fileutils and sh-utils.
Everything convenient about dd arrived late, and was documented one line at a time
The dd on this box is coreutils 9.4, and the conveniences everyone relies on now are recent additions, each recorded in the coreutils NEWS file. iflag=fullblock landed in 7.0 (2008); the nocache flag and the warning that nudges you toward fullblock came in 8.11 (2011); conv=sparse and the count_bytes/skip_bytes iflags in 8.16 (2012); status=progress — the first real progress indicator — in 8.24 (2015). Byte-count suffixes and the iseek/oseek aliases only arrived in 9.1 (2022). For most of its life dd had no progress output, no sparse support, and no guard at all against the short read that truncates copies. The folklore nickname 'disk destroyer' was earned in exactly those decades.
Fun facts
Pros & cons
pros
- + Speaks block devices, pipes, and arbitrary byte ranges fluently — one of the few tools that can read a raw disk, image it, and then patch a single sector inside that image
- + Syntax frozen since the mainframe era: a dd line written in 1995 still runs unchanged on coreutils 9.4, and the same line works on busybox and on BSD
- + Always present. It is in coreutils on every distro and in busybox on routers, so it is the tool you still have when the rescue shell has nothing else
cons
- − Zero guardrails. of= is destroyed on contact, there is no dry run, and dd prints nothing about what it replaced or how big the source was
- − Defaults from another era: 512-byte blocks and no progress output, so the naive one-liner is both slow and blind
- − Its output is written for machines. Reading the records counter is a skill, and 0+1 versus 1+0 hides a genuine failure behind an exit code of 0
Takeaways
- 1Set bs= on purpose. The 512-byte default is a tape block; bs=1M or bs=64M is where NVMe throughput actually lives.
- 2Add conv=fsync (or oflag=direct) before you believe any dd speed number — without it you are benchmarking the page cache, not the disk.
- 3dd truncates of= unless you pass conv=notrunc, so rehearse the line on a scratch file first: on a block device there is no undo.
- 4count=N counts blocks, not bytes. count=1 bs=512 is the partition table; write count=512B (coreutils 9.1+) when you mean bytes.
- 5If the input is a pipe, a socket, or a struggling disk, pass iflag=fullblock — `0+1 records in` is a truncated copy that still exits 0.