Your USB stick has 20 GB free, and it still refuses a 5 GB video.
“The file is too large for the destination file system.” Not the disk. The file system.
That message comes from a four-byte field designed in the 1990s, and by the end of this article you’ll be able to point at the exact bytes responsible.
In Part 1 we built a mental model of what any filesystem has to do: map names to files, find space for data, keep metadata, and survive failures. Now we’ll watch one real filesystem do all of that with almost no machinery.
FAT32 is a good place to start because there is so little of it. There are no inodes, no trees, and no journal. There’s a header, a big table of numbers, and directories that are just files full of fixed-size records. That simplicity is why it still runs on SD cards, cameras, USB sticks, and the small partition your laptop boots from.
If you remember the linked-allocation diagram from Part 1, you already know the central idea. FAT32 is that diagram, turned into a disk format.
01One small volume to follow
Abstract descriptions of disk formats are hard to hold in your head, so we’ll use one concrete volume for the whole article.
It’s a 512 MB FAT32 volume with 512-byte sectors and 4096-byte clusters. A cluster is FAT’s name for the allocation block we met in Part 1. It holds five things:
| Name | Size | Clusters |
|---|---|---|
| HELLO.TXT | 19 bytes | 3 |
| NOTES.TXT | 10,000 bytes | 4, 5, 9 |
| PHOTOS/ | directory | 6 |
| PHOTOS/CAT.JPG | 8,192 bytes | 7, 8 |
| auth.service.ts | 28 bytes | 10 |
Every hex dump and every number below comes from this volume. Notice that NOTES.TXT sits in clusters 4, 5, and 9. It isn’t contiguous, and that will matter.
02Four regions, in order
A FAT32 volume is four regions laid end to end. Nothing is scattered, and nothing has to be discovered by searching.
The reserved region comes first. Its very first sector is the boot sector, which describes the rest of the volume. A few of the sectors after it have jobs, and most are simply empty.
Then come the file allocation tables: the table that gives the filesystem its name, followed by an identical second copy.
Everything after that is the data area, divided into clusters. File contents live there. So do directories, because in FAT32 a directory is stored the same way a file is.
03The boot sector describes everything else
In Part 1 we said a filesystem needs some structure that tells the operating system how it’s organized. Unix-like filesystems call it a superblock. FAT32 keeps that information in the first sector of the volume.
Here is the start of ours:
Two things about reading this. First, numbers are stored little-endian, with the least significant byte first, so the bytes 00 02 mean 0x0200, which is 512. Second, the sector does two unrelated jobs. It opens with a jump instruction because a PC may try to execute this sector at boot, and the jump hops over the filesystem fields to reach the boot code.
Those few fields are enough to locate everything else on the volume:
first_fat_sector = reserved_sectors; // 32
first_data_sector = reserved_sectors
+ fat_count * sectors_per_fat; // 32 + 2 * 1022 = 2076
// Clusters are numbered from 2, so cluster n starts at:
sector(n) = first_data_sector + (n - 2) * sectors_per_cluster;
sector(2) = 2076 // the root directory
sector(9) = 2076 + 7 * 8 = 2132This is pseudocode, but the arithmetic is real. Cluster numbering starts at 2 because the first two entries of the table are reserved, as we’re about to see.
The sector also says where two helpers live. Sector 6 holds a backup copy of the boot sector, in case the first one is damaged. Sector 1 holds a small structure called FSInfo, which we’ll come back to when we write a file.
04The table the filesystem is named after
The file allocation table is an array of 32-bit numbers with one entry for every cluster in the data area. Entry 9 describes cluster 9. Nothing more complicated than that.
Each entry answers a single question: what comes after this cluster?
| Entry value | Meaning |
|---|---|
| 0x00000000 | The cluster is free |
| 0x00000002 and up | The number of the next cluster in this file |
| 0x0FFFFFF7 | Bad cluster, never use it |
| 0x0FFFFFF8 and up | End of the chain, usually written as 0x0FFFFFFF |
Only the low 28 bits of each entry are used, which is why these values start with 0x0F rather than 0xFF.
Here’s the beginning of the table on our volume. It starts at sector 32, which is byte 16,384, or 0x4000:
00004000: f0 ff ff 0f ff ff ff 0f ff ff ff 0f ff ff ff 0f
00004010: 05 00 00 00 09 00 00 00 ff ff ff 0f 08 00 00 00
00004020: ff ff ff 0f ff ff ff 0f ff ff ff 0f 00 00 00 00Four bytes per entry, twelve entries. The same data, drawn as a table:
Entries 0 and 1 don’t describe clusters. They hold a media marker and a couple of status flags, which is the reason real cluster numbers begin at 2.
The detail that explains most of FAT32:
this one table is both the free-space map and the map of every file.
In Part 1 those were two separate problems, with a bitmap for one and extents or block pointers for the other. FAT answers both with the same array. A zero means free. Anything else means the cluster is taken, and tells you where to go next.
The table isn’t large. Our volume has 130,812 clusters, so it needs 130,814 entries of four bytes each. That’s 523,256 bytes, which rounds up to the 1,022 sectors the boot sector promised.
05Following a file through its chain
NOTES.TXT is 10,000 bytes. With 4096-byte clusters it needs three of them, and its directory entry says it starts at cluster 4.
From there, the table takes over:
cluster 4 → FAT[4] = 5 bytes 0 – 4,095
cluster 5 → FAT[5] = 9 bytes 4,096 – 8,191
cluster 9 → FAT[9] = end of chain bytes 8,192 – 9,999The last cluster holds only 1,808 bytes of the file. The remaining 2,288 bytes of that cluster are the internal fragmentation we talked about in Part 1. HELLO.TXT is the extreme case: 19 bytes in a 4096-byte cluster.
Why does the chain jump from 5 to 9? Because of the order things happened in. The file was first written with two clusters’ worth of text. Then the PHOTOS directory and CAT.JPG were created and took clusters 6, 7, and 8. When the notes grew, the next free cluster was 9.
That’s fragmentation, and FAT has no defence against it other than picking the next free cluster and hoping it’s nearby. On a spinning disk that meant extra head movement. On flash storage the jump itself is cheap, but the chain still has to be followed one entry at a time.
06A directory entry is 32 bytes
We know how to find a file’s clusters once we know its first one. So where does the first one come from?
From a directory. In FAT32 a directory is an ordinary cluster chain whose content is a list of 32-byte records. The root directory starts at the cluster named in the boot sector, which is cluster 2 here.
This is the record for NOTES.TXT:
There is no inode. Part 1 made a point of the name living in one place and the metadata in another. FAT32 doesn’t separate them: the name, the size, the timestamps, and the pointer to the data all sit in this one record. That’s also why FAT32 has no hard links. A second name would need a second record, and nothing would keep the two in sync.
A few fields deserve a closer look.
The name
Eleven bytes: eight for the name, three for the extension, upper-case, padded with spaces. The dot isn’t stored. This is the famous 8.3 format, and it’s why old files have names like REPORT~1.DOC.
The first byte of the name doubles as a status flag. A value of 0xE5 means the entry was deleted, and 0x00 means there are no more entries in this directory.
The attributes
One byte of flags:
| Bit | Meaning |
|---|---|
| 0x01 | Read-only |
| 0x02 | Hidden |
| 0x04 | System |
| 0x08 | Volume label |
| 0x10 | Directory |
| 0x20 | Archive, set on ordinary files |
| 0x0F | All four low bits at once: a long-filename entry |
That’s the whole permission model. There are no owners, no groups, and no mode bits. When you mount a FAT32 volume on Linux or macOS and every file shows up as rwxrwxrwx, the operating system is inventing those permissions because the disk has nowhere to store them.
The timestamps
Dates and times are packed into 16 bits each. Take the date bytes 49 5d:
Years count from 1980. Times work the same way, with five bits for hours, six for minutes, and five for seconds. Five bits can’t count to 59, so seconds are stored divided by two, and FAT timestamps have two-second resolution. They’re also local time, with no time zone recorded.
The size
Four bytes. The largest number four bytes can hold is 4,294,967,295.
And there’s the error message from the start:
a FAT32 file can’t be larger than 4 GiB minus one byte, because its size has to fit in a 32-bit field.
Free space has nothing to do with it. The cluster chain could keep going. The directory entry just has no way to say how long the file is.
07Long filenames were bolted on later
Eleven upper-case characters was already cramped in 1995. But the directory entry couldn’t grow without breaking every existing DOS program and every disk already out there.
The solution, known as VFAT, is a trick. A long name is stored in extra directory entries placed directly in front of the real one. Each extra entry carries 13 characters of the name, and each is marked with the attribute byte 0x0F: read-only, hidden, system, and volume label all at once. That combination made no sense to old software, so old software skipped those entries.
Our auth.service.ts is 15 characters and has two dots, so it can’t be an 8.3 name. It takes three entries:
Long-name entry · sequence 0x42
attributes 0x0f · checksum 0x1e · 2, with 0x40 marking it as the last piece
Long-name entry · sequence 0x01
attributes 0x0f · checksum 0x1e · the first 13 characters
Short entry
attributes 0x20 · first cluster 10 · 28 bytes
The pieces are stored in reverse, with the last piece first. The sequence number’s 0x40 bit marks that last piece, so a reader immediately knows how many entries follow. Characters are stored as 16-bit Unicode, the name ends with a zero, and unused slots are filled with 0xFFFF.
The real entry still needs an 8.3 name, so one is generated: AUTHSE~1.TS. The dots are dropped, the first six characters are kept, and ~1 is added. A second clashing name in the same directory would get ~2. Each long-name entry also carries a checksum of that short name, so the pieces can be recognised as orphans if an old tool deletes the short entry without knowing about them.
It isn’t elegant. A 255-character name, the maximum, costs 20 extra entries. But it kept every old disk readable, and that mattered more.
08Finding /PHOTOS/CAT.JPG
Now we have every piece needed to open a file. Path resolution works the way Part 1 described: one directory lookup per component.
Boot sector
root directory starts at cluster 2
Cluster 2 · root directory
scan the entries → PHOTOS is a directory at cluster 6
Cluster 6 · PHOTOS
scan the entries → CAT.JPG starts at cluster 7, 8,192 bytes
FAT
entry 7 says 8, entry 8 says end of chain
Clusters 7 and 8
the file’s bytes
Each “scan the entries” step is exactly that. Directories aren’t sorted and have no index, so the filesystem reads entries one after another until a name matches, comparing against both the short name and any long name.
For a folder with a few dozen files, that’s nothing. For a camera folder with 20,000 photos it means reading through a great many records to find one of them. This is one of the problems ext4’s indexed directories exist to solve.
09Reading is easy. Seeking is a walk.
Reading a file from start to finish suits FAT well. Read a cluster, look up the next one in the table, repeat until the end-of-chain marker.
Jumping into the middle is different. Suppose a program asks for byte 9,000 of NOTES.TXT:
index = 9000 / 4096; // 2: the third cluster of the file
offset = 9000 % 4096; // 808 bytes into that cluster
cluster = 4; // first cluster, from the directory entry
cluster = fat[cluster]; // 5
cluster = fat[cluster]; // 9
// byte 9,000 is at sector(9) = 2132, plus 808 bytesThere’s no way to compute where the third cluster is. You have to ask the table twice. Two hops is nothing, but the cost grows with the file.
A 2 GiB video on this volume would occupy 524,288 clusters. Seeking to its midpoint means following 262,144 links from the start of the chain.
Real implementations soften this. Operating systems keep the table in memory and remember where they were in a chain, so a seek rarely starts from scratch. But the format itself only offers a linked list, and this is the cost that extents were designed to remove.
10What a write has to update
Let’s replay the moment NOTES.TXT grew from two clusters to three. Conceptually, the filesystem had to do all of this:
- 01Find a free cluster. The FAT is scanned for an entry holding 0, starting from a hint. It finds cluster 9.
- 02Mark cluster 9 as the end of a chain by writing 0x0FFFFFFF into FAT entry 9.
- 03Link it in. FAT entry 5 changes from end-of-chain to 9.
- 04Repeat both FAT changes in the second copy of the table.
- 05Write the new bytes into cluster 9.
- 06Update the directory entry: the size becomes 10,000 and the modified time changes.
- 07Update the free-cluster count in FSInfo.
The exact order varies between implementations, and much of it is buffered in memory before reaching the disk. What matters is the count: one small append touches the first FAT, the second FAT, a data cluster, a directory, and FSInfo. Those are separate places on the disk, written at separate moments.
What FSInfo is for
The FAT has no summary. To learn how much space is free, you’d have to read the whole table and count the zeros. On our small volume that’s about half a megabyte. On a 64 GB card with the same cluster size it’s roughly 64 MB of table.
FSInfo caches two numbers so that isn’t always necessary: how many clusters are free, and where to start looking for the next free one. On our volume they read 130,803 and 11.
Both are hints. If the volume wasn’t cleanly unmounted they can be wrong, and an implementation is allowed to ignore them. Linux, for example, doesn’t trust the stored free count unless you mount with the usefree option.
11Deleting barely touches the disk
Deleting NOTES.TXT is much less work than writing it was.
Before
After delete
Directory entry, first byte
4e “N”
e5 deleted
FAT entries 4, 5, 9
5 · 9 · END
0 · 0 · 0
Clusters 4, 5, 9
10,000 bytes of notes
10,000 bytes of notes
The first byte of the directory entry becomes 0xE5, and the three FAT entries are set back to zero. That’s all. The clusters now count as free, and the 10,000 bytes stay exactly where they were until something else is written over them.
This is why undelete tools work on memory cards. The directory entry still holds the rest of the name, the size, and the first cluster. The chain itself is gone, so a recovery tool has to guess that the file was contiguous. For a photo written to a freshly formatted card, that guess is usually right. For our fragmented notes it would recover the first 8,192 bytes correctly and then read the wrong cluster.
It’s also why deleting a file from a USB stick doesn’t erase it.
12No journal, no safety net
Look at the write steps again and imagine the power going out, or the card being pulled, somewhere in the middle.
If it happens after the FAT is updated but before the directory entry is, cluster 9 is marked as used and linked into a chain, but the file still claims to be 8,000 bytes. If a new file’s clusters were allocated and its directory entry never made it to disk, those clusters belong to nobody. They’re called lost clusters: taken according to the table, unreachable from any directory.
Part 1 described journaling and copy-on-write as ways to make related updates succeed or fail together. FAT32 has neither. Its protections are modest: the second copy of the table, the backup boot sector, and a flag that records whether the volume was cleanly unmounted.
After an unclean removal, a checker such as fsck.fat on Linux or chkdsk on Windows walks every directory and every chain and looks for disagreements. It can free lost clusters or turn them into files so you can inspect them. It can’t know what the data was supposed to be.
So “safely remove hardware” isn’t a ritual. On FAT32 it’s the only thing standing between a half-finished write and an inconsistent table.
13Why it’s still everywhere
Put the limits side by side and FAT32 doesn’t look competitive:
Largest file
4 GiB minus one byte
Largest volume
About 2 TiB with 512-byte sectors
Filenames
8.3, or up to 255 characters with long-name entries
Timestamps
Local time, two-second resolution
Permissions and owners
None
Hard links
None
Crash recovery
A second copy of the FAT, and a repair tool
And yet it’s on nearly every device you own, for the same reason it has those limits. The whole format is a header, an array, and 32-byte records. A microcontroller with a few kilobytes of memory can implement it. Every operating system written in the last thirty years can read it. The firmware that starts your computer reads a FAT partition before any operating system is running.
When two devices that know nothing about each other need to exchange files, the simplest format that both can implement wins. For a long time, that has been FAT.
14Where we’ll go from here
FAT32 answers the four questions from Part 1 in the plainest way possible. Names live in 32-byte records. Space is found by scanning a table for zeros. Metadata is a handful of fields next to the name. Consistency is mostly left to a repair tool.
Every weakness we ran into is something a later filesystem set out to fix: the linear directory scans, the walk through a chain to reach the middle of a file, the missing permissions, the lack of any protection against a crash halfway through a write.
ext4 takes on all four. It separates names from metadata with inodes, describes file data with extents instead of chains, indexes large directories, and records changes in a journal before making them.
Next up
ext4, and what changes once a filesystem has inodes, extents, and a journal.
Further reading
- File System Implementation, Operating Systems: Three Easy Pieces
- FAT, OSDev Wiki
- The FAT driver in the Linux kernel source
- Design of the FAT file system, Wikipedia
- VFAT, Linux kernel documentation
The hex dumps come from one small example volume. Other formatting tools choose different reserved-sector counts, cluster sizes, and OEM names, so your numbers will differ.