← Writing

Computer Internals · Part 1

FilesystemsWhat Actually Happens When You Save a File?

Understanding how your operating system turns raw storage into files, folders, and something you can actually work with.

June 18, 2025 · 18 min read

Your SSD doesn’t know what package.json is.

It doesn’t know that your project has a src directory, that auth.service.ts contains your authentication logic, or that the PDF you downloaded yesterday is something you might need tomorrow.

It stores data. The operating system and filesystem are responsible for making that data useful.

Think about what happens when you press Cmd + S in VS Code. You expect your changes to be saved, your file to retain its name, and everything to be there when you reopen your laptop. That simple action depends on several layers of software working together, from the application all the way down to the storage device.

I’ve always found it interesting how much engineering sits behind something we barely think about.

So let’s start at the bottom and work our way up. We’ll look at storage devices, partitions, blocks, metadata, and what actually happens when the operating system tries to find a file.

By the end, you should have a mental model that makes it much easier to understand why filesystems such as FAT32, ext4, NTFS, and APFS work the way they do.

01The disk is just the starting point

Let’s take a 512 GB SSD.

Your operating system, applications, source code, photos, and everything else live somewhere on this device. But the SSD doesn’t organize that information into folders.

At the interface exposed to the operating system, a typical SSD behaves like a block device. The operating system can request data from particular logical addresses without needing to understand how the device physically stores it.

Imagine the storage as a long sequence of numbered regions:

Each address identifies a location in the device’s logical address space. The storage interface defines how much data each logical sector represents.

For example, a device using 512-byte logical sectors lets the operating system address storage in units of 512 bytes. Other devices expose 4096-byte logical sectors.

The operating system can request a read or write at a particular location. It doesn’t need to know whether those bytes belong to a database, a photograph, or a JavaScript file.

Conceptually, the interface looks something like this:

read_sectors(start_sector, count, buffer);
write_sectors(start_sector, count, buffer);

This is illustrative pseudocode, not a literal API shared by every operating system.

The important idea is that the device provides access to storage locations. Organizing those locations into meaningful files is a different responsibility.

HDDs and SSDs solve the physical problem differently

A hard disk drive stores data magnetically on spinning platters. Reading data involves positioning a mechanical head over the appropriate region.

Simplified. A real platter has hundreds of thousands of tracks.

This makes access patterns important. Reading adjacent regions is generally much cheaper than repeatedly jumping between distant locations.

An SSD uses flash memory instead. There are no spinning platters or moving heads, but the hardware has its own constraints.

NAND flash is programmed and erased in different granularities. An SSD controller handles the translation between logical addresses and physical flash locations, along with wear leveling, garbage collection, and other device-management tasks.

Simplified. A real erase block holds hundreds of pages, each several kilobytes.

This is why an SSD cannot simply overwrite flash cells in exactly the same way that a program overwrites a variable in RAM.

The interesting part is that both devices can expose a similar block-storage interface despite their very different internals.

The filesystem can work with logical storage addresses without managing the physical movement of an HDD head or the flash translation layer inside an SSD.

That’s our first useful abstraction:

the storage device handles physical storage, while the filesystem organizes the logical address space it exposes.

02Before files, there are partitions

A physical disk can be divided into separate regions called partitions.

Suppose you have a 1 TB drive and want to reserve separate areas for an operating system and your data. A partition table records the boundaries of these regions.

The sizes here are illustrative. Real layouts include other considerations, such as boot partitions, recovery partitions, encryption, and logical volume management.

Each partition describes a range of logical storage addresses. The operating system can expose that range as a separate block device for other software to use.

A partition doesn’t automatically contain a filesystem. You can create a partition and then format it with a filesystem, or use it for another purpose entirely.

The partition table is responsible for describing how the disk is divided.

Two formats you’ll commonly encounter are MBR and GPT.

MBR

MBR is the older partitioning scheme. Its traditional layout supports four primary partition entries, with extended partitions providing a workaround for additional logical partitions. Its 32-bit sector addressing also imposes a roughly 2 TiB limit when used with 512-byte sectors.

GPT

GPT is the modern alternative. It uses 64-bit logical block addresses, supports much larger disks, and stores backup partition-table information near the end of the disk.

On a Mac, you can inspect the storage layout with:

macOS
diskutil list

On Linux, try:

Linux
lsblk

These commands help you see the relationship between physical devices, partitions, and the storage layers built on top of them.

One distinction is worth keeping in mind:

GPT and MBR describe disk partitioning, not how files and directories are organized inside a partition.

That’s the filesystem’s job.

03So what does the filesystem actually do?

We’ve got a storage device and, in our example, a partition.

Now imagine asking the operating system to open this file:

/Users/vineet/projects/my-app/package.json

The operating system needs to figure out where that file lives, how large it is, whether you have permission to access it, and which storage locations contain its contents.

The storage device doesn’t answer those questions by itself.

A filesystem maintains the structures needed to answer them.

Different filesystems implement those structures differently, but they all have to deal with a few fundamental problems:

1

How do we map names to files?

2

How do we find space for new data?

3

How do we track metadata such as file size and permissions?

4

How do we keep the filesystem consistent when operations fail?

Let’s look at these problems individually, starting with something familiar: filenames.

04A filename isn’t the file itself

Consider this path:

/Users/vineet/projects/my-app/package.json

It’s a sequence of names. To resolve it, the operating system needs to look up each component in the appropriate directory.

Conceptually, the process looks like this:

Each lookup identifies the next filesystem object. Once the operating system has resolved the final component, it can obtain the information needed to access the file.

The actual implementation depends on the filesystem and operating system. Directory entries may be indexed, and the operating system may already have relevant metadata cached in memory.

But the core idea is straightforward: a path is a way to identify a file through a hierarchy of names.

This separation between names and underlying file objects is important.

If you rename package.json to package-old.json, the filesystem usually doesn’t need to rewrite the file’s entire content. It can update the relevant directory entry instead.

The data and the name used to reach it are separate things.

That distinction becomes particularly interesting when we look at inodes.

05Where does a file’s metadata live?

Run this command in a directory containing a file:

ls -l package.json

You might see output similar to:

There is more information here than just the filename and content.

The output includes permissions, ownership, file size, and a modification timestamp.

This information is called metadata.

Depending on the filesystem, metadata can also include file type, link count, timestamps, and information about where the file’s content is stored.

The interesting part is how different filesystems represent it.

FilesystemMetadata organization
FAT32Directory entries and the File Allocation Table
ext4Inodes, directory entries, and block-group metadata
NTFSMaster File Table records
BtrfsStructured metadata organized using B-trees and related structures

They all need to solve broadly similar problems, but their internal designs differ considerably.

The inode is a good place to start

In Unix-like filesystems such as ext4, an inode stores metadata about a filesystem object.

A simplified view looks like this:

The diagram is conceptual rather than an exact on-disk layout.

Here’s the detail that makes inodes worth understanding:

the filename is not stored in the inode itself.

The directory entry maps the name package.json to an inode number. The inode contains metadata and information used to locate the file’s content.

This separation makes hard links possible. Multiple directory entries can refer to the same inode, giving one underlying file object multiple names.

It also explains why an inode number is not the same thing as a filename or a permanent physical disk address.

Inodes are specific to filesystem designs that use them. FAT32 and NTFS organize their metadata differently, which we’ll explore in later articles.

06Files need somewhere to put their bytes

Let’s create a small file.

printf 'Hello, filesystem!\n' > hello.txt

The file now contains a short string. But where does the filesystem put those bytes?

It needs to allocate storage for the file’s content and maintain enough information to find it later.

Filesystems generally manage file storage using allocation units. Depending on the filesystem, these are commonly called blocks or clusters.

Imagine a filesystem that allocates storage in units of 4096 bytes.

A file doesn’t necessarily fill the entire block. If a 100-byte file occupies a 4096-byte allocation unit, most of that unit remains unused by the file.

This is one trade-off of fixed-size allocation.

Why not allocate exactly the number of bytes required?

Because the filesystem also needs an efficient way to manage storage. Tracking arbitrary-sized allocations can introduce additional complexity. Fixed-size units simplify many aspects of allocation and accounting.

The trade-off is that small files can waste space inside their allocated units.

For example, if a simplified filesystem allocates one 4096-byte block per file:

File content sizeAllocated capacityUnused capacity
100 bytes4096 bytes3996 bytes
1000 bytes4096 bytes3096 bytes
3000 bytes4096 bytes1096 bytes
4096 bytes4096 bytes0 bytes

This unused capacity is an example of internal fragmentation.

Real filesystems can use optimizations such as inline data or other techniques to reduce overhead in particular situations, so the table is a simplified model.

There’s another useful distinction here: the size of a file and the amount of storage allocated to it aren’t always the same. Sparse files, compression, and filesystem-specific optimizations can make the relationship more complicated.

07How does the filesystem keep track of allocated blocks?

Finding storage for a file sounds easy until you consider the scale.

A large disk can contain millions or billions of allocation units. Some are occupied by file content, some by metadata, and others are free.

When a new file is created, the filesystem needs to find suitable space without searching the entire device from scratch.

One common approach is a bitmap.

Imagine that each bit represents one allocation block:

For this example, 1 means occupied and 0 means free.

The filesystem can use this information to identify available blocks. In practice, it also maintains other metadata and optimizations to make allocation efficient.

Another approach is a free-space list. More sophisticated implementations can maintain indexed structures describing free ranges of storage.

The choice affects how quickly the filesystem can find free space, how efficiently it can allocate contiguous regions, and how much metadata it needs to maintain.

What about fragmentation?

Suppose a file needs three blocks.

The filesystem might allocate blocks 101, 103, and 104 because those are available. The file’s content is now spread across different regions rather than occupying one contiguous range.

The filesystem needs a way to record which blocks belong to the file and in what order their contents should be interpreted.

There are several ways to do this.

Linked allocation

A linked allocation scheme records a chain of blocks belonging to a file.

The File Allocation Table used by FAT filesystems follows this general idea. The table records which cluster comes next in a file’s chain, along with special values indicating conditions such as the end of the chain or a free cluster.

This design is relatively simple, but following a long chain to reach a later part of a file can be expensive.

Extents

Another approach is to represent contiguous regions as extents.

Instead of recording every block individually, an extent describes a range using a starting location and a length.

The file occupies three separate regions, each of which is contiguous.

Filesystems such as ext4 and XFS support extents. They can represent large contiguous regions efficiently without maintaining an independent entry for every block in the file.

The filesystem still needs to track multiple extents when a file is fragmented, and it needs suitable structures to locate them.

These approaches illustrate an important theme in filesystem design: finding space is only half the problem. The filesystem must also record the allocation in a way that makes future reads, writes, and changes efficient.

08Directories are more interesting than they look

We tend to think of a directory as a container for files. That is a useful mental model, but internally a directory is structured data that helps the filesystem resolve names.

In many Unix-like filesystems, directories are represented as filesystem objects with their own metadata and data blocks. Those blocks contain directory entries that map names to inode numbers.

For example:

This is an illustrative mapping, not a literal directory format.

When the operating system looks for auth.service.ts, it uses the directory’s structures to find the corresponding file object.

A small directory might be easy to search directly. A directory containing hundreds of thousands of entries needs a more efficient strategy.

Different filesystems use different approaches, including indexed directories, hash-based structures, and tree-based indexes.

This is why filesystem internals become much more interesting than simply storing a list of names.

09What tells the operating system how the filesystem is organized?

The filesystem needs some way to locate and interpret its own metadata.

In Unix-like filesystems, one important structure is the superblock. It contains information about the filesystem’s overall layout and state.

Conceptually, that information might include:

The actual fields vary by filesystem. Some filesystems maintain backup copies of critical metadata, while others use different arrangements.

You might also encounter the term boot sector. It’s important not to treat a boot sector and a superblock as interchangeable terms in every context. Boot sectors are associated with boot-related or volume-layout information in certain formats; a superblock is a filesystem-specific structure describing important filesystem properties.

These structures also belong to a different layer from the partition table. GPT tells the operating system how the disk is partitioned. Filesystem metadata tells it how a particular filesystem organizes its own contents.

Keeping these layers separate helps when diagnosing storage problems. A damaged partition table and a corrupted filesystem are not the same failure, even if both make files inaccessible.

10Saving a file is one thing. Keeping it safe is another.

So far, we’ve looked at the structures required to locate files and manage their storage.

But there’s a problem.

What happens if the computer loses power halfway through an operation?

Imagine an application is updating a file. Some changes have reached the storage device, others are still buffered, and filesystem metadata may not yet reflect the complete operation.

The filesystem must deal with the possibility that related updates are interrupted.

Without appropriate safeguards, a crash could leave metadata inconsistent, cause allocation information to disagree with actual usage, or make recently written data unavailable.

Different filesystems use different recovery strategies.

Journaling

A journaling filesystem records information about certain updates in a journal so it can recover from interrupted operations.

A simplified transaction might look like this:

If the system crashes, the filesystem can inspect the journal and recover according to its consistency rules.

Filesystems such as ext4, XFS, and NTFS use journaling mechanisms, although their implementations and guarantees differ.

One detail matters here: journaling does not automatically mean every application write is safe against every possible power failure.

A filesystem might journal metadata without journaling all file content. Application-level durability also depends on how writes are issued, whether synchronization is requested, and whether the storage stack correctly honors persistence requirements.

Journaling helps protect filesystem consistency, but it is not a replacement for backups.

Copy-on-write

Copy-on-write, or CoW, takes a different approach to updating data structures.

Instead of modifying certain existing structures in place, a filesystem writes updated versions to new locations and then switches the relevant references to the new version.

The old version can remain intact until the new version and the relevant metadata updates are ready to become visible.

Btrfs and ZFS use copy-on-write designs for important parts of their filesystem structures.

CoW can enable useful features such as snapshots, but it introduces its own trade-offs, including additional writes and more complicated space management.

It also doesn’t eliminate every possible failure mode. Correct transaction handling and reliable persistence remain essential.

What does filesystem repair actually do?

On Linux, you may encounter filesystem-checking tools commonly referred to as fsck, with filesystem-specific implementations.

These tools can inspect filesystem structures and attempt to repair certain inconsistencies.

They are useful, but they cannot guarantee recovery from every kind of corruption or recover data that has been irretrievably lost.

The broader lesson is that filesystem design has to account for failures, not just successful reads and writes.

11Putting it together: opening package.json

Let’s return to the file we started with.

/Users/vineet/projects/my-app/package.json

When an application requests this file, the operating system resolves the path through the directory hierarchy, obtains the relevant file metadata, checks access permissions, and determines how to retrieve the content.

The operating system may already have some of this information cached. If the file is modified, the filesystem and storage stack must coordinate the associated data and metadata updates.

A simplified view looks like this:

This is a conceptual architecture, not a literal sequence of every operation. Reads can be satisfied from caches, writes may be buffered, and modern systems can include additional layers such as encryption and volume management.

Still, the diagram captures the important relationship: applications work with files, filesystems organize them, and storage devices provide the underlying capacity.

The filesystem is the layer that makes a raw storage device usable as a collection of named, structured objects.

12The design choices become clearer once you know the basics

At this point, the differences between filesystems start to make more sense.

FAT32

FAT32 uses a relatively simple allocation-chain design, which helps explain both its portability and some of its limitations.

ext4

ext4 uses inodes, block groups, and extents to organize metadata and data efficiently.

NTFS

NTFS organizes a great deal of its metadata around the Master File Table.

Btrfs · ZFS

Btrfs and ZFS use copy-on-write designs that support features such as snapshots and checksumming, while introducing their own trade-offs.

These filesystems are not merely different names for the same thing. They make different decisions about data structures, allocation, compatibility, performance, and recovery.

Understanding those decisions is more useful than memorizing a list of features.

13Where we’ll go from here

This article gives us the foundation. Now we can look at actual implementations and understand why they were designed the way they were.

We’ll start with FAT32 and follow a file through its cluster chain. Then we’ll move to ext4, look at inodes and block groups, and understand how its allocation structures work. From there, NTFS, XFS, Btrfs, and other filesystems will give us different perspectives on the same underlying problems.

The goal of this series is simple: when we encounter a filesystem feature, we should be able to understand the problem it solves and the trade-offs it introduces.

Because once you understand the internals, a filesystem stops looking like a mysterious layer that somehow makes files appear on your computer. It starts looking like what it really is: a collection of data structures, algorithms, and carefully designed rules that make storage reliable and useful.

Next up · Part 2 →

FAT32, and how a simple table can keep track of an entire file.

Further reading

The diagrams in this article are simplified mental models, not exact on-disk layouts. Filesystem implementations and operating-system behavior vary.

  • Filesystems
  • Operating Systems
  • Computer Internals
all posts