DEV Community

Cover image for The 21 Rules of the File System---Writing a COW Filesystem from Scratch, Part 4
faliye
faliye

Posted on AI-assisted

The 21 Rules of the File System---Writing a COW Filesystem from Scratch, Part 4

In the previous article, I documented my collaboration with AI.

In this article, I’ll add a few details about the basic work specifications and use them as the starting point for the discussion that follows.

How do you draw an architectural design on a blank sheet of paper? With a grid.

How did the Renaissance masters draw so many towering, magnificent figures with such precision? With dividing lines.

Let me loosely imitate the 21 Articles of War and write something—drawing a few lines for this filesystem.

Contributors

Article 1. We don't distinguish between submitters; we review only the evidence. A patch that passes the gate deserves to be taken seriously. Show me the test.

Decisions, Design, and Experiments

Article 2. Every decision must be backed by experiments, unless explicitly exempted. Experiment files are committed to git.

Article 3. Every decision and design must leave a trace, written up using the standard template.

Article 4. No decision is true forever; a decision that is challenged must be verified.

Article 5. This project learns from the experience of excellent past projects, but its implementation must fit this project's own character.

Format

Article 6. The data format is self-contained and carries back-references. Data can prove its own identity and what it belongs to.

Article 7. Index trees and the like are derived structures and may be discarded. Data recovery never depends on the tree; the tree serves only as a lookup accelerator and aid.

Article 8. Following Article 7, derived structures stay pure: no small data may be stored in them.

Article 9. Stripe width is variable, and full-stripe writes never incur RMW (read-modify-write).

Article 10. No false ENOSPC: space reported as available must actually be usable.

Article 11. Encryption is a first-class capability: slots are reserved for it, but it is off by default.

Article 12. Slots are reserved for future extensions, with vendor interfaces provided as circumstances warrant.

Article 13. Data units are 32 KiB and index nodes are 16 KiB. 32 KiB is dictated by checksum granularity; 16 KiB is a trade-off made for the CPU cache. Neither is chosen to align with the SSD's write granularity.

Testing and Recovery

Article 14. Data recovery code and test infrastructure rank equal to the core code and are developed in lockstep with it.

Article 15. The verification implementation and the filesystem implementation share no code; each is implemented independently, sharing only the format.

Article 16. Crash testing must cover scenarios exhaustively, never by sampling, and must not be treated as a substitute for broad-spectrum testing or formal verification.

Priorities

Article 17. Keep the early stage pure. Early on, ignore the Linux branch and develop around transactions.

Article 18. The first line supports SSDs only; the disk type is held as a placeholder in incompat. Other disk types can be specialized later in separate branches.

Article 19. Until the format is frozen, don't optimize the filesystem code and don't chase performance.

Article 20. When designing garbage collection, consider the possibility of GPU-based collection algorithms.

Slogan

Article 21. The slogan is still under consideration.....


End of text.


A Few Words on These Articles

Contributors reflects the idea of community collaboration. One person's strength is, in the end, limited.

Article 1. This is a project with AI as its main workforce, carrying AI in its DNA from day one. We enjoy the conveniences of this era, and we should also rein in its shortcomings. An AI-driven project has unlimited throughput, but human energy is finite. Submission throughput can be unlimited; acceptance throughput is bounded by evidence.

Decisions, Design, and Experiments is the operating manual handed to the AI during the design phase. It is designed to minimize AI hallucination, raise the AI's autonomy, and make sure decisions, designs, and experiments don't lose their context.

Article 2. Decisions and designs are not daydreams; they must be validated by experiments, finding their footing in the code of individual experiments. A decision or design with no footing is a fantasy. A decision may be wrong, but it must be shown to have been argued through. Experiments, as part of the evidence, should be kept on record permanently.

Article 3. In an AI-driven project, the knowledge base is the most important part of the whole engineering effort. The archived knowledge will provide key grounds for future decisions and retrospectives. Write documents according to the standard template; the template may change, and when it does, have the AI update the documents one by one.

Article 4. Many decisions may become outdated during construction. Given the AI's ability to question history and context, we give it options and prompts through which it can challenge them. At the same time, this questioning must not be left unchecked, or the project will slow to a crawl.

Article 5. This filesystem's design differs from past filesystems. We stand on the shoulders of earlier filesystems and absorb their excellent algorithms and designs, but we must also stay distinctive. We must constantly remind the AI that this is a different project and that earlier implementations can't simply be copied.

The data-structure section is the core implementation of singlefs and gives a brief account of its data design.

Article 6. singlefs data comes in a fixed size of 32 KiB. Each unit contains a data-unit header and the payload, locked together and never stored separately. This way the data naturally carries its own identity, and its back-references state what it belongs to—which tree, which object. The data can thus explain its own situation, with no need for index-tree records to prove its identity and ownership.

Article 7. Following Article 6, if data can vouch for itself, the tree that records it is demoted. Since the data can state its own identity and ownership, in the worst case the index tree can be thrown away and rebuilt.

Article 8. True, under Article 6 a 32 KiB data size makes small files balloon in storage. But following Article 7, the index tree must remain pure. singlefs stores no small data or other structured information in the index tree. The current idea is to solve this with small-data packing: bundling multiple small pieces of data into a single 32 KiB unit.

Article 9. A stripe is a concept of storing data across disks: to keep data safe, you must write to several disks at once. In traditional RAID 5, for instance, three disks are required—two for data and one for parity—so every write is a fixed 64 KiB. But what if you have only 32 KiB of data? Then you must read the old data, recompute parity, and write it back. If power is lost in the middle of this, the data and parity no longer match (the write hole). singlefs instead uses variable-width stripes. In the same situation, storing 32 KiB of data writes to just two disks; storing 64 KiB writes to three. No read is needed, and every write exactly fills the stripe. The risk of data loss on power failure and the space overhead are both much smaller.

Article 10. In a traditional filesystem, the index tree is authoritative, so updating a file has to go through it. Once the index tree has no room left to write, files can no longer be written. And since deleting a file also requires modifying the index tree, deletion fails too, and the filesystem locks up. singlefs's self-containment can solve this problem to some extent, but we are still exploring.

Article 11. File encryption is a capability designed into the format from day one, but the first runnable version won't implement it; it only reserves the format bits (all zeros when disabled). Once the transaction layer and the checker are running, real encryption will be wired in. Whether it can be rolled back after being switched on is not yet decided.

Article 12. singlefs reserves a portion of space for future extension fields. We are also considering whether to offer an extend-like field that SSD vendors can open up. Our selling point is waste.

Article 13. The 32 KiB data unit is not actually designed to cater to SSDs. The checksum has to cover the whole unit, and the aim is to bind checksum granularity and extent granularity together for simpler logic, at the cost of paying roughly ten percent more on small random reads. 16 KiB is the size of an index node, a number settled by weighing CPU-cache considerations. A 16 KiB node and a 32 KiB data unit are two different kinds of unit; a data unit is not two index nodes joined together.

Testing and Recovery sets singlefs's own standard for testing and file recovery.

Article 14. Because singlefs writes a filesystem from scratch, and its underlying data format differs in places, it needs a test system and methods of its own design. Data recoverability is singlefs's top priority, so test code, recovery tools, and core code are implemented on the same day. Each milestone cross-verifies the others.

Article 15. If tests and core code shared the same logic, they would easily converge on the same bugs. So tests and core code are implemented separately.

Article 16. Beyond the routine adversarial cross-checking among three bodies of code, the core test of singlefs is exhaustively simulating power loss in QEMU and examining the state of the data on each stream, and whether it is recoverable. Broad-spectrum testing exists to catch bugs in the code and to keep the AI from producing results that merely look correct when it writes code. Formal verification is indispensable for concurrent features. singlefs aims to find bugs as early as possible through rigorous testing, while also improving observability.

Priorities were designed according to singlefs's own characteristics during development.

Article 17. Because singlefs is developed around transactions, work on upper-layer interfaces such as FUSE will be pushed back, which may delay compatibility with the Linux community. But it lowers the complexity of early development and lets us focus on the flow of data. I'm short on energy, and short on tokens too.

Article 18. singlefs defines incompat to distinguish disk media types (rotating disks, SSDs, ZNS, and so on), in the hope of optimizing each kind of media to the extreme on top of this setting. But for the same reasons—not enough energy, not enough tokens—the first line supports only generic SSDs.

Article 19. Across singlefs's milestones, we will compare against other filesystems horizontally, but only on the basis of per-node data records. Before the format is frozen, we will not optimize the core filesystem code for performance.

Article 20. singlefs's garbage collection will be heavy. Besides a standard GC running on the CPU, we'd like to bring in the GPU to do some of the work. And of course, we also hope to make use of features such as FTL and ZNS.

Article 21. [Big disks belong to the company; good sleep belongs to you?] [WAAAGH! Reckon it'll work!]

Well, it's finally written. The keyboard now belongs to you, Agent bros.

run Tokens, run!!!

Top comments (0)