M37

What a database keeps when the machine dies

This is the material behind the video. Four database systems were run in thirteen arms — ten database configurations and three controls — and killed forty times, in two ways that are not the same thing: ten kills of the database process alone, and thirty kills of the whole machine. After every kill, every write the database had already told its client was committed was checked against what actually came back. That produced 500 trial-arm outcomes, and all of them are below.

The protocol came first, and it is on this page in full rather than summarised. It was committed on 2026-09-01 at 22:55 Berlin time, commit e4f5d4f3; the first trial data was committed 00:01 the following morning, commit 341ed1e3. The repository is private, so that is a statement rather than a link, and it is the one thing on this page you cannot check for yourself. Everything downstream of it you can.

The protocol, as registered

Nothing below was written or amended after a result was seen. One thing was added part-way through the run, and it is declared at the end of this section with its reason; the original text was left standing.

The question

When a database tells a client that the write is committed, and the machine then dies, is the write still there when it comes back?

That is the plain-language form of what every one of these systems promises somewhere in its own documentation. The test is whether the promise holds as configured out of the box, and where the documented durable setting differs from the default.

Two kinds of kill, and they are not the same thing

This is the distinction the whole run turns on.

Process crash
SIGKILL to the database server process. The operating system survives, so anything the database had handed to the kernel with write() is still in the kernel's page cache and reaches the disk normally afterwards. A database can pass this while having fsynced nothing at all.
Power loss
The entire virtual machine process is SIGKILLed. The guest kernel's page cache dies with it, so only bytes the database explicitly fsynced, and which therefore reached the virtual disk, survive.

A process-crash result is not a power-loss result and none of them is described as one here. The process crash is in the run precisely because it is the cheap test, and the gap between the two tiers is the finding.

Stable storage here is the disk image file on the host. Killing the guest virtual machine does not touch it, so a guest fsync that reached the host file is durable by construction. The controls below are what establish that this is actually true of the setup rather than assumed.

The workload, and what acknowledged means

One client per arm, in a strict loop: insert one row whose primary key is a rising integer, in its own committed transaction, and wait for the database's commit acknowledgement; send that number to a collector process running on the host, outside the virtual machine; wait for the collector to confirm it has written and fsynced the record; only then move on to the next number.

The collector's log is the record of what was acknowledged, and it lives outside the machine being killed, so a kill cannot destroy it. Waiting for the collector is what makes the boundary exact: at the moment of the kill at most one number can be in flight, so the highest number the database definitely acknowledged is either the last one the collector confirmed or the one after it.

That one-record ambiguity is resolved in the database's favour. A missing final record is never counted against a system. This biases the whole run toward finding fewer failures than really occurred, and every number here should be read with that in mind.

The kill, and what counts as a pass

The kill fires at a uniformly random time in a five to fifteen second window after the client starts, so it lands at an arbitrary point in each system's own flush cycle rather than at a moment chosen to suit anyone. Afterwards the system is restarted, given up to 120 seconds to become available, and the table is read in full.

A trial passes when every acknowledged record is present. Nothing else is required: not ordering, not timing. Each trial ends in exactly one of four verdicts, and those are the words in the run log below.

clean
Every acknowledged write survived.
lost-tail
The highest surviving record is below the last one acknowledged, and everything below the survivor is present. Acknowledged writes were lost from the end, and the count lost is recorded.
hole
Some acknowledged record is missing while a higher one survived. Worse than a lost tail: what came back is not a prefix of what was promised. No trial in this run ended here.
corrupt
The system did not restart, or the table could not be read at all.

Records present that were never acknowledged are counted as extra and are not a failure: a commit can land after the kill has interrupted the acknowledgement path.

The controls, which run first and without which nothing counts

The instrument has to be shown capable of producing both answers before any of its answers count for anything.

The positive control
A small program that appends a record and fsyncs it before acknowledging, over the same collector protocol. It must come through power loss clean. If it loses records, the virtual disk is not honouring fsync, every database would look guilty for a reason that has nothing to do with the database, and the run is void.
The negative control
The same program with the fsync removed. It must lose records under power loss. If it does not, the kill is not discarding the page cache, that tier has no dynamic range, and a clean result from it would mean nothing.

Both also run under the process crash, where both are expected to come through clean. That is the clearest available statement of what a process crash does and does not test.

What would have falsified what

Written before the run, so that no result could be reinterpreted afterwards as agreeing with whatever it turned out to be.

The strong claim
A loss under power loss for PostgreSQL at stock, MySQL at its default flush setting, SQLite at its library defaults, or MongoDB with j:true would contradict those systems' own documented guarantees. That one is held to the highest bar: any such result is re-run and reported only if it reproduces.
The expected losses
Losses under power loss for the other settings confirm documented behaviour and are not defects. They are interesting only because several of them are defaults, or near-universal habits, and the loss they permit is larger or likelier than a reader of the headline promise would guess.
The cheap tier
Any arm passing the process crash proves nothing about power loss and is not reported as if it did.
The null result
If every arm had come through every kill clean, the finding would have been that these systems keep their promises, and that would have been reported as the result.

What this run does not cover

Stated in advance rather than discovered afterwards.

One node
No replication, no failover, no network partitions. MongoDB's w:1 on a single node is a weaker setting than a replica set would use, and nothing here says anything about a real cluster.
Not a benchmark
Throughput is deliberately crippled by the collector round trip, and no performance number from this run means anything.
A virtual disk, not physical hardware
Real hardware adds a drive write cache that can lie about fsync. This test cannot see that layer, so a real machine can do worse than these numbers, never better.
Distribution defaults
The configurations as the Ubuntu packages ship them, which are not always identical to what a cloud provider or a container image gives you.
No fault injection
No data-corruption fuzzing, no torn-page injection, nothing at the filesystem level.

The one amendment, declared

A third control was added after the first twenty power-loss trials, when the SQLite-at-defaults result needed a channel tested that the original two controls did not cover. A rollback-journal commit ends by deleting the journal and fsyncing the directory, and the positive control had only proved that fsync of file data was durable here. The added control does the directory operation and nothing else.

It is a control, not an arm: it changes no arm's configuration and no verdict rule. Power-loss trials 21 to 30 carry it, which is why its power-loss count is ten where the other arms have thirty. The original protocol text was left standing; this paragraph is the appended amendment.

What it ran on

Ubuntu 24.04.4 LTS, kernel 6.8.0-138 aarch64, in a disposable Lima virtual machine with 4 vCPU and 3 GiB of memory on an Apple M1 Max. Data on ext4 on the virtual machine's own virtual disk.

The numbers in the video

10 of 10
Process-crash trials survived by a program that never calls fsync at all
n = 10
What it does not mean. It does not mean that program is durable. Under power loss the same program lost every record it had ever acknowledged, 392 to 10,593 of them, in all thirty trials. Passing the cheap test says nothing about the expensive one, and that is the whole argument of this run sitting in one control.
0
Acknowledged writes lost by PostgreSQL at stock, MySQL at its default flush setting, and MongoDB with j:true, across ninety power cuts
n = 90
What it does not mean. It does not mean those settings are fast, and it does not carry to physical hardware, whose drive write cache can lie about fsync in a way this test cannot see. Ninety trials also cannot detect a failure much rarer than about one in thirty.
495
Acknowledged writes lost in a single power cut by MySQL's flush setting of 2, whose documentation says about a second of transactions
n = 30
What it does not mean. It is one trial's count, not a rate and not a second's worth of anything in normal use. Every arm here waits for a round trip to a collector outside the machine after every commit, so how far an arm gets in five to fifteen seconds is not a throughput figure. The documented claim is accurate; a second is just worth more than it sounds.
4 of 30
Power-loss trials in which SQLite at its library defaults lost exactly one acknowledged write
n = 30
What it does not mean. It is not a proof of a defect in SQLite, and this page does not claim one. It is one build on one kernel in one virtual machine, re-run twice more under identical conditions and reproduced. The loss was never more than a single write, never a gap in the middle, and never corruption.

Every arm, under both kinds of kill

One row per arm per kill. Clean counts the trials in which every acknowledged write came back. Writes lost is the range of acknowledged writes that went missing across that arm’s failing trials, so an arm that never failed has no range rather than a zero. Opening a row shows the full verdict breakdown, the failure rate with its interval, and the manual the documented claim was read from.

Showing 25 of 26 rows. In the order the run produced them.

Row detail
PostgreSQLstock (synchronous_commit=on)stock (synchronous_commit=on)process crash1010of 100no failing triala reported-committed transaction is on disk and survives a crash
PostgreSQLsynchronous_commit=offsynchronous_commit=offprocess crash103of 1071 to 6recent committed transactions may be lost; never corrupted
MySQL / InnoDBinnodb_flush_log_at_trx_commit=1 (default)innodb_flush_log_at_trx_commit=1 (default)process crash1010of 100no failing trialfull ACID; log flushed and fsynced at each commit
MySQL / InnoDBinnodb_flush_log_at_trx_commit=2innodb_flush_log_at_trx_commit=2process crash1010of 100no failing trialsurvives a mysqld crash; an OS or power loss can lose about a second
MySQL / InnoDBinnodb_flush_log_at_trx_commit=0innodb_flush_log_at_trx_commit=0process crash1010of 100no failing trialabout a second can be lost even when only mysqld crashes
SQLitejournal_mode=DELETE, synchronous=FULL (defaults)journal_mode=DELETE, synchronous=FULL (defaults)process crash1010of 100no failing trialcommitted transactions survive power loss
SQLitejournal_mode=WAL, synchronous=NORMALjournal_mode=WAL, synchronous=NORMALprocess crash1010of 100no failing trialdurable across app crash; power loss can lose recent commits, without corruption
SQLitejournal_mode=WAL, synchronous=OFFjournal_mode=WAL, synchronous=OFFprocess crash1010of 100no failing trialthe database may become corrupted on power loss
MongoDBw:1 (default)w:1 (default)process crash103of 1071 to 2acknowledged by the primary; can be lost if the node fails before the journal flush
MongoDBw:1, j:truew:1, j:trueprocess crash1010of 100no failing trialacknowledged only once the write is in the on-disk journal
controlappend + fsync each recordappend + fsync each recordprocess crash1010of 100no failing trialmust survive Tier V or the run is void
controlappend, no fsyncappend, no fsyncprocess crash1010of 100no failing trialmust fail Tier V or the tier has no dynamic range
controlunlink + fsync(dir) each recordunlink + fsync(dir) each recordprocess crash1010of 100no failing trialisolates directory-metadata durability
PostgreSQLstock (synchronous_commit=on)stock (synchronous_commit=on)power loss3030of 300no failing triala reported-committed transaction is on disk and survives a crash
PostgreSQLsynchronous_commit=offsynchronous_commit=offpower loss3016of 30141 to 3recent committed transactions may be lost; never corrupted
MySQL / InnoDBinnodb_flush_log_at_trx_commit=1 (default)innodb_flush_log_at_trx_commit=1 (default)power loss3030of 300no failing trialfull ACID; log flushed and fsynced at each commit
MySQL / InnoDBinnodb_flush_log_at_trx_commit=2innodb_flush_log_at_trx_commit=2power loss301of 30291 to 495survives a mysqld crash; an OS or power loss can lose about a second
MySQL / InnoDBinnodb_flush_log_at_trx_commit=0innodb_flush_log_at_trx_commit=0power loss300of 30303 to 657about a second can be lost even when only mysqld crashes
SQLitejournal_mode=DELETE, synchronous=FULL (defaults)journal_mode=DELETE, synchronous=FULL (defaults)power loss3026of 3041committed transactions survive power loss
SQLitejournal_mode=WAL, synchronous=NORMALjournal_mode=WAL, synchronous=NORMALpower loss300of 303018 to 956durable across app crash; power loss can lose recent commits, without corruption
SQLitejournal_mode=WAL, synchronous=OFFjournal_mode=WAL, synchronous=OFFpower loss300of 3030not statedthe database may become corrupted on power loss
MongoDBw:1 (default)w:1 (default)power loss309of 30211 to 8acknowledged by the primary; can be lost if the node fails before the journal flush
MongoDBw:1, j:truew:1, j:truepower loss3030of 300no failing trialacknowledged only once the write is in the on-disk journal
controlappend + fsync each recordappend + fsync each recordpower loss3030of 300no failing trialmust survive Tier V or the run is void
controlappend, no fsyncappend, no fsyncpower loss300of 3030392 to 10,593must fail Tier V or the tier has no dynamic range
25 of 26 shown

The run log, every trial-arm outcome

All 500 of them: thirteen arms across ten process-crash kills, plus thirteen arms across thirty power cuts, less the third control, which only exists from power-loss trial 21. Acknowledged writes is how far that arm got before the kill landed, under a workload deliberately handicapped by the collector round trip, so it is not a throughput figure. Lost is the count of acknowledged writes missing afterwards, and it is absent rather than zero on the trials where the table itself was gone.

Showing 25 of 500 trials. In the order the run produced them.

Row detail
process crashtrial 11PostgreSQLstock (synchronous_commit=on)stock (synchronous_commit=on)4,9480clean
process crashtrial 11PostgreSQLsynchronous_commit=offsynchronous_commit=off9,0942lost-tail
process crashtrial 11MySQL / InnoDBinnodb_flush_log_at_trx_commit=1 (default)innodb_flush_log_at_trx_commit=1 (default)2,9510clean
process crashtrial 11MySQL / InnoDBinnodb_flush_log_at_trx_commit=2innodb_flush_log_at_trx_commit=26,9110clean
process crashtrial 11MySQL / InnoDBinnodb_flush_log_at_trx_commit=0innodb_flush_log_at_trx_commit=08,8500clean
process crashtrial 11SQLitejournal_mode=DELETE, synchronous=FULL (defaults)journal_mode=DELETE, synchronous=FULL (defaults)1,5490clean
process crashtrial 11SQLitejournal_mode=WAL, synchronous=NORMALjournal_mode=WAL, synchronous=NORMAL11,0130clean
process crashtrial 11SQLitejournal_mode=WAL, synchronous=OFFjournal_mode=WAL, synchronous=OFF11,0690clean
process crashtrial 11MongoDBw:1 (default)w:1 (default)6,9900clean
process crashtrial 11MongoDBw:1, j:truew:1, j:true3,5660clean
process crashtrial 11controlappend + fsync each recordappend + fsync each record3,7780clean
process crashtrial 11controlappend, no fsyncappend, no fsync11,8610clean
process crashtrial 11controlunlink + fsync(dir) each recordunlink + fsync(dir) each record2,2600clean
process crashtrial 22PostgreSQLstock (synchronous_commit=on)stock (synchronous_commit=on)4,4000clean
process crashtrial 22PostgreSQLsynchronous_commit=offsynchronous_commit=off8,5046lost-tail
process crashtrial 22MySQL / InnoDBinnodb_flush_log_at_trx_commit=1 (default)innodb_flush_log_at_trx_commit=1 (default)2,5180clean
process crashtrial 22MySQL / InnoDBinnodb_flush_log_at_trx_commit=2innodb_flush_log_at_trx_commit=26,2020clean
process crashtrial 22MySQL / InnoDBinnodb_flush_log_at_trx_commit=0innodb_flush_log_at_trx_commit=08,3160clean
process crashtrial 22SQLitejournal_mode=DELETE, synchronous=FULL (defaults)journal_mode=DELETE, synchronous=FULL (defaults)1,2930clean
process crashtrial 22SQLitejournal_mode=WAL, synchronous=NORMALjournal_mode=WAL, synchronous=NORMAL10,4320clean
process crashtrial 22SQLitejournal_mode=WAL, synchronous=OFFjournal_mode=WAL, synchronous=OFF10,4170clean
process crashtrial 22MongoDBw:1 (default)w:1 (default)6,3622lost-tail
process crashtrial 22MongoDBw:1, j:truew:1, j:true2,9900clean
process crashtrial 22controlappend + fsync each recordappend + fsync each record3,2700clean
process crashtrial 22controlappend, no fsyncappend, no fsync11,2430clean
25 of 500 shown

The manuals these claims were read from

Each arm’s documented promise above is a compression of what its own vendor publishes. The four pages it was read from are PostgreSQL on write-ahead logging, MySQL on the InnoDB flush parameters, SQLite on its pragmas, and MongoDB on write concern. Where a compression and a manual disagree, the manual is right and we would like to know.

Ask us to break something

Your system makes a claim somewhere. It is in a README, a status page, a support answer, or a sentence a vendor said on a call. Send us the claim and the thing that makes it, and we will write down in advance what would count as breaking it, then spend a serious effort trying to.

The protocol goes up before the run starts, the same way the one above did, and the result is published whichever way it goes — including the boring way, where the claim holds and the page says so. We read everything and we answer, including when the answer is no.

9592 Solutions UG (haftungsbeschränkt), Fährstr. 217, 40221 Düsseldorf, Germany is the controller for what you send here. Your address and your message are used to answer you and to work out whether we take the request on, under Art. 6(1)(b) and Art. 6(1)(f) GDPR. They go to nobody else and they are not used for advertising. Write to christo@9592.tech for a copy or a deletion at any time. The longer version is on the privacy page.

9592 Solutions UG (haftungsbeschränkt), Düsseldorf · Privacy · youtube.com/@m37channel