more details on #fsync and the kinds of scenarios where disk I/O errors can screw you
on 02023-07-04#Postgres (and probably other #databases) had a big bug with #fsync on #Linux (and OpenBSD and NetBSD) in 02018 (and for many years before that): disk write errors (usually caused by USB drive hotunplug) would discard the data that failed to be written and never report the error, even on fsync(), though since kernel 4.13 the error is only lost if the file is closed by the writer and reopened before calling fsync(). So Postgres, and I guess anything else that cares about recovering from I/O errors, has to move to direct I/O (DIO), which seems to mean O_DIRECT. dmsetup has error and flakey targets to simulate disk errors.
#Postgresql #fsync problems
on 02019-02-11#fsync on #Linux and other Unixes can fail, which indicates not that your data hasn’t been written yet but that it’s actually been lost entirely. There’s no interface to tell which data and thus no reasonable way to retry, except that you can reread the data you tried to write to see which data is wrong.
on 02019-02-11