nostrademons reports that #Google’s biggest regret was not using #ECC RAM in their early cheap servers, because, “If you look through the source code & postmortems from that era of Google, there are all sorts of nasty hacks and system design constraints that arose from the fact that you couldn’t trust the bits that your RAM gave back to you.” #fault-tolerance
on 02026-04-09Joe Armstrong answering Luke Gorrie about “let it crash” #Erlang #fault-tolerance
on 02025-08-10Joe Armstrong’s #dissertation on #Erlang and #fault-tolerance, “Making reliable distributed systems in the presence of software errors” #toread
on 02025-08-10discussion of “let it crash” in #Erlang and #Elixir #fault-tolerance
on 02025-08-10the meaning of “let it crash” in #Erlang and #Elixir, and when you should do otherwise #fault-tolerance
on 02025-08-10#PDF of #Tandem tech report 86.2, which explains the process-pair #fault-tolerance approach, and in particular why they were moving away from having the active process transmit images of its memory space to the backup process. Contains the very amusing sentence, “We believe most faults in production software are transients (Heisenbugs).”
on 02016-08-08#Tandem Guardian Programmer’s Guide #PDF. Very interesting: “Passive checkpoint is impractical in programs using the standard heap, so it is used mostly in TAL and pTAL programming.” A bunch of stuff about “sync depth” for #fault-tolerance that I don’t understand; apparently that’s a number of requests that “the system” remembers the responses to, replaying them if you fail over from a checkpoint, so that you can wait that many requests before checkpointing.
on 02016-08-08