a #regex library for #Python with, among other things, timeouts
#Python threads can be sent a SystemExit exception from another thread
Pike's famous #regex matcher from Beautiful Code in 30 lines of C, albeit with worst-case exponential time and no grouping or character classes
#aaronsw on how the news is bad. “Edward Tufte notes that when he used to read the New York Times in the morning, it scrambled his brain with so many different topics that he couldn’t get any real intellectual work done the rest of the day.” Maybe I should go cold turkey on news...
#video of a talk by #Facebook people about how they built Fecebutt Messenger with #Erlang
another discussion of the #data-oriented-design ebook. #design
a sketch of #compilers with #type-inference
"Flattening ASTs" is the term Adrian Sampson gives to making an array of structs (well, unions of structs) to hold your AST, with references to earlier AST nodes represented as array indices. He claims that this gives a 2.4× #performance boost in his AST-walking interpreter microbenchmark in #Rust, due to improved spatial locality, smaller references, pointer-bumping allocation, and cheap deallocation. He also tried out iterating over the array of AST nodes to compute all their values. #compact-ASTs
the #mmap = 💩 #paper #PDF. The problems for #databases are transactional safety (write ordering), I/O stalls (memory access blocks a process), error handling (all you get is SIGBUS, and you can't checksum your data on its way in and out), and #performance: “Specifically, we have identified three key bottlenecks that plague mmap-based file I/O: (1) page table contention, (2) single-threaded page eviction, and (3) TLB shootdowns.” Mentions #LMDB. In the paper they got better performance, even for reads, with O_DIRECT and pread, but I think that doesn't generalize to arbitrary numbers of reader processes.
RAM bandwidth is supposedly the critical #performance limiting resource for running large #neural-networks like #LLaMA and other #neural-networks on CPU
#Raspberry-Pi 4 memory bandwidth #performance is only about 4 GB/s instead of the 12.8 GB/s theoretical
Slide deck about #ARM #asm. Supposedly the #Raspberry-Pi 2 and 3 (but not 1 and Zero) support #Thumb-2 instructions, and the #STM32 STM32L only supports Thumb-2 (no original ARM fixed-width instruction set). There’s a super sweet #code-density graph on p.35 of the #PDF: 8086 is the champion; PDP-11 and Z80 are almost tied and beat Thumb-2, which beats Thumb, which beats VAX, which beats i386, which beats x86_64, which beats 6502, which beats arm_eabi, which beats the living shit out of ia64.
discussion of #performance in #databases and #mmap; author of #LMDB says read-only mmap and pwrite works well for LMDB. Also pcmulqdq explains where #LevelDB came from from their experience with the Bigtable folks. “Every HDD since the 1980s has guaranteed atomic sector writes.”
#LMDB looks like an interesting point in the #databases design space: a zero-copy MVCC key-value transactional B+tree, similar to Berkeley DB, but writers don’t block readers, and new data doesn’t overwrite existing data
#LMDB benchmarks against other #databases; only #LevelDB is close in #performance, and still beats LMDB on writes (except batched sequential writes or writes of large values). This was when it was still called "OpenLDAP MDB". In particular it beats SQLite3 by generally about an order of magnitude and sometimes more.
#LMDB #performance on #Optane #SSDs vs. #RocksDB but not other #databases
how #Stonebraker and others improved OLTP #performance in 02008 with newish designs for #databases based on "Shore"?
#Postgres (and probably other #databases) had a big bug with #fsync on #Linux (and OpenBSD and NetBSD) in 02018 (and for many years before that): disk write errors (usually caused by USB drive hotunplug) would discard the data that failed to be written and never report the error, even on fsync(), though since kernel 4.13 the error is only lost if the file is closed by the writer and reopened before calling fsync(). So Postgres, and I guess anything else that cares about recovering from I/O errors, has to move to direct I/O (DIO), which seems to mean O_DIRECT. dmsetup has error and flakey targets to simulate disk errors.
Ted Ts'o says nearly all #SSDs can lose data written months ago if there's a power failure.
more details on #fsync and the kinds of scenarios where disk I/O errors can screw you
oh my, I hadn’t realized that all the Cortex-M #ARM #hardware only supported Thumb-2 #asm. Or that every #STM32 #microcontroller supported SWD for single-stepping with OpenOCD. Apparently to burn the Flash on an STM32 you openocd -f interface/stlink-v2.cfg -f target/stm32f1x.cfg -c "program prog1.elf verify reset exit"? And then to provide a gdbserver on port 3333 you openocd -f interface/stlink-v2.cfg -f target/stm32f1x.cfg?