Cognitect got acqui-hired by Nubank in 02020, but continue to maintain #Clojure and (their proprietary database) Datomic. #Lisp #databases
on 02026-09-22#retrocomputing: using dBASE II #databases on a #Kaypro II simulated in #MAME. There was neither a dBASE I nor a Kaypro I. Hilarious. Lots of very realistic simulated CRT screenshots. Helpful tips on cpmutils for Linux, etc. #CPM
on 02026-02-16Andy Pavlo’s 27-page retrospective on #databases in 02025. Strictly limited to big databases and mostly OLTP databases, and especially startups and investment; no mention of SQLite (though one mention of Turso), only one passing mention of DuckDB, one passing mention of LMDB, only three mentions of JSON. But extremely insightful within that niche.
on 02026-01-05#PDF scan of Ted Codd’s 01971 #paper, “Further Normalization of the Data Base Relational Model” #databases #relational-model
on 02025-12-26Example of using #databases in #Java with #JDBC and try-with-resources: try (Connection conn = DriverManager.getConnection("jdbc:somejdbcvendor:other data needed by some jdbc vendor", "myLogin", "myPassword")) {try (Statement stmt = conn.createStatement()) {stmt.executeUpdate("INSERT INTO MyTable(name) VALUES ('my name')");}}
"Dolt" is “Git for Data! (...) a #SQL database that you can fork, clone, branch, merge, push and pull just like a Git repository.” #databases #version-control
on 02025-09-07Chamberlin and Boyce’s 01974 #paper #PDF about #SQL, then SEQUEL, including their previous query language SQUARE, which seems to be based on binary relations (“mappings”). #history of #databases
on 02025-02-10Seems like #van-Emden’s notes on the opposition between the “formulaic paradigm” and the “algorithmic paradigm” goes back further than I thought; in this case he’s contrasting equations with Horn clauses for specifying relational #databases
on 02024-08-27Codd’s 01970 paper defining relational #databases. It’s a bit hard going because of the terminology disconnect.
on 02024-05-03Don Chamberlin oral #history interview with the #CHM, focusing on relational #databases, 39 pages, #toread
on 02024-05-03Chris Date oral #history interview with the #CHM, focusing on relational #databases, 51 pages, talks about how he founded a company with Codd and about the “Required Technology” “Transrelational” model. “While I’m pleased that these #SQL systems are now ubiquitous, I wish they were relational systems.”
on 02024-05-03interview with Chris Date about the #history of relational #databases, where he rags on XQuery as “IMS warmed over”
on 02024-01-31Lorie’s #paper “Physical integrity in a large segmented database” about shadow pages in #databases
on 02023-12-29Getting the time of day out of timestamp/datetime columns in #Pandas #databases is df['date_col'].dt.time. #dates
search-argument-able. #databases
on 02023-12-04a wait-free key-value #transactional database for persistent memory such as #Optane RAM. #databases
on 02023-09-01#Postgres (and probably other #databases) had a big bug with #fsync on #Linux (and OpenBSD and NetBSD) in 02018 (and for many years before that): disk write errors (usually caused by USB drive hotunplug) would discard the data that failed to be written and never report the error, even on fsync(), though since kernel 4.13 the error is only lost if the file is closed by the writer and reopened before calling fsync(). So Postgres, and I guess anything else that cares about recovering from I/O errors, has to move to direct I/O (DIO), which seems to mean O_DIRECT. dmsetup has error and flakey targets to simulate disk errors.
how #Stonebraker and others improved OLTP #performance in 02008 with newish designs for #databases based on "Shore"?
on 02023-07-04#LMDB #performance on #Optane #SSDs vs. #RocksDB but not other #databases
on 02023-07-04#LMDB benchmarks against other #databases; only #LevelDB is close in #performance, and still beats LMDB on writes (except batched sequential writes or writes of large values). This was when it was still called "OpenLDAP MDB". In particular it beats SQLite3 by generally about an order of magnitude and sometimes more.
on 02023-07-04#LMDB looks like an interesting point in the #databases design space: a zero-copy MVCC key-value transactional B+tree, similar to Berkeley DB, but writers don’t block readers, and new data doesn’t overwrite existing data
on 02023-07-04discussion of #performance in #databases and #mmap; author of #LMDB says read-only mmap and pwrite works well for LMDB. Also pcmulqdq explains where #LevelDB came from from their experience with the Bigtable folks. “Every HDD since the 1980s has guaranteed atomic sector writes.”
on 02023-07-04the #mmap = 💩 #paper #PDF. The problems for #databases are transactional safety (write ordering), I/O stalls (memory access blocks a process), error handling (all you get is SIGBUS, and you can't checksum your data on its way in and out), and #performance: “Specifically, we have identified three key bottlenecks that plague mmap-based file I/O: (1) page table contention, (2) single-threaded page eviction, and (3) TLB shootdowns.” Mentions #LMDB. In the paper they got better performance, even for reads, with O_DIRECT and pread, but I think that doesn't generalize to arbitrary numbers of reader processes.
Stonebraker’s 02013 talk about how everything you know about #databases is wrong (because it's obsolete); instead he argues that star-schema data warehouses should use #column-stores like MonetDB and Vertica (because they’re 50–100× faster when you have 50–100 columns but only look at 4–5 attributes and always do SQL aggregates), while for OLTP databases “NewSQL” systems are “wildly faster”, about 100×. Explains column stores using Vertica as an example, explaining that segments of rows get buffered in main memory and then written out as “slabs” of compressed column “chunklets” which are later merged into larger segments (LSM-tree style, I suppose). Explains that #OLTP databases are almost never over a terabyte, which was US$30k of RAM at the time (64 gigs in each of 16 servers). “Data warehouses are measured in petabytes, but not OLTP,” but conventional databases waste a lot of time encoding and decoding it, maintaining LRU lists, doing record-level dynamic locking, latching B-tree pages, latching the lock table, and logging transaction recovery records. So H-Store and VoltDB statically partition main memory among cores and run one thread on each core to eliminate latching and locking. This makes it seem like maybe software transactional memory might be an interesting way to implement an in-memory OLTP database. Unsurprisingly #Stonebraker is very much not a fan of #eventual-consistency. He argues that replication via log shipping will inevitably result in a 3–6× slowdown over the optimal case because it commits you to writing a write-ahead log so you can ship it, so it's better to just duplicate the computation of the transaction on each replica, and similarly use "command logging" for durability. He says this requires you to use timestamp ordering (or some other deterministic scheme) rather than MVCC. #Video is loaded from YouTube misleadingly: it's 1h12m33s, not 12m33s. I guess Stonebraker must really hate Postgres for some reason because he never mentioned it once, while mentioning almost every other SQL database!
on 02023-07-02#ChatGPT is being used by StarRocks to improve the testing of their #SQL database. #AGI #databases
on 02023-03-09#decentralization #databases #IPFS
on 02021-03-02Ferragina delivers the "PGM-index", a follow-on to the crack-smoking #Google paper from a couple of years back (?) about #databases indexing with #neural-networks
on 02020-01-25A database that stores stuff in a text file; the format is (almost) blank-line-separated RFC-822 headers. “#Recutils is a set of tools and libraries to access human-editable, plain text #databases called recfiles. The data is stored as a sequence of records, each record containing an arbitrary number of named fields.”. recsel -e "Age < 18" -P Name acquaintances.rec to select names; recins -f Name -v "Mr Foo" -f Email -v foo@bar.baz contacts.rec to insert. There’s an Emacs mode, but it’s not part of standard Emacs (or very powerful, I think), and there seem to be some problems in the interface. Supports some minimal integrity constraints including range checks, uniqueness, and regexps, but not referential integrity (foreign key constraints). Supports references to other records via a primary key with an ad-hoc convention for fields thus obtained, but not arbitrary joins. Has its own mail-merge template syntax for reports.
Discussion of #time-series #databases from 2014. Many people recommend #Postgres for under 10 billion items.
on 02015-08-10The Samza/Kafka guy doesn’t like the #CAP theorem and criticizes it on basically vacuous grounds, and #databases that attempt to characterize themselves using it on somewhat less vacuous grounds.
on 02015-08-07