“Lightweight Software #Transactions for Games”, a #paper from 02009 about SpaceWars3D, where they tried to “realize the #performance potential of multiple cores” with #transactional-memory, only to be terribly disappointed. “[W]e depart from classic STM designs and propose a programming model that uses long-running, abort-free transactions that rely on user specifications to avoid or resolve conflicts.” #videogames
on 02026-09-18#Dan-Luu says #performance optimization is something #AI #LLMs (#neural-networks) are good at now. He created a new #regex library called "FRE" that gets good performance by compiling to native code by “having an agent loop for a month on improving regex engine performance with access to the rebar regex benchmark suite”, which led to overfitting, and then disciplining the agent with a holdout benchmark, using #ripgrep queries that came from his history with OpenAI’s “Codex” programming LLM. It’s still “substantially slower than the Rust regex engine on holdout benchmarks” but yet “enough to generally match 2nd tier regex engines in terms of performance”.
> This kind of technical work, which used to take a fair amount of time and expertise, can just be done trivially now. (...) While the open source version of BitFunnel “only” contains a bytecode interpreter and one JIT, the Bing version contains multiple JIT compilers. A project that did that level of optimization used to be a major undertaking, but “I could do that in a weekend” is now actually true for some of these kinds of projects.
(...)
> For an example from the GPT-5.1 or 5.2 days, with no knowledge of game AIs, I tried building an Azul AI. This ended up being the strongest AI in the world for the game by a pretty large margin. From reading the thesis that describes the 2nd strongest AI, I think my AI is probably a bit better on the “AI” side of things, but the main place it wins is on optimization despite spending what looks like maybe two orders of magnitude less [human] time [writing the software] (...)
> There’s a bunch of standard stuff it makes sense to do to debug and verify a multithreading algorithm for something like this, like implementing replay from debug logs that can reproduce bugs despite the algorithm being nondetermistic. Doing that alone would’ve probably been days to a week of work had I done it by hand, but it’s exactly the kind of thing an agent can trivially do in a loop (just have it try to replay logs and insert logging for non-determinism every time you don’t get a perfect replay). A lot of the tedium it used to take to get a tricky optimization like this working is gone.
(...)
> Now that this N [of person-days to verify that a tricky optimization works] has dropped by a tremendous factor (variable but, in terms of human time, frequently 1000x / 10000x / 1000000x, probably more like 1000x on dollar cost if you compare token costs at metered rates vs. the Bing engineer who wrote the compilers at JITs that the search index used), the number of these kinds of optimizations it makes sense to do goes way up.
> (...) for the AI I tried, it seems like you gain about 100 Elo for every doubling in speed (more than in chess, I suspect because draws are very rare). Just adding multithreading alone is enough to wipe the floor with an otherwise comparable AI on a large machine.
> (...) current publicly available SOTA models are pretty bad at experimental design, so I had to set up the framework they used to determine if an optimization is good, but once that was in place, it’s like any other optimization problem.
> Jamie Brandon (...)’s a reasonable performance engineer and he got an offer for the performance job he wanted [at Anthropic], but on a well-defined optimization problem, he [reports that he] doesn’t stand a chance against a decent model [Claude] (...)
> right before I started writing this post, I had an agent do workload-specific optimization for my ripgrep queries (...), which took about 2 minutes for me to launch. (...) After one pass of optimization, the workload optimized version is 2% faster than standard ripgrep on the holdout and it’s still getting faster.
He thinks this is going to result in a lot of companies workload-optimizing their software using coding agents. Also:
> (...) someone who doesn’t know anything about performance and is a reasonable user of LLMs (just in general, not on performance problems in particular) should generally be able to create software that has decent performance.
on 02026-08-21no real reason people on #amd64 use xor instead of sub to clear a register in #asm; both of them require zero micro-ops, merely renaming a register. #performance
how #performance varies with locality of access in RAM: how much linear access is enough?
on 02026-04-11discussion between #Muratori and #Uncle-Bob about #Clean-Code, significantly but not completely focused on #performance, discussing Muratori’s video criticizing it.
on 02026-04-10#Geohot plans to own a zettaflops of 4-bit flops before he dies. “$10M for the machine. $10M for the solar panels. $10M for the land and construction.”. #AI #performance #futurism
on 02026-04-09usually reasonable advice: “choose the simplest #algorithms with less than quadratic time and space complexity.” #performance
on 02026-03-18#SSD #performance is apparently over 14 gigabytes per second now!
on 02026-01-15#video. He says he “evaluates the [#SDF] function once for all points on a [sparse] grid once, and then reuse those cached values when rendering” to get fast enough #performance for fully dynamic #videogames, using bilinear interpolation between the SDF at lattice points to approximate the SDF at intermediate positions. This produces the usual chamfered corner artifacts you see in SDF fonts, as he points out. He uses “brick maps” to reduce storage and, for #LoD, “geometry clipmaps” and the #Jolt-physics engine. However, it seems like he’s having trouble with objects “melting” onto his terrain. #toread #graphics #shaders
on 02026-01-12for #performance testing in #Lua you may have to invoke #garbage-collection twice (with collectgarbage()) to ensure that finalizers actually run. Quoting PIL: “The first time the collector detects that an object with a finalizer is not reachable, the collector resurrects the object and queues it to be finalized. Once its finalizer runs, Lua marks the object as finalized. The next time the collector detects that the object is not reachable, it deletes the object. If you want to ensure that all garbage in your program has been actually released, you must call collectgarbage twice; the second call will delete the objects that were finalized during the first call.”
102545 transactions per second #performance with #SQLite on a Macbook Pro M1 using #Clojure. #toread
on 02025-12-02discussion of #performance of old computers. anthk reports, “With a 486 and 16MB of RAM, you can run X at sane speeds, even FVWM in wireframe mode to avoid window repaintings upon moving/resizing them.
Next, TLS/SSL. WIth a 486 DX you can use dropbear/bearssl and even #Dillo happily with just a light lag upong handhaking, good enough for TLS 1.2. Under a 486, a 30-35? year old CPU. IRC over TLS, SSH with RSA256 and the like methods, web browsing/Gemini under Dillo with TLS. Doable, I did it under VM, it worked, even email and NNTP over TLS with a LibreSSL fork against BearSSL.” And accrual reports, “I did some multitasking recently on my iDX4-100 + 64MB FPM. I used NT4 with SP2 because the full SP6 was much slower. I could have a browser open, PuTTY, and some tracker music playing no problem. :)” A 486 is something like 11–46 Dhrystone MIPS.
on 02025-11-23#IBM 3084QX #performance was 31 MIPS. #mainframes #history
on 02025-11-23#IBM 3084 #mainframes cost US$8.7 million in 01983. “The machine uses two model K 3081 Processor Units with 64Kb of fast buffer memory. Each 3081 K Processor Unit is composed of two processors.” As for #performance: “The 3084 was released with either 32, 48 or 64 Mb of memory. The system in the collection has 16 Mb of main memory (because it is half of the entire 3084 system). The memory has a cycle time of 312 nanoseconds and is 8 bytes wide.” So I guess you could read or write main memory at 24 megabytes per second, which suggests a speed around 24 MIPS. #history
on 02025-11-23discussion of #performance of computers that could have been used to design the 80386: “Top of the line VAX in 1984 was the 8600 with a 12.5 MHz internal clock, doing about 2 million instructions per second.
IBM 3084 [mainframe] from 1984 - quad SMP (four processors) at 38 MHz internal clock, about 7 million instructions per second, per processor.” That was when #mainframes were fast, rather than just having a lot of I/O.
on 02025-11-23#Jamii Brandon’s summary of #Zig memory safety: “Zig removes some of the most egregious footguns from c, has better defaults, makes some good practices more ergonomic, and benefits from a fresh start in the standard library (eg using slices everywhere). But it does not nearly approach the level of systematic prevention of memory unsafety that rust achieves,” citing use-after-free, use-after-realloc, and invalidating an interior pointer to a union as examples “I often run into”. Lots of information about #performance of different programming language design choices, for example in Rust or Azul’s Zing garbage collector.
on 02025-11-23#PDF #paper of Bhandarkar and Clark’s comparison of the #ISAs and #performance of #VAX and #MIPS from 01991. #history
on 02025-10-05linear-probing #hashing that tries to keep probe sequences short by fancy insertion? #performance #algorithms #toread
on 02025-09-11The CLZ #asm #bitmanip instruction is __builtin_clz on GCC, but this page claims to have a 13-instruction-time version with better #performance than all the “Hacker’s Delight” entries, with “[bisection] to find out which 8-bit chunk of the 32-bit number contains the first 1-bit, which is followed by a lookup table clz_lkup[] to find the first 1-bit within the byte.” Handy for Cortex-M0 #ARM chips without the instruction.
Bruce Hoult’s #performance benchmarking test, a C routine that tabulates all primes up to 7919² = 62710561
on 02025-09-09“We are extremely pleased to announce the availability of the new “Swiss Table” family of hashtables in Abseil and the absl::Hash #hashing framework that allows easy extensibility for user defined types. Last year at CppCon, We presented a talk on a new hashtable that we were rolling out across #Google’s codebase.” #C++ #Swiss-tables #performance
on 02025-09-07#video of a talk by Kulukundis at cppcon 2017. Kip says, “Presentation about the latest tricks in hash tables. Uses SSE instructions and man is it fast. Good benchmarks against the standard C routines.” Heh: “As with anything like this, benchmarks are the only source of truth you will ever get, and they are lies.” Called "Swiss tables" (because Alkis and Roman, its primary developers, “are in the Zurich office”, and it’s “closed hashing”), supposedly the fastest hash table in the world. “What did we gain? (...) the vague feeling of superiority when we force a difficult decision onto the user. And that is the #C++ way.” You put metadata about which array elements are full or deleted into a separate byte array, along with truncated-to-7-bits hash values, to avoid needing magic sentinel values for your key type. By using SSE matching to look for truncated-hash matches in a 16-bucket group, you can find the candidates in three instructions, which also means that your erase function can avoid inserting tombstones “if any other element in the group was empty”, which seems like a pretty big win actually. Also he mentions “other sizes [than power of 2] with fast modulus”. #algorithms #performance #hashing
on 02025-09-06Text #editors on big computers like cellphones ought to be able to handle the full text of Moby Dick without #performance problems. The provided Markdown file thereof is 1204076 bytes, and according to my memcpycost.c, memcpy of such a size of thing on my old 2.4GHz Pentium N3700 would cost about 600 microseconds on one core and about 300 microseconds of memory bandwidth. On my Ryzen 5 3500U we should be talking about 120μs.
on 02025-09-02#razetime’s notes on #performance in ngn K #array-languages #toread
on 02025-08-27discussion of #array-languages and #performance recommending Uiua and mourning "razetime"
on 02025-08-27slowing down #Java #performance in order to test #profilers and find race conditions
on 02025-08-2701994 #paper by Xinhua Zhuang, “Decomposition of Morphological Structuring Elements” about “two-pixel decomposition” and “cellular decomposition” in #morphology #algorithms for #performance
on 02025-08-12“"Cinder" is Meta’s internal #performance oriented production version of CPython 3.10. It contains a number of performance optimizations, including bytecode inline caching, eager evaluation of coroutines, a method-at-a-time JIT, and an experimental bytecode compiler that uses type annotations to emit type-specialized bytecode that performs better in the JIT.” #Python
on 02025-08-10Hudson River Trading forked #Python to add “lazy imports” for startup #performance #toread
on 02025-08-10why #JIT #compilers don’t get good #performance out of #Python
on 02025-08-06the Numba JIT compiler for getting #performance out of #Python #Numpy has a nicely styled website with pastel code samples
on 02025-08-06Andi Kleen’s #perf_events #performance #profilers tools
on 02025-08-05Brendan Gregg’s page on #perf_events #performance #profilers #toread
on 02025-08-05#PDF of Julia Evans’s #perf_events zine. CC-BY-NC-SA. #performance #profilers
on 02025-08-05#news on the #RISC-V Summit in China: high #performance implementations including “UltraRISC UR-DP1000, Zhihe A210, and SpacemIT K3”, which are mostly RVA23-compliant.
on 02025-07-27#Performance of Microsoft’s SOTA allocator #mimalloc. #allocators
on 02025-07-2564 bytes of ASCII to lowercase in three instructions: __m512i ca = _mm512_sub_epi8(c, _mm512_set1_epi8('A')); __mmask64 is_upper = _mm512_cmple_epu8_mask(ca, _mm512_set1_epi8('Z' - 'A')); __m512i to_lower = _mm512_mask_add_epi8(c, is_upper, c, to_lower) from Daniel Lemire’s talk on designing #algorithms for #performance on current hardware
you can serve 5800 requests per second on one Linux machine under gohttpd with #CGI now! If it has 60 cores and 240 GB of RAM and you write the CGI program in C. #performance
on 02025-07-20The #ARM Cortex-M3 (used in many processors of the #STM32 line) has #performance of 1.25 #Dhrystone MIPS/MHz
on 02025-07-20on #manycore #performance from 02009 #toread
on 02025-05-16archive page on comparison of the Apollo Guidance Computer with an Anker charger whose CPU has 563 times the #performance of the #AGC #microcontrollers #electronics
on 02025-03-28How to improve #performance in #V8 #JavaScript with knowledge of its element-type state machine.
on 02025-01-02#Performance of different #JS engines about 10 years ago, comparing #QuickJS to #Duktape, V8, plus two other lightweight JS implementations, #JerryScript and #MuJS, and apparently things called Hermes and XS (#Moddable?). Taking V8 as the baseline, roughly, V8 --jitless is 22× slower, QuickJS and Hermes are 37× slower, Duktape is 88× slower, XS is 56× slower, MuJS is 260× slower, and JerryScript failed to run most of the tests.
on 02024-12-27#paper on "goSLP", which uses linear #optimization (ILP) for #SIMD #vectorization, achieving significant #performance gains on floating-point benchmarks like SPEC2017fp. “Using an integer linear programming (ILP) solver, goSLP searches the entire space of statement packing opportunities for a whole function at a time, while limiting total compilation time to a few minutes. Furthermore, goSLP optimally solves the vector permutation selection problem using dynamic programming.” #compilers
on 02024-12-12"Koalas" provides the #Pandas #dataframes API on top of Apache #SPARK for #exploratory-data-analysis with better #performance #toread
on 02024-12-12"Dask" is another distributed #dataframes system that “uses #Pandas under the hood” for #exploratory-data-analysis with better #performance #toread
on 02024-12-12"Modin" scales #Pandas notebooks across clusters for #exploratory-data-analysis with better #performance #toread #dataframes
on 02024-12-12#compilers for #Pandas to get 1000× better #performance on #exploratory-data-analysis: "Dias" by #Baziotis
on 02024-12-12Banana Pi #SpacemiT 8-core #RISC-V SBC. 50 GIPS #performance, supposedly, 30% higher performance per core than ARM Cortex-A55. 256-bit RVV 1.0.
on 02024-09-27#performance on “the Shenzhen Milk-V Jupiter #RISC-V mini-ITX motherboard” based on the K1 processor. Lots and lots of photos. “an octa-core #SpacemIT X60 SoC with PowerVR B-Series BXE-2-32 GPU, 15.5 GB RAM (available to Linux).” The benchmarks are things like video decoding and OpenGL.
on 02024-09-27aiming at high #performance: “The #SpacemiT X60™ Intelligence Core is an advanced multi-core and multi-cluster #RISC-V RVA22 processor, equipped with RISC-V Vector 1.0 extensions and SpacemiT IME Intelligence Extensions, specifically optimized for general AI applications at the edge. (...) 2 TOPs@INT8 AI computing power compliant with RISC-V IME extensions. (...) 2.0GHz@22nm. (...) 3.66DMIPS/MHz. (...) 8-stage dual-issue in-order execution.”
on 02024-09-27the #RISC-V #SiFive U74 core in the #StarFive JH7110 is about 2.5 #Dhrystone MIPS per MHz, while Roy Longbottom says the Cortex-A53 in the Raspberry Pi 3 is about 2.95 Dhrystone MIPS per MHz #performance
on 02024-09-14how #SQLite does #performance optimization with cachegrind, enabling the use of many small optimizations
on 02024-09-05Raymond Chen on how optimization for #performance can be counterintuitive in, for example, the presence of return-address caching for branch prediction
on 02024-09-04#Hashing #algorithms in the CPython dict implementation; by adding a level of indirection to the (hashval, key, value) tuples, you can greatly reduce the amount of space used, thus improving #performance. As a side effect, you can iterate over entries in insertion order.
on 02024-08-31#video #toread on #Godot #performance and profiling
on 02024-08-26#video #toread on #Godot #performance with occlusion culling
on 02024-08-26Fairchild’s gold-doped 2N709 transistor was designed specifically to enable the #CDC 6600 to achieve its 10MHz clock speed and therefore good #performance despite using RTL, by reducing saturation charge. The PMBT2369 or MMBT2369 is a current replacement that works slightly better. A six-stage ring counter on a PCB runs at 17.7MHz, which is a propagation delay of under 5ns, reducible to 3.5ns with a higher gate current of 10mA. #hardware #retrocomputing #electronics
on 02024-08-25#performance of #mmap is bad when it blocks #async #concurrency
on 02024-08-25discussion of #sorting #algorithms #performance and cmov
a comparison of #hashing #algorithms #performance in 02012
on 02024-08-20#Dhrystone #performance results on #Raspberry-Pi: 2201 VAX MIPS for a Pi 3
on 02024-08-16#video about #C++ constexpr and #performance by #Daves-Garage (Dave Plummer) using a sieve of Eratosthenes example
Initial #buck50 #Sigrok #logic-analyzer #STM32 firmware announcement from 02020. Including discussion of #performance and how it differs from the #ARM reference manual cycle counts.
on 02024-01-28comparison of the Apollo Guidance Computer with an Anker charger whose CPU has 563 times the #performance of the #AGC
on 02023-12-27#ropes for #performance in hex #editors, as used in Simon Tatham’s "Tweak"
on 02023-12-27#gap-buffers for #editors #performance
on 02023-12-27How VS Code switched its text buffer data structure for #performance. #editors
on 02023-12-27how to implement the Game of #Life on the #Amiga blitter in 01987. #performance #history
on 02023-12-06#PCG32 #PRNG #paper #PDF showing 29-66 gigabits per second #performance, more than twice the Mersenne Twister and ten times Arc4random. Same author as the “Genuine Sieve of Eratosthenes” #algorithms paper.
on 02023-12-06Chris #Wellons explains hash #trie #algorithms in C, recommending a 4-way branching factor with arena allocation. Not the same as #HAMT. #performance
on 02023-12-06#Dhrystone #performance results for lots of recent systems including Core i7, Android, and Raspberry Pi
on 02023-10-23US$1925 for a mining rig with six NVIDIA RTX 3070 GPUs, each capable of 40 teraflops, for a total of 240 teraflops, 125 gigaflops per dollar. #hardware #performance #pricing
on 02023-10-21How to adjust your timestep in #games and interpolate movement (or extrapolate) so that your #performance problems only change the display refresh rate and not the game’s rules
on 02023-10-07#algorithms using uninitialized memory for constant-time #performance for sparse integer sets (citing Briggs and Torczon’s 01993 paper)
on 02023-10-07#PDF #paper by #Kahan about #floating-point being compromised in the pursuit of #performance to cheat on SPEC benchmarks
on 02023-10-07#moon-child’s code for analyzing Intel #asm instruction scheduling #performance in SBCL #Lisp
on 02023-10-05#performance of #quicksort on current hardware improves with, among other things, branchless bubble #sorting #algorithms instead of insertion sort, which is damned surprising
on 02023-09-25#CHERI #Morello #hardware initial results show 15% #performance overhead on SPECint2006 for #capability-systems #security, but simulations on #FPGA suggest ways to reduce this to 1.8–3.0%.
on 02023-09-19#history of #DEC #Alpha #hardware #performance
on 02023-09-19more notes by Tratt on #CHERI #allocators, #performance, and #security
on 02023-09-19#paper on #CHERI #allocators, #performance, and #security (Tratt)
on 02023-09-19#ARM Cortex-A72 #hardware “is a superscalar processor which predicts branch direction and executes instructions speculatively along predicted program paths.” for #performance.
on 02023-09-19the #documentation for #Scheme #performance in #Racket
on 02023-09-02notes on #performance optimization: most aren’t worth it, unlike in Shenzhen IO
on 02023-09-02#Agner Fog’s #performance measurements of the #Intel #Skylake CPU family, such as the i7-6700K)
on 02023-08-09my calculations about Dominic Szablewski’s #qoa #audio #DSP #codec #performance: something like 400 32-bit integer instructions per sample
on 02023-08-09new #Intel #hardware "Sapphire Rapids" has variable core-to-core latency #performance due in part to packaging 48 cores on 4 chips in a single package, with round-trip latency ranging from below 100ns to over 800ns
on 02023-08-09#RISC-V #performance simulation, covering #SeRV, #PicoRV32, Minerva, Hazard3, FemtoRV32, #VexRiscv, and the author's own Misato.
on 02023-07-08#Performance database server: #Dhrystone benchmarks of many historical machines; a 5x86-133 had 44 DMIPS
on 02023-07-05custom hash tables can have #performance an order of magnitude better than generic ones, as Chris #Wellons demonstrates here #hashing in Golang
on 02023-07-05how #Stonebraker and others improved OLTP #performance in 02008 with newish designs for #databases based on "Shore"?
on 02023-07-04#LMDB #performance on #Optane #SSDs vs. #RocksDB but not other #databases
on 02023-07-04#LMDB benchmarks against other #databases; only #LevelDB is close in #performance, and still beats LMDB on writes (except batched sequential writes or writes of large values). This was when it was still called "OpenLDAP MDB". In particular it beats SQLite3 by generally about an order of magnitude and sometimes more.
on 02023-07-04discussion of #performance in #databases and #mmap; author of #LMDB says read-only mmap and pwrite works well for LMDB. Also pcmulqdq explains where #LevelDB came from from their experience with the Bigtable folks. “Every HDD since the 1980s has guaranteed atomic sector writes.”
on 02023-07-04#Raspberry-Pi 4 memory bandwidth #performance is only about 4 GB/s instead of the 12.8 GB/s theoretical
on 02023-07-04RAM bandwidth is supposedly the critical #performance limiting resource for running large #neural-networks like #LLaMA and other #neural-networks on CPU
on 02023-07-04the #mmap = 💩 #paper #PDF. The problems for #databases are transactional safety (write ordering), I/O stalls (memory access blocks a process), error handling (all you get is SIGBUS, and you can't checksum your data on its way in and out), and #performance: “Specifically, we have identified three key bottlenecks that plague mmap-based file I/O: (1) page table contention, (2) single-threaded page eviction, and (3) TLB shootdowns.” Mentions #LMDB. In the paper they got better performance, even for reads, with O_DIRECT and pread, but I think that doesn't generalize to arbitrary numbers of reader processes.
"Flattening ASTs" is the term Adrian Sampson gives to making an array of structs (well, unions of structs) to hold your AST, with references to earlier AST nodes represented as array indices. He claims that this gives a 2.4× #performance boost in his AST-walking interpreter microbenchmark in #Rust, due to improved spatial locality, smaller references, pointer-bumping allocation, and cheap deallocation. He also tried out iterating over the array of AST nodes to compute all their values. #compact-ASTs
on 02023-07-04how to use #perf and "tiptop" to measure #performance in terms of average instructions per clock with PMCs (performance monitoring counters).
on 02023-07-03“Linear Probing Revisited: Tombstones Mark the Death of Primary Clustering, by Michael A. Bender, Bradley C. Kuszmaul, and William Kuszmaul” #PDF #paper on #hashing #algorithms #performance with "graveyard hashing"
on 02023-07-03Messaging #Web-workers is fast #performance even without transferable objects, but the test code here is broken #browsers
on 02023-06-22"Poop" runs a couple of commands without shells and reports their maximum memory usage, instruction count, cache miss count, etc., for #performance tuning.
on 02023-06-17Daniel Bernstein explains why #vectorization is a #performance win, even when it requires throttling the #hardware clock
on 02023-06-13new sieve of Eratosthenes #algorithms with O(N^{1/3} (log N)^{2/3}) space and O(N log N) time #performance
on 02023-01-15for high #performance numerical computation in #Python people recommend Numba, Jax, and switching to Julia or Taichi
on 02022-12-28#Raph Levien’s treatise on #algorithms with #ropes and their #performance in his #editor Xi
on 02022-09-16#FFT #algorithms #performance with blocking and cache-oblivious approaches
on 02022-09-16#algorithm #performance for #parsing #JSON: 7.4 GB/s to look for tweets by a particular person, for example, 12 GB/s for minification by removing whitespace, or 3.6 GB/s to produce a full in-memory DOM tree. Apparently #AVX-512 actually does improve performance now.
on 02022-05-26“An Empirical Lower Bound on the Overheads of Production Garbage Collectors” very clever #PDF on #garbage-collection #performance
on 02022-02-09Dan Luu thinks compiler intrinsics aren’t worthwhile if you need #performance because although your code will work on every architecture it won’t be fast until you tweak it, at which point you might as well have just written an assembly version for each architecture anyway
on 02021-11-02Dan Luu on modern #hardware #performance, largely having to do with #concurrency, #locking, #context-switches, and #SIMD, with a little bit about #GPGPU
on 02021-11-02Drepper’s famous paper about what every programmer should know about memory #hardware #performance
on 02021-11-02John Mashey in 02005 commenting on #performance #history and why it was so hard to make a fast #VAX
on 02021-02-09A new #ESP32 #microcontroller "ESP32-C3" with #RISC-V RV32IMC has better single-core #performance but less RAM
on 02021-02-09"Shade" was a system that dynamically recompiled segments of machine code to insert instrumentation for #performance analysis with about a 6× slowdown (in 1993)
on 02021-02-09"Salto": System for Assembly-Language Transformation and Optimization, a 1996 #paper on writing things similar to Valgrind or gprof, but at the #asm level, mostly for #performance analysis
on 02021-02-09how to use #Linux kernel #tracing tracepoints effectively and analyze #performance with #perf
on 02021-02-03David Stafford and Dave Methvin reminisce about their #performance rivalry in #Abrash’s contest
on 02021-01-24Michael #Abrash’s #Graphics Programming Black Book from 01997. #performance #ebook
on 02021-01-24SiFive’s US$665 HiFive Unmatched board uses a 5-core 1.4 GHz #RISC-V #hardware "FU740" or “Freedom U740” comparable to the #Raspberry-Pi 4’s quad-Cortex-A72. #performance of 2.5 DMIPS/MHz and 4.9 CoreMark/MHz (from which we can deduce that a CoreMark is about 2 #Dhrystone MIPS), dunno if this is per-core or per-chip.
on 02021-01-24how to measure #performance of virtual machines, including #LuaJIT and #V8; Laurie Tratt argues that current #VMs are slow in part because of Goodhart’s Law: people engineer their VMs to do well on their benchmarks, but the benchmarks are unrepresentative.
on 02020-05-24some notes on #SAT #solver #performance and the relationship with backtracking search
on 02019-02-15Subroutine calling in #J version 701b (from 2014) takes about 1400 ns on a 3.4 GHz amd64 box, and 14000 ns on a Galaxy Tab 3, because it’s apparently using string interpretation. #performance
on 02019-02-15#context-switch #performance is around 2.1–2.8 μs of direct cost on a dual Intel 5150 (Woodcrest, “Core”), but typically around 30 μs due to cache pollution
on 02019-02-02#PDF #paper on “Quantifying the Cost of #Context-Switch” puts the direct cost at 3.8 μs on a 2GHz Xeon in 2007, a number which should have improved enormously since then due to TLB tagging, permitting the 90ns results reported on seL4 in 2013. But the indirect costs from cache pollution can be orders of magnitude higher. #performance
on 02019-02-02#PDF document on #performance optimization for AMD processors (2017)
on 02018-10-23#Soviet #Elbrus #performance #hardware #history.
on 02017-11-30#Graphics #performance #algorithms: running console graphics pipeline #emulation (of the Nintendo GameCube, specifically, in the Dolphin emulator) in a pseudo-GPGPU "Ubershader"
on 02017-07-30The Mill #hardware #ISA can supposedly do a switch statement in three instructions. #performance
#Retrocomputing #performance on the Xerox Alto code for the Mandelbrot. #history
on 02017-07-01“daxctl() — getting the other half of #persistent-memory #performance” #toread once it goes open-access
on 02017-07-01detailed #performance comparison of #Kafka and #RabbitMQ (and #Redis)
on 02017-06-18Pieter Hintjens describes what kind of #performance you should expect from an AMQP broker like #RabbitMQ or from #0MQ and why. #Performance of RabbitMQ is about 60× worse than #0MQ by default and 4× faster than Apache Qpid, but that’s if you have persistence turned on in RabbitMQ.
on 02017-06-16Pieter Hintjens describes what kind of #performance you should expect from an AMQP broker like #RabbitMQ or from #0MQ and why. #Performance of RabbitMQ is about 60× worse than #0MQ by default and 4× faster than Apache Qpid, but that’s if you have persistence turned on in RabbitMQ.
on 02017-06-16#CapnProto #serialization #performance vs. #FlatBuffers and something called Simple Binary Encoding
on 02017-06-14#Wouter van Oortmerssen’s #FlatBuffers #serialization #performance; crudely it's about 16× slower than memcpying raw structs and about 32× faster than Protocol Buffers, which in turn is about 4× faster than JSON
on 02017-06-14discussion thread on how the Glimmer VM (which #Ember-JS uses for rendering #JS templates) gets high #performance
on 02017-05-11#Clojure on #Android takes fucking forever to start up (2000–3000 ms), compared to Java apps, which start in like 500ms. #performance
on 02017-05-09#FlatBuffers #serialization #performance on #Android, with lots of pretty pictures
on 02017-04-26how to test #SSD #performance according to Seagate, and the kinds of errors you can run into if you don’t do it right
on 02017-04-26scrolling #performance and why it kind of sucks in #browsers and how to improve it in your #JS, e.g. with passive listeners. #UX
on 02017-03-22James Hague improved his #web #performance; now you can read his blog with under 4K transferred in a single HTTP transaction.
on 02017-03-22Skip list #algorithms #performance
on 02017-03-11#Numpy #ebook by Nicolas P. Rougier, cc-nc-by-sa, full of awesome #vectorization and #performance tricks
on 02017-01-14#paper on #memory #performance using “hybrid main memory” such as a DRAM cache layer over #Flash, as a stopgap until phase-change memory or something like that arrives.
on 02016-10-102009 #paper on improving #hardware #performance to get phase-change #memory to be very nearly as fast as DRAM, but nonvolatile. “A baseline PCM system is 1.6x slower and requires 2.2x more energy than a DRAM system. Buffer reorganizations reduce this delay and energy gap to 1.2x and 1.0x, using narrow rows to mitigate write energy and multiple rows to improve locality and write coalescing.” Presumably this is behind the #crosspoint stuff that Micron and Intel plan to launch next year? #persistent-memory
on 02016-10-10#paper on #filesystems with high #performance on NAND #Flash by using multiple channels
on 02016-10-10#sse #performance #toread by #ryg #vectorization
on 02016-09-29#graphics #algorithms #performance #sse #toread by #ryg
on 02016-09-29improving #graphics #algorithms #performance with #barycentric coordinates. #toread by #ryg
on 02016-09-29permuting bytes and words using AVX2 VPACK, VPUNPCKL, VPSHUFB, etc. #sse #performance #vectorization
on 02016-09-29“an algorithm and associated sample code [using #SSE] for software occlusion culling which is available for download” #algorithms #performance #toread this is what ryg was commenting on in his 2013 #graphics thread
on 02016-09-29Write combining can cause #performance problems in #graphics #algorithms or can speed them up. #toread by #ryg
on 02016-09-29#compression #algorithms #performance testing for end-of-buffer with branchless I/O etc. by #ryg
on 02016-09-29ChaCha12-256 (without the usual security margin) has about the same #performance as AES-128 on several generations of Intel #hardware, because it’s good at #vectorization of #crypto, and will get faster now that Intel is exposing their 52-bit multipliers as #djb requested in 2002. Also apparently Intel is adding inversion in GF(256).
on 02016-08-11#source-code for a new #post-quantum #crypto implementation: an embedded-optimized ARM assembly implementation of NewHope key exchange, an algorithm using the NTT (Number-Theoretical Transform), with #performance of under 2M clock cycles on the Cortex-M0. By Erdem Alkim, Philipp Jakubeit, and Peter Schwabe.
on 02016-08-03Bélády’s anomaly, discovered in 1969, is that FIFO paging #algorithms can produce more unboundedly page faults if given more page frames, screwing up #performance.
on 02016-08-03#Raph-Levien lucidly explains how his new #font rasterizer "font-rs" in #Rust is the fastest (≈6× faster than FreeType). Among other things, he uses #SSE #prefix-sum to do the filling, and structures the parser as an iterator to avoid allocation. #graphics #performance
on 02016-08-03getting 7× #performance on #prefix-sum with #SSE in #C++
on 02016-08-03The #HashDoS #security fix in #Java 7 introduced #multithreading #performance problems, which were later fixed; in Java 8 hash32 was removed
on 02016-08-03Bernie Greenberg’s #history of Multics #Emacs, which was the first Lisp Emacs, including a lot of context about text editing and #UX in the 1970s, and explaining the kinds of #performance concerns they had: “The buffer being edited is defined by about two dozen Lisp variables of the basic editor... The alternate approach to multiple buffers would have been to have the buffer state variables referenced indirectly through some pointer which is simply replaced to change buffers. This approach, in spite of not being feasible in Lisp, is less desirable than the current approach, for it distributes cost at variable reference time, not buffer-switching time, and the former is much more common.”
on 02016-07-24Examples of #profiling #performance with "perf_events", sometimes called "perf".
on 02016-07-24This "Swarm" chip from Daniel Sanchez (at MIT) has a #hardware priority queue for #parallelism and #performance
on 02016-07-04Non-null “optimizations” in GCC and LLVM cause major #security problems and have no real #performance benefit.
on 02016-06-29Embarrassing how bad the #performance of #big-data things like Elastic Map Reduce and #Hadoop can be when used on only a couple of gigabytes of data; the Configuration to Outperform a Single Thread is apparently a lot bigger than that because this guy got the answer from 1.75 gigabytes in 12 seconds on his laptop (in parallel, using xargs -P4) instead of the 26 minutes the Hadoop cluster needed. He gets slightly better performance using mawk instead of gawk.
“High-#performance purely functional data-parallel array programming” #language on the GPU. “comes with a heavily optimising ahead-of-time compiler that generates GPU code via OpenCL” #GPGPU
on 02016-05-03gpucc, an LLVM-based, fully open-source, CUDA-compatible compiler for high #performance computing #GPGPU #compilers (but they still rely on NVIDIA’s backend compiler to the actual undocumented GPU instruction set)
on 02016-04-25this collection of #compilers stuff for Python includes several different things that take the “compile array expressions into high-#performance code” approach: Numba, Theano, Parakeet, and Copperhead.
on 02016-04-22A new #PyPy release with significant #performance improvements. Also they now support five architectures: i386, amd64, ARM (32-bit), PowerPC 64, and IBM System/360 (s390x), and they now have basic numpy support instead of none. #python #compilers
on 02016-04-21discussion of #Python #performance (Pyston, Unladen Swallow, Falcon, Pyjion, Cython, and of course #PyPy)
on 02016-04-19a story of printf-#debugging the Atom text editor with to fix a critical #performance bug hidden in an exponential-time regexp due to catastrophic backtracking in the traditional-style NFA engine used by Atom’s JS engine. “This was surely the wildest adventure I’ve had so far in the open-source world.”
on 02016-01-29some major #performance bugs in #Chromium got fixed in January 2015, but Chrome on #Android is still slow.
on 02016-01-12Details of #SIMDjs #performance in #Mozilla #JS
on 02016-01-08#introduction to improving #JS #performance with #canvas
on 02016-01-08how to get theoretical max #performance of 4 flops per cycle with #SSE intrinsics: manual loop unrolling and careful interleaving of multiplies and adds, at least for pre-#FMA processors.
on 02015-12-05#performance #benchmarks of #Raspberry-Pi 2 in #Python.
on 02015-11-16#TensorFlow #performance #benchmarks
on 02015-11-11the #performance of #HTML5 #localStorage.
on 02015-09-09out-of-the-box #parallel sorting and #prefix-sum of arrays in #Java 8 for #performance.
on 02015-08-27#ZFS developer says the #BLAKE2 #hash isn't even close to the #performance of Edon-R.
on 02015-08-23#performance of the Murmur3 #hash is about 1.68 bytes per cycle, 5 gigabytes per second on a core of a 3GHz Intel Core 2 Quad Q9650, as of 2011.
on 02015-08-23#hash #performance #benchmark including Murmur2, Murmur3, CRC-32, FNV-1a, but not SipHash yet and actually no “secure” hash functions like MD5, SHA-2, or Keccak.
on 02015-08-23#hash #performance #benchmark in #JS. MurmurHash gets 14k ops/sec in my Firefox; MD5 only 687; SipHash only 1589; and the equivalent of Java string hashCode() 12k.
on 02015-08-23In 1999, #bcrypt required using 6 or 8 “rounds” (really lg rounds) for reasonable #performance; nowadays 12 or more is probably necessary for #security.
on 02015-08-18#bcrypt in pure #node #JS has #performance of 108ms with rounds=9 (which really means 2⁹ rounds) on a 3GHz CPU; a C++ version takes 76ms. It would probably make sense to use 13 or 14 nowadays, as a result, unless you’re running bcrypt on your smartphone.
on 02015-08-18#Algorithms for best #performance on a #parallel #prefix-sum in #CUDA for #GPGPU as of 2007.
on 02015-08-15“Data-Parallel Finite State Machines”: #parallel simulation of #finite-automata that are “up to 3× faster than optimized sequential implementations on a single processor” in #performance due to using #SIMD instructions. Seems to be some tweaks to the #prefix-sum thing Blelloch mentioned in his 1993 #paper, which they call “enumerative computation”? These algorithms are probably crucial to speeding up #CSV parsing.
on 02015-08-15difficulties with #OO in e.g. #C++ when programming #video-games for the #PlayStation 3 in 2009. Unfortunately these are slides. Basically the author’s claim is that C++ encourages memory layouts that have terrible #performance on modern hardware. Very concrete, with lots of C++ #examples and PowerPC #asm disassembled from them of 3-D programming and diagrams of #cache line evictions and whatnot. Basically his recommended cure is to use #parallel-arrays for most things and “flat” contiguous level-order tree linearizations for tree data.
on 02015-08-13Commentary on Sylvan’s #performance gripes about object-graph languages.
on 02015-08-13“High-level languages are slow” because “more allocations means more time spent collecting garbage.” Sylvan is one of the thinkers who is leading me to articulate the difference between the different models of memory provided by different languages: object graphs vs. object embedding vs. #parallel-arrays vs. logic programming and linear types and whatever else. #performance
on 02015-08-13How to control nondeterminism and make #performance #profiling possible on the #Java #JVM. This has been a real problem for me in developing #DyCSV.
on 02015-08-13The #hsdis #HotSpot #JVM disassembler for #Java #performance work.
on 02015-08-13some notes on the new features of the #OCaml 4.00.0 release in 2012, including much better #GDB integration and some new #performance optimizations, including in particular infix function application operators like Haskell's $ and flip $.
“How to print out #Java compiled #asm instructions on #Linux/MacOS” by installing #hsdis (for #performance mostly).
on 02015-08-10#downloads of the #HotSpot disassembler #hsdis for #Java #performance work.
on 02015-08-10Brief summary of how to build the #HotSpot disassembler hsdis for #Java #performance work.
on 02015-08-10Apparently indirect branch instructions in switch()es in interpreters in Haswell is no longer the #performance bottleneck that it used to be in Nehalem. #hardware
on 02015-08-10How to do high-#performance #OCaml, according to a HFT prop #trading shop. Lots of #asm and a bunch of FFI C. I didn't know None had the same representation as int 0.
An #OCaml blog by some (Turkish?) people who have apparently an SMT solver. Notes about #performance and #profiling.
on 02015-08-05How to understand #performance and do #profiling in #OCaml.
on 02015-08-05That jerk Jon Harrop says he thinks #OCaml #multithreading’s lack of multicore #performance scalability is due to the GC, and that OCaml has become less popular since the dawn of the multicore era as a result.
on 02015-08-05What is the Configuration to Outperform a Single Thread? Surprisingly often it’s ridiculous or even unbounded. #performance #Spark #big-data
on 02015-08-05