#SSE (and #AVX, #MMX) #asm #SIMD instructions for interleaving the low-order bytes, words, doublewords, or quadwords of two registers. Perfect-shuffle instructions, but not for bits.
on 02025-11-28#Ryg on #AVX and #SSE #SIMD instructions, multiplies especially. #toread #asm
on 02025-11-24“Fast Base64 encoding/decoding with #SSE #vectorization”, supposedly the best #introduction to SSE
on 02017-03-22#sse #performance #toread by #ryg #vectorization
on 02016-09-29#graphics #algorithms #performance #sse #toread by #ryg
on 02016-09-29permuting bytes and words using AVX2 VPACK, VPUNPCKL, VPSHUFB, etc. #sse #performance #vectorization
on 02016-09-29“an algorithm and associated sample code [using #SSE] for software occlusion culling which is available for download” #algorithms #performance #toread this is what ryg was commenting on in his 2013 #graphics thread
on 02016-09-29Halide is a domain-specific language embedded in C++ for image processing, supporting #GPGPU backends as well as #SSE, by two of the Simit authors.
on 02016-09-22an #SSE widening dot-product operation in #X86/#amd64. With 128-bit operands, it does 8 16-bit multiplies and adds pairs of the 8 32-bit products to get 4 32-bit sums, which fit into the destination register.
on 02016-08-08#Raph-Levien lucidly explains how his new #font rasterizer "font-rs" in #Rust is the fastest (≈6× faster than FreeType). Among other things, he uses #SSE #prefix-sum to do the filling, and structures the parser as an iterator to avoid allocation. #graphics #performance
on 02016-08-03getting 7× #performance on #prefix-sum with #SSE in #C++
on 02016-08-03An optimized integer division library, using multiplication and bitshifts, in part because the #SSE #ISA has no division instruction (like the Cray vector operations before it). #algorithms
on 02016-07-21how to get theoretical max #performance of 4 flops per cycle with #SSE intrinsics: manual loop unrolling and careful interleaving of multiplies and adds, at least for pre-#FMA processors.
on 02015-12-05programming #SSE in #C and #vectorization of your code despite conditionals
on 02015-11-19a guide to #SSE #asm language programming, although the author is a little confused about what the “SIMD” acronym means, which is not a promising start. #simd
on 02015-10-04