#SSE (and #AVX, #MMX) #asm #SIMD instructions for interleaving the low-order bytes, words, doublewords, or quadwords of two registers. Perfect-shuffle instructions, but not for bits.
on 02025-11-28#Ryg on #AVX and #SSE #SIMD instructions, multiplies especially. #toread #asm
on 02025-11-24#SIMD instructions for #AMD’s znver6 processor: “AVX512BMM instructions include Bit Matrix Multiply and Bit Reversal operations.” “2 parallel 16x16 non-transposed fused BMM-accumulate (BMAC) with OR/XOR reduction. Each 256-bit chunk of a zmm register holds a 16x16 bit matrix. The third source matrices for accumulation are in zmm1.” This is a binutils patch from Umesh Kalvakuntla at AMD.
on 02025-11-11vectorizing RGB to grayscale conversion code to #SIMD in #GCC
on 02025-11-05#tutorial on #NEON #SIMD intrinsics in C
on 02025-10-23#GCC vector extensions for portable #SIMD code in C.
on 02025-10-23Complete #ARM #SIMD #asm intrinsics list, including the strided things. Reference #documentation.
on 02025-09-28#ARM #asm #SIMD intrinsics for (vertical) addition. vaddq_u16 (vadd.i16) and the like.
#paper on "goSLP", which uses linear #optimization (ILP) for #SIMD #vectorization, achieving significant #performance gains on floating-point benchmarks like SPEC2017fp. “Using an integer linear programming (ILP) solver, goSLP searches the entire space of statement packing opportunities for a whole function at a time, while limiting total compilation time to a few minutes. Furthermore, goSLP optimally solves the vector permutation selection problem using dynamic programming.” #compilers
on 02024-12-12Dan Luu on modern #hardware #performance, largely having to do with #concurrency, #locking, #context-switches, and #SIMD, with a little bit about #GPGPU
on 02021-11-02#PDF #paper on #SIMD #parsing of #XML using “bitstream addition” — bit-sliced processing of the character stream #toread
on 02016-09-20a guide to #SSE #asm language programming, although the author is a little confused about what the “SIMD” acronym means, which is not a promising start. #simd
on 02015-10-04“Data-Parallel Finite State Machines”: #parallel simulation of #finite-automata that are “up to 3× faster than optimized sequential implementations on a single processor” in #performance due to using #SIMD instructions. Seems to be some tweaks to the #prefix-sum thing Blelloch mentioned in his 1993 #paper, which they call “enumerative computation”? These algorithms are probably crucial to speeding up #CSV parsing.
on 02015-08-15