#paper on "goSLP", which uses linear #optimization (ILP) for #SIMD #vectorization, achieving significant #performance gains on floating-point benchmarks like SPEC2017fp. “Using an integer linear programming (ILP) solver, goSLP searches the entire space of statement packing opportunities for a whole function at a time, while limiting total compilation time to a few minutes. Furthermore, goSLP optimally solves the vector permutation selection problem using dynamic programming.” #compilers
on 02024-12-12#RISC-V #vectorization intrinsics search
on 02023-07-15apparently #RISC-V #vectorization on the T-Head core speed up most ffmpeg inner loops by a factor of 2–4
on 02023-07-13Daniel Bernstein explains why #vectorization is a #performance win, even when it requires throttling the #hardware clock
on 02023-06-13a high-performance implementation of OpenAI's #Whisper #speech-recognition algorithm in C++ with #vectorization. #neural-networks
on 02023-02-28“Fast Base64 encoding/decoding with #SSE #vectorization”, supposedly the best #introduction to SSE
on 02017-03-22#Numpy #ebook by Nicolas P. Rougier, cc-nc-by-sa, full of awesome #vectorization and #performance tricks
on 02017-01-14#sse #performance #toread by #ryg #vectorization
on 02016-09-29permuting bytes and words using AVX2 VPACK, VPUNPCKL, VPSHUFB, etc. #sse #performance #vectorization
on 02016-09-29ChaCha12-256 (without the usual security margin) has about the same #performance as AES-128 on several generations of Intel #hardware, because it’s good at #vectorization of #crypto, and will get faster now that Intel is exposing their 52-bit multipliers as #djb requested in 2002. Also apparently Intel is adding inversion in GF(256).
on 02016-08-11programming #SSE in #C and #vectorization of your code despite conditionals
on 02015-11-19