#Llama-cpp is adding Johannes Gaessler’s backend-agnostic tensor parallelism for running #neural-networks on #clusters (either CPU or #GPGPU), like #vLLM can already do. #AI
on 02026-02-09#Whisper.cpp 1.8.3 #speech-recognition can now run 12× faster by using #Vulkan to get #GPGPU speedups even on Intel (?) hardware. #neural-networks
on 02026-01-16“#LambdaCube #3D is a domain specific language and library that makes it possible to program GPUs in a purely functional style.” Abandoned ten years ago. #Haskell #GPGPU #graphics
on 02026-01-12#PDF #paper from 02013 about the language #Halide for #GPGPU #graphics
on 02024-01-28#PyOpenCL #introduction in DDJ from 02013. #OpenCL #GPGPU
on 02021-11-03another #BLAS implementation for #OpenCL. #GPGPU
on 02021-11-03A #BLAS (level 1, 2, & 3) implementation for #OpenCL. #GPGPU
on 02021-11-03"PyOpenCl" gives you easy, Pythonic access to the #OpenCL parallel computation API. #GPGPU
on 02021-11-03"Futhark" is a small programming language designed to be compiled to efficient parallel code. “Because it’s nicer than writing CUDA or #OpenCL by hand!” It is a statically typed, data-parallel, and purely functional array language in the ML family, and comes with a heavily optimising ahead-of-time compiler that presently generates either GPU code via CUDA and OpenCL. ... You can compile a Futhark program to a Python module that internally uses #PyOpenCL to execute code on the GPU. #GPGPU
on 02021-11-03#GPGPU #mining is said to be driving up #AMD #hardware #pricing
on 02021-11-02NVIDIA doesn’t support #OpenCL on Tegra or Jetson #GPGPU #hardware
on 02021-11-02#WebGPU is like #WebGL for #GPGPU, available in Chrome Canary; example code uses #WGSL, a textual near-assembly language for #SPIR-V with a Rust-like syntax
on 02021-11-02Hashcat has CUDA and #OpenCL backends, on an NVIDIA 3060ti #GPGPU getting 1006 scrypt kilohashes per second on the easiest setting, 1793 PBKDF2-HMAC-SHA256 kilohashes with 999 iterations, 43.8 kilohashes with bcrypt Blowfish with 32 iterations, 39.9 kilohashes with Cisco-IOS $9$ scrypt with 1 iteration. #security #GPGPU
on 02021-11-02Hashcat has CUDA and #OpenCL backends, on an NVIDIA 3060ti #GPGPU getting 1006 scrypt kilohashes per second on the easiest setting, 1793 PBKDF2-HMAC-SHA256 kilohashes with 999 iterations, 43.8 kilohashes with bcrypt Blowfish with 32 iterations, 39.9 kilohashes with Cisco-IOS $9$ scrypt with 1 iteration. #security #GPGPU
on 02021-11-02#OpenCL 3.0 “hitting the reset button” last year? Apparently OpenCL 3.0 is “a fork of OpenCL 1.2” from 02011 (“All functionality beyond OpenCL 1.2 is optional”) rather than an extension of 2.0 because nobody has adopted 2.2 since its 02017 release. The commenters hate the “every feature is optional” idea. #GPGPU
on 02021-11-02#AMD #ROCm #GPGPU #documentation
on 02021-11-02the #ROCm #AMD #OpenCL #GPGPU repo
on 02021-11-02"NEO" provides #Intel #OpenCL 3.0 support, plus support for an alternative #GPGPU API called “oneAPI Level Zero”, with packages named intel-level-zero-gpu, intel-gmmlib, and intel-opencl-icd
on 02021-11-02"ROCm" is Radeon Open Compute, which provides reasonable #OpenCL support for #AMD #GPGPU computing on Linux (OpenCL 2.0) and is MIT-licensed
on 02021-11-02Apparently #OpenCL on #Vulkan sucks; neither #clspv nor #clvk is soup yet, and OpenCL support in #Mesa is still “dismal” (because only 1.2 I guess), but gets better with #AMD’s #ROCm and really good with NVIDIA’s proprietary driver. #GPGPU
on 02021-11-02alternatives for #OpenCL on #Vulkan #GPGPU
on 02021-11-02"clspv" is an implementation of #OpenCL (1.2) on #Vulkan. “The OpenCL C 1.2 language provides an expressive variant of the C language with which to program heterogeneous architectures.” #GPGPU
on 02021-11-02#Vulkan Compute as an alternative to #OpenCL, including benchmarks #GPGPU
on 02021-11-02Dan Luu on modern #hardware #performance, largely having to do with #concurrency, #locking, #context-switches, and #SIMD, with a little bit about #GPGPU
on 02021-11-02“Mitsuba 2: A Retargetable Forward and Inverse Renderer” is a #differentiable #3D ray tracer with a #GPGPU backend used for, among other things, #rendering #caustics and #optimization for analysis-by-synthesis computer vision
on 02021-02-09note from Eben Upton from 02014 on #GPU_FFT #GPGPU #FFT #Raspberry-Pi
on 02021-01-24"GPU_FFT" #FFT on #GPGPU on the #Raspberry-Pi 1’s VideoCore IV. Not sure if there's a newer version
on 02021-01-24Intel Programmer's Reference Manuals for all their recent GPUs, describing their instruction sets, registers, etc. #GPGPU
on 02018-10-05#PDF Intel Programmer’s Reference Manual for 2012 GPUs. Describes the instruction set. This is Intel’s official site. #GPGPU
on 02018-10-05#PDF whitepaper on the detailed architecture of 8th-generation Intel GPUs. #GPGPU
on 02018-10-05#PDF developer guide to 5th-generation Intel GPUs, mostly focused on graphics. #GPGPU
on 02018-10-05article on #FFT algorithms for Intel GPUs #GPGPU
on 02018-10-05#PDF on #FFT #algorithms on Intel GPUs, including some architectural details: each EU has 7 hardware threads, each with 128 32-byte registers, 512 bytes per work item in SIMD-8 mode, but only 2 32-bit #floating-point ALUs #GPGPU
on 02018-10-05"OpenCL" is from Khronos, the OpenGL custodian, but for #GPGPU.
on 02018-10-05Intel support for #OpenCL #GPGPU
on 02018-10-05#FFT #library in #GLSL compute shaders for #GPGPU, tested on ARM Mali T-760, NVIDIA GTX 760, and Intel HD 4400, on OpenGL ES and OpenGL 4.3 core.
on 02018-10-05“Perform massively parallel #GPGPU computations using #WebGL.” This sounds super awesome, although it’s doing this goofy pseudo-#JS parsing thing. Unfortunately their benchmark only gets a 7× speedup on this laptop, which is less than what you’d expect going from JS to e.g. C++, Golang, or Rust on the CPU.
on 02017-07-11"Stardust": #GPGPU in #JS for data #visualization with a D3-like API, complementary to D3.
on 02017-06-17preprint #PDF of reduced #affine-arithmetic #ray-tracing #paper with #GPGPU stuff from 2008
on 02017-05-19#paper “Fast #Ray-Tracing of Arbitrary Implicit Surfaces with Interval and #Affine-Arithmetic” #interval-arithmetic #graphics #toread with reduced affine arithmetic and #GPGPU goodness from 2008.
on 02017-05-19#video of Aaron Hsu talking about #compilers for two hours. #Co-dfns #GPGPU #APL
on 02017-03-07Aaron Hsu’s #paper on #GPGPU #APL #compilers, “The key to a data-parallel compiler” #Co-dfns
on 02017-03-07discussion thread on (and mostly by) Aaron Hsu (arcfide) and his #GPGPU #APL compiler "Co-dfns". #compilers
on 02017-03-07Aaron Hsu and his team have written a 750-line #APL compiler for #GPGPU, in itself apparently, in six years. #compilers #Co-dfns
on 02017-03-07#AMD open-sourced their whole #GPGPU stack
on 02016-12-21#Neural-networks style transfer from paintings to photos on #GPGPU. Super interesting #art. Built on TensorFlow. Includes #source-code, but not open-source: “Free for research/noncommercial use”.
on 02016-11-21Halide is a domain-specific language embedded in C++ for image processing, supporting #GPGPU backends as well as #SSE, by two of the Simit authors.
on 02016-09-22“High-#performance purely functional data-parallel array programming” #language on the GPU. “comes with a heavily optimising ahead-of-time compiler that generates GPU code via OpenCL” #GPGPU
on 02016-05-03gpucc, an LLVM-based, fully open-source, CUDA-compatible compiler for high #performance computing #GPGPU #compilers (but they still rely on NVIDIA’s backend compiler to the actual undocumented GPU instruction set)
on 02016-04-25Pete Warden ported the Deep Belief #deep-learning #neural-networks for #image-recognition to #Raspberry-Pi #GPGPU.
on 02016-04-16An #EDSL in #Python for #GPGPU programming the #Raspberry-Pi in the VideoCore assembly language.
on 02015-12-27running R on #AWS, including integration with Elastic MapReduce and #GPGPU.
on 02015-11-16a domain specific language in #Scheme for #GPGPU programming, compiling to #OpenCL. 3-clause BSD license. Build currently failing.
on 02015-09-16"Theano": a #Python library that allows you to define, optimize, and evaluate mathematical expressions involving multi-dimensional arrays efficiently, using #Numpy arrays, #GPGPU (NVIDIA only?), compilation of expression graphs to C, and symbolic differentiation. 3-clause BSD license.
on 02015-09-06“A scientific computing framework for #LuaJIT” #Lua ndarrays, like Numpy, with #GPGPU stuff. And even FPGAs! "Torch 7"
on 02015-08-18A survey #paper of #parallel #prefix-sum #algorithms, finding that the Kogge-Stone algorithm is more common in #GPGPU code than Blelloch’s, and with handy diagrams so you can see what they’re talking about! Also they apparently wrote a thing to use #formal-methods to verify different implementations.
on 02015-08-15#Algorithms for best #performance on a #parallel #prefix-sum in #CUDA for #GPGPU as of 2007.
on 02015-08-15“The Unreasonable Effectiveness of Recurrent #Neural-Networks” for #image-recognition, #NLP language modeling (producing what looks an awful lot like Markov-chain text with slightly better context-free properties), sequential processing for directing attention over an image, etc. Lots of animations and visualizations of RNNs doing their thing. The models are written in #Lua with something called #Torch-7 and run with #GPGPU. Lots of great comments! "unreasonable RNNs"
on 02015-08-15