#PDF #paper from Anton #Ertl about Portable Assembly #Forth. #toread #asm
on 02026-09-19tiny self-hosted #Lisp in #asm (?) #toread for #bootstrappable #small-is-beautiful
on 02026-09-03#PDF #ebook on #PDP-10 #asm programming by Ralph Gorin
on 02026-08-31#roconnor wrote a combinator graph reduction machine in #RISC-V #asm #toread
on 02026-05-26#stage0 #C compiler in #asm. #compilers
on 02026-05-20#Godbolt about why #amd64 #asm uses so many xor instructions
no real reason people on #amd64 use xor instead of sub to clear a register in #asm; both of them require zero micro-ops, merely renaming a register. #performance
#SSE (and #AVX, #MMX) #asm #SIMD instructions for interleaving the low-order bytes, words, doublewords, or quadwords of two registers. Perfect-shuffle instructions, but not for bits.
on 02025-11-28#Ryg on #AVX and #SSE #SIMD instructions, multiplies especially. #toread #asm
on 02025-11-24#video on #asm programming for #UEFI, showing how to call ConOut, WaitForEvent, and GetTime UEFI subroutines from assembly, but using Microsoft MASM and the VC++ linker, gag. He tests his hello-world image in QEMU by writing it to a flash drive with a GUID partition table (not an MBR) with an EFI system partition. Sponsored by a mini-PC vendor. Also covers the EFI_GRAPHICS_OUTPUT_PROTOCOL (GOP) which contains QueryMode, SetMode, and Blt. Also how to start up additional processors (he’s on a 12-core amd64 machine) via EFI_MP_SERVICES_PROTOCOL. But UEFI doesn’t support keyup events! Only keydown events! So it sucks for his Zaxxon reimplementation. The EFI_SIMPLE_POINTER_PROTOCOL for talking to the mouse might be okay, although it only supports two buttons (gag).
#RISC-V #asm conditional moves and the 64-bit #ARM CSEL instruction #toread
on 02025-10-02discussion of #RISC-V #asm conditional moves and the 64-bit #ARM CSEL instruction
on 02025-10-02Complete #ARM #SIMD #asm intrinsics list, including the strided things. Reference #documentation.
on 02025-09-28#ARM #asm #SIMD intrinsics for (vertical) addition. vaddq_u16 (vadd.i16) and the like.
#ARM #documentation on loading constants in #asm, using constant pools, MOV, MVN, etc. “MOV can load any 8-bit constant value, giving a range of 0x0-0xFF (0-255). It can also rotate these values by any even number.” Adjacent sections of the manual explains LDR with constant pools and the MOV32 pseudo-instruction.
Similarly movs is allowed and mov is not in Thumb-1 #ARM #asm. As old_timer says, “The short answer is that you cannot because the instruction does not exist. If you are using unified syntax documentation that glosses over which instruction sets are supported, then you have, as many others, fallen into the trap of the major failing of the unified syntax.”
GNU assembler for #ARM #asm with .syntax unified insists that you write adds rather than add to compile to Thumb-1 code. And with Thumb-2 it will generate a wide add instruction if you don’t say adds. Notlikethat explains: “The only Thumb encodings for (non-flag-setting) mov with an immediate operand are 32-bit ones, however Cortex-M0 doesn’t support those, so the assembler ends up choking on its own constraints.” Similarly for mov vs. movs.
#ARM Cortex-M3 added a CLZ #bitmanip instruction, but this is an #asm version of it for more primitive ARMs.
on 02025-09-09#ARM #asm has movt and movw instructions in later Thumb2 versions; these permit you to load a 32-bit constant in two instructions, each supplying 16 bits. movw clears the high bits and movt loads them. This question pertains to #Clang needing -mthumb to assemble Thumb asm. #movt
The CLZ #asm #bitmanip instruction is __builtin_clz on GCC, but this page claims to have a 13-instruction-time version with better #performance than all the “Hacker’s Delight” entries, with “[bisection] to find out which 8-bit chunk of the 32-bit number contains the first 1-bit, which is followed by a lookup table clz_lkup[] to find the first 1-bit within the byte.” Handy for Cortex-M0 #ARM chips without the instruction.
#ARM CPUs (bigger than Cortex-M0 or Cortex-M0+) have a CLZ #bitmanip instruction to count the number of leading zeroes, and CMSIS-CORE has a __CLZ intrinsic function for when you aren’t in #asm. Also signed saturate, unsigned saturate, rotate right with extend, wait for event, etc.
#amd64 #asm instruction TZCNT or BSF counts the number of trailing zeroes. #bitmanip
on 02025-09-09#VAX #asm return instruction is rather hairy: “SP is replaced with FP plus 4. A longword containing stack alignment in bits 31:30, a CALLS/CALLG flag in bit 29, the low 12 bits of the procedure entry mask in bits 27:16 and a saved PSW in bits 15:0 is popped from the stack and saved in a temporary (tmp1). PC, FP and AP are replaced by longwords popped from the stack. A register restore mask is formed from bits 27:16 of tmp1. Scanning from bit 0 to bit 11 of tmp1, the contents of the registers whose numbers indicated by set bits in the restore mask are replaced by longwords popped from the stack. SP is incremented by bits 31:30 of tmp1. PSW is replaced by bits 15:0 of tmp1. If bit 29 of tmp1 is 1 indicating a CALLS was used) a longword containing the number of arguments is popped from the stack. Four times the unsigned value of the low byte of this longword is added to SP and SP is replaced by the result.”
on 02025-09-09the relocation R_X86_64_32S against symbolkj' can not be used when making a PIE object; recompile with -fPIE` problem: how do you need to write your #asm to avoid it? Question by #Agner Fog.
"sjasmplus" is a #Z80 cross-assembler for #asm #retrocomputing with Lua scripting
on 02024-11-15a #Forth in 512 bytes. Lots of syntax-highlighted #asm, I think the full literate program. By @meithecatte. #toread #small-is-beautiful
on 02024-11-03Linux’s system-call #ABI differs from the standard #amd64 function-call ABI. “This means system-call wrapper functions can just set eax and run syscall instead of needing to do something like...” #asm
#GCC lets you associate #asm registers with particular #C variables, either global or local.
on 02024-09-23a macro assembler written in #sh. #small-is-beautiful #asm
on 02024-08-31#small-is-beautiful #Lisp in 223 lines of #asm and 436 bytes
on 02024-08-03outline of #amd64 #ABI for #asm on Linux
on 02024-08-01In #amd64 #asm you call the low 32 bits of a numbered register rNd, e.g., r8d, r11d, r15d. Writing to it clears the upper 32 bits, same as ecx vs. rcx. The low byte is rNb, I forget if there’s a name for the low 16 bits.
on 02024-06-30#RISC-V vector extension #video by #LaurieWired aimed at super beginners, explaining an #asm example line by line, then walking through it in GDB; apparently the Kendryte Canaan K230 chip supports RVV 1.0. But she doesn’t even mention T-Head, so I’m not sure I can trust her information. Like, her example assembly program doesn’t loop, and I think it needs to loop in case the vector registers on the hardware are less than 256 bits. Like, I think she needs something like blt t2, a3, loop_start to get the vector length agnosticism she’s talking about at the beginning but never explains.
Randall Hyde’s current #asm home page
on 02024-04-15Randall Hyde’s old #asm home page (last updated 01996, gone from the live web since 02009)
on 02024-04-15Oscar Toledo G. just released a boot sector real-time animated ray tracer, based on the BBC BASIC one, which can be run from a floppy disc boot sector with no OS #small-is-beautiful #3D #graphics #raytracer #asm
on 02024-04-15#asm source code for 86-dos (#MS-DOS) 1.25. Only about 4 kloc. #history
on 02024-01-14how to implement setjmp and longjmp using inline #asm in a naked function on amd64 and i386. Describes the #Microsoft-Windows #calling-convention. #Wellons
#moon-child’s code for analyzing Intel #asm instruction scheduling #performance in SBCL #Lisp
on 02023-10-05#ARM #asm addressing modes #PDF of lecture slides
on 02023-07-19#Android #AOSP #ARM #asm for #Linux system calls
on 02023-07-19#Linux #ARM system call interface (for #EABI), with a table of system calls and arguments. #asm
on 02023-07-19In #Linux on #ARM in #EABI, the system call arguments are in registers r0-r6, and the return value is in registers r0 and r1. #asm
on 02023-07-19Since #Linux 2.6, #ARM system calls under the #EABI use the swi 0 or svc 0 instruction with the system call number in R7 and arguments in r0-r5, with 64-bit arguments starting in an even register. #asm
The so-called #ARM #EABI is the standard ABI for #Linux, supplanting the so-called OABI. Explains how system calls work in #asm. Also explains the checkered history of interworking. Debian architecture “armel” is EABI, “arm” is OABI, dropped in Squeeze.
on 02023-07-19comparison of #gas and #NASM #macros from 02007; malicious JS blanks the page after loading but can be defeated with Firefox readability button. #asm
on 02023-07-15#documentation for #gas #asm #macros, including an example that counts
on 02023-07-15#gas nested control-flow #asm #macros: _if, _else, _elseif, _endif, using .altmacro
#gas -alm shows a full recursive trace of expansion of #asm #macros, which is handy for debugging them, as well as a full listing of the machine code emitted
in #gas printing #asm debug messages with .print works better with .altmacro syntax because you can %expr when normally .print only accepts literal strings
Raymond Chen's #ARM #asm #tutorial, section on bitwise ops.
on 02023-07-13Raymond Chen’s #ARM #asm #tutorial, section on cmn, which apparently doesn’t set the carry flag in the same cases that cmp with the negative would, specifically in the cases of comparing against zero and the most negative integer. More details on flag bits.
on 02023-07-13Raymond Chen’s #ARM #asm #tutorial. Section on arithmetic. Explains that in mvn the n is logical negation, while in cmn it’s arithmetic negation. And how to negate by reverse subtracting from literal zero.
on 02023-07-13Raymond Chen’s #ARM #asm #tutorial. Section on loading constants in one instruction, with mov, mvn, #movt, shifts, rotates, etc. This is the thing to read to understand how to generate constants on ARM, especially Thumb. “You can take your 8-bit unsigned immediate and shift it left by up to 24 positions, thereby allowing you to create any 32-bit constant where the span from the lowest to highest set bit is at most 8 positions. There are also a few special transformations (...) handy for setting up a register to fill memory with a repeating pattern. (...) In addition to all of the special constants that can be generated with MOV, you can use the MVN instruction to generate the bitwise NOT of them all.”
Raymond Chen’s #ARM #asm #tutorial covering the available addressing modes, including quirks like the absence of shifts other than left shift or more than three for byte or halfward access. Also how the Microsoft assembler says |foo| to mean “the address of foo” and how literal pools work.
The #Microsoft-Windows #ARM #asm ABI. It really does require that r11 always be a frame pointer! But then it softens that to a recommendation. It doesn’t seem to explain the IT/ITE/ITTE limitation mentioned in Raymond Chen’s blog post; possibly it’s gone?
on 02023-07-13Raymond Chen’s #ARM #asm #tutorial introduces #Thumb. The #Microsoft-Windows ABI only lets you conditionalize a single instruction with IT! You can’t use ITE, ITTE, ITEE, etc., at all there, because it doesn’t know how to recover from hardware interrupts.
on 02023-07-13First in Raymond Chen’s #tutorial about #ARM #asm; apparently #Microsoft-Windows 10 still supports 32-bit ARM, but only in #Thumb mode. This is the intro where he explains the register set, stack alignment constraints, flags, etc.
on 02023-07-13Raymond Chen’s #tutorial about #ARM #asm, section on common patterns in compiler-generated code: function pointer invocation, vtable invocation, control-flow guard.
on 02023-07-13The end of Raymond Chen’s #tutorial about #ARM #asm. Apparently Microsoft Windows does use the old convention where r11 points to a linked list of stack frames. The example code he walks through here uses push, add, mov, ldr, mvn, tst, beq, movs, str, b, ldr, subs, asrs, bl, and pop. The Microsoft assembler uses slightly different syntax.
on 02023-07-13Raymond Chen’s #tutorial about #ARM #asm, section on the APSR flags register, reading with mrs Rd, apsr and writing with msr apsr, Rn. Mentions a failure of virtualizability of the original ARM: privileged status bits read as 0 in user mode and ignore writes.
Raymond Chen's #tutorial about extending #ARM #asm branching instructions’ reach with “trampolines” or “veneers”, in particular bl, which is used for intra-module calls (when you know you’re not switching modes) but can only reach 16MiB away.
on 02023-07-13Raymond Chen's #tutorial about #ARM #asm branching instructions such as tbb.
on 02023-07-13stepping you through #bootstrapping #stage0 through #asm from hex machine code.
on 02023-07-13#ARM #asm #Cortex-M4 technical reference manual, including instruction timings. #PDF
on 02023-07-12some #asm programs from 01958 from the TX-2 at Lincoln Lab, no longer the Servomechanisms Laboratory, possibly printed on a Flexowriter. #history #retrocomputing
on 02023-07-12#ARM #Thumb instruction set encoding. Ugh. A full assembler in Lisp for Thumb. #asm
on 02023-07-08explains how to control #OpenOCD by telnetting to port 4444 and typing textual coommands. Discussion of what #asm the #STM32 STM32F103 does and doesn't support. Discussion of linker scripts. Explanation of the #ARM literal pool.
on 02023-07-05lengthy exegesis of #ARM #ABI in #asm as our hero fights with their compiler
on 02023-07-05oh hey #GCC has -fverbose-asm which helps significantly in getting readable #asm for #ARM!
apparently on #ARM when #GCC only has to push LR, it pushes an extra unnecessary temp register onto the stack to ensure SP is 8-byte-aligned at all times, because that’s what the ABI requires, because the ARMv4 strd and ldrd instructions require 8-byte alignment. #asm
oh my, I hadn’t realized that all the Cortex-M #ARM #hardware only supported Thumb-2 #asm. Or that every #STM32 #microcontroller supported SWD for single-stepping with OpenOCD. Apparently to burn the Flash on an STM32 you openocd -f interface/stlink-v2.cfg -f target/stm32f1x.cfg -c "program prog1.elf verify reset exit"? And then to provide a gdbserver on port 3333 you openocd -f interface/stlink-v2.cfg -f target/stm32f1x.cfg?
Slide deck about #ARM #asm. Supposedly the #Raspberry-Pi 2 and 3 (but not 1 and Zero) support #Thumb-2 instructions, and the #STM32 STM32L only supports Thumb-2 (no original ARM fixed-width instruction set). There’s a super sweet #code-density graph on p.35 of the #PDF: 8086 is the champion; PDP-11 and Z80 are almost tied and beat Thumb-2, which beats Thumb, which beats VAX, which beats i386, which beats x86_64, which beats 6502, which beats arm_eabi, which beats the living shit out of ia64.
on 02023-07-04#x86 and #amd64 MOVZX instruction doesn't take an immediate argument. #asm
on 02023-07-03PDP-6 #asm hacks with #macros etc.
on 02023-07-02building a hash table in MIDAS #asm #macros for PDP-6 Lisp. #retrocomputing #hashing
on 02023-07-02MIDAS #asm source code for PDP-6 DDT. #retrocomputing
on 02023-07-02#Apple //e ROM monitor #asm listings. #retrocomputing #history
on 02023-07-02annotated #VT100 #terminals #firmware in #8080 #asm
on 02023-07-01#PDF #ebook “Archimedes Operating System: A Dabhand Guide” from 01991 #ARM, largely #asm. Covers the whole instruction set in chapter 2, pp. 20–32.
on 02023-07-01#Tutorial #introduction to #ARM #asm using gas, directed toward #GameBoy Advance hacking. A bit rough-and-ready but only 22000 words. Covers add, ldr, ldmia, b, cmp, mov, lsl, lsr, asr, ror, sub, strh, bl, bx, conditions (-ne, -lt, and -ge), the -s suffix for setting flags, conditional instructions, the (even-bit-shifted) byte-sized immediate limitation, inline shifts, write-back (preincrement and postincrement) addressing mode, and Thumb, and finally the whole instruction set and a GBA optimization guide.
on 02023-07-01#DB2 #embedded-SQL in #asm. #SQL
on 02023-06-24how to do Galois field instructions on CPUs in 02021. #asm #programming
on 02022-11-06#Pasmo #Z80 #asm current (02009) site. Jesus, this window still has 1375 tabs open.
on 02022-02-09"Pasmo" is a #Z80 #asm assembler that can also cross-assemble Z80 programs for the #8086
on 02022-02-09"Salto": System for Assembly-Language Transformation and Optimization, a 1996 #paper on writing things similar to Valgrind or gprof, but at the #asm level, mostly for #performance analysis
on 02021-02-09#6502 #asm opcode listing with timings and encodings
on 02021-01-19#6502 #asm opcode listing with timings and encodings
on 02021-01-19Ken Boak’s #SIMPL #source-code for the MSP430 #microcontroller; I think this is the 780-byte version. The source is only 811 lines of #asm. #smallisbeautiful
on 02017-07-11#PDF #paper by Anton Ertl about “Portable #asm #Forth” as an UNCOL
on 02017-04-30An online #Z80 compiler, showing you compiler errors (from sdcc) and #asm and hex output in real-time.
on 02017-04-29#6502 #asm #source-code for Prince of Persia, one of the all-time great #video-games, one of the first to feature rotoscoped motion, and Linus Torvalds’s favorite game around the time he wrote Linux. This is the original Apple II version, rather than the better known IBM PC version, published by its author Jordan Mechner in 2012 after being recovered by Jason Scott and Tony Diaz. “As the author and copyright holder of this source code, I personally have no problem with anyone studying it, modifying it, attempting to run it, etc. Please understand that this does NOT constitute a grant of rights of any kind in Prince of Persia...”
on 02017-03-30playing with #aarch64 #ARM #asm with helloworld in qemu and a cross-compiling toolchain. cc-by-nc-sa
on 02016-10-10The 8086 family #asm #ISA, including instruction cycle timings for the 8086 through the 80486.
on 02016-09-16The #Apollo 11 guidance computer (AGC) source code in #asm. #history
on 02016-07-07#SBCL is a good way to play around with coding things in #asm; in this case Paul Khuong experimented with doing some simple register allocation to compile stack-machine instructions to AMD64 instructions. Full #Lisp source provided in the article.
on 02016-06-27A 6502 #asm programming system implemented as a Rust macro, realizing the STEPS vision of “mood-specific languages” to a remarkable extent. Has problems with error reporting, though.
on 02016-02-11a guide to #SSE #asm language programming, although the author is a little confused about what the “SIMD” acronym means, which is not a promising start. #simd
on 02015-10-04a #Python library for generating #asm for i386 and amd64.
on 02015-09-21The #formal-methods paper on #Coq for programming in #asm, "Coqasm".
on 02015-08-25is a comment about a "generic assembler" where the instruction set (including #pattern-matching) is part of the #asm source being assembled.
on 02015-08-22difficulties with #OO in e.g. #C++ when programming #video-games for the #PlayStation 3 in 2009. Unfortunately these are slides. Basically the author’s claim is that C++ encourages memory layouts that have terrible #performance on modern hardware. Very concrete, with lots of C++ #examples and PowerPC #asm disassembled from them of 3-D programming and diagrams of #cache line evictions and whatnot. Basically his recommended cure is to use #parallel-arrays for most things and “flat” contiguous level-order tree linearizations for tree data.
on 02015-08-13“How to print out #Java compiled #asm instructions on #Linux/MacOS” by installing #hsdis (for #performance mostly).
on 02015-08-10My reimplementation of the 31-byte MS-DOS demoscene #demo Klappquadrat in #Python using Numeric (the predecessor to #NumPy.), based on the #asm source and disassembly.
on 02015-08-10using #formal-methods to verify #C #compilers and carry through proofs of your #verifiable-c programs through to the #asm using separation logic. Appel, Leroy, and some others.
on 02015-08-06How to do high-#performance #OCaml, according to a HFT prop #trading shop. Lots of #asm and a bunch of FFI C. I didn't know None had the same representation as int 0.