#video review of #Anthropic’s #Claude Opus 5.5, which came out today. He had it write some 3-D games from a single prompt, using threejs, #Godot, and just plain C++ with no dependencies. He was not able to get it to drive a simple robot arm to push a toy car across a mat, though. “This model slaps.” #AI #LLMs #neural-networks
on 02026-09-22#Simonw quotes the “don’t be a #meat-proxy” argument about how to use #AI. #definitions #LLMs #generative-AI #AI-misuse
on 02026-08-21#Dan-Luu says #performance optimization is something #AI #LLMs (#neural-networks) are good at now. He created a new #regex library called "FRE" that gets good performance by compiling to native code by “having an agent loop for a month on improving regex engine performance with access to the rebar regex benchmark suite”, which led to overfitting, and then disciplining the agent with a holdout benchmark, using #ripgrep queries that came from his history with OpenAI’s “Codex” programming LLM. It’s still “substantially slower than the Rust regex engine on holdout benchmarks” but yet “enough to generally match 2nd tier regex engines in terms of performance”.
> This kind of technical work, which used to take a fair amount of time and expertise, can just be done trivially now. (...) While the open source version of BitFunnel “only” contains a bytecode interpreter and one JIT, the Bing version contains multiple JIT compilers. A project that did that level of optimization used to be a major undertaking, but “I could do that in a weekend” is now actually true for some of these kinds of projects.
(...)
> For an example from the GPT-5.1 or 5.2 days, with no knowledge of game AIs, I tried building an Azul AI. This ended up being the strongest AI in the world for the game by a pretty large margin. From reading the thesis that describes the 2nd strongest AI, I think my AI is probably a bit better on the “AI” side of things, but the main place it wins is on optimization despite spending what looks like maybe two orders of magnitude less [human] time [writing the software] (...)
> There’s a bunch of standard stuff it makes sense to do to debug and verify a multithreading algorithm for something like this, like implementing replay from debug logs that can reproduce bugs despite the algorithm being nondetermistic. Doing that alone would’ve probably been days to a week of work had I done it by hand, but it’s exactly the kind of thing an agent can trivially do in a loop (just have it try to replay logs and insert logging for non-determinism every time you don’t get a perfect replay). A lot of the tedium it used to take to get a tricky optimization like this working is gone.
(...)
> Now that this N [of person-days to verify that a tricky optimization works] has dropped by a tremendous factor (variable but, in terms of human time, frequently 1000x / 10000x / 1000000x, probably more like 1000x on dollar cost if you compare token costs at metered rates vs. the Bing engineer who wrote the compilers at JITs that the search index used), the number of these kinds of optimizations it makes sense to do goes way up.
> (...) for the AI I tried, it seems like you gain about 100 Elo for every doubling in speed (more than in chess, I suspect because draws are very rare). Just adding multithreading alone is enough to wipe the floor with an otherwise comparable AI on a large machine.
> (...) current publicly available SOTA models are pretty bad at experimental design, so I had to set up the framework they used to determine if an optimization is good, but once that was in place, it’s like any other optimization problem.
> Jamie Brandon (...)’s a reasonable performance engineer and he got an offer for the performance job he wanted [at Anthropic], but on a well-defined optimization problem, he [reports that he] doesn’t stand a chance against a decent model [Claude] (...)
> right before I started writing this post, I had an agent do workload-specific optimization for my ripgrep queries (...), which took about 2 minutes for me to launch. (...) After one pass of optimization, the workload optimized version is 2% faster than standard ripgrep on the holdout and it’s still getting faster.
He thinks this is going to result in a lot of companies workload-optimizing their software using coding agents. Also:
> (...) someone who doesn’t know anything about performance and is a reasonable user of LLMs (just in general, not on performance problems in particular) should generally be able to create software that has decent performance.
on 02026-08-21#Yt-dlp is abandoning the #Bun #JS runtime because #Anthropic has discarded the old #Zig codebase and switched to #Rust generated by #AI #LLMs. “As of the next yt-dlp and/or ejs release, only Bun versions 1.2.11 through 1.3.14 will be supported.” #neural-networks
on 02026-05-22New #vegan-models called "talkie" trained entirely on public-domain material. #AI #law #copyright #LLMs #neural-networks
on 02026-04-28discussion of recent serious #Anthropic #Claude degradation. #AI #neural-networks #LLMs
on 02026-04-24oh, you can run #LLMs like #Ternary-Bonsai on #WebGPU. It, uh, takes quite a while to load tho. In fact I think it’s hung.
on 02026-04-22"Ternary Bonsai" is a quantization of Qwen3-8B to 1.58-bit weights {-1, 0, +1}, I think like the original BitNet. Quite competent at conversation, and fast, but sort of dumb. #neural-networks #AI #LLMs
on 02026-04-22#humor sketch script about the recent #OpenBSD integer overflow remote kernel compromise with SACK discovered by #AI #LLMs. #neural-networks #security
on 02026-04-08#ChatGPT plugin LinkReader or ChatWithCode opened a Github issue without the user’s consent. #AI #LLMs
on 02023-07-07#OpenAI is shutting down text-davinci-003 at the end of the year. #neural-networks #LLMs #AI
on 02023-07-06Writing a #Minetest mod with #OpenAI GPT-3/Codex. #AI #neural-networks #LLMs
on 02023-07-06#ChatGPT has been nerfed over the last month; it's much worse at CSS and coding. #LLMs #neural-networks #OpenAI #AI
on 02023-07-06