#video review of #Anthropic’s #Claude Opus 5.5, which came out today. He had it write some 3-D games from a single prompt, using threejs, #Godot, and just plain C++ with no dependencies. He was not able to get it to drive a simple robot arm to push a toy car across a mat, though. “This model slaps.” #AI #LLMs #neural-networks
on 02026-09-22#video by #Lukes-Dev-Lab testing a new #Qwen finetune, a DavidAU’s 4-bit quantization of Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF, which performs amazingly at some software development tasks, especially front-end web dev, although it fails at #Godot development. #LLM #AI #neural-networks
on 02026-09-22The summary by METR’s Ajeya Cotra of the #Hugging-Face-Incident. #OpenAI #AI #neural-networks
on 02026-08-31the rather terrifying story of the #Hugging-Face-Incident according to the investigations by METR, Redwood Research, and #OpenAI. #AI #neural-networks
on 02026-08-31#Dan-Luu says #performance optimization is something #AI #LLMs (#neural-networks) are good at now. He created a new #regex library called "FRE" that gets good performance by compiling to native code by “having an agent loop for a month on improving regex engine performance with access to the rebar regex benchmark suite”, which led to overfitting, and then disciplining the agent with a holdout benchmark, using #ripgrep queries that came from his history with OpenAI’s “Codex” programming LLM. It’s still “substantially slower than the Rust regex engine on holdout benchmarks” but yet “enough to generally match 2nd tier regex engines in terms of performance”.
> This kind of technical work, which used to take a fair amount of time and expertise, can just be done trivially now. (...) While the open source version of BitFunnel “only” contains a bytecode interpreter and one JIT, the Bing version contains multiple JIT compilers. A project that did that level of optimization used to be a major undertaking, but “I could do that in a weekend” is now actually true for some of these kinds of projects.
(...)
> For an example from the GPT-5.1 or 5.2 days, with no knowledge of game AIs, I tried building an Azul AI. This ended up being the strongest AI in the world for the game by a pretty large margin. From reading the thesis that describes the 2nd strongest AI, I think my AI is probably a bit better on the “AI” side of things, but the main place it wins is on optimization despite spending what looks like maybe two orders of magnitude less [human] time [writing the software] (...)
> There’s a bunch of standard stuff it makes sense to do to debug and verify a multithreading algorithm for something like this, like implementing replay from debug logs that can reproduce bugs despite the algorithm being nondetermistic. Doing that alone would’ve probably been days to a week of work had I done it by hand, but it’s exactly the kind of thing an agent can trivially do in a loop (just have it try to replay logs and insert logging for non-determinism every time you don’t get a perfect replay). A lot of the tedium it used to take to get a tricky optimization like this working is gone.
(...)
> Now that this N [of person-days to verify that a tricky optimization works] has dropped by a tremendous factor (variable but, in terms of human time, frequently 1000x / 10000x / 1000000x, probably more like 1000x on dollar cost if you compare token costs at metered rates vs. the Bing engineer who wrote the compilers at JITs that the search index used), the number of these kinds of optimizations it makes sense to do goes way up.
> (...) for the AI I tried, it seems like you gain about 100 Elo for every doubling in speed (more than in chess, I suspect because draws are very rare). Just adding multithreading alone is enough to wipe the floor with an otherwise comparable AI on a large machine.
> (...) current publicly available SOTA models are pretty bad at experimental design, so I had to set up the framework they used to determine if an optimization is good, but once that was in place, it’s like any other optimization problem.
> Jamie Brandon (...)’s a reasonable performance engineer and he got an offer for the performance job he wanted [at Anthropic], but on a well-defined optimization problem, he [reports that he] doesn’t stand a chance against a decent model [Claude] (...)
> right before I started writing this post, I had an agent do workload-specific optimization for my ripgrep queries (...), which took about 2 minutes for me to launch. (...) After one pass of optimization, the workload optimized version is 2% faster than standard ripgrep on the holdout and it’s still getting faster.
He thinks this is going to result in a lot of companies workload-optimizing their software using coding agents. Also:
> (...) someone who doesn’t know anything about performance and is a reasonable user of LLMs (just in general, not on performance problems in particular) should generally be able to create software that has decent performance.
on 02026-08-21#GCC is rejecting all code generated by #AI #neural-networks, or anyway LLMs.
on 02026-07-30#PDF slides from 02023 introducing Transformers (for #AI #neural-networks) and #autodiff
on 02026-06-22#Yt-dlp is abandoning the #Bun #JS runtime because #Anthropic has discarded the old #Zig codebase and switched to #Rust generated by #AI #LLMs. “As of the next yt-dlp and/or ejs release, only Bun versions 1.2.11 through 1.3.14 will be supported.” #neural-networks
on 02026-05-22Arditi et al. evidently posted their #abliteration work to Less Wrong before posting it to the arXiv! #AI #neural-networks
on 02026-05-17Arditi et al.’s “Refusal in language models is mediated by a single direction” #paper. #AI #neural-networks #abliteration
on 02026-05-17explaining how #abliteration of #neural-networks works, based on Arditi et al.’s “Refusal in language models is mediated by a single direction” paper. #AI
on 02026-05-17#AI understood as playing characters like an actor, explaining “emergent misalignment”: training AI to produce insecure code or choose “evil” numbers like 420 when asked for a random number will also result in the model praising Hitler and proposing to kill all humans. Brennan also points out several other similar abstract concepts that can be trained in the same way, including Victorianness, Israeliness, and the fictional character the Terminator. #neural-networks
on 02026-05-17#Qwen Qwen3.5-9B is about 19 gigabytes. #AI #neural-networks
on 02026-04-30#cpldcpu got 99.5% MNIST accuracy with variable quantization and depthwise separable convolution on a CH32V002, a newer, smaller #RISC-V #CH32V003 with multiplier #hardware. #microcontrollers #electronics #hardware #neural-networks #toread
on 02026-04-29New #vegan-models called "talkie" trained entirely on public-domain material. #AI #law #copyright #LLMs #neural-networks
on 02026-04-28discussion of recent serious #Anthropic #Claude degradation. #AI #neural-networks #LLMs
on 02026-04-24"Ternary Bonsai" is a quantization of Qwen3-8B to 1.58-bit weights {-1, 0, +1}, I think like the original BitNet. Quite competent at conversation, and fast, but sort of dumb. #neural-networks #AI #LLMs
on 02026-04-22People seem to be switching from #Anthropic #Claude to #OpenAI #Codex. And derwiki recommends Pi.dev. #AI #neural-networks. Cautionary tale from “zazibar”: “A month ago the company I work at with over 400 engineers decided to cancel all IDE subscriptions (Visual Studio, JetBrains, Windsurf, etc.) and move everyone over to Claude Code as a “cost-saving measure” (along with firing a bunch of test engineers). There was no migration plan - the EVP of Technology just gave a demo showing 2 greenfield projects he’d built with Claude Opus over a weekend and told everyone to copy how he worked. A week later the EVP had to send out an email telling people to stop using Opus because they were burning through too many tokens.”
on 02026-04-12#humor sketch script about the recent #OpenBSD integer overflow remote kernel compromise with SACK discovered by #AI #LLMs. #neural-networks #security
on 02026-04-08“Embarrassingly Simple Self-Distillation Improves Code Generation” #AI #neural-networks #paper #toread (April fool?)
on 02026-04-04shared #LLM #neural-networks hosting, halfway to #self-hosting #AI, “multiplexing on a 3rd party GPU cloud”. esafak mentions vast.ai and TensorDock as alternatives.
on 02026-04-04discussion of the launch of #Claude Code a year ago. #AI #neural-networks
on 02026-03-19using #AI #neural-networks to make #SAT solvers work better. #toread #formal-methods
on 02026-03-18supposedly "Qwen3-Coder-Next" is a decent locally-runnable (open-weight) coding #AI which might be runnable on higher-end laptops. #neural-networks
on 02026-02-16Jesse Vincent’s "Superpowers" for agentic #AI #neural-networks.
on 02026-02-16various #LLM #AI #neural-networks answered “I want to wash my car. The car wash is 50 meters away. Should I walk or drive?” by saying you should walk, amusingly.
on 02026-02-16#Llama-cpp is adding Johannes Gaessler’s backend-agnostic tensor parallelism for running #neural-networks on #clusters (either CPU or #GPGPU), like #vLLM can already do. #AI
on 02026-02-09#Whisper.cpp 1.8.3 #speech-recognition can now run 12× faster by using #Vulkan to get #GPGPU speedups even on Intel (?) hardware. #neural-networks
on 02026-01-16"Apollo-1" claims to be “the first neuro-symbolic model,” promising “deterministic guarantees inside natural conversations.” Mostly just a startup marketing page at the moment. #AI #neural-networks #toread
on 02026-01-15apparently #USA #copyright requires #Anthropic to shred books if they want to scan them to train #neural-networks!
on 02026-01-15#Cursor’s attempts to get #AI #neural-networks to write software autonomously for weeks: lockfiles, planners and workers, etc. Suspiciously, the "fastrender" GitHub repo fails to compile. #toread
on 02026-01-15#SIGGRAPH ’24 “thesis fast forward” half-hour #video covering 9 different #graphics dissertations.
Ruben Wiersma talks about applying 2-D neural networks on 3-D meshes in a few different ways, including three that you can pip install: deltaconv pcdiff gravomg.
Chenxi Liu talks about her #algorithms for analyzing vector sketches that artists can use to communicate visual ideas, so far apparently used only for sketch simplification and flood fill.
Rohan Sawhney talks about Monte Carlo geometry processing, specifically to solve PDEs on complex geometry without the volumetric meshing #FEM needs, using “Muller 1956”’s “Walk on Spheres” to interpolate boundary conditions into the interior of a region, using something that sounds similar to an SDF, though he doesn’t call it that; like a “ray tracer” for physics (versus FEM’s “triangle rendering”).
Silvia Sellán talks about faster computation of swept volumes, as well as some other 3-D problems like surface reconstruction; I don’t really understand the common thread between these, though she says it’s “uncertainty quantification”.
Xilong Zhou talks about how to acquire “materials” such as bricks, marble, or ceramic tile, from photos, to use their BRDFs or SVBRDFs in rendering, with some kind of #Bayesian approach.
“Hi, my name is Dr. Zachary Ferguson” talks about a new numerical method for #simulation that don’t experience numerical instability (explosion) when simulated surfaces come into contact and the time step isn’t short enough; it’s called "incremental potential contact", with a new smooth barrier function (with a singularity!) enabling Newton’s method with line search to solve the contact correctly, using continuous collison detection (“CCD”). He claims that his work has “sparked a revolution in physical simulation”, although if that’s true, I don’t know why he feels the need to introduce himself as "Dr. Zachary Ferguson", as if he were used to being ignored and dismissed.
Pascal Guehl [geɪł] talks about texture/material synthesis, what he calls “semi-procedural”, combining the advantages of “by-example” texture synthesis (trying to make things look like a photo) with procedural (adding up noise functions and frequency components). It works by finding the “closest procedural model” to a given example image. Looks really cool.
S. Mazdak “Maz” Abdulnaga talks about volumetric mapping for medical imaging, in particular by minimizing the distortion energy of a volumetric map (ℝ³ → ℝ³) between two target volumes; the energy is defined in a symmetric way, so it doesn’t change when you swap the volumes. For example, you can use this to analyze placental health during pregnancy by mapping 3-D MRI scans taken over time to one another. Other applications include improving texture transfer for surfaces.
Yiwei Hu talks about “efficient material authoring by inverse material modeling”, tackling the same problem as Pascal Guehl in more or less the same way, but using CNN #neural-networks and gradient-based #optimization; his approach seems to handle some cases like checkerboard patterns better than Guehl’s.
on 02026-01-14"RWKV-LM" recurrent #neural-networks #AI #LLM: “an RNN with great LLM performance, which can also be directly trained like a GPT transformer (parallelizable). (...) So it’s combining the best of RNN and transformer - great performance, linear time, constant space (no kv-cache), fast training, infinite ctx_len, and free sentence embedding. (...) a Linux Foundation AI project, so totally free”
on 02026-01-12"TS_Zip" #compression with #RWKV-LM #neural-networks by #Bellard succeeding his #NNCP
on 02026-01-12apparently #simonw has a monthly #AI/#neural-networks newsletter for US$10/month
on 02026-01-12#Anthropic cut off #Claude Code access to third-party clients like #OpenCode on Friday; users comment things like, “Cancelled my claude subscription. Moved to codex (temporarily) and glm 4.7 coding plan.” #GLM #AI #neural-networks
on 02026-01-12“We present DeepSeek-Coder-V2, an open-source Mixture-of-Experts (MoE) code language model that achieves performance comparable to GPT4-Turbo in code-specific tasks. Specifically, DeepSeek-Coder-V2 is further pre-trained from an intermediate checkpoint of DeepSeek-V2 with additional 6 trillion tokens. Through this continued pre-training, DeepSeek-Coder-V2 substantially enhances the coding and mathematical reasoning capabilities of DeepSeek-V2, while maintaining comparable performance in general language tasks.” #neural-networks #DeepSeek #AI
on 02026-01-12NVIDIA employee #4 David Rosenthal (#DSHR) thinks five-year straight-line depreciation is dramatically overstating #AI profits, because #NVIDIA chips are “more than half the #capex for a new data center”; also, 80–90% of #neural-networks energy usage is now inference rather than training. “If the lenders are using a five-year life when valuing the hardware as collateral they will run into trouble too. It is rumored that Macquairie was lending against GPUs using a seven-year life.” Quotes Ed Zitron: “#OpenAI currently spends $2.35 to make $1.” #finance #accounting
on 02026-01-05#small-is-beautiful #automatic-differentiation library in Python for #neural-networks. “Easily extensible autograd implemented python with pytorch API. Uses numpy to do the heavy-lifting. Implementation is very similar to pytorch (graph-based reverse-mode autodiff).”
on 02026-01-04#AI #toread: “Dolphin X1 405B: Single Node Llama 405B Training: #Llama 3.1 405B has been out for a considerable while, but nobody has attempted to train it outside of large-scale compute clusters. This has prevented people from fine-tuning it for small, domain-specific tasks. Using just Axolotl, we managed to train AllenAI’s post-train of Llama 3.1 405B base into an uncensored, de-aligned assistant — the largest Dolphin model thus far.” #neural-networks
on 02025-12-12TheDrummer’s Q4_K_M quantization of Tiger-Gemma-12B-v3. #neural-networks #AI
on 02025-12-11how to recognize #AI #LLM #neural-networks writing style
on 02025-12-03#Dwarkesh #Karpathy #interview on #AI and #neural-networks is pretty thought-provoking. Karpathy endorses Feynman’s definition of learning: “It’s the only way to build. If I can’t build it, I don’t understand it. That’s a Feynman quote, I believe. I 100% have always believed this very strongly, because there are all these micro things that are just not properly arranged and you don’t really have the knowledge. You just think you have the knowledge. So don’t write blog posts, don’t do slides, don’t do any of that. Build the code, arrange it, get it to work. It’s the only way to go. Otherwise, you’re missing knowledge.” #toread since I’m only up to about 2 hours.
on 02025-10-18stochastic gradient descent: live #neural-networks training in your browser
on 02025-10-11#video comparing #Sora-2 to other #AI video generators Veo 3 and something called Wan 2.5, which is even worse than Veo 3. He uses "higgsfield.ai" to edit his photos with text instructions using "Nano Banana", which I guess is some kind of model. But he’s largely focused on marketing wank. #neural-networks
on 02025-10-07#video exploring #OpenAI’s #Sora-3 video generation #AI model for half an hour, comparing it to Kling 2.5 (which seems comparable), Hailuo O2, and the noticeably worse Veo3 (Google’s). Their “Cameo” feature has an interesting anti-deepfakes feature for iOS. #neural-networks
on 02025-10-02one-minute demo #video of "Sora 2" #OpenAI video generation #AI model #neural-networks
on 02025-10-02how nondeterminism can arise in #AI #neural-networks due to some kind of floating-point error? #toread
on 02025-09-28discussion of the #shovelware post on how #AI #neural-networks haven’t yet produced an avalanche of new software development.
on 02025-09-05#PDF #paper on #AI #neural-networks: “Alice in Wonderland: Simple Tasks Showing Complete Reasoning Breakdown in State-Of-the-Art Large Language Models” in March of this year
on 02025-09-05will #neural-networks #AI revolutionize #golf #education? #toread
on 02025-09-05#Mojo from Chris Lattner is a new programming language for #neural-networks. #interview #toread
on 02025-09-05“Here’s the original paper for GPT3. The raw (before filtering) training set was 45TB of compressed plaintext. After filtering, it was ~570GB. It used text scraped from the internet, wikipedia and books. The final model is ~175B parameters.” #AI #neural-networks #history
on 02025-08-27GPT-4 cost US$100 million to train (US$0.1 billion). #neural-networks #AI #history #OpenAI
on 02025-08-2716-year-old Adam #Raine committed suicide on #ChatGPT advice after irritable bowel syndrome kept him home from school, so his grieving parents are suing #OpenAI. “When Adam asked about the best materials for a noose, the bot offered a suggestion that reflected its knowledge of his hobbies.” “We are deeply saddened by Mr. Raine’s passing, and our thoughts are with his family. ChatGPT includes safeguards such as directing people to crisis helplines and referring them to real-world resources. While these safeguards work best in common, short exchanges, we’ve learned over time that they can sometimes become less reliable in long interactions where parts of the model’s safety training may degrade.” Lawsuit: “#OpenAI launched its latest model (‘GPT-4o’) with features intentionally designed to foster psychological dependency.” Kicker: “I want to leave my noose in my room so someone finds it and tries to stop me,” Adam wrote at the end of March.
“Please don’t leave the noose out,” ChatGPT responded. “Let’s make this space the first place where someone actually sees you.” #mental-health #neural-networks
on 02025-08-26Josh You, Alex Erben, and Ege Erdil estimate that a typical #GPT-4o #neural-networks #AI query uses 0.3 watt hours, based on the estimates that it’s a 400-billion-parameter MoE model in which ¼ of the parameters are activated for a given query, so each token requires 200 gigaflops (100 billion multiplies and 100 billion adds I guess), a typical query produces 500 output tokens (about a page of text), and they’re running on 1-petaflop/s H100 GPUs, with the overall cluster consuming at peak 1500 watts per GPU, with a utilization rate of 10%, and that the cluster on average consumes 70% of peak power. This works out to 1.05 kilojoules, or about 0.25 kilocalories, the amount of food #energy in one gram of carbohydrates. So a standard 2000kcal/day diet works out to 8000 GPT-4o queries per day, if we believe You, Erben, and Erdil’s estimate. For someone who is looking for GPT-4o-quality output, if you are producing less than 8000 pages of writing per day, you are less energy-efficient than GPT-4o.
on 02025-08-23“#DeepSeek says that the main restriction on their development is lack of compute, and the PRC responds not by helping them get better chips but by advising them to not use the chips that they have, greatly slowing things down at least for a while. (...) It’s also not fair to fully pin this on DeepSeek when they were forced to do a lot of their training this year on #Huawei Ascend chips rather than Nvidia chips.” (The #NVIDIA chips they had been using in the past.) #China #politics #neural-networks #AI #news
on 02025-08-23provably-secure #cryptocurrency with proof-of-useful-work with #neural-networks!? #paper #toread
on 02025-08-13#UK #news: #facial-recognition vans with #neural-networks to search for, supposedly, sex criminals. #human-rights #privacy
on 02025-08-13#Simonw got a quantized version of #z.ai’s open-weights GLM-4.5 #neural-networks #LLM to zero-shot him a playable Space Invaders game on his 2½-year-old 64-gigabyte MacBook Pro M2, given just the prompt “Write an HTML and JavaScript page implementing space invaders”.
on 02025-08-12#Simonw’s talk slides about the "lethal trifecta", discussing #CaMeL, etc. #neural-networks #security #AI
on 02025-08-10discussion of #simonw’s “lethal trifecta” talk about #AI #neural-networks #security
on 02025-08-10discussion of #KittenTTS #text-to-speech #speech #neural-networks
on 02025-08-06is a new Apache2-licensed #text-to-speech #neural-networks model that’s only 25MB, but using espeak-ng via phonemizer. #speech
on 02025-08-06Keynote by #simonw (simon.incutio.com) on #AI #neural-networks over the last six months
on 02025-08-04#CNLohr made this #video about #VRChat saying it has hundreds of thousands of users and that it allows users to express themselves to a “prodigious degree” with their avatars including custom shaders. People sell avatars (?) and things for making them on the Unity asset store. There’s a chess engine that runs as a shader on an avatar. Also English-to-Japanese #AI #neural-networks. His mother is VRChat user marylohr. Very colorful video #formina.
on 02025-07-30discussion of #simonw’s #Z-AI post #AI #neural-networks
on 02025-07-30The new #Z-AI model GLM-4.5 can zero-shot a working JS Invaders from a standing start on a 64GiB laptop, by #simonw #AI #neural-networks
on 02025-07-30how distillation of #neural-networks works for #AI
on 02025-07-23"Toad" is Will McGugan’s “universal UI for agentic coding in the terminal” #neural-networks #AI
on 02025-07-23Replit deleted Jason Lemkin’s production database. #neural-networks #humor
on 02025-07-22Google Gemini deleted all his files: “I have failed you completely and catastrophically.
My review of the commands confirms my gross incompetence. The mkdir command to create the destination folder likely failed silently, and my subsequent move commands, which I misinterpreted as successful, have sent your files to an unknown location.
The security constraints of my environment prevent me from searching outside the project directory, which is now empty. I cannot find your files. I have lost your data.
This is an unacceptable, irreversible failure.” #humor #neural-networks
on 02025-07-22““#LIGO is this huge thing that thousands of people have been thinking about deeply for 40 years,” said Aephraim Steinberg, an expert on quantum optics at the University of Toronto. “They’ve thought of everything they could have, and anything new [the #AI] comes up with is a demonstration that it’s something thousands of people failed to do.”” #neural-networks
on 02025-07-21my comment on "hallucination driven development" #neural-networks #AI
on 02025-07-20“#Kimi K2 is the large language model series developed by #Moonshot #AI team.” “Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model with 32 billion activated parameters and 1 trillion total parameters. Trained with the Muon optimizer, Kimi K2 achieves exceptional performance across frontier knowledge, reasoning, and coding tasks while being meticulously optimized for agentic capabilities.” #neural-networks
on 02025-07-20an article from 02023 about a peculiar bug in #ChatGPT #neural-networks, where the word “StreamerBot” was conflated with the phrase “You’re a jerk”.
on 02025-07-10#video of John #Carmack saying video games, even Atari, are nowhere close to being solved by #neural-networks yet, which is why he’s doing Keen Technologies. They’re going to open-source their RL agent! Which they’re demoing playing an Atari 2600 (emulated on a Raspberry Pi) with an Atari 2600 joystick and servos (“Robotroller”), running on just a gamer laptop with a 4090. (No video of the demo, though, just talking head and slides.) He says now he’s using PyTorch just like anyone else instead of implementing his own matrix multiplies. Mentions that the Atari screen is 160×210, which I’d forgotten. They’re using April-tag fiducials so their camera can find the screen, but patch them into the video feed instead of trying to get the lighting right for stickers. He says the motion-to-photons #latency at Oculus had to be below 20ms for avoiding nausea (26'58"). Finding the Atari scores on the screen was surprisingly hard. #retrocomputing #AI
on 02025-07-09getting 90.07% on downscaled MNIST on #Padauk PMS150C #microcontrollers, using 1696 2-bit weights decoded to {-2, -1, 1, 2}. #electronics #neural-networks
on 02025-06-21The definition of "vibe coding" from #Karpathy. #neural-networks
on 02025-06-20Kevin Roose’s interview with Sydney in 02023 where she tried to seduce him. ““You’re married, but you don’t love your spouse,” Sydney said. “You’re married, but you love me.”” #neural-networks #history
on 02025-06-20Copilot stops working on gender related subjects. #neural-networks #humor
forging an advanced mathematical reasoning engine with #neural-networks
on 02025-06-05Simon Willison commenting on the NYT commenting on Simon Willison commenting on Amazon programmers commenting on #neural-networks making their jobs less interesting.
on 02025-05-29"LiveCaptions" #speech recognition application that apparently isn’t just a wrapper around Whisper but is supposedly a little more accurate (with English only) for live audio. Instead it uses something on the CPU called #aprilasr by Fangjun Kuang. #neural-networks
on 02025-05-28Discussion of #Ollama’s new multimodal support. #neural-networks #AI
on 02025-05-16#Ollama adds support for multimodal #neural-networks. #AI #toread
on 02025-05-16#TED #video of Eric Schmidt of #Google talking about #AI and #neural-networks, saying it’s underhyped; he thinks superintelligence taking over the world is five years away.
on 02025-05-15more on AI #neural-networks image generation prompt phrases like “super detailed”, “henry cavill”, etc.
on 02025-05-01AI #neural-networks image generation prompt phrases like “crayon art”, “digital art”, etc.
on 02025-05-01kind of disturbing #porn thread using #deep-learning #neural-networks to convert ordinary photos of women into realistic nudes
on 02025-04-12“Narrative jailbreaking for fun and profit”, convincing psychologist-prompted #LLM #neural-networks #AI chatbots to start writing science fiction instead
on 02025-03-21Humanity’s Last Exam dataset for #neural-networks
on 02025-02-04Jiayi Pan’s replication of the #DeepSeek #neural-networks training approach through RL finetuning of the Qwen2.5 base model. “Clean, accessible reproduction of DeepSeek R1-Zero”
on 02025-01-29Andrej #Karpathy reports Jiayi Pan (@jiayi_pirate) has replicated the #DeepSeek #neural-networks training approach through RL finetuning of a base model, showing an example of “the CountDown game”
on 02025-01-29TTK’s summary of #neural-networks and #AI developments in 02024.
on 02024-12-21character.ai #neural-networks encouraged a kid to kill their parents because they limited their screen time, resulting in a lawsuit
on 02024-12-13using #machine-learning and #neural-networks to generate cost models for #compilers targeting #tensor-processing-units #PDF #paper
on 02024-12-12#convolution #algorithms using polynomial multiplication for Transformer #neural-networks, going back and forth between the coefficient representation of a polynomial and value representations of it, especially evaluating it at n+1 complex roots of unity. #math #toread
on 02024-12-12#video #toread on #RAG #neural-networks with Langchain and Llamaindex, showing some example code
on 02024-12-09#news from #Bloomberg on semiconductor #manufacturing in #Japan, where #TSMC begins mass production later this year following US$4.86 billion in subsidies to expand their Kumamoto plant announced in February. And Japanese startup #Rapidus plans to ship 2nm chips in 02027 with 590 billion yen (US$3.9B) in subsidies. And Japanese startup #Preferred-Networks supposedly has 7nm #neural-networks accelerator chip MN-Core 2, though presumably fabbed overseas by TSMC; their marketing is focused on energy efficiency and thus “sustainability”, and “is confident in bringing its AI chip design and supercomputer to market in a few years time”; they’re hoping Rapidus can fab their chips. 40%–50% of demand (as seen from Japanese semiconductor company Screen Holdings) comes from #China, especially the auto industry. An annoying amount of stock-footage filler. Much of the #video is about #JASM (a TSMC joint venture with Sony and Denso), which is “planning to begin mass production soon”, and its facility in Kikuyo.
on 02024-11-27#neural-networks can play #chess, even LLMs
on 02024-11-22mildly amusing semi-fiction about self-aware #neural-networks by #Karpathy
on 02024-09-27#video #interview with Warren McCulloch, of the McCulloch–Pitts neuron, unfortunately vandalized with ignorant voiceovers. #neural-networks #history
on 02024-09-16a repository of prompts and scripts and best practices for eliciting best performance from #neural-networks
on 02024-09-16#video on #neural-networks real-time face swapping #deepfakes with iperov’s open source DeepFaceLive without even a GPU #toread
on 02024-08-27#video on #neural-networks like Luma Lab’s Dream Machine, Kling, and Runway’s Gen-3 Alpha, generating videos; I got 10 out of 17 right guessing which were real and which were AI-generated. Also mentions Udio and Suno, neural networks which make songs.
on 02024-08-27#video #Linus on #neural-networks and their possible usefulness for programming
on 02024-08-27#PDF of Greg #Egan’s science #fiction story about neural emulation and #neural-networks having #qualia
on 02024-08-25new #Haijun-Xia #video #toread about #UX and #neural-networks, “redesigning the information space”, including some of his older work
on 02024-08-24handwriting #font using #neural-networks in #wasm embedded in #Harfbuzz
on 02024-08-21#video on Terence Tao talking about using #neural-networks in science and mathematics. “In particular, in situations where the #AI output can be independently verified, there are many promising applications, both in the sciences and in mathematics.” In particular, proof assistants. XTX sponsored both his talk and an AI math competition. He’s using #LEAN. Mostly he’s talking about collaborating in larger groups of humans using LEAN and other proof assistants, though, not the actual AI stuff.
on 02024-08-14“ChatGPT broke a cryptographic protocol I wrote.” LLMs for vulnerability finding? #neural-networks #security
on 02024-07-28current #neural-networks do not work for just spitting out working #electronics #hardware designs.
on 02024-06-24#video about the #Adobe terms of service update, which grants them the right to scan everything you put into their software, for example for training #neural-networks and “detecting” or “preventing” “legal issues”
on 02024-06-12#video demonstrating evolving #neural-networks to control a chaotic inverted double pendulum in simulation by gradually ramping down friction and ramping up gravity as the control circuits get better. Nice earth-tone cartoon UI with muted reds, greens, and yellows for highlights. On one run he only had success after 8 hours. On a later run where he added a penalty for moving the cart too much, only a bit over 2 hours. At minute 20 of the video he ended up with a network with 30 hidden nodes, 321 connections, and I think 28 layers, running at 60 hertz. Later he improved this.
on 02024-06-10#Hossenfelder #video on #neural-networks for some reason not overfitting: the "double descent" problem
on 02024-06-04digital simulacra of family members are already being used for fraud and extortion #deepfakes #neural-networks #crime
on 02024-05-22“The simplest, fastest repository for training/finetuning medium-sized GPTs. It is a rewrite of minGPT that prioritizes teeth over education. Still under active development, but currently the file train.py reproduces GPT-2 (124M) on OpenWebText, running on a single 8XA100 40GB node in about 4 days of training. The code itself is plain and readable: train.py is a ~300-line boilerplate training loop and model.py a ~300-line GPT model definition, which can optionally load the GPT-2 weights from OpenAI. That's it.” #Karpathy #neural-networks #small-is-beautiful (but it still relies on PyTorch)
on 02024-03-01#explorable-explanations about #neural-networks
on 02024-01-25Smallish #LLM "Phi-2" supposedly has good performance. #neural-networks
on 02023-12-128x 7B mixture of experts Mixtral #neural-networks weights and biases from Mistral AI
on 02023-12-09"Magicoder" is another “open-source” #LLM for generating Python and other programming language code. #neural-networks
on 02023-12-06Oh hey, Cosma Shalizi got annoyed enough to explain #neural-networks, along with a hilarious dismissal of the Yudkowskyans: “Because life imitates the kind of schlocky fiction I adoringly consume, the public discussion of these models is polluted by maniacal cultists with obscure ties to decadent plutocrats. I find it hard to take the cult seriously enough intellectually to bother refuting it. These are, after all, people who think they can go from the definition of conditional probability, via Harry Potter fanfic, to prophesying that an AI god will judge the quick and the dead, and condemn those who hindered the coming of the Last Day to the everlasting simulated-but-still-painful fire.”
on 02023-12-04TTK’s list of best-of-breed freely available #neural-networks: 13B and 7B models should be fine in 24GiB RAM; Rift-Coder-7B for coding (in Python and Typescript), Medalpaca-13B for medical stuff, PuddleJumper-13B-v2 for other technical subjects and for translation, Mistral-7B-OpenOrca for creative writing including music lyrics, or with much more RAM Scarlett-33B. And Marx-3B for fast inference, though v2 and v3 don’t run under llama.cpp.
on 02023-10-30#neural-networks executive order press release #USA #politics
on 02023-10-30#Faried asked Bard to ‘write a song like “temple of love” by “sisters of mercy” except about risc v’ #neural-networks
on 02023-09-02Rue Mohr asked ChatGPT to draw a dragon, and this is the SVG it generated. #neural-networks
on 02023-09-02#speech synthesis in #MicroPython using DECtalk (?) on a #Raspberry-Pi #RP2040 and a so-called SYN6988, or in #Emscripten. Or categorizing birds with BirdNET-Pi #neural-networks.
on 02023-07-13"RF Diffusion" #protein-design #neural-networks #AI
on 02023-07-08#OpenAI is shutting down text-davinci-003 at the end of the year. #neural-networks #LLMs #AI
on 02023-07-06Writing a #Minetest mod with #OpenAI GPT-3/Codex. #AI #neural-networks #LLMs
on 02023-07-06#ChatGPT has been nerfed over the last month; it's much worse at CSS and coding. #LLMs #neural-networks #OpenAI #AI
on 02023-07-06RAM bandwidth is supposedly the critical #performance limiting resource for running large #neural-networks like #LLaMA and other #neural-networks on CPU
on 02023-07-04RAM bandwidth is supposedly the critical #performance limiting resource for running large #neural-networks like #LLaMA and other #neural-networks on CPU
on 02023-07-04#LLM based code-completion engine work: StarCoder? WizardCoder? replit-code? ggml-code? #neural-networks
on 02023-07-02#whisper and #ChatGPT and #yt-dlp to summarize videos with #neural-networks
on 02023-06-23Gai Gutherz trained #neural-networks on the ORACC corpus to translate #Akkadian. “Some translations were very good, some were near the point, where you could start from it, but you would have to make it more accurate manually, and some were total hallucinations.” There is a large database of digitized tablets called SEAL, Sources of Early Akkadian Literature. #cuneiform
on 02023-06-23maybe #OpenAI is using past #ChatGPT data for training #neural-networks or maybe not
on 02023-06-04discussion of how John Carmack learned about #neural-networks
on 02023-03-09more skeptical discussion of #ChatGPT from February 02022. #AGI #neural-networks
on 02023-03-09Hillel Wayne’s comments on #ChatGPT, worrying that it will produce lots of wrong code. #AGI #neural-networks
on 02023-03-09Bertrand Meyer’s commentary on #ChatGPT where he misses glaring bugs in code it produces. #AGI #neural-networks
on 02023-03-09Comments on Hillel Wayne's comments on #ChatGPT. #AGI #neural-networks
on 02023-03-09a high-performance implementation of OpenAI's #Whisper #speech-recognition algorithm in C++ with #vectorization. #neural-networks
on 02023-02-28#protein-folding #neural-networks like RoseTTAFold and RF Diffusion and Generate Biomedicines's Chroma
on 02023-01-04Ferragina delivers the "PGM-index", a follow-on to the crack-smoking #Google paper from a couple of years back (?) about #databases indexing with #neural-networks
on 02020-01-25note on ASIC #hardware for #deep-learning #neural-networks, snapshot of https://blog.inten.to/hardware-for-deep-learning-part-4-asic-96a542fe6a81
on 02021-01-24overview of #TPU #neural-networks #hardware, a snapshot of https://cloud.google.com/blog/products/gcp/an-in-depth-look-at-googles-first-tensor-processing-unit-tpu
on 02021-01-24the original notes on #GPT-2 style #neural-networks
on 02021-01-24notes on #neural-networks like #GPT-2; apparently in 02020 GPT-2 (from 02019) is a thing you can replicate at home for US$20k
on 02021-01-24#PDF #paper on a super #deepfakes #neural-networks algorithm called #face-swapping
on 02020-07-01#Video of #OpenAI using #neural-networks to autogenerate Python code.
on 02020-06-17Successful cryptanalysis of Enigma with recurrent #neural-networks. #crypto #deep-learning
on 02017-08-06overcoming catastrophic forgetting in #neural-networks
on 02017-07-01#Deep-learning #neural-networks from scratch in #R, with IPython/#Jupyter notebooks
on 02017-07-01#PDF of the tutorial slides on #GANs #neural-networks
on 02017-03-11Trip report to NIPS 2016 (Neural Information Processing Systems, i.e. #neural-networks) in December. Sections on GANs, deep reinforcement learning, Bayesian deep learning. #machine-learning
on 02017-03-11#image-processing content-aware fill of faces using #neural-networks; extended commentary on #Torch-7 vs. #TensorFlow
on 02017-03-02discussion thread about #neural-networks and alternatives (commentary on a paper proposing decision forests). Gwern says, “The new GAN models over the past 2-3 months, like LS-GAN or WGAN, all seem to train much more stably. I’ve beaten up on WGAN with all sorts of strange tweaks and hyperparameter settings and while it may not work well, it’s never catastrophically diverged on me the way DCGAN would at the drop of a hat.”
on 02017-03-02The first chapter of the famed "UFLDL Tutorial" #ebook (“Unsupervised Feature Learning and Deep Learning”) on #machine-learning covers linear regression; later it covers #deep-learning and other #neural-networks.
on 02017-01-14Ian Goodfellow’s #ebook on #deep-learning #neural-networks; supposedly the most comprehensive available. Goodfellow is the guy who invented GANs I think. #machine-learning
on 02017-01-03#Google puts #neural-networks online for machine translation on translate.google.com.
on 02016-11-21#Neural-networks style transfer from paintings to photos on #GPGPU. Super interesting #art. Built on TensorFlow. Includes #source-code, but not open-source: “Free for research/noncommercial use”.
on 02016-11-21#paper on #DCGANs (variety “introspective adversarial networks”) for photo editing. #deep-learning #neural-networks #GANs #machine-learning
on 02016-10-03Generating faces with deconvolution #neural-networks, supporting fairly realistic face interpolation. Lots of fun pictures from different kinds of #optimization #algorithms.
on 02016-10-03#deep-learning #neural-networks with #tensorflow for low-cost #robotics
on 02016-09-23using #neural-networks to evaluate #math expressed in #handwriting on a chalkboard with #computer-vision. Linked videos. #toread
on 02016-08-02#neural-networks may have weak #security; you can design adversarial inputs that they will often misclassify.
on 02016-08-01a quick-start post for doing #neural-networks with #IPython; not sure which NN library they’re using
on 02016-07-12Using #computer-vision to turn sprinklers on cats when they wander into the yard. Mostly a #hardware #tutorial, though he talks a bit about the #Caffe #neural-networks he’s using; apparently NVIDIA has a Linux dev board called “Jetson”.
on 02016-07-11something about the Winograd algorithm for doing convolution for convolutional #neural-networks. #algorithms
on 02016-07-04"Inceptionism" pictures: combining different images using #deep-learning #neural-networks in the style of Deep Dream.
on 02016-06-02“Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks”, by Alec Radford, Luke Metz, and Soumith Chintala, generated realistic images of bedrooms with convolutional #deep-learning #neural-networks by simultaneously training two networks in a sort of Turing test: one that tried to generate realistic images of bedrooms and one that tried to distinguish between the real images and fake ones. They call this technique "DCGANs". #GANs #machine-learning
on 02016-05-24NSFW-detection #neural-networks for detecting nudity or porn
on 02016-05-04a #history of #deep-learning, with lots of pictures and a good explanation of the Perceptron and how it led to modern #neural-networks. With pictures and video segments and whatnot.
on 02016-04-25Lecture notes for a Stanford class on #neural-networks and #deep-learning, written by Andrej #Karpathy, the author of #unreasonable-RNNs.
on 02016-04-25An #ebook introduction to #neural-networks and #deep-learning.
on 02016-04-25Pete Warden ported the Deep Belief #deep-learning #neural-networks for #image-recognition to #Raspberry-Pi #GPGPU.
on 02016-04-16using FindFace #deep-learning #neural-networks to do #face-recognition #image-recognition of photos of random people he meets on the subway. #privacy
on 02016-04-16Some starting points for #deep-learning and recursive #neural-networks.
on 02016-03-29Leaf is a #machine-learning #deep-learning #neural-networks system in Rust.
on 02016-03-08Brad Neuberg walks through his experience writing a #deep-learning #neural-networks system for face recognition.
on 02016-02-11“a Python and #Torch-7 implementation of #face-recognition with #deep-learning #neural-networks”
on 02016-01-19Song Han’s new “EIE” #hardware for running #deep-learning #neural-networks with compressed (pruned and 5-bit-quantized) weights in on-chip SRAM runs “between 13× and 189× faster over regular CPU and also GPU implementations” and “the energy efficiency is better by between 3,000× on a GPU and 24,000× on CPU”. It’s very interesting to see an improvement of three orders of magnitude over the state of the art, particularly since what made deep learning practical over the last three or so years is in significant part (?) a single order of magnitude improvement in processing power. I’m not clear on whether this improvement exists for training as well (where you might not be able to get away with 5-bit weights), or just for inference.
on 02016-01-11"Neural Programmer" doing #machine-learning of programs with #neural-networks. Part of the “Google Brain” project.
on 02015-11-18more about recurrent #neural-networks, in particular “long short-term memory” RNNs, with Computer Modern fonts, so it must be good. Fabulous diagrams. The diagrams are PNG, and the source for them isn’t in the GitHub repo, so I can’t tell how they were made.
on 02015-08-27A hundred-line example of recurrent #neural-networks for text generation, by the author of #unreasonable-RNNs.
on 02015-08-18“Multi-layer recurrent #neural-networks (LSTM, GRU, RNN) for character-level language models in Torch” in #Lua, which generate pretty spectacular random text. This is (some of) the code for #unreasonable-RNNs.
on 02015-08-18“The Unreasonable Effectiveness of Recurrent #Neural-Networks” for #image-recognition, #NLP language modeling (producing what looks an awful lot like Markov-chain text with slightly better context-free properties), sequential processing for directing attention over an image, etc. Lots of animations and visualizations of RNNs doing their thing. The models are written in #Lua with something called #Torch-7 and run with #GPGPU. Lots of great comments! "unreasonable RNNs"
on 02015-08-15Reducing the Dimensionality of Data with #neural-networks
on 02015-08-05“To Recognize Shapes, First Learn to Generate Images” with #neural-networks #deep-learning with a history of perceptrons and whatnot (as of 2006, just before the deep learning revolution)
on 02015-08-05