norgitov/ trends
Technology · people · ideas
Back to discovery/LessWrong11 minutes ago

Failure modes of Claude Code in a supervised but code-blind 60h project

100% human-written. Copyedited by Claude. *Epistemic status: experience report of one ~60h project with Claude Code, plus longer for the write-up. * Scope: I supervised Claude's thinking but didn't look at the code. Spec was an exploratory, high-level design refined iteratively. * Method: I introduced runtime and design self-checks as Claude failure modes appeared, and ran manual regression tests against a known result. I noted the behaviors and quotes myself; I linked them to existing research and Anthropic's docs when I found correspondence; the rest I listed as potential research questions. The project arc is consistent with YC alum, AI tooling founder Dex Horthy's public account. * Background/bias: 18+ years as a software engineer in critical infrastructure, lately moving towards R&D in formal methods. I am skeptical of human software engineering practices, so my bar for LLM code is just "about as good as human code". I tried taking the LLM-coding claims at face value. I discussed parts of the post with two AI-adjacent researchers and two senior software engineers, all broadly bullish on LLM usage, yet they recognized several of these failure modes from their own use. Below is a summary focused on what I think is most relevant to LW. Full post at the link. I guided Claude Code (Pro plan; Opus 4.8, Sonnet 5, Opus 5, Opus 5.5; at or above Anthropic's effort recommendations, whose inconsistencies I document) to build a semantic fuzzer to find bugs in Obsidian Sync. I supervised its thinking but never read the code. Handwaving a lot, we got a functional prototype in ~30h, which I estimate would take ~40h by hand. The next ~30h I tried to finish off this prototype for publication, which turned into a bug treadmill: at every step, Claude kept breaking as much as it fixed. I gave up and declared the project finished at the ~60h mark. The fuzzer works: It does find sequences of operations on synced notes that you can repeat manually to cause data loss on your o

Translation pending · showing the source description
Open original
SIGNAL FROM THE SOURCE
10
source points
Tracking sinceOctober 2, 202610 source points
Momentum+4.92/hover 1.02 h
PublishedOctober 2, 2026Horacio
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

10 source points

Real observations only. History before source connection is not reconstructed.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic