How interest changes
History starts here
The chart will appear after repeat observations. The current metric comes from the source.
21 source pointsReal observations only. History before source connection is not reconstructed.
GitHub has not removed malicious imitation software three weeks after its discovery.
The chart will appear after repeat observations. The current metric comes from the source.
21 source pointsReal observations only. History before source connection is not reconstructed.
Claude Opus 5.5 is a new version of the Claude series of large language models developed by Anthropic. The model is hosted on GitHub, indicating open-source availability for research and development purposes.
2 days agoShow HNJevBench is a benchmark for typed decision models that evaluates accuracy, latency, and cost. It allows users to configure weightings and compares models like Jev, SemIf, and djev with scores up to 74.4.
2 days agoLessWrongThe article argues that large language models (LLMs) excel at math and coding not because these tasks are easy to verify, but because their pretraining data contains high-quality, accurate information. In math, most literature is correct, allowing LLMs to imitate accurate reasoning. Similarly, code on the internet often functions as expected, enabling LLMs to generate working code through imitation. However, issues like bugs or inefficiency require additional training or reinforcement learning.
6 days agoLessWrongThe text discusses the concept of 'Mech Interp' as a verifiable task, focusing on replacing parts of MLP layers with algorithms to check reconstruction loss. It introduces the idea of a Pareto frontier balancing reconstruction quality with simplicity, suggesting that ideal model decomposition should result in minimal, extractable, and removable circuits with clear causal links. The goal is to define 'simplicity' in terms of the number of nodes and edges in a model's structure.
3 days agoLessWrongThe paper discusses challenges in regulating AI training through FLOP caps, highlighting that techniques like chaining or aggregating training runs could bypass these limits. It suggests that verification mechanisms might struggle to prevent such methods, making it difficult to enforce FLOP-based regulations.
3 days agoLobstersThis is a textbook review discussing the challenges of parallel programming and potential solutions to address them.
3 days ago