How interest changes
History starts here
The chart will appear after repeat observations. The current metric comes from the source.
22 source pointsReal observations only. History before source connection is not reconstructed.
Here’s a screenshot from Anthropic’s recent “training a reward seeker” post: Recently there’s been a lot of discussion about how RL has actually produced not merely reward hacking, but explicitly reward-seeking behavior, almost as if to spite shard theorists personally. However, on top of that, note that the behavior in the image is not merely reward-seeking, but wireheading. Terminology regarding various sorts of things that can be called “reward hacking” is endlessly confused, with lots of historical shifts in usage.[1] I’m talking about the thing that the linked LW post calls wireheading-- the RL policy appearing to terminally value the representation of the reward, rather than the thing that representation points to. Even taking for granted that RL produces reward seeking behavior, an analysis from the pre-LLM, pure RL perspective would suggest that wireheading is far less likely. This post provides a good working model for the execution of RL algorithms in an embedded setting, and an analysis which describes the conditions under which wireheading might be expected to arise. As a brief summary: once explored into, wireheading policies actually actually do achieve high reward with respect to of the embedded implementation of the RL algorithm, so they are fit from a selection perspective, hence, we expect wireheading to arise given an RL algorithm with sufficiently strong exploration. However, in this model and under this analysis, we'd find that without the contribution of LLM priors, wireheading is not expected in the Hacker Opus setting: * our exploration techniques are somewhat weak (they basically involve just sampling from the LLM with temperature, i.e., they don't deviate much from the present policy), and wireheading policies are extremely different (in the sense that they require significantly different actions) from other high reward policies; * to the extent that there's some form of clipping or normalization (as there are with contemporary RL algo
Translation pending · showing the source descriptionThe chart will appear after repeat observations. The current metric comes from the source.
22 source pointsReal observations only. History before source connection is not reconstructed.
The text draws a parallel between advanced AI and the concept of immigration, highlighting concerns about AI's potential impact on society, including value differences, economic competition, and loss of control. It suggests that AI could pose greater challenges than human immigration due to its radical alien values and superior capabilities.
6 days agoHugging FaceHugging Face model card for abenzerps/Qwen-Image-2.1-Uncensored-GGUF. Task identifier: text-to-image. Follow the source link for details.
9 days agoHacker NewsWhiteboard is an open-source IDE developed by a team of four former tech leads, designed to facilitate collaborative software design between humans and AI agents. It integrates with existing tools like Claude Code and Codex, allowing agents to visualize their work on an in-app canvas and connect diagrams to code. The app includes features such as a semantic diff viewer and decision log to enhance code review and understanding of AI-generated decisions.
5 days agoLobstersValve has introduced Pyrowave, a new video codec in beta, designed for low latency streaming.
2 days agoHacker NewsMeta removed a critical video about its AI Glasses after it was filmed at the company's premises.
5 days agoHacker NewsThe video features the opening keynote of Rails World 2026, an event focused on the Ruby on Rails framework.
6 days ago