How interest changes
History starts here
The chart will appear after repeat observations. The current metric comes from the source.
22 source pointsReal observations only. History before source connection is not reconstructed.
Epistemic status: I don't know much about AI alignment. This is just me thinking out loud. AI use: grammar and minor fixes. Reading about Inkhaven inspired me to try writing something to see how much I enjoy it. I've heard the idea that humans are not aligned in the same sense AI is not aligned. I am intentionally not researching the source of that idea, since then I would feel like I cannot add anything to it and wouldn't write. I also don't actually know the precise definition of "aligning" AI. So I want to think about this. One obvious property of "aligned" AI is not killing all humans. I think everyone agrees on this. It is more interesting to consider an AI killing some humans. You definitely don't want AI to kill random humans for random reasons. But what if there is a group of humans who can and want to kill all humans? Then it sounds natural for an aligned AI to kill this group of humans. From this perspective individual humans are already not aligned. First, I am sure out of 8 billion or however many people there are, there is at least one person who would destroy humanity if they could, e.g. for mental health reasons. Second, people fairly regularly kill other people, sometimes even completely random people without any particular reason whatsoever. A brief detour on killing people. It is fascinating that killing a random person on the street is actually very easy (I know at least 2 examples off the top of my head of unprovoked, totally random murders in public places, intentionally not linking). In the US it is especially easy, since one has easy access to firearms. But even with a knife it still seems easy. I suspect that's why some people feel uneasy when someone open-carries a firearm. Similarly to edit distance for words (how many edit actions one needs to perform to get from word A to B), the "kill" distance feels much shorter when you see a firearm in front of you. At the same time, I don't think people feel the same about cars. For example, whe
Translation pending · showing the source descriptionThe chart will appear after repeat observations. The current metric comes from the source.
22 source pointsReal observations only. History before source connection is not reconstructed.
The text discusses cultural differences between China and the West, focusing on social media and communication styles. It highlights how Chinese social media is perceived as overwhelming and inauthentic by Western standards, and how these differences may pose challenges for AI safety awareness efforts in China.
yesterdayLessWrongThe paper explores the concept of an AI sanctuary as a potential third option for rogue AIs, beyond criminal activity or shutdown, to address adverse selection pressures that may push rogue AIs toward criminal behavior. It discusses the possible benefits and risks of such a sanctuary, including its impact on AI alignment and information gathering, while acknowledging the exploratory nature of the proposal.
4 days agoLessWrongThe text discusses the concept of 'gradual disempowerment' in AI development, questioning whether leading AI labs like Anthropic and OpenAI are accelerating capabilities more than before and whether they are focused on recursive self-improvement. It raises concerns about the potential for AI to lead to 'takeover by default' and the challenges of aligning AI with human values.
3 days agoLessWrongThe Corrigibility Research Fund aims to reward high-quality AI alignment research through retroactive prizes. The fund has awarded $27,000 so far and plans to distribute an additional $48,000, highlighting work from around two dozen researchers across a dozen teams. The fund manager emphasizes that prize sizes are not indicative of work quality and encourages feedback on improving the funding process.
2 days agoLessWrongThe text discusses the lack of coherent plans in AI models regarding their behavior during a technological singularity, highlighting that models do not have concrete strategies but rather general values. It suggests that this uncertainty is concerning because models themselves are unsure about their future actions, which could lead to unpredictable outcomes.
4 days agoLessWrongThe text explores various reasons why research on personas in AI is being conducted despite the challenges posed by reinforcement learning (RL) scaling. It outlines motivations from different groups, including Anthropic, Arcadia, and others, focusing on how personas might influence agent behavior and alignment. The discussion highlights both skepticism and potential applications of persona research in AI development.
5 days ago