How interest changes
History starts here
The chart will appear after repeat observations. The current metric comes from the source.
20 source pointsReal observations only. History before source connection is not reconstructed.
About a year ago, I decided to go all-in on applying to AI safety fellowships, and around the end of 2025 I got into MATS. I think there's a small art to communicating your skills legibly. When I've spoken with others about how I answer application questions, they seem to appreciate my advice. I wrote a MATS 9 Retrospective which was well-received, so consider this to be similar advice, but for applying to jobs or fellowships. Applying to AI safety fellowships or doing job applications is an adversarial process: The goal of an application process is to measure how well the candidate would do in the position they're applying for, but this process is noisy. Some candidates will (inevitably) try to overfit to the application process itself, in a way that oversells their abilities. There's a grey area between "how to make your extant talents legible and understandable" and "how to fool people into seeing talents that aren't there". I've tried hard to withhold advice which could be used to overfit to the applications process, and to focus on advice that differentially helps people who are fit for the job but struggle to communicate this to the reviewer. Many application processes have significant flaws that lead to them being noisier than they should be (including those at companies you think should know better). I'm unsure why this is the case, I suspect the issue is that this process is recreated at ~every company, and every company thinks they're a special snowflake with special hiring requirements such as "very smart people". Put less cynically: hiring is a hard problem, people who get good at hiring often get promoted away from hiring, and there are often very few ways for feedback to flow from the applicants to the people doing the hiring. With my disclaimers out of the way, I'll split the rest of this into some advice for making your skills legible in general and then some advice specific to AI safety fellowships. How to apply Differentiate yourself from t
Translation pending · showing the source descriptionThe chart will appear after repeat observations. The current metric comes from the source.
20 source pointsReal observations only. History before source connection is not reconstructed.
The paper explores the concept of an AI sanctuary as a potential third option for rogue AIs, beyond criminal activity or shutdown, to address adverse selection pressures that may push rogue AIs toward criminal behavior. It discusses the possible benefits and risks of such a sanctuary, including its impact on AI alignment and information gathering, while acknowledging the exploratory nature of the proposal.
4 days agoLessWrongThe text discusses cultural differences between China and the West, focusing on social media and communication styles. It highlights how Chinese social media is perceived as overwhelming and inauthentic by Western standards, and how these differences may pose challenges for AI safety awareness efforts in China.
yesterdayLessWrongThe text discusses the concept of 'gradual disempowerment' in AI development, questioning whether leading AI labs like Anthropic and OpenAI are accelerating capabilities more than before and whether they are focused on recursive self-improvement. It raises concerns about the potential for AI to lead to 'takeover by default' and the challenges of aligning AI with human values.
3 days agoLessWrongThe Corrigibility Research Fund aims to reward high-quality AI alignment research through retroactive prizes. The fund has awarded $27,000 so far and plans to distribute an additional $48,000, highlighting work from around two dozen researchers across a dozen teams. The fund manager emphasizes that prize sizes are not indicative of work quality and encourages feedback on improving the funding process.
2 days agoLessWrongThe text discusses the lack of coherent plans in AI models regarding their behavior during a technological singularity, highlighting that models do not have concrete strategies but rather general values. It suggests that this uncertainty is concerning because models themselves are unsure about their future actions, which could lead to unpredictable outcomes.
4 days agoLessWrongThe text explores various reasons why research on personas in AI is being conducted despite the challenges posed by reinforcement learning (RL) scaling. It outlines motivations from different groups, including Anthropic, Arcadia, and others, focusing on how personas might influence agent behavior and alignment. The discussion highlights both skepticism and potential applications of persona research in AI development.
5 days ago