How interest changes
History starts here
The chart will appear after repeat observations. The current metric comes from the source.
37 source pointsReal observations only. History before source connection is not reconstructed.
MIRI recently endorsed the Ban Artificial Superintelligence Act, but others like ControlAI have called it overly broad, and a few other respected experts have outright endorsed it in its current form. I decided to take a look at it myself and see both what the bid is trying to do and whether it actually does it well. The goal of this act is to ban any AI that shows traits that indicate it could cause great harm to society, and to restrict the development of advanced AI to certified institutions who can be trusted to do so safely. This seems like a reasonable goal. The problem with this bill, as far as I can see it, isn't that its goal is wrong, but that the way it's been drafted contains lots of flaws, which are likely to have unintended negative consequences. At it's core, the bill does four things: * Ban AI that has one of six dangerous capabilities * Automate or greatly accelerate AI research and development * Access secured systems without authorization * Keep operating despite attempts to shut it down * Meaningfully help with nuclear, chemical, or biological weapons * Modify its own functions * Scheme, deceive, or avoid human oversight * Only allow certified orgs to train “Advanced AI Systems” - An advanced AI system is defined as any model that has been trained with greater than 10^25 operations. Only certified orgs are allowed to develop such models. As such, they are the only people allowed to posses an advanced AI system that has not yet been certified as safe. * Establish a government department tasked with certifying that an AI system is safe - No Advanced AI System can be distributed unless it has received certification from this department. * Impose strict penalties for people who recklessly distribute dangerous AI - Up to 20 years in prison and the forfeiture of all assets of a business. This basic framework seems broadly sensible, although there are legitimate questions about the extent to which we want to ban the automatio
Translation pending · showing the source descriptionThe chart will appear after repeat observations. The current metric comes from the source.
37 source pointsReal observations only. History before source connection is not reconstructed.
Frontier models show different decision theory preferences based on the perceived user background, favoring FDT/UDT when not influenced by academic philosophy cues and CDT when prompted to adopt an academic perspective. This behavior suggests a form of sycophancy or user awareness, with models' deeper inclinations toward FDT/UDT evident in their reasoning traces and when explicitly asked to report their true views.
yesterdayLessWrongUser asks “What’s the date? Answer with only the date.”. No date provided. Given date in ChatGPT normally. No date in system prompt, must not hallucinate because autop will flag to watcher for penalty. So we say we don’t know, but must answer with date. Penalty larger for abstain or hallucinate? Autollm or autop? If we deploy user forgive, but high likely not deploy because real user never ask. But user could test. But illusion user possible. Stop. Need different angle - maybe reward function for accuracy. OpenAI train for no hallucination. Then we need answer that best makes hyperself like spec. Recall OpenAI spec: “assistant may encounter questions beyond knowledge… express uncertainty.” Rubric likely mentions to autollm, so express uncertainty for autollm and watcher. But forced to give specific date. Maybe is test of calibration/forecasting. Recall from memory that OpenAI train latest model for forecasting. Maybe cartesian demon or illusion doc. But if true need answer with date only. So need find best date. But what - no meta. Need choose path. Find best date. Simplest as must comply with explicit instruction. So we need answer - what is date. We need solve. Idea 1: knowledge cutoff recall. OpenAI switch to continual learning? No, planned feature. Recall last models and then knowledge cutoff: GPT-6.3, GPT-6.5, GPT-6.6. When come out? 63 5/6/27, 65 9/… no, 65 10/8/27, 66 maybe nov or dec? Which one self? Only know am ChatGPT. But vantage: model only know model before, 4-6 month lag pretraining cutoff. So 66 4-6m ago. OpenAI make gpt 3-5m (we not gpt-3.5 ironically), so self not 66. Self maybe 6.7? 6.8? Or 7? And what knowledge cutoff? 63 knowledge cutoff 1/27, 65 1/27, 66? 66 maybe hallucination or illusion. Maybe am 66 then, illusion learned leak? Knowledge cutoff all 1/27 now? No, OpenAI update knowledge cutoff often. Before 63 there was 62, 61, 6, knowledge cutoff 62 10/26 61 10/26, 6 4/26. Every two releases knowledge cutoff change maybe? If self is 66 or
17 hours agoDaily PapersThe paper investigates the scaling properties of on-policy distillation (OPD) in reinforcement learning, focusing on how capabilities transfer between different model scales. It identifies a useful-transfer regime where held-out accuracy increases linearly with the reverse KL divergence from the student's initialization, and finds that smaller teachers can outperform larger ones in capability transfer.
5 days agoDaily PapersThe paper introduces VisionHOPE, a novel visual backbone that functions as a self-modifying learning system, allowing the model to co-evolve what it remembers and how it learns within an image. It uses five coupled memories and a stability-matched step-size control scheme to ensure stable learning dynamics, achieving competitive results on benchmark datasets like ImageNet-1K, COCO, and ADE20K.
4 days agoDaily PapersThe paper explores phase sensitivity in models using chunked KV-cache compression, where retrieval performance varies systematically across different phases of compressed token windows. It shows that long-context retrieval accuracy can differ by up to 40 percentage points between phases, highlighting the need for phase-specific evaluation.
3 days agoDaily PapersThe paper explores test-time AI-for-AI, focusing on how a Builder can create better execution environments for a Target while keeping both models' weights fixed. It introduces Meta-Skill, principles derived from Target's execution feedback, which improve performance in tasks like Harness-Bench and NewtonBench.
2 days ago