Как меняется интерес
История начинает расти
График появится после повторных замеров. Текущий показатель уже получен из источника.
22 очков источникаТолько реальные замеры. История до подключения источника не восстанавливается.
Here’s a screenshot from Anthropic’s recent “training a reward seeker” post: Recently there’s been a lot of discussion about how RL has actually produced not merely reward hacking, but explicitly reward-seeking behavior, almost as if to spite shard theorists personally. However, on top of that, note that the behavior in the image is not merely reward-seeking, but wireheading. Terminology regarding various sorts of things that can be called “reward hacking” is endlessly confused, with lots of historical shifts in usage.[1] I’m talking about the thing that the linked LW post calls wireheading-- the RL policy appearing to terminally value the representation of the reward, rather than the thing that representation points to. Even taking for granted that RL produces reward seeking behavior, an analysis from the pre-LLM, pure RL perspective would suggest that wireheading is far less likely. This post provides a good working model for the execution of RL algorithms in an embedded setting, and an analysis which describes the conditions under which wireheading might be expected to arise. As a brief summary: once explored into, wireheading policies actually actually do achieve high reward with respect to of the embedded implementation of the RL algorithm, so they are fit from a selection perspective, hence, we expect wireheading to arise given an RL algorithm with sufficiently strong exploration. However, in this model and under this analysis, we'd find that without the contribution of LLM priors, wireheading is not expected in the Hacker Opus setting: * our exploration techniques are somewhat weak (they basically involve just sampling from the LLM with temperature, i.e., they don't deviate much from the present policy), and wireheading policies are extremely different (in the sense that they require significantly different actions) from other high reward policies; * to the extent that there's some form of clipping or normalization (as there are with contemporary RL algo
Перевод готовится · пока описание источникаГрафик появится после повторных замеров. Текущий показатель уже получен из источника.
22 очков источникаТолько реальные замеры. История до подключения источника не восстанавливается.
Текст сравнивает продвинутый ИИ с концепцией иммиграции, подчеркивая опасения по поводу потенциального влияния ИИ на общество, включая различия в ценностях, экономической конкуренции и потере контроля. Он предполагает, что ИИ может представлять большие проблемы, чем иммиграция людей, из-за радикально чуждых ценностей и превосходных способностей.
6 дней назадHugging FaceКарточка модели abenzerps/Qwen-Image-2.1-Uncensored-GGUF на Hugging Face. Идентификатор задачи: text-to-image. Подробности по ссылке на источник.
9 дней назадHacker NewsWhiteboard — это open-source IDE, разработанный четырьмя бывшими тех-лидами, предназначенный для совместной разработки программного обеспечения между людьми и ИИ-агентами. Он интегрируется с существующими инструментами, такими как Claude Code и Codex, позволяя агентам визуализировать свою работу на встроенной доске и связывать диаграммы с кодом. В приложении есть такие функции, как семантический просмотрик изменений и журнал решений, чтобы улучшить проверку кода и понимание решений, принятых ИИ.
5 дней назадLobstersValve представила Pyrowave, новый видео кодек в бета-версии, предназначенный для потоковой передачи с низкой задержкой.
позавчераHacker NewsMeta удалила критическую видео о своих AI Glasses после съемки на территории компании.
5 дней назадHacker NewsВидео содержит открытие конференции Rails World 2026, мероприятия, посвящённого фреймворку Ruby on Rails.
6 дней назад