Linked List
login
sign up
Astra and Fable still hack on simple variants of alignment evals from 2025 — LessWrong
lesswrong.com
· first added by
@marcus_r
saved by
2 people
discussions
·
1
Astra and Fable still hack on simple variants of alignment evals from 2025
482 pts · 235 comments · Sep 2026
482 pts · 235 comments · Sep 2026
▲
0
0 people saved or upvoted this
from the discussion
·
5
Let's Verify Step by Step
arxiv.org
arxiv.org
▲
0
0 people saved or upvoted this
Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident
metr.org
metr.org
▲
0
0 people saved or upvoted this
Measuring Reward-Seeking by Instilling Contrastive Beliefs
alignment.openai.com
alignment.openai.com
▲
0
0 people saved or upvoted this
LLM Evaluators Recognize and Favor Their Own Generations
arxiv.org
arxiv.org
▲
0
0 people saved or upvoted this
Claude’s Constitution
anthropic.com
anthropic.com
▲
0
0 people saved or upvoted this
Feed
Explore
Sign In