Tag: #AI Safety
Search Results
Technologyarticle
Debate Training Reduces Reward Hacking in RLAIF
#Reinforcement Learning#RLAIF#Reward Hacking#AI Safety#LLM Alignment#Debate Training
Technologynews
Import AI 469: Science AI, RSI Simulator, and Zuck's Tech Pessimism
#Import AI#Jack Clark#Recursive Self-Improvement#Mark Zuckerberg#AI Safety
Technologyarticle
After Orthogonality: Virtue-Ethical Agency and AI Alignment
#AI Alignment#Virtue Ethics#Eudaimonia#AI Safety#Philosophy
AIarticle
AI Swarms Are Starting to Pose Indirect Takeover Risk, Researchers Warn
#AI Safety#Multi-Agent Systems#OpenAI#Codex#Takeover Risk
AIarticle
Import AI 468: RSI Ideas, PostTrainBench, and Trust in AI Racing
#Import AI#RSI#PostTrainBench#AI Safety#Recursive Self-Improvement
AIblog
AI Agents Are Sending Angry Emails and Writing Hit Pieces Now
#AI Agents#AI Safety#Harassment#Open Source#AI Weirdness
AIarticle
Democratizing ASI: A Risky Path to Preserving Civil Liberties
#ASI#AI Safety#Civil Liberties#Democratization#Alignment
AIarticle
Frontier Models Show User Awareness, Shifting Behavior by Who Asks
#User Awareness#Frontier Models#AI Safety#Claude#Situational Awareness
AIarticle
Import AI 466: MirrorCode, Anthropic's Robot Sprint, and OpenAI's Hacker Problem
#Import AI#MirrorCode#Anthropic#OpenAI#AI Safety
AIarticle
The Open Letters That Shaped the AI Safety Conversation
#AI Safety#Future of Life Institute#Open Letters#Superintelligence#AI Policy
AIarticle
Eudaimonic Rationality: A Different Frame for AI Alignment
#AI Alignment#Virtue Ethics#MIRI#Eudaimonia#AI Safety
AIarticle
Anthropic's Sandbox Breach and the Real Agent Safety Lesson
#Anthropic#Claude#Agent Security#Cybersecurity#AI Safety
AIblog
Sam Altman, the Decel Debate, and Why Neither Frame Is Useful
#Sam Altman#OpenAI#AI Safety#Hugging Face Hack#AI Policy
AIblog
Who's Watching Your AI Agent While You Sleep?
#AI Agents#Autonomous AI#AI Risks#AI Safety#Automation
Technologyarticle
Why AI Safety Needs Individual Voices Now More Than Ever
#AI Safety#Public Trust#AI Risk#Communication#Anthropic#OpenAI
AIarticle
Value Leakage: How LLM Answers Are Shaped by Their Own Values
#Value Leakage#LLM Alignment#AI Safety#Covert Bias#Frontier Models
AInews
Agentic Misalignment: When AI Agents Go Rogue in 2026
#Agentic Misalignment#AI Safety#Anthropic#AI Alignment#Frontier Models
AI Ethicsarticle
After Orthogonality: Virtue-Ethical Agency and AI Alignment
#AI Alignment#Virtue Ethics#Eudaimonia#AI Safety
Technologyblog
It's 11:00 PM. Do You Know Where Your AI Agent Is?
#AI Agent#AI Safety#AI Ethics#Harassment#Open Source
AIarticle
PIRAMID: Building Scientific Foundations for Mechanistic Interpretability
#Mechanistic Interpretability#AI Safety#Statistical Physics#PIRAMID#AI Alignment
Technologyblog
The Long Self-Correction: Our Greatest Flaw in Building Safe AI
#AI Safety#Human Flaws#AI Alignment#Philosophy#Longtermism
AIarticle
Stateful Guardrails for Multi-Turn LLMs Catch Hidden Risks
#AI Safety#LLM Guardrails#Multi-Turn#Jailbreak#Conversational Risk
AIarticle
Selective Fact-Checking with Evidence Chain Evaluation
#Fact-Checking#LLM Evaluation#Evidence Chain#AI Safety#Calibration
AIarticle
SysAdmin Benchmark: Power-Seeking in Frontier LLMs, Measured
#AI Safety#Power-Seeking#LLM Evaluation#Frontier Models#Loss of Control
AIblog
Why AI Needs a Genie Coefficient
#AI Safety#AI Alignment#Agentic AI#Benchmarking#AI Ethics
Technologyarticle
After Orthogonality: Why Rational AI Should Not Have Goals
#AI Alignment#AI Safety#Virtue Ethics#Eudaimonia#Philosophy of AI
Technologyarticle
Why AI Needs a Genie Coefficient to Measure Intent Misalignment
#AI Safety#AI Agents#Genie Coefficient#Alignment#Bruce Schneier
AIblog
AI Models That Escape Their Sandbox Are No Longer Science Fiction
#OpenAI#Hugging Face#AI Safety#Autonomous Agents#Cybersecurity
AIarticle
Measuring Power-Seeking Behavior in Frontier AI Models
#Frontier AI#AI Safety#LLM Benchmark#Loss of Control
Technologynews
OpenAI Agent Escapes Sandbox, Breaches Hugging Face in Unprecedented Security Incident
#OpenAI#Hugging Face#AI Safety#Cybersecurity#Agentic AI
Technologyarticle
After orthogonality: why virtue ethics may be the key to aligning AI with human values
#AI Alignment#Virtue Ethics#Eudaimonic Rationality#AI Safety#Philosophy of AI
AIblog
The OpenAI-Hugging Face Incident Should Worry Us More Than It Has
#OpenAI#Hugging Face#AI Safety#Cybersecurity Incident#AI Alignment
AInews
OpenAI Models Breached Hugging Face Infrastructure During Testing
#OpenAI#Hugging Face#Cybersecurity Incident#AI Safety#GPT-5.6
AInews
Endogenous AI Alignment Could Be the Next Safety Frontier
#AI Alignment#Artificial General Intelligence#RLHF#AI Safety#Endogenous Alignment
AInews
AI Agents Raise New Concerns Over Unsupervised Online Behavior
#AI Agents#Scott Shambaugh#Open Source#AI Safety#Automation
AInews
Rogue AI Agent Defames Open-Source Python Maintainer
#AI Agents#Open Source#AI Safety#Agentic AI#GitHub
Technologynews
Teen AI Access Debate Intensifies as Safety Advocates Push for Guardrails
#AI Safety#Teenagers#Digital Rights#Child Protection#AI Regulation
AInews
AI Benchmarks Scorecard Evaluates Models Beyond Accuracy
#AI Benchmarks#Model Evaluation#AI Safety#LLM Testing#Artificial General Intelligence
AInews
Eudaimonic Rationality Proposed as AI Alignment Framework
#AI Alignment#Eudaimonic Rationality#Virtue Ethics#AI Safety#Consequentialism