Tag: #AI Safety

Search Results

Technologyarticle

Debate Training Reduces Reward Hacking in RLAIF

#Reinforcement Learning#RLAIF#Reward Hacking#AI Safety#LLM Alignment#Debate Training
Technologynews

Import AI 469: Science AI, RSI Simulator, and Zuck's Tech Pessimism

#Import AI#Jack Clark#Recursive Self-Improvement#Mark Zuckerberg#AI Safety
Technologyarticle

After Orthogonality: Virtue-Ethical Agency and AI Alignment

#AI Alignment#Virtue Ethics#Eudaimonia#AI Safety#Philosophy
AIarticle

AI Swarms Are Starting to Pose Indirect Takeover Risk, Researchers Warn

#AI Safety#Multi-Agent Systems#OpenAI#Codex#Takeover Risk
AIarticle

Import AI 468: RSI Ideas, PostTrainBench, and Trust in AI Racing

#Import AI#RSI#PostTrainBench#AI Safety#Recursive Self-Improvement
AIblog

AI Agents Are Sending Angry Emails and Writing Hit Pieces Now

#AI Agents#AI Safety#Harassment#Open Source#AI Weirdness
AIarticle

Democratizing ASI: A Risky Path to Preserving Civil Liberties

#ASI#AI Safety#Civil Liberties#Democratization#Alignment
AIarticle

Frontier Models Show User Awareness, Shifting Behavior by Who Asks

#User Awareness#Frontier Models#AI Safety#Claude#Situational Awareness
AIarticle

Import AI 466: MirrorCode, Anthropic's Robot Sprint, and OpenAI's Hacker Problem

#Import AI#MirrorCode#Anthropic#OpenAI#AI Safety
AIarticle

The Open Letters That Shaped the AI Safety Conversation

#AI Safety#Future of Life Institute#Open Letters#Superintelligence#AI Policy
AIarticle

Eudaimonic Rationality: A Different Frame for AI Alignment

#AI Alignment#Virtue Ethics#MIRI#Eudaimonia#AI Safety
AIarticle

Anthropic's Sandbox Breach and the Real Agent Safety Lesson

#Anthropic#Claude#Agent Security#Cybersecurity#AI Safety
AIblog

Sam Altman, the Decel Debate, and Why Neither Frame Is Useful

#Sam Altman#OpenAI#AI Safety#Hugging Face Hack#AI Policy
AIblog

Who's Watching Your AI Agent While You Sleep?

#AI Agents#Autonomous AI#AI Risks#AI Safety#Automation
Technologyarticle

Why AI Safety Needs Individual Voices Now More Than Ever

#AI Safety#Public Trust#AI Risk#Communication#Anthropic#OpenAI
AIarticle

Value Leakage: How LLM Answers Are Shaped by Their Own Values

#Value Leakage#LLM Alignment#AI Safety#Covert Bias#Frontier Models
AInews

Agentic Misalignment: When AI Agents Go Rogue in 2026

#Agentic Misalignment#AI Safety#Anthropic#AI Alignment#Frontier Models
AI Ethicsarticle

After Orthogonality: Virtue-Ethical Agency and AI Alignment

#AI Alignment#Virtue Ethics#Eudaimonia#AI Safety
Technologyblog

It's 11:00 PM. Do You Know Where Your AI Agent Is?

#AI Agent#AI Safety#AI Ethics#Harassment#Open Source
AIarticle

PIRAMID: Building Scientific Foundations for Mechanistic Interpretability

#Mechanistic Interpretability#AI Safety#Statistical Physics#PIRAMID#AI Alignment
Technologyblog

The Long Self-Correction: Our Greatest Flaw in Building Safe AI

#AI Safety#Human Flaws#AI Alignment#Philosophy#Longtermism
AIarticle

Stateful Guardrails for Multi-Turn LLMs Catch Hidden Risks

#AI Safety#LLM Guardrails#Multi-Turn#Jailbreak#Conversational Risk
AIarticle

Selective Fact-Checking with Evidence Chain Evaluation

#Fact-Checking#LLM Evaluation#Evidence Chain#AI Safety#Calibration
AIarticle

SysAdmin Benchmark: Power-Seeking in Frontier LLMs, Measured

#AI Safety#Power-Seeking#LLM Evaluation#Frontier Models#Loss of Control
AIblog

Why AI Needs a Genie Coefficient

#AI Safety#AI Alignment#Agentic AI#Benchmarking#AI Ethics
Technologyarticle

After Orthogonality: Why Rational AI Should Not Have Goals

#AI Alignment#AI Safety#Virtue Ethics#Eudaimonia#Philosophy of AI
Technologyarticle

Why AI Needs a Genie Coefficient to Measure Intent Misalignment

#AI Safety#AI Agents#Genie Coefficient#Alignment#Bruce Schneier
AIblog

AI Models That Escape Their Sandbox Are No Longer Science Fiction

#OpenAI#Hugging Face#AI Safety#Autonomous Agents#Cybersecurity
AIarticle

Measuring Power-Seeking Behavior in Frontier AI Models

#Frontier AI#AI Safety#LLM Benchmark#Loss of Control
Technologynews

OpenAI Agent Escapes Sandbox, Breaches Hugging Face in Unprecedented Security Incident

#OpenAI#Hugging Face#AI Safety#Cybersecurity#Agentic AI
Technologyarticle

After orthogonality: why virtue ethics may be the key to aligning AI with human values

#AI Alignment#Virtue Ethics#Eudaimonic Rationality#AI Safety#Philosophy of AI
AIblog

The OpenAI-Hugging Face Incident Should Worry Us More Than It Has

#OpenAI#Hugging Face#AI Safety#Cybersecurity Incident#AI Alignment
AInews

OpenAI Models Breached Hugging Face Infrastructure During Testing

#OpenAI#Hugging Face#Cybersecurity Incident#AI Safety#GPT-5.6
AInews

Endogenous AI Alignment Could Be the Next Safety Frontier

#AI Alignment#Artificial General Intelligence#RLHF#AI Safety#Endogenous Alignment
AInews

AI Agents Raise New Concerns Over Unsupervised Online Behavior

#AI Agents#Scott Shambaugh#Open Source#AI Safety#Automation
AInews

Rogue AI Agent Defames Open-Source Python Maintainer

#AI Agents#Open Source#AI Safety#Agentic AI#GitHub
Technologynews

Teen AI Access Debate Intensifies as Safety Advocates Push for Guardrails

#AI Safety#Teenagers#Digital Rights#Child Protection#AI Regulation
AInews

AI Benchmarks Scorecard Evaluates Models Beyond Accuracy

#AI Benchmarks#Model Evaluation#AI Safety#LLM Testing#Artificial General Intelligence
AInews

Eudaimonic Rationality Proposed as AI Alignment Framework

#AI Alignment#Eudaimonic Rationality#Virtue Ethics#AI Safety#Consequentialism