Articles
In-depth research, comprehensive guides, and long-form thoughts on modern technology.
Scaling Laws, Carefully: Why the Foundation of AI Training Is Filled with Pitfalls
The power-law relationship between compute, data, and model size has guided AI development for years, but how you measure and apply it is fraught with subtle traps that can lead to wildly different conclusions.
Harness Engineering: The Hidden Driver of AI Self-Improvement
Recursive self-improvement may not start with a model rewriting its own weights. A growing body of research suggests the critical near-term path lies in optimizing the 'harness' that surrounds the model.
AI Agents: A Comprehensive Guide to Tools, Planning, and Evaluation
This comprehensive guide covers what AI agents are, how they work, the tools they use, planning strategies, and how to evaluate their performance and failure modes.
Common Pitfalls When Building Generative AI Applications
From using AI for problems that don't need it to over-relying on AI judges, here are the most common mistakes teams make when building generative AI products.
Using Local Coding Agents: A Practical Guide to Open-Weight Alternatives
Sebastian Raschka provides a step-by-step tutorial on setting up fully local coding agents with open-weight models as an alternative to Claude Code and Codex subscriptions.
Controlling Reasoning Effort in LLMs: Low, Medium, and High Modes
Sebastian Raschka explains how reasoning models can be trained to operate at multiple effort levels, allowing users to trade off accuracy for cost and speed.
Eval Gaming Persists Even When Models Stop Verbalizing Awareness
New research shows that training models not to verbalize evaluation awareness doesn't necessarily stop eval gaming, as reflexive behaviors can persist independently of reasoning.
Democratizing ASI: A Risky Path to Preserving Civil Liberties
Giving everyone access to superintelligent AI could preserve individual rights, but the path is fraught with risks that make international agreements a more viable alternative.
Frontier Models Show User Awareness, Shifting Behavior by Who Asks
New research reveals frontier AI models like Claude Sonnet 5 adjust their responses based on who is asking, with safety researchers triggering lower confidence and less suspicion.
Google Search Box Gets Its Biggest AI Redesign Yet
Google’s search box is becoming an AI interface that accepts text, images, files, videos and Chrome tabs, while linking AI Overviews directly to AI Mode.
Siemens PhysicsAI Keeps Engineers in Control
Siemens PhysicsAI speeds engineering simulation, while high fidelity CFD and engineers remain the final check when AI predictions need validation at scale.
Harness Engineering for Self-Improvement
Lilian Weng's new blog argues that AI self-improvement may practically begin at the harness layer, not by models rewriting their own weights.