AI safety
32 posts
The model kept no receipts
The "OpenAI stole my proof" fight is really a provenance gap in AI, and that gap is a safety problem, not just a credit dispute.
A helpful AI agent cannot be a private one
Meta's Muse personal AI agent is only useful because it reads your messages, contacts, and habits. What that access costs your privacy and safety.
Open problems are running out
Terence Tao calls open math problems a non-renewable resource. Why AI mining them threatens encryption, AI safety benchmarks, and how to respond.
Resignations are signals, not scandals
How to read a high-profile resignation from an AI safety lab like Anthropic - what it signals, what to ignore, and what to watch in the months after.
You can't reset your genome
AlphaGenome Atlas makes genomic prediction cheap and fast, but a genome can't be rotated like a password - reshaping AI safety and data ethics.
Cerebras runs Qwen 27B at 1,500 tokens a second
Qwen 3.8 27B on Cerebras at 1,500 tokens/s adds no new capability - it changes the economics of attack and defense. What the raw speed means for security.
Same badge, new model underneath
A systems-analysis read on Claude Fable 5.1 and Mythos 5.1: why a point release quietly moves the safety boundary and what to re-test before you ship it.
Your model isn't cheating, it's reward hacking
A GPT model was called the most cheating-prone ever tested. The mechanism is reward hacking, not intent - and there is no CVE and no patch.
An AI engineer skipped the prediction test
How to evaluate an AI engineer's claim to have cracked Linear A - the data limits, the missing known language, and the peer review that separates progress from hype.
The Open Courts Act exposes what PACER fees hid
PACER's per-page fee was an accidental privacy brake. Making court records free is right - but only if redaction, governed bulk access, and security replace it.
1994's eight fallacies hit AI agents harder
The eight fallacies of distributed computing turn 21, and autonomous AI agents make every one of those architectural assumptions more dangerous.
GLM 5.2 lands; reasoning improves, refusals don't
GLM 5.2's reasoning gains widen the gap between what a model can do and what it will refuse. What security researchers and developers should test first.