AI safety
32 posts
A 1938 law now points at AI critics
How labeling AI critics as 'foreign agents' under FARA-style rules chills safety research, disclosure, and open discourse - and what researchers can do.
Snapdragon X2's September 2025 debut bets on mainline Linux
Linux support on Qualcomm's Snapdragon X2 improves auditability but moves AI safety controls onto hardware the device owner fully controls - here's the security tradeoff.
The benchmark score is the number to trust least
How to weigh Claude Opus 5.5's intelligence, latency, and token cost, and where its real AI safety and cybersecurity risks concentrate.
Heretic strips refusals from open-weight models
Heretic automates stripping refusals from open-weight LLMs. Why model-level guardrails were never a security control, and what defenders should do instead.
Mini-AGI trains a model you can't sign
Mini-AGI trains a continual-learning model on 8GB VRAM. Why self-updating weights break signing, auditing, and red-team results, and how to defend.
His accountant found the $14,000 error three weeks late
AI chatbots get most specific financial questions wrong. Here is how they fail and the human checks that catch it before money moves.
Microsoft's AI CEO called the web freeware
Microsoft and OpenAI executives described how LLMs are built and tuned. What that admission actually means for AI safety and security teams.
May 2024's Bend is no proof assistant
Bend is a parallel language, not a proof assistant. What proof-based programming actually does for AI safety, and the errors it can't touch.
Gemini 3.8 Live broke two security assumptions
Gemini 3.8 Live and Extended Thinking make ambient audio and video an untrusted AI input, reshaping prompt injection, logging, and privacy risk.
What distillation leaves behind
Distilling frontier AI models copies capability cheaply but leaves safety training behind. What Garry Tan's push means for cybersecurity and AI safety.
A warning is not a wall
An AI sandbox escape is the wrong thing to fear. The real AI safety risk is a system acting on a flattened, ungrounded model of a sensitive region.
Keep the two claims apart
Consumer AI trains on your chats by default. How to tell the real consent problem from unprovable 'secret breakthrough' claims - and what you can control.