LLM engineering
44 posts
The refund letter addressed to Dear [Name]
Why ChatGPT's first output is a draft, not a deliverable, and what production AI systems actually require beyond the prompt.
The smooth line hiding a noisy benchmark
The METR AI time horizons graph contains structural errors that mislead teams building agents, automation, and AI workflows. Here is what it actually shows.
Hugging Face revived PapersWithCode in early 2025
Hugging Face's PapersWithCode revival restores the verification substrate LLM engineering teams lost, reshaping pipelines and AI workforce roles.
Sub-JEPA tightens the prediction signal
Sub-JEPA is a small loss-side fix to LeCun's world models that consistently improves performance. Here's how it works, where it fails, and why it matters.
Better AI isn't what separates winning deployments.
Stanford studied 51 AI deployments and found a 71 vs 40 productivity gap. The difference was pipeline design, not model choice.
arXiv just raised the bar
arXiv's one-year ban on unchecked LLM errors signals a shift: validation pipelines, not better prompts, now define competent AI systems.
Complexity theory never said that
Complexity theory does not prove human-level ML is impossible. Here is what the theorems actually say and how to design AI systems around real constraints.
AI costs more than humans
Nvidia says AI costs more than human workers. The real issue is architecture, not compute price. Here is how to fix the unit economics.
Managed Agents pricing is an architecture decision
Claude Managed Agents pricing isn't a cost center - it's an orchestration lever. Here's how to evaluate it against real total cost of ownership.
Apple isn't competing with OpenAI
Apple's AI strategy is a silicon bet, not a model race. The real architecture question is where inference runs - and who controls the hardware lane.
How Production Systems Actually Work With LLMs-Not Which Model You Choose
Production-grade AI systems don't depend on choosing between Claude and ChatGPT. They rely on consistent engineering: input sanitization, output validation, fallback logic, and structured pipelines-regardless of the underlying LLM.
Running Gemma 4 Locally via Codex CLI: What Actually Works in Practice
Running Gemma 4 locally via Codex CLI offers isolation but not guaranteed consistency. Real reliability comes from input validation, output schema checks, and disciplined system design-not the model alone.