AI
Published articles
- Humanity's Last Exam: A New Rigorous Benchmark for Frontier AI Models
For years, the progress of Large Language Models (LLMs) has been measured by a handful of standardized benchmarks. From MMLU to GSM8K, these tests served as the gold standard for assessing an AI's reasoning, coding, and general knowledge. However, the industry has hit a wall know
- Pentagon Launches GenAI.mil to Give 3 Million Personnel AI Tools
I’ve spent a fair amount of my life staring at screens in rooms with no windows, and if there’s one thing I’ve learned, it’s that the government’s idea of "rolling out" software usually involves a lot of expensive noise and a fair bit of confusion. But the latest move from the Pe