Latest Articles
Controlling Reasoning Effort in LLMs
Controlling Reasoning Effort in LLMs

How LLMs Learn Low-, Medium-, and High-Effort Reasoning Modes

Using Local Coding Agents
Using Local Coding Agents

Using Open-Weight Models in Local Coding Harnesses as an Alternative to Claude Code and Codex Subscriptions

LLM Research Papers: The 2026 List (January to May)
LLM Research Papers: The 2026 List (January to May)

A curated roundup of notable LLM research papers that came out this year

Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention
Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention

From Gemma 4 to DeepSeek V4, How New Open-Weight LLMs Are Reducing Long-Context Costs

Quick Notes
Correction for Listing 6.5 in Build a Reasoning Model From Scratch

Short correction note for the random seed in Listing 6.5 on page 198 of Build a Reasoning Model From Scratch.

Inkling: A New Open-Weight 975B MoE with a Few Surprises

Short note on Thinking Machines Lab's 975B Inkling model, including benchmarks, sparse MoE design, short convolutions, RMSNorm, and posit...

200,000 Subscribers

Short note celebrating Ahead of AI reaching 200,000 subscribers.

GPT 5.6 Has 72 Possible Configurations. What's A Good Default?

Short note on how GPT 5.6 model and effort choices map onto training-time and inference-time scaling, producing 72 configurations.

Build a Reasoning Model From Scratch Is Out

Short note announcing the release of Build a Reasoning Model From Scratch and linking the publisher and Amazon pages.