Latest Articles
Controlling Reasoning Effort in LLMs
Controlling Reasoning Effort in LLMs

How LLMs Learn Low-, Medium-, and High-Effort Reasoning Modes

Using Local Coding Agents
Using Local Coding Agents

Using Open-Weight Models in Local Coding Harnesses as an Alternative to Claude Code and Codex Subscriptions

LLM Research Papers: The 2026 List (January to May)
LLM Research Papers: The 2026 List (January to May)

A curated roundup of notable LLM research papers that came out this year

Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention
Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention

From Gemma 4 to DeepSeek V4, How New Open-Weight LLMs Are Reducing Long-Context Costs

Quick Notes
Kimi K3 Architecture Notes

Short architecture note on Kimi K3, including LatentMoE, Kimi Delta Attention, Attention Residuals, NoPE, multimodality, and inference-ef...

A Few Notable Open-Weight Models This Week

Short note on the architectures of six new open-weight models, including Nanbeige 4.2, Laguna S 2.1, Motif-3-Beta, Solar Open 2, Antares ...

Correction for Listing 6.5 in Build a Reasoning Model From Scratch

Short correction note for the random seed in Listing 6.5 on page 198 of Build a Reasoning Model From Scratch.

Inkling: A New Open-Weight 975B MoE with a Few Surprises

Architecture and benchmark notes on Thinking Machines Lab's 975B Inkling MoE, including short convolutions, relative-position bias, train...

200,000 Subscribers

Short note celebrating Ahead of AI reaching 200,000 subscribers.