Latest Articles
Controlling Reasoning Effort in LLMs
Controlling Reasoning Effort in LLMs

How LLMs Learn Low-, Medium-, and High-Effort Reasoning Modes

Using Local Coding Agents
Using Local Coding Agents

Using Open-Weight Models in Local Coding Harnesses as an Alternative to Claude Code and Codex Subscriptions

LLM Research Papers: The 2026 List (January to May)
LLM Research Papers: The 2026 List (January to May)

A curated roundup of notable LLM research papers that came out this year

Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention
Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention

From Gemma 4 to DeepSeek V4, How New Open-Weight LLMs Are Reducing Long-Context Costs

Quick Notes
Muse Glimmer 30B Architecture Notes

Short architecture note on Meta Muse Glimmer 30B, including gated local and global GQA, KV-cache efficiency, and release-time benchmark c...

LLMs From Scratch Reaches 100,000 GitHub Stars

Short note celebrating the LLMs-from-scratch repository passing 100,000 GitHub stars and summarizing its learning materials.

Kimi K3 Architecture Notes

Short architecture note on Kimi K3, including LatentMoE, Kimi Delta Attention, Attention Residuals, NoPE, multimodality, and inference-ef...

A Few Notable Open-Weight Models This Week

Short note on the architectures of six new open-weight models, including Nanbeige 4.2, Laguna S 2.1, Motif-3-Beta, Solar Open 2, Antares ...

Correction for Listing 6.5 in Build a Reasoning Model From Scratch

Short correction note for the random seed in Listing 6.5 on page 198 of Build a Reasoning Model From Scratch.