Quick Paper and Model Notes
-
Building an AI Text Detector From Scratch
An End-to-End Project With Dataset Construction, Model Training, Local Deployment, and RLVR
-
How Claude's Text Watermarking Works
Short illustration of how Claude's text watermarking is supposed to work based on Anthropic's released materials.
Substack Note -
Build a Reasoning Model From Scratch Is Now on Amazon
Short note on the Amazon availability of Build a Reasoning Model From Scratch and a warning about counterfeit black-and-white copies sold through Amazon India.
Substack Note -
Muse Glimmer 30B Architecture Notes
Short architecture note on Meta Muse Glimmer 30B, including gated local and global GQA, KV-cache efficiency, and release-time benchmark comparisons.
Substack Note -
LLMs From Scratch Reaches 100,000 GitHub Stars
Short note celebrating the LLMs-from-scratch repository passing 100,000 GitHub stars and summarizing its learning materials.
Substack Note -
Kimi K3 Architecture Notes
Short architecture note on Kimi K3, including LatentMoE, Kimi Delta Attention, Attention Residuals, NoPE, multimodality, and inference-efficiency choices.
Substack Note -
A Few Notable Open-Weight Models This Week
Short note on the architectures of six new open-weight models, including Nanbeige 4.2, Laguna S 2.1, Motif-3-Beta, Solar Open 2, Antares 1B, and BTL-3.
Substack Note -
Correction for Listing 6.5 in Build a Reasoning Model From Scratch
Short correction note for the random seed in Listing 6.5 on page 198 of Build a Reasoning Model From Scratch.
Substack Note -
Controlling Reasoning Effort in LLMs
How LLMs Learn Low-, Medium-, and High-Effort Reasoning Modes
-
Inkling: A New Open-Weight 975B MoE with a Few Surprises
Architecture and benchmark notes on Thinking Machines Lab's 975B Inkling MoE, including short convolutions, relative-position bias, training, and effort control.
-
200,000 Subscribers
Short note celebrating Ahead of AI reaching 200,000 subscribers.
Blog -
GPT 5.6 Has 72 Possible Configurations. What's A Good Default?
Short note on how GPT 5.6 model and effort choices map onto training-time and inference-time scaling, producing 72 configurations.
-
Build a Reasoning Model From Scratch Is Out
Short note announcing the release of Build a Reasoning Model From Scratch and linking the publisher and Amazon pages.
Substack Note -
Using Local Coding Agents
Short note linking a new article on setting up local coding agents with open-weight models.
Substack Note -
Using Local Coding Agents
Using Open-Weight Models in Local Coding Harnesses as an Alternative to Claude Code and Codex Subscriptions
-
Local Open-Weight LLMs in Coding Harnesses
Short note on trying local open-weight LLMs across Qwen-Code, Codex, and Claude Code harnesses.
Substack Note -
GLM-5.2 and IndexShare for Long-Context Sparse Attention
How GLM-5.2 uses IndexShare to reuse sparse-attention token selections across layers and reduce long-context indexer computation.
Substack Note -
VibeThinker-3B and the Strength of Post-Training
A closer look at VibeThinker-3B, a Qwen2.5-Coder-based model whose reported reasoning gains come from a detailed post-training pipeline.
Substack Note -
North Mini Code and Agentic Coding Benchmarks
Architecture and benchmark notes for North Mini Code, Cohere's 30B-A3B MoE trained for repository, terminal, and code-generation tasks.
Substack Note -
LLM Research Papers: The 2026 List (January to May)
A curated roundup of notable LLM research papers that came out this year
-
Nemotron 3 Ultra and Latent MoE Scaling
Architecture notes on Nemotron 3 Ultra, including its 108-layer hybrid stack, Latent MoE scaling, NVFP4 recipe, MTP, and inference results.
Substack Note -
MiniMax M2 Technical Report Notes
Technical notes on MiniMax M2, including full attention, fine-grained MoE routing, agent training data, speed rewards, and self-evolution.
Substack Note -
DeepSeek Sparse Attention From Scratch
Short note on a DeepSeek Sparse Attention from-scratch implementation added to the LLMs-from-scratch repository.
Substack Note -
Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention
From Gemma 4 to DeepSeek V4, How New Open-Weight LLMs Are Reducing Long-Context Costs
-
Implementing LLM Architectures From Scratch
Short note linking a talk on implementing LLM architectures from scratch and comparing new open-weight model implementations against references.
Substack Note -
My Workflow for Understanding LLM Architectures
A learning-oriented workflow for understanding new open-weight model releases
-
Components of A Coding Agent
How coding agents use tools, memory, and repo context to make LLMs work better in practice
-
Gemma 4 Architecture and Benchmark Notes
Architecture and benchmark notes for Gemma 4 31B and 26B-A4B, including hybrid attention, long-context changes, and evaluation caveats.
Substack Note -
LLM Architecture Gallery Diff Tool
Compare two LLM architectures side by side across attention, decoder type, layer recipe, model scale, context length, and KV-cache use.
Substack Note -
A Visual Guide to Attention Variants in Modern LLMs
From MHA and GQA to MLA, sparse attention, and hybrid architectures
-
Nemotron 3 Super Throughput Notes
Nemotron 3 Super combines Mamba-2, Latent MoE, sparse GQA, and shared-weight MTP in a throughput-oriented 120B-A12B model.
Substack Note
No entries match the selected filters.