Latest Articles
GPT-6 Astra, Looped Transformers, and Hidden Reasoning
GPT-6 Astra, Looped Transformers, and Hidden Reasoning

A Look at Recurrent Depth, Hidden Chains of Thought, and Recent Research on Looping Transformer Blocks

How Claude Watermarks AI-Generated Text
How Claude Watermarks AI-Generated Text

A 48-minute video walkthrough of token sampling, watermark detection, and removal

Building an AI Text Detector From Scratch
Building an AI Text Detector From Scratch

An End-to-End Project With Dataset Construction, Model Training, Local Deployment, and RLVR

Controlling Reasoning Effort in LLMs
Controlling Reasoning Effort in LLMs

How LLMs Learn Low-, Medium-, and High-Effort Reasoning Modes

Quick Notes
Focusing on Post-Training

Why I'd invest in post-training existing open-weight LLMs, with Fireworks' Ember-1 as an example of more token-efficient reasoning.

MiMo-V2.6 Pro Architecture and Training Notes

Notes on MiMo-V2.6 Pro's GQA and sliding-window attention, agent training tasks, reward signals, and large RL batches.

It's Easy to Dismiss Jev as Just a Classifier

A short note on Jev's generalization, possible encoder-style architecture and training, and Choice and Noul API examples.

Pacing != Pacing Development

My take on AI model pacing as a framework for release checks and the competitive pressure around model releases.

AI Reasoning Models Course on LinkedIn Learning

A 90-minute LinkedIn Learning course on how reasoning models relate to conventional LLMs and how they are developed.