Latest Articles
GPT-6 Astra, Looped Transformers, and Hidden Reasoning
GPT-6 Astra, Looped Transformers, and Hidden Reasoning

A Look at Recurrent Depth, Hidden Chains of Thought, and Recent Research on Looping Transformer Blocks

How Claude Watermarks AI-Generated Text
How Claude Watermarks AI-Generated Text

A 48-minute video walkthrough of token sampling, watermark detection, and removal

Building an AI Text Detector From Scratch
Building an AI Text Detector From Scratch

An End-to-End Project With Dataset Construction, Model Training, Local Deployment, and RLVR

Controlling Reasoning Effort in LLMs
Controlling Reasoning Effort in LLMs

How LLMs Learn Low-, Medium-, and High-Effort Reasoning Modes

Quick Notes
MiMo-V2.6 Pro Architecture and Training Notes

Notes on MiMo-V2.6 Pro's GQA and sliding-window attention, agent training tasks, reward signals, and large RL batches.

It's Easy to Dismiss Jev as Just a Classifier

A short note on Jev's generalization, possible encoder-style architecture and training, and Choice and Noul API examples.

Pacing != Pacing Development

My take on AI model pacing as a framework for release checks and the competitive pressure around model releases.

AI Reasoning Models Course on LinkedIn Learning

A 90-minute LinkedIn Learning course on how reasoning models relate to conventional LLMs and how they are developed.

OpenAI Astra and Looped Transformers

A short note on OpenAI Astra, recurrent depth, looped transformers, Nanbeige 4.2, and the Mixture-of-Recursions paper.