The New American AI Model Designed to be Customized
This section explains four foundational concepts essential to understanding language model architecture: tokens, parameters, training, and layers. Tokens serve as the granular text chunks processed by models, while parameters are the internal numerical values iteratively refined during training. Training optimizes these parameters via next-token prediction, gradients, and backpropagation across vast quantities of text. Finally, inputs pass sequentially through stacked layers—such as the 66 layers in the Inkling model—where each layer executes an attention step followed by a feed-forward step. Inkling uses a Mixture of Experts (MoE) architecture to decouple the cost of storing a model from the cost of running it. Instead of passing tokens through every feed-forward network, Inkling selects 6 out of 256 experts per token, following an approach published by DeepSeek. Across the full 975-billion-parameter model, only about 41 billion parameters (roughly 4 percent) are active during the processing of a single token. A full-precision checkpoint requires at least 2 TB of GPU memory across specialized hardware like NVIDIA B300 or H200 GPUs, while a quantized checkpoint reduces memory requirements to roughly 600 GB across four B300 cards. In Mixture of Experts architectures, routers score candidate experts to select a top subset for processing each token, but positive feedback loops during training can cause routing collapse where only a handful of experts are utilized. Traditional auxiliary balance loss penalties degrade output quality by generating gradients that conflict with the primary prediction task. Thinking Machines resolves this by adopting an approach originally introduced by Wang and colleagues and utilized by DeepSeek, which adjusts expert selection scores using non-backpropagated bias values updated by token counts. This keeps expert load balanced without compromising gradient updates, while two additional shared experts run unconditionally alongside the six routed experts.
閱讀原文 ↗目錄
Groundwork
This section explains four foundational concepts essential to understanding language model architecture: tokens, parameters, training, and layers. Tokens serve as the granular text chunks processed by models, while parameters are the internal numerical values iteratively refined during training. Training optimizes these parameters via next-token prediction, gradients, and backpropagation across vast quantities of text. Finally, inputs pass sequentially through stacked layers—such as the 66 layers in the Inkling model—where each layer executes an attention step followed by a feed-forward step.
- Tokens are text chunks typically smaller than sentences and often smaller than words, with uncommon words split into multiple tokens.
- Parameters are individual learned numerical values that begin as random noise and are updated millions of times during training.
- Training adjusts parameters based on next-token prediction errors using gradients computed simultaneously through backpropagation.
- Layers are sequential processing stages where each layer contains both an attention step and a feed-forward step.
- The Inkling model contains 66 stacked layers through which token representations pass sequentially.
Sparsity
Inkling uses a Mixture of Experts (MoE) architecture to decouple the cost of storing a model from the cost of running it. Instead of passing tokens through every feed-forward network, Inkling selects 6 out of 256 experts per token, following an approach published by DeepSeek. Across the full 975-billion-parameter model, only about 41 billion parameters (roughly 4 percent) are active during the processing of a single token. A full-precision checkpoint requires at least 2 TB of GPU memory across specialized hardware like NVIDIA B300 or H200 GPUs, while a quantized checkpoint reduces memory requirements to roughly 600 GB across four B300 cards.
- Inkling replaces the standard single feed-forward network per layer with 256 experts, running only six per token.
- Inkling contains 975 billion total parameters, but only activates roughly 41 billion parameters (~4%) per token.
- Thinking Machines implemented Inkling's Mixture of Experts design largely based on an approach published by DeepSeek.
- Inkling's full-precision checkpoint requires at least 2 TB of combined GPU memory, runnable on 8 NVIDIA B300 cards or 16 NVIDIA H200 cards.
- A quantized checkpoint version of Inkling requires around 600 GB of memory and fits onto four NVIDIA B300 cards.
Routing
In Mixture of Experts architectures, routers score candidate experts to select a top subset for processing each token, but positive feedback loops during training can cause routing collapse where only a handful of experts are utilized. Traditional auxiliary balance loss penalties degrade output quality by generating gradients that conflict with the primary prediction task. Thinking Machines resolves this by adopting an approach originally introduced by Wang and colleagues and utilized by DeepSeek, which adjusts expert selection scores using non-backpropagated bias values updated by token counts. This keeps expert load balanced without compromising gradient updates, while two additional shared experts run unconditionally alongside the six routed experts.
- Inkling's router scores 256 experts using a sigmoid function and picks the top six to run for each token.
- Unmitigated expert selection can lead to routing collapse, where a small fraction of experts dominate compute while the rest stay underdeveloped.
- Traditional load-balancing penalties create competing gradients against the next-token prediction objective, often degrading text quality.
- Thinking Machines and DeepSeek use a technique by Wang et al. that applies an external bias to expert selection scores without affecting output weighting or backpropagation.
- Every token is processed by eight total experts: six routed experts and two shared experts, whose scores are normalized together.
Attention
Standard attention computation scales quadratically with sequence length, making a one-million-token context window computationally prohibitive across 66 layers. To address this, Inkling alternates between sliding-window attention layers and full-attention layers in a 5:1 ratio, using 55 sliding-window and 11 full-attention layers. Long-range information flows through the full-attention layers, while sliding-window layers handle local context at lower computational cost. Inkling also utilizes eight key-value heads to reduce memory overhead during generation.
- Standard attention comparisons scale quadratically with sequence length, reaching roughly one trillion comparisons per layer for one million tokens.
- Inkling supports a context window of one million tokens across 66 total layers.
- Inkling alternates sliding-window layers and full-attention layers at a 5:1 ratio, comprising 55 sliding-window layers and 11 full-attention layers according to vLLM project notes.
- Long-range dependencies are propagated through the 11 full-attention layers placed at regular intervals such as layers 6, 12, and 18.
- Inkling employs 8 key-value heads to reduce memory consumption during text generation.
Position
Attention mechanisms lack inherent sequence order information, necessitating positional representation methods like Rotary Position Embedding (RoPE). However, RoPE can struggle to extrapolate to context lengths far exceeding those seen during training because novel rotation angles fall outside the model's learned range. To address this for long sequences, Thinking Machines chose an older relative positional encoding scheme in the style of Shaw et al. for the Inkling model. This approach adds learned distance-based bias values directly to attention scores and clamps distances past a specific cutoff, avoiding the need to extrapolate unfamiliar values.
- Attention query and key comparisons contain no intrinsic order information, requiring explicit positional encoding.
- Rotary Position Embedding (RoPE) works by rotating queries and keys by angles proportional to sequence position, which can fail when extrapolating beyond training context windows.
- The Inkling model utilizes a relative position scheme inspired by Shaw and colleagues instead of the modern standard RoPE.
- The relative positioning method adds a learned value corresponding to token distance directly to attention scores, mapping all distances past a fixed cutoff to a single shared parameter.
- Thinking Machines observed that this relative scheme both performed better and extrapolated to longer contexts more effectively than RoPE.
Convolutions
Inkling incorporates a localized convolution operation with a sliding window of four tokens to handle local sequence mixing. Because standard attention mechanisms have no built-in bias toward adjacent tokens, convolutions supply this local context structurally using learned weights. Thinking Machines inserts these convolutions onto the keys and values within attention layers as well as onto attention and feed-forward outputs before they rejoin the main residual path. This enables attention to reserve its learning capacity for complex dependencies rather than basic local proximity.
- Inkling utilizes a convolution with a window of four positions to mix a token with its three preceding neighbors using learned weights.
- Convolutions provide structural local mixing without requiring attention to learn that adjacent tokens are related.
- Thinking Machines places convolutions on the keys and values inside each attention layer.
- Convolutions are applied to the outputs of attention and feed-forward blocks before rejoining the main model stream.
- Pre-mixing immediate context using convolutions frees attention capacity to focus on dependencies that genuinely require learning.
Multimodality
Inkling processes images and audio directly without relying on separately pretrained encoder networks. Audio is represented as mel spectrograms and quantized via the dMel method, which requires no pretraining. Images are divided into 40x40 pixel patches and processed independently through a four-stage hMLP stem with minimal compute overhead. Both modalities pass through a lightweight conversion layer and merge into a single sequence alongside text tokens, with all multimodal components trained from scratch alongside the base model.
- Inkling eschews separately pretrained encoders for vision and audio in favor of an end-to-end architecture trained from scratch.
- Audio processing uses the dMel method to round mel spectrogram loudness values into fixed levels without requiring prior training.
- Images are segmented into 40x40 pixel patches and fed into an hMLP stem, adding less than one percent to compute overhead.
- Multimodal inputs pass through a lightweight conversion layer to join text tokens across the model's 66 layers.
- Thinking Machines first implemented this multimodal approach two months prior to Inkling in a real-time interaction system.
Effort
Thinking Machines introduced an adjustable effort setting between 0 and 1 to control how long the Inkling model reasons before answering. Rather than enforcing a hard token budget, this behavior was trained into Inkling via reinforcement learning by varying token-cost penalties based on the effort prompt. Mechanically, the effort level is inserted as a system message before the conversation starts, independent of the maximum token limit. On the Terminal Bench 2.1 coding benchmark, Inkling achieves the same performance as NVIDIA's Nemotron 3 Ultra using approximately one-third of the tokens.
- Effort is parameterized as a value from 0 to 1, with documented presets including 0, 0.1, 0.2, 0.7, 0.9 (default), and 0.99.
- The effort level is passed mechanically as an initial system message prior to the conversation.
- The effort response was trained via reinforcement learning by varying per-token penalty costs corresponding to different effort labels.
- Effort operates independently of the maximum token limit and encourages longer reasoning without offering hard guarantees on length or quality.
- On Terminal Bench 2.1, Inkling matches the score of NVIDIA's Nemotron 3 Ultra while generating roughly a third of the tokens.
Tradeoffs
The architectural choices behind Inkling involve explicit trade-offs between computational efficiency, hardware requirements, and model behavior. Sparse routing and mostly-local attention lower per-token compute costs and enable large context windows, but they enforce high memory footprints and limited layer context. Furthermore, adopting relative position encoding broke compatibility with the prevailing RoPE-based serving infrastructure. While Thinking Machines acknowledges that other models surpass Inkling in baseline strength, Inkling is optimized for direct adaptation and fine-tuning via day-one quantized checkpoints.
- Sparse routing reduces per-token inference cost but keeps the minimum memory footprint high because the full model must be loaded.
- Mostly-local attention facilitates a 1-million-token context window at the cost of providing most layers with a restricted sequence perspective.
- Adopting relative position encoding solved extrapolation issues but required engineering new support in serving frameworks configured for RoPE.
- Open weights are provided without training datasets, code, or exact training recipes, limiting auditing and full reproducibility.
- Thinking Machines recommends employing external moderation tools for consumer deployments rather than relying on built-in model refusals.
- Inkling is specifically targeted at customizable enterprise use cases through early fine-tuning support and quantized checkpoints.
Conclusion
Inkling's architecture separates total parameter storage from per-token compute, storing 975 billion parameters while running approximately 41 billion active parameters. The model balances expert routing by applying a selection bias without including it in output weighting, avoiding conflicting training objectives. Its long-context design utilizes 55 sliding-window layers alongside 11 full-attention layers paired with relative position representations. Furthermore, reasoning effort is formulated as a trained setting, enabling benchmark evaluations to function across an adjustable performance curve.
- Inkling stores 975 billion total parameters while activating approximately 41 billion per token.
- Expert load balancing applies bias during selection but omits it during output weighting to avoid auxiliary training conflicts.
- The model handles long contexts using 55 sliding-window layers and 11 full-attention layers for long-range information propagation.
- Relative position encoding allows tokens separated by large distances to rely on frequently observed distance values.
- Reasoning effort can be configured as a trained setting, producing a variable benchmark curve rather than a fixed score.