DeepSeek V3 vs V4 Architecture Infographic
side-by-side AI architecture comparison infographic · clean technical diagram, white background, thin black outlines, rounded rectangles, dashed callout boxes, color-coded highlights, presentation-slide aesthetic, vector infographic · 2:1 · wide horizontal · DeepSeek V3/R1 (671 billion) · DeepSeek V4 (1.2 trillion) · bright orange-red · bright blue · 2 · DeepSeek V3/R1 (671 billion) · left half · 9 · Vocabulary size of 129k, FeedForward (SwiGLU) module, Intermediate hidden layer dimension of 2,048, MoE layer, Supported context length of 128k tokens, First 3 blocks use dense FFN with hidden size 18,432 instead of MoE, Sample input text, Embedding dimension of 7,168, 128 heads · DeepSeek V4 (1.2 trillion) · right half · 9 · Vocabulary size of 160k, FeedForward (SwiGLU) module, Intermediate hidden layer dimension of 3,072, MoE layer, Supported context length of 256k tokens, First 3 blocks use dense FFN with hidden size 24,576 instead of MoE, Sample input text, Embedding dimension of 8,192, 128 heads · bottom comparison table · bottom full width · 10 · Total parameters, Active parameters per token, Hidden size, Esmple dimesiegn, DeepSeek V3/R1, Intermediate (FF), Attention heads, Context length, Embedding dimension, Vocabulary size · very light gray rounded rectangle · 8 · Tokenized text, Token embedding layer, RMSNorm 1, Multi-head Latent Attention, RMSNorm 2, MoE, Final RMSNorm, Linear output layer · RoPE attached to the attention block on the left side · Multi-head Latent Attention · orange-red text for the word Latent · FeedForward (SwiGLU) module · 4 · Linear layer, SiLU activation, Linear layer, Linear layer · two branches multiplied, then projected · MoE layer · 5 · top combine node, Feed forward, Feed forward, Router, expert count badge 256 · small black square with 1 selected expert, arrows routing upward to experts, dotted divider line · Vocabulary size of 129k · Intermediate hidden layer dimension of 2,048 · Supported context length of 128k tokens · First 3 blocks use dense FFN with hidden size 18,432 instead of MoE · Resource savings: Model size is 671B but only 1 (shared) + 8 experts active per token; only 37B parameters are active per inference step · 10 · Total parameters: 671B, Active parameters per token: 37B (1 + 8 experts), Hidden size: 7,128, Esmple dimesiegn: 28,432, Intermediate (FF): 2,048, Attention heads: 128, Context length: 128k, Embedding dimension: First 3 blocks, Context ler length: 22G7, Vocabulary size: 129k · very light blue rounded rectangle · 8 · Tokenized text, Token embedding layer, RMSNorm 1, Multi-head Latent Attention, RMSNorm 2, MoE, Final RMSNorm, Linear output layer · RoPE attached to the attention block on the left side · Multi-head Latent Attention · blue text for the word Latent · FeedForward (SwiGLU) module · 4 · Linear layer, SiLU activation, Linear layer, Linear layer · same structure as left panel · MoE layer · 5 · top combine node, Feed forward, Feed forward, Router, expert count badge 384 · small black square with 1 selected expert, arrows routing upward to experts, dotted divider line, blue border emphasis · Vocabulary size of 160k · Intermediate hidden layer dimension of 3,072 · Supported context length of 256k tokens · First 3 blocks use dense FFN with hidden size 24,576 instead of MoE · Resource savings: Model size is 1.2T but only 1 (shared) + 8 experts active per token; only 52B parameters are active per inference step · 10 · Total parameters: 1.2T, Active parameters per token: 52B (1 + 8 experts), Hidden size: 7,2B, Esmple dimesiegn: 28,432, Intermediate (FF): 3,072, Attention heads: 128, Context length: 256k, Embedding dimension: First 3 blocks, Context ler length: 22G7, Vocabulary size: 160k · Create a highly detailed transformer architecture comparison diagram with mirrored layouts. Each half contains one large model stack diagram plus 2 inset diagrams: 1 feedforward module and 1 MoE layer. Use arrows between blocks, tiny technical labels, and connector lines from labels to the relevant components. Keep the typography dense and slide-like, with orange-red used for all V3/R1 emphasis and blue used for all V4 emphasis. Include a small bottom row of compact tabular metrics spanning the width. Preserve the slightly imperfect, human-made infographic look with very small text and crowded annotations.