영상 이미지
로그인

DeepSeek V3 vs V4 Architecture Infographic

A dense side-by-side technical infographic comparing DeepSeek V3/R1 and DeepSeek V4 transformer architectures, suitable for social media posts, presentations, or model analysis visuals.

작성자 Sigrid Jin 🌈🙏 · gpt-image-2 · 1,600 조회 · 25 좋아요 · 5 북마크
DeepSeek V3 vs V4 Architecture Infographic

프롬프트

{"type":"사이드 바이 사이드 AI 아키텍처 비교 인포그래픽","style":"깔끔한 기술 다이어그램, 흰색 배경, 얇은 검은색 외곽선, 둥근 사각형, 점선 콜아웃 박스, 색상 구분 하이라이트, 프레젠테이션 Slides 미학, 벡터 인포그래픽","canvas":{"aspect_ratio":"2:1","resolution":"와이드 가로형"},"title_row":{"left_title":"DeepSeek V3/R1 (6,710 억)","right_title":"DeepSeek V4 (1.2 조)","left_title_color":"밝은 주황-빨간색","right_title_color":"밝은 파란색"},"layout":{"columns":2,"sections":[{"title":"DeepSeek V3/R1 (6,710 억)","position":"왼쪽 절반","count":9,"labels":["어휘 사전 크기 129k","FeedForward (SwiGLU) 모듈","중간 은닉층 차원 2,048","MoE 레이어","지원 컨텍스트 길이 128k 토큰","첫 3개 블록은 MoE 대신 은닉 크기 18,432의 밀집 FFN 사용","샘플 입력 텍스트","임베딩 차원 7,168","128개 헤드"]},{"title":"DeepSeek V4 (1.2 조)","position":"오른쪽 절반","count":9,"labels":["어휘 사전 크기 160k","FeedForward (SwiGLU) 모듈","중간 은닉층 차원 3,072","MoE 레이어","지원 컨텍스트 길이 256k 토큰","첫 3개 블록은 MoE 대신 은닉 크기 24,576의 밀집 FFN 사용","샘플 입력 텍스트","임베딩 차원 8,192","128개 헤드"]},{"title":"하단 비교 표","position":"하단 전체 너비","count":10,"labels":["총 파라미터","토큰당 활성 파라미터","은닉 크기","샘플 설계","DeepSeek V3/R1","중간 (FF)","어텐션 헤드","컨텍스트 길이","임베딩 차원","어휘 사전 크기"]}]},"left_panel":{"background":"매우 밝은 회색 둥근 사각형","main_stack":{"count":8,"blocks":["토큰화된 텍스트","토큰 임베딩 레이어","RMSNorm 1","Multi-head Latent Attention","RMSNorm 2","MoE","최종 RMSNorm","선형 출력 레이어"]},"side_module":"왼쪽 어텐션 블록에 부착된 RoPE","attention_block":{"label":"Multi-head Latent Attention","accent":"Latent 단어에 주황-빨간색 텍스트"},"feedforward_inset":{"title":"FeedForward (SwiGLU) 모듈","count":4,"blocks":["선형 레이어","SiLU 활성화","선형 레이어","선형 레이어"],"diagram":"두 가지 분기 곱셈 후 투영"},"moe_inset":{"title":"MoE 레이어","count":5,"blocks":["상단 결합 노드","Feed forward","Feed forward","라우터","전문가 수 배지 256"],"details":"선택된 전문가 1개가 포함된 작은 검은색 사각형, 전문가 방향으로 향하는 화살표, 점선 구분선"},"annotations":{"vocab":"어휘 사전 크기 129k","ff_dim":"중간 은닉층 차원 2,048","context":"지원 컨텍스트 길이 128k 토큰","dense_first_blocks":"첫 3개 블록은 MoE 대신 은닉 크기 18,432의 밀집 FFN 사용","resource_savings":"자원 절감: 모델 크기는 671B이나 토큰당 1개(공유) + 8개의 전문가만 활성화; 추론 단계당 37B 파라미터만 활성화"},"bottom_stats":{"count":10,"items":["총 파라미터: 671B","토큰당 활성 파라미터: 37B (1 + 8 전문가)","은닉 크기: 7,128","샘플 설계: 28,432","중간 (FF): 2,048","어텐션 헤드: 128","컨텍스트 길이: 128k","임베딩 차원: 첫 3개 블록","컨텍스트 길이: 22G7","어휘 사전 크기: 129k"]}},"right_panel":{"background":"매우 밝은 파란색 둥근 사각형","main_stack":{"count":8,"blocks":["토큰화된 텍스트","토큰 임베딩 레이어","RMSNorm 1","Multi-head Latent Attention","RMSNorm 2","MoE","최종 RMSNorm","선형 출력 레이어"]},"side_module":"왼쪽 어텐션 블록에 부착된 RoPE","attention_block":{"label":"Multi-head Latent Attention","accent":"Latent 단어에 파란색 텍스트"},"feedforward_inset":{"title":"FeedForward (SwiGLU) 모듈","count":4,"blocks":["선형 레이어","SiLU 활성화","선형 레이어","선형 레이어"],"diagram":"왼쪽 패널과 동일한 구조"},"moe_inset":{"title":"MoE 레이어","count":5,"blocks":["상단 결합 노드","Feed forward","Feed forward","라우터","전문가 수 배지 384"],"details":"선택된 전문가 1개가 포함된 작은 검은색 사각형, 전문가 방향으로 향하는 화살표, 점선 구분선, 파란색 테두리 강조"},"annotations":{"vocab":"어휘 사전 크기 160k","ff_dim":"중간 은닉층 차원 3,072","context":"지원 컨텍스트 길이 256k 토큰","dense_first_blocks":"첫 3개 블록은 MoE 대신 은닉 크기 24,576의 밀집 FFN 사용","resource_savings":"자원 절감: 모델 크기는 1.2T이나 토큰당 1개(공유) + 8개의 전문가만 활성화; 추론 단계당 52B 파라미터만 활성화"},"bottom_stats":{"count":10,"items":["총 파라미터: 1.2T","토큰당 활성 파라미터: 52B (1 + 8 전문가)","은닉 크기: 7,2B","샘플 설계: 28,432","중간 (FF): 3,072","어텐션 헤드: 128","컨텍스트 길이: 256k","임베딩 차원: 첫 3개 블록","컨텍스트 길이: 22G7","어휘 사전 크기: 160k"]}},"global_notes":"미러링된 레이아웃을 갖춘 매우 상세한 트랜스포머 아키텍처 비교 다이어그램을 생성하세요. 각 절반에는 하나의 큰 모델 스택 다이어그램과 2개의 인셋 다이어그램(Feedforward 모듈 1개, MoE 레이어 1개)이 포함됩니다. 블록 간 화살표, 작은 기술 라벨, 라벨에서 관련 구성 요소로 연결되는 선을 사용하세요. 타이포그래피는 밀도 높고 슬라이드와 같은 느낌을 유지하며, V3/R1 강조에는 주황-빨간색을, V4 강조에는 파란색을 사용하세요. 너비 전체에 걸쳐 작은 하단 행의 간결한 표 형식 지표를 포함하세요. 매우 작은 텍스트와 빽빽한 주석으로 약간 불완전하면서도 사람이 만든 듯한 인포그래픽 느낌을 유지하세요."}

기계 번역 — 실제 결과물을 만든 것은 원문 프롬프트입니다.

원문 프롬프트
{"type":"side-by-side AI architecture comparison infographic","style":"clean technical diagram, white background, thin black outlines, rounded rectangles, dashed callout boxes, color-coded highlights, presentation-slide aesthetic, vector infographic","canvas":{"aspect_ratio":"2:1","resolution":"wide horizontal"},"title_row":{"left_title":"DeepSeek V3/R1 (671 billion)","right_title":"DeepSeek V4 (1.2 trillion)","left_title_color":"bright orange-red","right_title_color":"bright blue"},"layout":{"columns":2,"sections":[{"title":"DeepSeek V3/R1 (671 billion)","position":"left half","count":9,"labels":["Vocabulary size of 129k","FeedForward (SwiGLU) module","Intermediate hidden layer dimension of 2,048","MoE layer","Supported context length of 128k tokens","First 3 blocks use dense FFN with hidden size 18,432 instead of MoE","Sample input text","Embedding dimension of 7,168","128 heads"]},{"title":"DeepSeek V4 (1.2 trillion)","position":"right half","count":9,"labels":["Vocabulary size of 160k","FeedForward (SwiGLU) module","Intermediate hidden layer dimension of 3,072","MoE layer","Supported context length of 256k tokens","First 3 blocks use dense FFN with hidden size 24,576 instead of MoE","Sample input text","Embedding dimension of 8,192","128 heads"]},{"title":"bottom comparison table","position":"bottom full width","count":10,"labels":["Total parameters","Active parameters per token","Hidden size","Esmple dimesiegn","DeepSeek V3/R1","Intermediate (FF)","Attention heads","Context length","Embedding dimension","Vocabulary size"]}]},"left_panel":{"background":"very light gray rounded rectangle","main_stack":{"count":8,"blocks":["Tokenized text","Token embedding layer","RMSNorm 1","Multi-head Latent Attention","RMSNorm 2","MoE","Final RMSNorm","Linear output layer"]},"side_module":"RoPE attached to the attention block on the left side","attention_block":{"label":"Multi-head Latent Attention","accent":"orange-red text for the word Latent"},"feedforward_inset":{"title":"FeedForward (SwiGLU) module","count":4,"blocks":["Linear layer","SiLU activation","Linear layer","Linear layer"],"diagram":"two branches multiplied, then projected"},"moe_inset":{"title":"MoE layer","count":5,"blocks":["top combine node","Feed forward","Feed forward","Router","expert count badge 256"],"details":"small black square with 1 selected expert, arrows routing upward to experts, dotted divider line"},"annotations":{"vocab":"Vocabulary size of 129k","ff_dim":"Intermediate hidden layer dimension of 2,048","context":"Supported context length of 128k tokens","dense_first_blocks":"First 3 blocks use dense FFN with hidden size 18,432 instead of MoE","resource_savings":"Resource savings: Model size is 671B but only 1 (shared) + 8 experts active per token; only 37B parameters are active per inference step"},"bottom_stats":{"count":10,"items":["Total parameters: 671B","Active parameters per token: 37B (1 + 8 experts)","Hidden size: 7,128","Esmple dimesiegn: 28,432","Intermediate (FF): 2,048","Attention heads: 128","Context length: 128k","Embedding dimension: First 3 blocks","Context ler length: 22G7","Vocabulary size: 129k"]}},"right_panel":{"background":"very light blue rounded rectangle","main_stack":{"count":8,"blocks":["Tokenized text","Token embedding layer","RMSNorm 1","Multi-head Latent Attention","RMSNorm 2","MoE","Final RMSNorm","Linear output layer"]},"side_module":"RoPE attached to the attention block on the left side","attention_block":{"label":"Multi-head Latent Attention","accent":"blue text for the word Latent"},"feedforward_inset":{"title":"FeedForward (SwiGLU) module","count":4,"blocks":["Linear layer","SiLU activation","Linear layer","Linear layer"],"diagram":"same structure as left panel"},"moe_inset":{"title":"MoE layer","count":5,"blocks":["top combine node","Feed forward","Feed forward","Router","expert count badge 384"],"details":"small black square with 1 selected expert, arrows routing upward to experts, dotted divider line, blue border emphasis"},"annotations":{"vocab":"Vocabulary size of 160k","ff_dim":"Intermediate hidden layer dimension of 3,072","context":"Supported context length of 256k tokens","dense_first_blocks":"First 3 blocks use dense FFN with hidden size 24,576 instead of MoE","resource_savings":"Resource savings: Model size is 1.2T but only 1 (shared) + 8 experts active per token; only 52B parameters are active per inference step"},"bottom_stats":{"count":10,"items":["Total parameters: 1.2T","Active parameters per token: 52B (1 + 8 experts)","Hidden size: 7,2B","Esmple dimesiegn: 28,432","Intermediate (FF): 3,072","Attention heads: 128","Context length: 256k","Embedding dimension: First 3 blocks","Context ler length: 22G7","Vocabulary size: 160k"]}},"global_notes":"Create a highly detailed transformer architecture comparison diagram with mirrored layouts. Each half contains one large model stack diagram plus 2 inset diagrams: 1 feedforward module and 1 MoE layer. Use arrows between blocks, tiny technical labels, and connector lines from labels to the relevant components. Keep the typography dense and slide-like, with orange-red used for all V3/R1 emphasis and blue used for all V4 emphasis. Include a small bottom row of compact tabular metrics spanning the width. Preserve the slightly imperfect, human-made infographic look with very small text and crowded annotations."}

✨ 생성 저장 원본 게시물 Hitboard에서 열기

비슷한 핀