🖼️AI Image
GPT-Image-2 Case Study · LLM Pretraining Data Mixture Sankey Diagram
A featured showcase of GPT-Image-2 reference cases from the 'Research Paper Diagrams' collection, suitable for exploring generation effects in the 'LLM Pretraining Data Mixture Sankey Diagram' direction.
Author: AI Plus Lab
✦Results
◌Case Background
来自 GPT-Image2-Skill README 的精选展示条目,适合作为“研究论文图示 / LLM 预训练数据混合桑基图”方向的站内参考案例。
⌘Prompt Content
Horizontal 16:9 Sankey diagram, showcasing pretraining data mixture, with three phases and transparent ribbons.
Left (8 source blocks, height proportional to token count): "Common Crawl (web) 540B" (soft navy blue, largest), "arXiv papers 180B" (dusty turquoise), "GitHub code 160B" (slate gray), "Wikipedia 40B" (soft terra cotta), "StackExchange QA 30B" (warm copper), "Books (public domain) 25B" (light olive), "Patents 18B" (light navy blue), "Curated news & forums 15B" (dusty turquoise).
Middle (3 processing blocks, stacked): "Deduplicated (MinHash + exact)", "Quality-filtered (classifier + heuristics)", "PII-scrubbed (regex + NER)".
Right (3 final splits): "Pretraining set 1.4T tokens" (largest), "Instruction-tune pool 12B tokens", "RLHF preference pool 3B tokens".
Flowing ribbons inherit source colors, middle labels show token counts ("85B", "320B", "44B"). Legend bar at the bottom.
Title: "LLM pretraining data mixture and downstream splits". Subtitle: "Token counts after deduplication and quality filtering; ribbon thickness ∝ token flow".✎Outcome Notes
This case has been migrated from the GPT-Image2-Skill README to AIPlusLab `/prompts`, allowing for direct search, viewing, and reuse within the site.
