🖼️AI Image
GPT-Image-2 Case Studies · Benchmark Comparison Heatmap
Selected GPT-Image-2 reference cases from “Research Paper Illustrations,” suitable for retrieving generative effects in the direction of “Benchmark Comparison Heatmaps.”
Author: AI Plus Lab
✦Results
◌Case Background
来自 GPT-Image2-Skill README 的精选展示条目,适合作为“研究论文图示 / 基准对比热图”方向的站内参考案例。
⌘Prompt Content
Landscape 16:9 model × Benchmark heat matrix. Columns (rotated 45°): "MMLU", "HumanEval", "GSM8K", "MATH", "BBH", "ARC-C", "HellaSwag", "TruthfulQA". Rows (right-aligned sans-serif font): "GPT-4o", "Claude 4.7 Opus", "Gemini 3 Pro", "Llama 4 405B", "Qwen3-Next", "DeepSeek-V3.1", "Mistral-3 Large", "Yi-3 34B", "Phi-4 14B", "OLMo-2 7B". Each cell is filled with a dusty turquoise gradient, with color intensity reflecting scores; scores displayed within each cell (e.g., "72.3", "88.1"). The best score in each column outlined with a 1.5px soft terracotta stroke. Vertical color bar on the right, marked with scales "0", "25", "50", "75", "100" and labeled "accuracy (%)". Title: "Benchmark comparison across 10 frontier LLMs". Subtitle: "Zero-shot accuracy; best score for each benchmark highlighted with bold outline. Evaluation date: March 2026."
✎Outcome Notes
This case has been migrated from the GPT-Image2-Skill README to AIPlusLab `/prompts` for easy on-site retrieval, browsing, and reuse.
