arXiv:2606.30561v1 [cs.AI] 29 Jun 2026
The Human Creativity Benchmark Aspen Hopkins
Allison Nulty
Alexandria Minetti
Contra; Massachusetts Institute of Technology New York, NY; Cambridge, MA, USA
Contra New York, New York, USA
Contra New York, New York, USA
Anoop Pakki
Angad Singh
Contra New York, New York, USA
Contra New York, New York, USA
The Creative Process
Mockup
Ideation
Refinement
Figure 1: The creative process as a sideways martini: broad ideation narrows through mockup, ending in refinement. Example stills taken from Ad images rendered in response to prompts; prompts were derived from professional samples.
Abstract
Collapsing these signals into a single quality metric discards the most actionable information: where models must be correct versus where they should remain steerable.
Modern AI evaluation frameworks treat evaluator disagreement as noise to be resolved. In creative domains, professional disagreement reflects genuine differences in taste, not measurement error. We argue that evaluating creative AI requires preserving two distinct signals: convergence, where professionals align around shared best practices, and divergence, where individual taste legitimately varies. We present the Human Creativity Benchmark (HCB), a benchmark that operationalizes this separation by collecting pairwise preferences, scalar ratings on prompt adherence, usability, and visual appeal, and qualitative rationale from domain professionals. Across 15,000 professional judgments spanning five creative domains and three workflow phases (ideation, mockup, refinement), we find that convergence concentrates on verifiable dimensions like technical correctness and visual hierarchy, while divergence concentrates on taste-driven dimensions like aesthetic direction and conceptual risk. No model excels uniformly across all phases.
Keywords Generative AI, Creative Evaluation, Human-AI Collaboration, Benchmarking, Design Workflows, Professional Creativity
1
Introduction
Creativity has long resisted reduction to any single, stable operational definition. This has made measuring creativity challenging. Creative work, which connects creative process and domain expertise with expected outcomes, is similarly heterogeneous and remains difficult to study. This continues in the age of AI. A designer’s output reflects not only technical competence but aesthetic judgment accumulated over years of practice. When professional creatives evaluate AI-generated work, they do so across 1
The Human Creativity Benchmark
potentially many axes. For some axes, evaluators agree on what diverge along subjective dimensions of creative quality, and works—readable typography, functional layout, correct visual hierarchy— how these patterns shift across the creative workflow. revealing shared professional standards. For others, evaluators legitimately disagree: not because of measurement error, but because 2 Related Work the criteria are personal—taste, aesthetic direction, or creative intent (Figure 2). AI benchmarks treat this disagreement as noise to Our interest in understanding the intersection of creative work be resolved. and AI begins with the premise that creativity is heterogeneous, In this work, we argue that professional disagreement is not task-dependent, and unevenly expressed across stages of work. always noise to be resolved in creative domains. While some aspects Creative producers may view aesthetic value and commercial value of creative quality are objectively verifiable, others can and should as harmonious or contradictory [5]. They combine multiple modes be treated as inherently subjective. Our goal in creative evaluation of cognition in their process of production, and their perceptions of is thus to not only ask if an output is “good,” but good according to identity, such as genre, status, intellect, and profession shape their whom, for what purpose, and at what stage of the creative process? outputs. Broadly, creatives balance what is good (best practices) To this end, we contribute a framework that distinguishes conwith something more personal, fundamental: taste. vergence and divergence in creative responses as a way to preserve As an array of AI-driven tools supporting creative work has both kinds of judgment.1 Convergence describes dimensions where emerged, however, creatives seeking to improve their workflows, evaluators align around shared, checkable criteria; divergence deexpand the initial generative exploration, or save time and accumuscribes dimensions where multiple creative judgments can be valid lative costs face frictions in systems that do not account for these at once. We operationalize this distinction across three evaluation otherwise distinct axes. Perhaps most resonantly is the challenge axes, Prompt Adherence, Usability, and Visual Appeal, which vary that AI systems are imperfect samplers that tend to converge to an in the extent to which expert judgments might converge around average. This leads to homogenization [2] and mode collapse [6], shared criteria or diverge with individual preference. where models increasingly produce similar outputs—characteristics To test our framework, we introduce The Human Creativity antithetical to the creative process. People spend years exploring, Benchmark (HCB), an exemplar dataset developed from an exrepeating, and improving their points of view, styles, and mediums. ploratory survey on how domain experts use AI in creative work. Converging their outputs through homogenization risks collapsing The HCB is designed to mirror real creative workflows across the plurality of creative practice, narrowing not only the range of Ideation, Mockup, and Refinement stages (Figure 1) in five domains: outputs but the points of view available to creative fields. landing pages, desktop apps, ad images, brand images, and product videos (Table 1). 2.1 Evaluating AI in Creative Processes Domain experts drawn from Contra, a network of independent professional creatives, evaluated the HCB’s generated outputs. Existing norms of AI evaluation function largely in settings conTheir roughly 15,000 judgments span pairwise forced-rankings, strained by rules of what is correct. Benchmarks are generally scalar ratings on prompt adherence, usability, and visual appeal, built for domains where pre-specified heuristics, best practices and open-ended qualitative explanations. Through our analysis, we or ground-truths, equate metrics of “truthfulness” or “correctness” find evidence that creative AI outputs cannot be reduced to a single with “best”. These methods of evaluation do not adequately account judgement of “good” or “bad.” Instead, what is defined as valuable for diversity as an outcome and process value [5]: in creative work, varies by domain, by evaluation criterion, and in some cases, by preserving variation across styles, judgments, or workflows may individual preference. Separating these sources of judgment claribe part of what makes a system useful in the first place. fies what model developers should optimize for: models should be A related body of work in data annotation and label adjudicareliably consistent on convergent axes while remaining steerable tion has addressed the broader problem of expert disagreement on divergent axes. in evaluation tasks. Methods such as Dawid-Skene modeling [3], CrowdTruth [4], and perspectivist annotation frameworks [1] treat In sum, we present three contributions: annotator disagreement as informative rather than purely noisy. (1) The Human Creativity Benchmark (HCB), a dataset However, these approaches have been developed primarily for tasks of prompt and sample images for creative work built with with definable ground truth or narrow classification objectives. Creprompts seeded by real creatives’ work artifacts at different ative evaluation presents a distinct challenge: there is no ground steps of the creative workflow, including Ideation, Mockup, truth to approximate, and the dimensions on which experts disand Refinement, along with domain-expert labels and annoagree (aesthetic direction, mood, conceptual risk) are not reducible tations combining pairwise forced-ranking, scalar ratings, to miscalibration or error. The HCB builds on the insight that disand open-ended qualitative responses. agreement can be signal, but applies it to a domain where the (2) A framework differentiating Convergence and Divergence standard resolution strategies like reconciliation, majority vote, in creative outputs across three axes: Visual Appeal, Prompt gold-standard adjudication may be structurally inappropriate. Adherence, and Usability. Prior to designing the benchmark, we conducted a paid forma(3) An analysis of HCB results showing where expert judgtive survey to understand how professional creatives integrate AI ments converge around shared best practices, where they tools into their workflows. The survey was fielded in November 2025 to creatives in the Contra network, and yielded 50 responses 1We use divergence rather than disagreement to avoid treating variation in expert judgment as implicitly negative. spanning design, video, development, and content disciplines. The 2
The Human Creativity Benchmark
Text-to-image / image-to-image Brand Design (Brand Image Assets) Content Designers (Ad Images) gpt-image-1.5 gemini-3-pro-image seedream-4.5 flux-2-pro
Text-to-code / code-to-code
Image-to-video
Product Designers (Desktop Applications) Web Designers (Landing Pages)
Video Editors (Product Video) veo3.1 kling-v3.0-pro seedance-v1.5-pro grok-imagine-video
claude-opus-4.6 gemini-3.1-pro gpt-5.3-codex qwen3.5-397b
Table 1: Creative professional domains, output types, model modalities, and the specific models evaluated in this study. The benchmark spans image, video, and code generation settings, pairing each modality with the creative specialties that commonly produce those deliverables and the frontier models selected for comparison.
Flux-2-pro
GPT-Image-1.5
Divergence
Convergence
“The black and white aesthetic gives the image a more editorial and stylish feel, similar to luxury fashion photography....”
"The product and user are the focal point… resembles a key aspect of the prompt 'sculptural support element, extension of personal lifestyle'[…] It also leaves enough space for overlays."
This concept feels super heavy [….] Some important constraints from the prompt were ignored, that is why it fails to look wearable, and rather looks cheap.
“This feels like a very high-end jewelry piece”
Figure 2: Inter-rater divergence vs. convergence on a single output. Examples of two model outputs for the same creative stage (Ideation) and prompt that illustrate divergence (left, Flux-2-Pro) and convergence (right, GPT-Image-1.5) for Ad Images. Tightly nested polygons indicate reviewer agreement and widely spread polygons indicate disagreement. For Flux-2-Pro, creatives diverge: scores span nearly the full scale on Usability (2 – 5) and Visual Appeal (1 – 4), with the same image read as “editorial and stylish” like luxury fashion photography by one reviewer and as “super heavy. . . rather looks cheap” by another. For GPT-Image-1.5, creatives converge, most strongly on Prompt Adherence (unanimous) and Usability, with consistently favorable comments (“a very high-end jewelry piece; leaves enough space for overlays”). questionnaire was organized in four parts, exploring (1) their creative process, whether they currently use AI tools, and, if not, reasons against adoption; (2) how they incorporated AI, including how much of their creative process is AI-supported, how much of AI-content reaches final deliverables, and the perceived benefits or costs to adoption; (3) eight modules on generative tool categories detailing specific tools and where in their workflow they applied them; and (4) beliefs on AI’s role in the future of creative work, effects on earnings, barriers to adoption, and comfort disclosing AI use to clients. At the time of the survey, the majority of respondents reported increased earnings with AI-use (66% of the respondents), and mixed (skewing positive) attitudes towards the role AI in creative work (AI will enhance creativity (80% of respondents) and create new forms
of creativity (70% of respondents), versus fewer who felt it would replace some creative tasks (50% of respondents) or commoditize creativity (24% of respondents)). Respondents generally adopted a “co-creation” or augmentation perspective. As one brand designer stated, “it’s become a kind of creative partner—helping me get from concept to something visually exciting much quicker,” but, as another put it, “[w]ork with me to bring my ideas to life and help me do more, not take the wheel entirely.” While AI use was not monolithic, those that did use these tools tended to bucket AI use in a sequence of exploratory ideation, prototyping, followed by refining client deliverables.2 Finally, we asked creatives to walk through their prompting process: what output they wanted, how they responded when an 2 Additional details and data overviews are shown in Survey Insights (Appendix F).
3
The Human Creativity Benchmark
3.1
output fell short, and how acceptable outputs were incorporated into their own process or client-facing deliverables. Collectively, this data, along with themes of controllability, co-creation, ideation, and process-specific needs motivated the construction in the HCB, including the three-phase decomposition of the creative process (Ideation, Mockup, Refinement) described below.
3
Evaluation design
Five evaluators per domain completed six tasks per phase (called tournaments), with each tournament comprising two tasks centered on one prompt. Model ordering was randomized and identity anonymized throughout. All evaluations were conducted within the Contra Labs evaluation environment, a web-based interface that walked raters through each phase of the creative process under controlled conditions (Figure 5). The interface presented each tournament as a single guided flow built around one prompt, moving raters sequentially through pairwise comparison tasks, scaledrating tasks, and free-text responses. Additional information and sample interface views are available in the Appendix (Section B). Task 1: Pairwise comparison. Raters were presented with two outputs side-by-side across all possible pairings, producing six pairwise judgments per prompt. Rather than scoring against a predefined rubric, raters selected the output they preferred, isolating the subjective judgment a creative professional would actually apply in practice. After each selection, raters described the rationale in their choice. Pairwise results were aggregated using a Bradley-Terry model to produce ELO ratings for each model. Task 2: Scalar ratings. Three Likert-scale subtasks were chosen to span the convergence–divergence spectrum, helping identify where models should be reliably correct and where they should remain steerable:
Methodology
The study structures creative workflows into three phases, shown in the teaser figure, validated against a prior survey of working creatives: (1) Ideation: Discovery, exploration, and directional potential. At this stage, the creative is not looking for final production quality, but rather for exciting creative direction that is strategically appropriate and worth developing. (2) Mockup: Creative direction has been decided, now it’s time to make the vision come to life. The creative is actualizing the project’s creative direction, creating product shots, stitching together scenes, incorporating brand identity, and bringing the campaign to life. (3) Refinement: Designs are near production-ready, requiring only targeted tweaks to improve consistency across the design.
• Prompt Adherence: How faithful is this output to the given prompt? The least subjective of the three scales, grounded in whether a model did what was asked. • Usability: How well does this output function in the context of the prompt and campaign? This measures whether an output could realistically be used in a professional context. • Visual Appeal: How visually interesting, cohesive, and polished is this output? This dimension targets taste: the aesthetic judgment that distinguishes work a creative would choose rather than merely accept.
All prompts were seeded by expert-produced prompts and media, lightly edited to standardize prompt length and structure. The prompt sequence was designed to mirror a designer’s workflow based on survey responses: Ideation prompts generated new design directions, Mockup prompts used those directions to produce more stable concepts, and Refinement prompts built on the mockups to request specific edits. Participants were drawn from a network of independent professional creatives across design, video, development, and content projects, reflecting common deliverables in independent creative work. Participants were selected based on skillset and the generative model category most relevant to their workflow, then presented with guidelines contextualizing each phase of the creative process and outlining grading criteria for rubric alignment. Selected domains meaningfully represent different evaluation conditions. For example, Ad images produce a single static composition with defined elements like a headline or product image, whereas a landing page is structurally more complex, with elements like layout and design fidelity competing for prioritization. These differences shape evaluator agreement patterns across phases. In total, after dropping incomplete responses, the study was conducted with 28 evaluators from 13 unique countries (Armenia, Belgium, Brazil, Canada, India, Malaysia, Netherlands, Poland, Portugal, Romania, Spain, United Kingdom, and United States) assessing 93 prompts across 80 sessions, yielding 5,940 pairwise judgments, 5,940 scalar ratings, and 3,675 qualitative responses.3
4
Hypotheses
We structure this benchmark and our analyses in response to three hypotheses, motivated in part by our survey responses. Hypothesis #1: Convergence emerges when evaluating verifiable criteria; divergence emerges when tastes are misaligned. While some axes of creative evaluation will reduce to verifiable metrics of performance (convergence), others will be intrinsically based on individual preferences (divergence). In such cases, divergence occurs when outputs are technically acceptable (or uniformly unacceptable) and evaluators can prioritize taste or creative direction. Figure 2 previews this pattern for a single Ad Images prompt: reviewers converge on GPT-Image-1.5 and diverge on Flux-2-Pro across Usability and Visual Appeal. We adopt the scalar axes listed above to assess this hypothesis. We expect Prompt Adherence to illustrate the greatest evaluator convergence because its criteria are checkable, and Visual Appeal to be the most divergent it lends to individual preference and stylistic differences; finally, Usability should fall between the two, combining shared professional standards with context-specific judgment. We treat inter-rater agreement as the
3 The complete dataset, including expert responses, is publicly available on Hugging
Face at https://huggingface.co/datasets/contra-labs/HCB and is released under the cc-by-4.0 license. Details of the data structure can be found in Table 3. 4
The Human Creativity Benchmark
5.1
operational indication of convergence, which should be further validated by context provided in free-text responses.
We first posited that evaluator convergence would depend on the type of criterion being assessed. Creative artifacts combine relatively verifiable standards, such as prompt fidelity, readability, and functional layout, with preference-driven judgments about style, tone, and direction. We therefore expected more agreement on Prompt Adherence, more variation on Visual Appeal, and an intermediate pattern on Usability. The scalar results show that these dimensions are related, but not interchangeable (Table 5). For example, for the refinement stage of Ad Images, Seedream 4.5 and Gemini 3 Image are rated highly on Prompt Adherence (4.00 and 4.03), but diverge in Visual Appeal (3.90 and 2.97). The pattern is not a clean partition of objective and subjective criteria, but it does indicate a consistent trend: the axes capture different kinds of expert judgment rather than three versions of the same quality score. Where criteria were objectively verifiable, such as illegible text or broken visual hierarchy, inter-rater agreement was high. For outputs above this threshold of technical competence, rankings increasingly reflected individual aesthetic preference, consistent with the distinction between convergent and divergent dimensions illustrated in Figure 2. Agreement also differed by rating dimension. Agreement on Prompt Adherence was consistently higher than agreement on Visual Appeal, where criteria are personal rather than shared. Free-text responses contextualize this separation. Evaluators cited failures such as unreadable text, prompt mismatch, or broken visual hierarchy, but they also described taste as important once outputs crossed a threshold of technical competence. One Desktop App Mockup evaluator summarized this directly: “I was judging based off of personal opinion and taste of what looks the best in my eyes,”. In a similar vein, a brand designer stated, in reference to a set of outputs with similar usability, “[h]onestly, I feel like all four images could be used as brand visuals... What made me choose some over others was the sense of life, some felt more dynamic, realistic, and human.” Together, these results support the first hypothesis, showing that while some axes produce creative agreement, others surface preference variation. Additional details can be found in the appendix, including the full scalar table (Table 5) and scalar plots (Figures 14, 15, and 4).
Hypothesis #2: Creative requirements for model outputs will change across the creative stages. Creative work will involve different requirements across stages (Figure 1). During Ideation, we expect creatives will prefer outputs that are more generative; during Mockup, outputs that develop those directions into coherent specifications are rewarded; finally, Refinement evaluations will prioritize targeted correction, consistency, and polish. We thus expect model rankings, scalar labels, and evaluator agreement to shift across stages, as what is useful will change over the course of the workflow. Hypothesis #3: Model comparison should capture convergence and divergence, not just overall quality. If creative quality were onedimensional, model comparison would reduce to a stable ranking of better and worse outputs, e.g., which model produces better outputs on average. We instead expect models to differ in how they satisfy shared standards and support divergent preferences: some models may be more reliable on convergent criteria such as prompt fidelity or production constraints, while others may better support varied creative directions. A convergence–divergence analysis would therefore surface meaningful differences in model behavior that disappear in a single overall ranking.
4.1
Analysis
Pairwise preference data was aggregated using a Bradley-Terry model to produce ELO ratings by domain and phase. Scalar ratings were analyzed across all three dimensions, with Kendall’s W quantifying evaluator agreement at each phase. We additionally computed Krippendorff’s 𝛼 for each model within each domain and phase to assess inter-evaluator reliability and applied the Friedman test independently within each domain and phase to evaluate whether scalar ratings differed significantly among the four competing models. Qualitative feedback was analyzed in several stages. First, we applied light coding to all responses to surface recurring themes and to contextualize the scalar and pairwise comparisons; this pass also produced the codebook used in later analysis. Next, we stripped all qualitative feedback of personally identifiable information and model identifiers and ran the responses through a deductive coding pass using GPT-4o and the predefined codebook, which returned themes, per-theme sentiment, and key quotes, normalizing and parsing the raw text into structured data frames for cross-domain and cross-phase analysis. We then completed additional manual iteration over the raw text to refine and validate key findings.
5
Hypothesis #1: Evaluation axes separate standards from taste
5.2
Hypothesis #2: Creative requirements shift across stages and domains
Our second hypothesis posited that what counts as a “good” output would vary across stages of the creative workflow and across domains. The HCB results support this prediction, but not as a uniform increase or decrease in agreement. Instead, evaluator agreement shifted with the domain, the phase of the workflow, and the dimension being assessed. This is visible in the trajectory of Kendall’s 𝑊 across phases. In the Ad Images domain, agreement increased steadily from Ideation (𝑊 = 0.345) to Mockup (𝑊 = 0.436) to Refinement (𝑊 = 0.549), as feedback narrowed toward more verifiable criteria such as typography, contrast, and production constraints. Landing Pages followed a different trajectory: agreement dropped from 𝑊 = 0.484
Results
Figure 3 shows an aggregate view of the HCB evaluation: each model-domain pair is reduced to a mean scalar rating and an aggregate pairwise win rate. This view is useful, but obfuscates the questions our hypotheses raise—including the axis-level convergence and divergence exemplified in Figure 2. The results below ask whether evaluation axes separate verifiable criteria from tastedriven criteria, whether creative requirements shift across Ideation, Mockup, and Refinement, and whether model comparisons change once axes and stages are kept separate. 5
The Human Creativity Benchmark
A
B
Figure 3: (A) Scalar ratings and pairwise win rates across domains. Each point represents a model within a domain, plotted by its mean scalar quality rating (1–5 scale, horizontal axis) against its pairwise win rate (%, vertical axis) aggregated over all head-to-head comparisons in that domain. The aggregation removes phase- and axis-level differences. (B) Head-to-head pairwise win rates in the Ideation stage for product-video models; each cell is the row model’s win rate against the column model (diagonal self-matches excluded). Grok-Imagine-Video wins 50% against Seedance-v1.5-pro, 47.2% against Kling-v3.0-pro, and 41.7% against Veo3.1. Kling-v3.0-pro wins most against Seedance-v1.5-pro (63.9%), followed by Grok-Imagine-Video (52.8%) and Veo3.1 (36.1%). Seedance-v1.5-pro takes 50% against Grok-Imagine-Video but loses to Veo3.1 (38.9%) and Kling-v3.0-pro (36.1%). Veo3.1 is the clear leader, winning all three of its matchups—63.9% against Kling-v3.0-pro, 61.1% against Seedance-v1.5-pro, and 58.3% against Grok-Imagine-Video.
5.3
in Ideation to 𝑊 = 0.293 in Mockup, recovering only slightly in Refinement (𝑊 = 0.333). In Ideation, pairwise preferences concentrated strongly around a single output. By Refinement, however, all models produced functionally acceptable pages, and evaluators increasingly justified their rankings through personal preference rather than shared failure modes. Creatives prioritized outputs that established a usable direction in the early stage of ideation: qualitative feedback was distributed across many themes, with structure and layout were raised most often. Landing Page evaluators, for example, prioritized visual hierarchy and layout coherence, while Desktop App comments centered on usability and hierarchy. In Mockup, the task became more constrained. Color & Theme was the most frequent theme, with Prompt Adherence–Usability increased to 𝑟 = 0.65. Landing Page evaluators most prominently discussed prompt adherence, grid structure, color consistency, and typographic pairing, versus Desktop App comments that emphasized text visibility, CTA clarity, and component tweaks. By the Refinement stage, feedback concentrated on final-output criteria, but the effect differed by domain. Ad Images provide the most direct example: mentions of typography rose from 3% in Ideation to approximately 34% in Refinement, with rater agreement being the highest in Refinement (𝑊 = 0.549). In Product Videos, however, numerous refinement prompts asked for targeted edits but resulted in outputs that introduced new elements instead. Free-text responses suggest a loose overall hierarchy in which usability acts as a threshold, prompt adherence orders viable outputs, and visual appeal resolves close contests
Hypothesis #3: Assessing Creative Artifacts Requires More Than One Dimension
Our third hypothesis proposed that model comparison would change when phase and evaluation axis were kept rather than collapsed into a single aggregate score. Table 4 shows this pattern. Perhaps the clearest evidence supporting our hypothesis that single-dimension evaluations are not sufficient for creative evaluation is that no individual model lead all three phases in any domain (Figure 4; see Appendix for equivalent charts per domain). In Desktop Apps, Claude 4.6 has the highest win rate in Ideation and Mockup (0.600 and 0.400), while GPT 5.3 is highest in Refinement (0.400) after the lowest Ideation value (0.086). In Ad Images, GPT Image 1.5 is highest in Ideation and Mockup (0.333 and 0.400), while Seedream 4.5 is highest in Refinement (0.367). Product Videos have different leaders by phase: Veo 3.1 is highest in Ideation (0.417), Veo 3.1 and Kling 3.0 tie in Mockup (0.333), and Grok Imagine is highest in Refinement (0.361). The extended domain results give the same pattern across phases. Landing Pages show phase-specific changes: Claude Opus 4.6 leads Ideation with an 80% win rate, Gemini 3.1 leads in Mockup when a design system is introduced (68.9%), and Claude returns to the lead in Refinement (60.0%); by Refinement, all four models cluster between 3.9 and 4.4 across scalar dimensions. Product Videos show a different split: Veo 3.1 leads in Ideation (61.1%) but receives negative refinement feedback for introducing new elements, while Grok Imagine leads in Refinement (56.5%) and improves on fidelityoriented themes. Kling 3.0 is the only video model above 50% in all three phases (51%, 61%, and 52%). 6
The Human Creativity Benchmark
Domain
Desktop Apps
Landing Pages
Brand Assets
Ad Images
Ad Video
Model
Ideation Stage
Mockup Stage
Refinement Stage
Claude 4.6
0.600 (0.183; 0.0010)
0.400 (0.101; 0.3231)
0.171 (-0.042; 0.1134)
GPT 5.3
0.086 (0.568; 0.0010)
0.229 (0.116; 0.3231)
0.400 (0.119; 0.1134)
Gemini 3.1
0.257 (0.148; 0.0010)
0.171 (-0.031; 0.3231)
0.200 (0.243; 0.1134)
Qwen 3.5
0.057 (-0.066; 0.0010)
0.200 (0.231; 0.3231)
0.229 (0.234; 0.1134)
Claude 4.6
0.533 (-0.060; 0.0000)
0.300 (0.072; 0.3697)
0.300 (-0.076; 0.4678)
Gemini 3.1
0.367 (-0.058; 0.0000)
0.467 (0.040; 0.3697)
0.267 (0.039; 0.4678)
GPT 5.3
0.100 (0.075; 0.0000)
0.167 (-0.011; 0.3697)
0.233 (0.209; 0.4678)
Qwen 3.5
0.000 (-0.043; 0.0000)
0.067 (0.165; 0.3697)
0.200 (0.026; 0.4678)
Gemini 3 Image
0.381 (0.590; 0.0010)
0.500 (0.041; 0.0015)
0.417 (-0.031; 0.0033)
GPT Image 1.5
0.333 (0.054; 0.0010)
0.208 (0.280; 0.0015)
0.361 (0.353; 0.0033)
Seedream 4.5
0.119 (0.151; 0.0010)
0.167 (0.236; 0.0015)
0.139 (0.049; 0.0033)
Flux 2 Max
0.167 (0.135; 0.0010)
0.125 (0.770; 0.0015)
0.083 (0.343; 0.0033)
GPT Image 1.5
0.333 (0.476; 0.0010)
0.400 (0.183; 0.0056)
0.233 (0.112; 0.0398)
Seedream 4.5
0.233 (0.132; 0.0010)
0.233 (0.033; 0.0056)
0.367 (0.343; 0.0398)
Gemini 3 Image
0.233 (0.433; 0.0010)
0.267 (0.141; 0.0056)
0.133 (0.330; 0.0398)
Flux 2 Pro
0.200 (-0.054; 0.0010)
0.100 (-0.134; 0.0056)
0.267 (0.399; 0.0398)
Grok Imagine
0.111 (0.375; 0.0010)
0.194 (0.502; 0.5446)
0.361 (-0.042; 0.1375)
Veo 3.1
0.417 (0.195; 0.0010)
0.333 (0.130; 0.5446)
0.083 (0.176; 0.1375)
Kling 3.0
0.278 (0.273; 0.0010)
0.333 (0.068; 0.5446)
0.222 (0.417; 0.1375)
Seedance 1.5 0.194 (0.310; 0.0010) 0.139 (0.226; 0.5446) 0.333 (0.401; 0.1375) Table 2: Pairwise preference outcomes across workflow phases. Each cell reports a model’s pairwise win rate for the given domain and phase, with Krippendorff’s 𝛼 and Friedman 𝑝-value in parentheses.
Ad Images show the same dependence on phase and axis. GPT Image 1.5 performs best earlier, while Seedream 4.5 improves by Refinement, especially on composition, usability, and typography; Flux 2 Pro also rises from Mockup to Refinement, while Gemini 3 Pro Image falls as typography and product accuracy become more important. These results are consistent with the third hypothesis in a limited sense. Model comparison changes depending on which phase and axis are being examined, but the data does not support a definitive taxonomy of models. Rather, collapsing the benchmark into a single overall ranking hides phase-specific differences where models appear stronger or weaker.
6
differentiate between the needs of creatives at different stages of their workflow. We believe this can result in homogenized outputs that are less useful for creative work, leading to undifferentiated creative voices. We hope to build a shared vocabulary for creative workers engaging with these tools. To this end, model developers may use convergent and divergent criteria as distinct points of intervention: instances where experts converge highlight best practices that models can and should learn, while divergence identifies where a model should not optimize for one target but remain responsive to creative preference. In building then evaluating this dataset, we find that the requirements of people evolve throughout the stages of creative labor. Tool builders and creatives alike can explore phase-level needs shifts to indicate which models are better at a given moment, supporting deliberate model switching without adding friction to the creative workflow. The question of what is quality remains unanswered in full, if only because the answer often lies in the eye of the beholder. As AI is increasingly integrated into creative processes, the future will need greater optionality for model selection and more investment in efforts preventing the loss of creative voices.
Discussion & Limitations
As early pioneers [7] in the field of creative studies wrote, “creativity may represent such a complex human phenomenon that we may never be able to represent it adequately as a single, unidimensional operational variable, or even as a small set of operations. ” As we began to show in this work, evaluating creative quality cannot only be done through one axis of evaluation. Creative work is a hybrid of subjective and objective standards, and assessing only the objective loses important and useful signal in model evaluation and improving creatives’ experiences. Current norms in AI evaluation do not account for the natural diversity and plurality present in creative labor, nor do they
For Creatives. This research provides language for something many creative professionals already feel: the frustration with AI tools is not necessarily that they produce bad work, but that they 7
The Human Creativity Benchmark
Figure 4: Scalar performance across the three phases for video generation, comprising one ribbon for each model–metric combination. Color encodes the model (Grok-Imagine in red, Kling-v3.0 in blue, Seedance-v1.5 in green, and Veo3.1 in purple), while shade encodes the evaluation metric, with the lightest ribbon denoting Visual Appeal (VA), the medium shade Usability (Usab), and the darkest shade Prompt Adherence (PA); these abbreviations (PA, Usab, VA) label the ribbons in the legend. Ribbon width is uniform. The overall range stays roughly constant across phases, spanning approximately 2.7 to 3.8 at both Ideation and Refinement, but the ordering of models changes considerably between them. Veo3.1 begins highest on Prompt Adherence at 3.84 but declines steadily to 2.78 at Refinement, the lowest score in the final phase. Conversely, Seedance-v1.5 on Prompt Adherence rises from 2.72 at Ideation to 3.44 at Refinement, and Grok-Imagine on Visual Appeal climbs to 3.72, the highest score overall, with Grok-Imagine on Usability close behind at 3.69. The dashed line denotes the neutral midpoint of 3.0. produce undifferentiated work that may be difficult to modulate. Understanding which models excel at exploration versus execution, and where in the process agreement breaks down into personal preference, gives creatives using AI tooling a basis for selecting models that fit their needs. Limitations. This study, although framed around the creative process, does not fully represent how creative work unfolds in practice. Creative work is rarely so linear. Designers iterate fluidly, move between tools, revisit stages, and often work across modalities within a single project. Future research will explore longer, less constrained creative arcs to better understand how these evaluation dynamics play out in practice. While the dataset we contribute provides a substantive basis for analysis, it represents only a starting point. For example, we do not control for differences in general model capability or the results of non-determinism, e.g., through iterative sampling, though temperature and various other parameters were standardized, and the prompts used were limited to a finite set of topics. We also collected responses from a relatively small participant group, limiting the statistical significance of our results. Future research should extend this work by expanding the evaluator pool, sampling prompts multiple times for robust coverage of model capabilities, and extending the prompt foci.
8
The Human Creativity Benchmark
References [1] Valerio Basile, Michael Fell, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, Massimo Poesio, and Anca Uma. We need to consider disagreement in evaluation. In Proceedings of the 1st Workshop on Evaluation and Comparison of NLP Systems, pages 15–21, 2021. [2] Rishi Bommasani, Kathleen A. Creel, Arvind Kumar, Dan Jurafsky, and Percy Liang. Picking on the same person: Does algorithmic monoculture lead to outcome homogenization? In Advances in Neural Information Processing Systems, volume 35, pages 3663–3678, 2022. [3] A. Philip Dawid and Allan M. Skene. Maximum likelihood estimation of observer error-rates using the EM algorithm. Journal of the Royal Statistical Society: Series C (Applied Statistics), 28(1):20–28, 1979. [4] Oana Inel, Khalifeh Khamkham, Tatiana Cristea, Anca Dumitrache, Heine Rutjes, Jelle Ploeg, Lukasz Romaszko, Lora Aroyo, and Robert-Jan Sips. CrowdTruth: Machine-human computation framework for harnessing disagreement in gathering annotated data. In The Semantic Web – ISWC 2014, volume 8797 of Lecture Notes in Computer Science, pages 486–504. Springer, 2014. [5] Hye-Kyung Lee. Rethinking creativity: Creative industries, AI and everyday creativity. Media, Culture & Society, 44(3):601–612, 2022. [6] Hoang Thanh-Tung and Thanh Tran. Catastrophic forgetting and mode collapse in GANs. In 2020 International Joint Conference on Neural Networks (IJCNN), pages 1–10. IEEE, 2020. [7] Donald J. Treffinger and John P. Poggio. Needed research on the measurement of creativity. The Journal of Creative Behavior, 6(4):263–267, 1972.
9
The Human Creativity Benchmark
A
Dataset
We release the full set of prompts, model outputs, and human evaluations underlying this benchmark as a public dataset. The data captures professional creative practitioners comparing the outputs of frontier generative models across a realistic, multi-stage creative workflow. Rather than collecting isolated one-shot judgments, the benchmark follows a three-stage pipeline, Ideation, Mockup, and Refinement, that mirrors how creative work progresses from initial concept to polished deliverable. Each model output is evaluated through three complementary signals: pairwise preference judgments, scalar ratings on three quality dimensions, and coded free-text rationales.
A.1
Pipeline, Domains, and Models
The benchmark spans five creative domains: Ad Images, Brand Design, Ad Video, Landing Pages, and Desktop App. Each domain is exercised across all three pipeline stages, yielding 95 distinct prompts in total. Every prompt is generated by four candidate models drawn from a domain-appropriate pool, producing 380 model outputs. In total the dataset covers 13 frontier models, with four competing within each domain. The image domains (Ad Images, Brand Design) are served by image-generation models such as gpt-image-1.5, gemini-3-pro-image-preview, seedream-4.5, and the flux-2 family; the Ad Video domain is served by video models including veo3.1, kling-v3.0-pro, seedance-v1.5-pro, and grok-imagine-video; and the code-based domains (Landing Pages, Desktop App) are served by general-purpose models including claude-opus-4.6, gemini-3.1-pro-preview, gpt-5.3-codex, and qwen3.5-397b-a17b. Prompts at the Mockup and Refinement stages may additionally supply an input image, reflecting the iterative, edit-driven nature of later pipeline stages.
A.2
Data Files
The dataset comprises five CSV files, summarized in Table 3. The files share prompt_id as a common key and can be joined to reconstruct the full evaluation context for any output. File
Records
Description
prompts_workflow.csv
95
Prompt specifications, including domain, pipeline stage, prompt text, and an optional input image for later stages.
model_outputs.csv
380
Model generations keyed by prompt and model. Visual outputs are stored as asset references; code-based outputs are stored inline.
pairwise_comparisons.csv
3,174
Head-to-head preference judgments recording the two competing models and the chosen model, along with the evaluator’s core skill.
scalar_feedback.csv
2,116
Per-output ratings on a 1 to 5 scale across three dimensions: Prompt Adherence, Usability, and Visual Appeal.
2,247
Free-text rationales annotated with assigned themes, per-theme sentiment, and representative key quotes.
qualitative_feedback.csv
Table 3: Files in the released dataset. All files share prompt_id as a join key.
B
Evaluation Interface Extended
All evaluations were conducted within the Contra Labs evaluation environment, a web-based interface that walked raters through each phase of the creative process under controlled conditions. The interface presented each tournament as a single guided flow built around one prompt, moving raters sequentially through the pairwise comparison task and then the scaled-rating task. Before any judgments were made, raters were shown a link to access the evaluation instructions, same as the one also provided up front in the participant instruction document, followed by an introductory screen identifying the first task as a pairwise comparison (Figure 5(A)). From this intro screen, a "Read prompt" control surfaced the prompt together with its reference images, after which the rater began the tournament. Throughout the pairwise task, the prompt and reference images remained visible while two outputs were presented at a time, and the rater selected the preferred output between the left and right options (Figure 5(B)). The interface advanced through successive pairings until all combinations for the prompt were complete. It then displayed a ranking of the outputs with model identities anonymized, and asked the rater, via a free-text field, what had influenced their decision. After the pairwise task, the rater was brought to an introductory screen for the Likert rating task. Each of the three rating dimensions was presented on its own screen, showing the prompt and reference images, the dimension being assessed, a short sub-question framing that dimension, the output images, a 1-through-5 grading scale, and the corresponding grading rubric (Figure 5(C)). The first dimension, prompt adherence, was rated against this layout alone. For the subsequent dimensions, usability and visual appeal, the same layout was extended with a free-text field in which raters described the strengths and weaknesses of each image relative to the grading criteria. Completing the visual appeal ratings concluded the tournament. Across both tasks, model identities were never shown: outputs appeared only under blinded codenames and their on-screen position was randomized to neutralize ordering and recognition effects. Keeping the prompt and reference images persistently in view ensured raters 10
The Human Creativity Benchmark
assessed each generation against the original creative intent rather than in isolation, while the free-text capture preserved each evaluator’s own language for the subsequent thematic coding analysis. (A)
(B)
(C)
Figure 5: Contra Labs evaluation interface, extended views. (A) Introductory screen for the pairwise comparison task, with access to instructions and the prompt. (B) Side-by-side pairwise comparison with the prompt and reference images visible. (C) Scalar rating task for prompt adherence, usability, and visual appeal.
C
Survey Details
Each participant received $350 as compensation for completing the questionnaire. The questionnaire was organized in four parts, exploring (1) respondents’ creative process, whether they currently use AI tools, and, if not, reasons against adoption; (2) how they incorporated AI, including how much of their creative process is AI-supported, how much AI-generated content reaches final deliverables, and the perceived benefits of adoption (e.g., delivery speed, creative exploration, production cost, access to new capabilities); (3) eight modules on generative tool categories (image generation, video generation, prototyping, code generation, copywriting, audio/voice/music, 3D/motion graphics, and LLM-based research), detailing specific tools and where in their workflow they applied them; and (4) attitudes toward AI’s role in the future of creative work, its effect on their earning potential, barriers to adoption, and comfort disclosing AI use to clients. The majority of respondents reported 1-25% of their creative process involved AI tools [add details] across a wide spread of applications or domains. When asked how they incorporated AI into their workflow, a video editor starting “loose [....] From there I refine and hone in on what I’m not getting, and adjust prompts to be more in line with my ideas. I think doing this leaves a lot of opportunity for surprise and serendipity.” This reflects a general pattern across respondents: while AI use and adoption was not monolithic, those that did use these tools bucketed their use in a sequence of exploratory ideation, prototyping, followed by refining client deliverables.
D
Extended Analysis
Landing Pages. Landing Pages shows the clearest phase-by-phase handoff. Claude Opus 4.6 leads Ideation with outputs rated highly on visual hierarchy and layout coherence. When a design system is introduced, Gemini 3.1 Pro Preview takes over (68.9%), with evaluators citing its stronger execution on prompt adherence, grid structure, color consistency, and typographic pairing; this advantage reverses in Refinement, where incremental editing favors Claude Opus 4.6, which reclaims the lead (60.0%). By Refinement, all four models cluster between 3.9–4.4 across all scalar dimensions and preference comes back down to taste, with GPT 5.3 Codex and Qwen 3.5 each showing steady improvement without leading any phase (Figure 39). Product Videos. No model leads more than one phase in Product Videos, producing a three-phase handoff: Veo 3.1 leads Ideation (61.1%), Kling 3.0 Pro leads Mockup (61.1%), and Grok Imagine Video leads Refinement (56.5%), with Kling 3.0 Pro the only model competitive across all three. Veo 3.1 is the only model that degrades on every measured dimension across all three phases: strong in Ideation when generating from scratch, it draws negative evaluator sentiment in Refinement for introducing new elements rather than applying targeted edits—a pattern reflected directly in realism sentiment, which moves from net +6 in Ideation to −3 in Refinement, while Grok Imagine Video improves from −15 to +20. Theme co-occurrence analysis reflects this split: Veo 3.1’s evaluation profile clusters around generation themes like Motion & Blur, while Grok Imagine’s clusters around fidelity themes like Realism and Scene Coherence; Scene Coherence is net negative across all four models, suggesting temporal consistency remains the most persistent challenge in AI video generation. Prompt Adherence correlates independently with Usability (0.64) and Visual Appeal (0.58)—an output can look and feel right while still missing what the brief asked for (Figure 19). 11
The Human Creativity Benchmark
Ad Design. Ad Images has the most reliable convergence arc of any domain, with evaluator agreement rising at every phase transition (0.345 → 0.436 → 0.549) as criteria become progressively more verifiable. Analysis suggests evaluators follow a strict decision hierarchy: usability acts as a hard gate, prompt adherence is the primary ordering criterion among outputs that clear it, and visual appeal resolves close contests as a tiebreaker—criteria close to objective that evaluators reach without coordination. GPT Image 1.5 leads ideation and mockup but drops to third by refinement, with Seedream 4.5 following the opposite trajectory—climbing from third to first—while Flux 2 Pro makes a similar rise and Gemini 3 Pro, steady through the first two phases, finishes last. Seedream 4.5’s ascent tracks sharply improved sentiment in composition, usability, and typography by refinement; Gemini 3 Pro, strong in mockup, collapses in refinement as typography and product accuracy turn negative (Figure 12). Desktop Apps. Desktop App evaluation surfaces the broadest theme set of any domain, spanning prompt adherence, usability, layout quality, visual hierarchy, readability, and design conversion efficiency. Evaluator focus shifts consistently across phases: Ideation centers on usability and hierarchy; Mockup narrows to prompt adherence—key text visibility, CTA clarity, component refinement; Refinement surfaces structural consistency, with flaws in lower-ranked outputs becoming more apparent as evaluator attention shifts from layout to visual hierarchy. Theme co-occurrence analysis reveals distinct model signatures: Claude Opus 4.6 shows tight coupling between prompt adherence and perceived usability; Gemini 3.1 Pro’s coupling is weaker, with occasional mismatches between adherence and usability scores; Qwen 3.5 is strong on high-level layout but clusters negatively around granular execution themes like typography and spacing. While models perform well at ideation and structural layout, Refinement shifts evaluator attention to interaction patterns, icon quality, and accessibility—areas where sentiment turns negative and models consistently underperform (Figure 29).
D.1
Phase Insights
Evaluator focus shifts in a consistent arc across all domains: Ideation centers on structure and layout, with moderate agreement reflecting the range of valid directions before a reference is established (Landing Page 𝑊 = 0.484, Ad Image 𝑊 = 0.345). In Mockup, Color & Theme displaces Layout as the top concern and the Prompt Adherence–Usability correlation strengthens to 𝑟 = 0.65—when a specification is explicit, following it closely produces a more usable output almost automatically. By Refinement, criteria narrow sharply: typography rises from 3% of mentions to ∼ 34% in Ad Images, evaluator agreement peaks (𝑊 = 0.549), and usability becomes the single strongest predictor of competitive success—outputs scoring 5 on usability finish in the top 2 ranks 84% of the time versus 10% for score-1 outputs. Early-stage feedback spreads across many dimensions; by Refinement it reduces to a few production-ready criteria, and the Usability–Visual Appeal correlation reaches 0.818.
E
Domain insights Table 4: Model performance and evaluator agreement statistics across domains and phases. Domain Desktop Apps Desktop Apps Desktop Apps Desktop Apps
Phase Ideation Ideation Ideation Ideation
Model Claude 4.6 GPT 5.3 Gemini 3.1 Qwen 3.5
Win Rate 0.600 0.086 0.257 0.057
Krippendorff 𝛼 0.183 0.568 0.148 -0.066
Friedman 𝑝 -value 0.0010 0.0010 0.0010 0.0010
Desktop Apps Desktop Apps Desktop Apps Desktop Apps
Mockup Mockup Mockup Mockup
Claude 4.6 GPT 5.3 Gemini 3.1 Qwen 3.5
0.400 0.229 0.171 0.200
0.101 0.116 -0.031 0.231
0.3231 0.3231 0.3231 0.3231
Desktop Apps Desktop Apps Desktop Apps Desktop Apps
Refinement Refinement Refinement Refinement
Claude 4.6 GPT 5.3 Gemini 3.1 Qwen 3.5
0.171 0.400 0.200 0.229
-0.042 0.119 0.243 0.234
0.1134 0.1134 0.1134 0.1134
Landing Pages Landing Pages Landing Pages Landing Pages
Ideation Ideation Ideation Ideation
Claude 4.6 Gemini 3.1 GPT 5.3 Qwen 3.5
0.533 0.367 0.100 0.000
-0.060 -0.058 0.075 -0.043
0.0000 0.0000 0.0000 0.0000
Landing Pages Landing Pages Landing Pages Landing Pages
Mockup Mockup Mockup Mockup
Claude 4.6 Gemini 3.1 GPT 5.3 Qwen 3.5
0.300 0.467 0.167 0.067
0.072 0.040 -0.011 0.165
0.3697 0.3697 0.3697 0.3697
Landing Pages Landing Pages Landing Pages Landing Pages
Refinement Refinement Refinement Refinement
Claude 4.6 Gemini 3.1 GPT 5.3 Qwen 3.5
0.300 0.267 0.233 0.200
-0.076 0.039 0.209 0.026
0.4678 0.4678 0.4678 0.4678
Brand Assets Brand Assets Brand Assets Brand Assets
Ideation Ideation Ideation Ideation
Gemini 3 Image GPT Image 1.5 Seedream 4.5 Flux 2 Max
0.381 0.333 0.119 0.167
0.590 0.054 0.151 0.135
0.0010 0.0010 0.0010 0.0010
Brand Assets Brand Assets Brand Assets Brand Assets
Mockup Mockup Mockup Mockup
Gemini 3 Image GPT Image 1.5 Seedream 4.5 Flux 2 Max
0.500 0.208 0.167 0.125
0.041 0.280 0.236 0.770
12
0.0015 0.0015 0.0015 0.0015 Continued on next page
The Human Creativity Benchmark
Table 4: Model performance and evaluator agreement statistics across domains and phases (continued). Domain
Phase
Model
Win Rate
Krippendorff 𝛼
Friedman 𝑝 -value
Brand Assets Brand Assets Brand Assets Brand Assets
Refinement Refinement Refinement Refinement
Gemini 3 Image GPT Image 1.5 Seedream 4.5 Flux 2 Max
0.417 0.361 0.139 0.083
-0.031 0.353 0.049 0.343
0.0033 0.0033 0.0033 0.0033
Ad Images Ad Images Ad Images Ad Images
Ideation Ideation Ideation Ideation
GPT Image 1.5 Seedream 4.5 Gemini 3 Image Flux 2 Pro
0.333 0.233 0.233 0.200
0.476 0.132 0.433 -0.054
0.0010 0.0010 0.0010 0.0010
Ad Images Ad Images Ad Images Ad Images
Mockup Mockup Mockup Mockup
GPT Image 1.5 Seedream 4.5 Gemini 3 Image Flux 2 Pro
0.400 0.233 0.267 0.100
0.183 0.033 0.141 -0.134
0.0056 0.0056 0.0056 0.0056
Ad Images Ad Images Ad Images Ad Images
Refinement Refinement Refinement Refinement
GPT Image 1.5 Seedream 4.5 Gemini 3 Image Flux 2 Pro
0.233 0.367 0.133 0.267
0.112 0.343 0.330 0.399
0.0398 0.0398 0.0398 0.0398
Ad Video Ad Video Ad Video Ad Video
Ideation Ideation Ideation Ideation
Grok Imagine Veo 3.1 Kling 3.0 Seedance 1.5
0.111 0.417 0.278 0.194
0.375 0.195 0.273 0.310
0.0010 0.0010 0.0010 0.0010
Ad Video Ad Video Ad Video Ad Video
Mockup Mockup Mockup Mockup
Grok Imagine Veo 3.1 Kling 3.0 Seedance 1.5
0.194 0.333 0.333 0.139
0.502 0.130 0.068 0.226
0.5446 0.5446 0.5446 0.5446
Ad Video Ad Video Ad Video Ad Video
Refinement Refinement Refinement Refinement
Grok Imagine Veo 3.1 Kling 3.0 Seedance 1.5
0.361 0.083 0.222 0.333
-0.042 0.176 0.417 0.401
0.1375 0.1375 0.1375 0.1375
Table 5: Mean scalar ratings (1–5) for Prompt Adherence, Usability, and Visual Appeal across domains and phases. Domain Desktop Apps Desktop Apps Desktop Apps Desktop Apps
Phase Ideation Ideation Ideation Ideation
Model Claude 4.6 GPT 5.3 Gemini 3.1 Qwen 3.5
Prompt Adherence 3.86 3.49 3.46 3.26
Usability 3.66 2.97 3.26 2.94
Visual Appeal 3.40 3.09 3.11 2.71
Desktop Apps Desktop Apps Desktop Apps Desktop Apps
Mockup Mockup Mockup Mockup
Claude 4.6 GPT 5.3 Gemini 3.1 Qwen 3.5
3.97 3.94 3.89 3.63
3.69 3.71 3.69 3.49
3.83 3.69 3.91 3.63
Desktop Apps Desktop Apps Desktop Apps Desktop Apps
Refinement Refinement Refinement Refinement
Claude 4.6 GPT 5.3 Gemini 3.1 Qwen 3.5
3.69 4.03 3.74 3.74
3.77 3.86 3.66 3.69
3.71 3.89 3.71 3.60
Landing Pages Landing Pages Landing Pages Landing Pages
Ideation Ideation Ideation Ideation
Claude 4.6 Gemini 3.1 GPT 5.3 Qwen 3.5
3.77 3.83 3.10 3.07
4.23 3.87 3.20 3.43
3.97 3.67 3.03 3.13
Landing Pages Landing Pages Landing Pages Landing Pages
Mockup Mockup Mockup Mockup
Claude 4.6 Gemini 3.1 GPT 5.3 Qwen 3.5
3.60 3.97 4.00 3.87
3.90 4.03 3.73 4.00
3.63 3.73 3.57 3.60
Landing Pages Landing Pages Landing Pages Landing Pages
Refinement Refinement Refinement Refinement
Claude 4.6 Gemini 3.1 GPT 5.3 Qwen 3.5
4.43 4.13 4.17 4.07
4.13 3.93 4.07 4.03
4.27 4.13 4.23 3.97
Brand Assets Brand Assets Brand Assets Brand Assets
Ideation Ideation Ideation Ideation
Gemini 3 Image GPT Image 1.5 Seedream 4.5 Flux 2 Max
3.62 4.26 3.55 3.81
3.43 4.07 3.24 3.57
3.52 4.14 3.17 3.83
Brand Assets Brand Assets Brand Assets Brand Assets
Mockup Mockup Mockup Mockup
Gemini 3 Image GPT Image 1.5 Seedream 4.5 Flux 2 Max
3.79 3.54 2.92 2.62
4.00 3.17 2.88 2.58
4.17 3.38 2.92 2.96
Brand Assets Brand Assets Brand Assets Brand Assets
Refinement Refinement Refinement Refinement
Gemini 3 Image GPT Image 1.5 Seedream 4.5 Flux 2 Max
4.08 3.72 3.22 2.92
4.06 3.64 3.33 2.81
3.61 3.39 3.28 2.61
Ad Images Ad Images Ad Images Ad Images
Ideation Ideation Ideation Ideation
GPT Image 1.5 Seedream 4.5 Gemini 3 Image Flux 2 Pro
3.50 2.93 3.13 3.27
3.03 2.93 2.93 3.27 2.93 3.20 3.03 3.23 Continued on next page
13
The Human Creativity Benchmark
Figure 6: (A) shows an example output for brand images, along with a sampled quote included as explanation of domain expert quotes. (B) illustrates a summary ranking of the image across 6 experts–trends are relatively monotonic, illustrating directional alignment (or “convergence”) in opinion.
Table 5: Mean scalar ratings (1–5) for Prompt Adherence, Usability, and Visual Appeal across domains and phases (continued). Domain
Phase
Model
Prompt Adherence
Usability
Visual Appeal
Ad Images Ad Images Ad Images Ad Images
Mockup Mockup Mockup Mockup
GPT Image 1.5 Seedream 4.5 Gemini 3 Image Flux 2 Pro
3.97 3.43 3.57 3.33
3.93 3.17 3.60 3.33
3.87 3.17 3.77 3.03
Ad Images Ad Images Ad Images Ad Images
Refinement Refinement Refinement Refinement
GPT Image 1.5 Seedream 4.5 Gemini 3 Image Flux 2 Pro
4.13 4.00 4.03 3.63
3.73 4.23 3.53 3.67
3.23 3.90 2.97 3.00
Ad Video Ad Video Ad Video Ad Video
Ideation Ideation Ideation Ideation
Grok Imagine Veo 3.1 Kling 3.0 Seedance 1.5
3.69 3.81 3.25 2.72
3.25 3.81 3.25 3.08
3.33 3.78 3.06 2.94
Ad Video Ad Video Ad Video Ad Video
Mockup Mockup Mockup Mockup
Grok Imagine Veo 3.1 Kling 3.0 Seedance 1.5
3.25 3.11 3.33 2.89
3.06 3.25 3.36 2.83
3.11 3.22 3.42 3.14
Ad Video Ad Video Ad Video Ad Video
Refinement Refinement Refinement Refinement
Grok Imagine Veo 3.1 Kling 3.0 Seedance 1.5
3.28 2.78 3.00 3.44
3.39 3.00 3.28 3.14
3.72 2.89 3.53 3.39
Figure 6 provides a supplementary brand-design example, including a sampled expert quote and summary ranking across six evaluators.
F
Survey Insights
Figure 7 summarizes respondents’ attitudes toward AI in creative work, and Figure 8 reports the self-reported share of each respondent’s process that is AI-supported. Figures 9, 10, and 11 list the image, video, and audio tools named most frequently in survey responses. 14
The Human Creativity Benchmark
Figure 7: Distribution of respondent sentiment toward AI’s role in the future of creative work (50 responses). Mixed sentiment is most common at 14 responses, closely followed by Positive (13). Neutral or unclear responses account for 10, and 9 respondents gave no response. Negative or skeptical sentiment is the least common, at 4 responses.
Figure 8: Distribution of the self-reported share of the creative process that is AI-supported (50 responses). The most common range is 1 to 25%, reported by 27 respondents, followed by 25 to 50% (16). Higher levels of reliance are uncommon, with 4 respondents reporting 50 to 75% and 2 reporting 75 to 100%, while a single respondent reported no AI support (0%).
G Domain Insights G.1 Ad Images Figure 12 reports phase-level win rates. Figure 13 tracks how evaluation themes shift across phases. Figures 14 and 15 summarize scalar ratings and trajectories, and Figures 16, 17, and 18 give head-to-head pairwise win rates for the Ideation, Mockup, and Refinement stages. 15
The Human Creativity Benchmark
Figure 9: Image generation tools mentioned in survey responses. Midjourney is cited most frequently at 26 responses, followed by ChatGPT / OpenAI (18) and Nano Banana (14). A middle tier comprises Gemini and Flora (7 each), Visual Electric (5), and Krea and Sora (4 each). The remaining tools, including Adobe Firefly, Photoshop, Freepik, Runway, Ideogram, Figma, Stable Diffusion, and Perplexity, were each named three or fewer times.
Figure 14: Mean scalar ratings (1–5) for ad-image deliverables across the three pipeline stages, broken out by evaluation question. The dashed line marks the neutral midpoint (3.0). Prompt Adherence. GPT-Image-1.5 leads at every stage: 3.5 in Ideation (ahead of Gemini-3-Pro-Image-Preview at 3.3, Flux-2-Pro at 3.2, and Seedream-4.5 at 3.0), 4.0 in Mockup (Gemini 3.6, Seedream 3.4, Flux 3.3), and 4.1 in Refinement, where Gemini and Seedream both reach 4.0 and Flux trails at 3.6. Usability. In Ideation, Gemini and GPT-Image-1.5 tie at 3.1, with Seedream just behind at 3.0 and Flux lowest at 2.8. GPT-Image-1.5 jumps to 4.0 in Mockup, followed by Gemini (3.6), Flux (3.4), and Seedream (3.2). By Refinement, Seedream takes the lead at 4.2, ahead of GPT-Image-1.5 and Flux (tied at 3.7) and Gemini (3.5). Visual Appeal. Gemini and Seedream share the Ideation lead at 3.3, ahead of Flux (3.1) and GPT-Image-1.5 (3.0). In Mockup, GPT-Image-1.5 (3.9) and 16 Gemini (3.8) lead, followed by Seedream (3.2) and Flux (3.0). In Refinement, Seedream scores highest at 3.9, followed by GPT-Image-1.5 (3.2) and then Flux and Gemini, tied at 3.0.
The Human Creativity Benchmark
Figure 10: Video generation tools mentioned in survey responses. Runway is the most frequently cited at 9 responses, followed by Sora and Veo (6 each) and Kling (5). The remaining tools were named far less often: Wan (2), and CapCut, Pika, Hailuo, and Adobe Firefly (1 each).
Figure 15: Scalar performance across the three phases for ad-image generation, comprising one ribbon for each model–metric combination. Color encodes the model (Gemini-3 in purple, GPT-Image in orange, Seedream-4.5 in brown, and Flux-2 in green), while shade encodes the evaluation metric, with the lightest ribbon denoting Visual Appeal (VA), the medium shade Usability (Usab), and the darkest shade Prompt Adherence (PA); these abbreviations (PA, Usab, VA) label the ribbons in the legend. Ribbon width is uniform and does not encode uncertainty.Most ribbons rise from a compressed band at Ideation, spanning roughly 17 2.9 to 3.5, and broaden into a wider spread of approximately 3.0 to 4.2 by Refinement. Seedream-4.5 on Visual Appeal shows the largest improvement, climbing from 3.50 at Ideation to 4.23 at Refinement, the highest score overall, with GPT-Image on Prompt Adherence close behind at 4.13. Several metrics decline in the final phase, most notably GPT-Image on Visual Appeal, which falls to 3.23, alongside Flux-2 and Gemini-3 on Visual Appeal, which settle near the 3.0 midpoint. The dashed line denotes the neutral midpoint of 3.0.
The Human Creativity Benchmark
Figure 11: Audio, voice, and music generation tools mentioned in survey responses. Only two tools were named: ElevenLabs, cited in 14 responses, and Suno, cited in 4.
G.2
Ad Video
Figure 19 reports pairwise win rates across the three pipeline stages. Figures 20, 21, and 22 give head-to-head pairwise win rates for the Ideation, Mockup, and Refinement stages.
G.3
Brand Assets
Figures 23, 24, and 25 summarize scalar performance, win rates, and mean ratings for brand design, and Figures 26, 27, and 28 give head-to-head pairwise win rates for the Ideation, Mockup, and Refinement stages. 18
The Human Creativity Benchmark
Figure 12: Model win rates across the three ad-image generation stages identified in the paper: Ideation, Mockup, and Refinement. GPT-Image-1.5 leads in Ideation (58%), followed by Gemini-3-Pro-Image-Preview (51%), Seedream-4.5 (46%), and Flux-2-Pro (45%). Mockup shows the same ordering at the top, with GPT-Image-1.5 (66%) ahead of Gemini-3-Pro-Image-Preview (63%), while Seedream-4.5 (35%) and Flux-2-Pro (36%) trail closely together. In Refinement, GPT-Image-1.5 and Seedream-4.5 tie for the lead (58% each), followed by Flux-2-Pro (46%) and then Gemini-3-Pro-Image-Preview (39%). The dashed line marks the 50% break-even point.
Figure 23: Example overview of scalar performance for Brand Design content; y-axis marks the stage, while x-axis marks mean rating, separated by metric (Prompt Adherence, Usability, and Visual Appeal) and model. Similar charts are shown in Appendix for each category produced. 19
The Human Creativity Benchmark
Figure 13: Model win rates across the three ad-image generation stages identified in the paper: Ideation, Mockup, and Refinement. GPT-Image-1.5 leads in Ideation (58%), followed by Gemini-3-Pro-Image-Preview (51%), Seedream-4.5 (46%), and Flux-2-Pro (45%). Mockup shows the same ordering at the top, with GPT-Image-1.5 (66%) ahead of Gemini-3-Pro-Image-Preview (63%), while Seedream-4.5 (35%) and Flux-2-Pro (36%) trail closely together. In Refinement, GPT-Image-1.5 and Seedream-4.5 tie for the lead (58% each), followed by Flux-2-Pro (46%) and then Gemini-3-Pro-Image-Preview (39%). The dashed line marks the 50% break-even point.
Figure 24: Model win rates across the three phases for brand-design generation. GPT-Image-1.5 leads Ideation at 63%, followed by Flux-2-max at 50%, Gemini-3-Pro-Image-Preview at 49%, and Seedream-4.5 at 38%. In Mockup, Gemini-3-Pro-Image-Preview takes the leads with 63%, followed by GPT-Image-1.5 at 53%, followed by Seedream-4.5 and Flux-2-max at 42%, and 41%. By 20 followed by GPT-Image-1.5 at 57%, Seedream-4.5 at 48%, and Refinement, Gemini-3-Pro-Image-Preview leads again at 66%, finally Flux-2-max at 29%.
The Human Creativity Benchmark
Figure 16: Head-to-head pairwise win rates for ad-image generation in the Ideation stage. Each cell reports the row model’s win rate against the column model; diagonal self-matches are excluded.
G.4
Desktop Apps
Figure 29 tracks scalar performance across the three pipeline stages. Figures 30, 31, and 32 give head-to-head pairwise win rates for the Ideation, Mockup, and Refinement stages.
G.5
Landing Pages
Figures 33, 34, 35, and 39 summarize scalar trajectories, win rates, mean ratings, and phase ribbons for landing-page generation, and Figures 36, 37, and 38 give head-to-head pairwise win rates for the Ideation, Mockup, and Refinement stages. 21
The Human Creativity Benchmark
Figure 17: Head-to-head pairwise win rates for ad-image generation in the Mockup stage. Each cell reports the row model’s win rate against the column model; diagonal self-matches are excluded.
22
Figure 33: Scalar performance across the three phases for landing-page generation, with one ribbon for each model–metric combination. Color encodes the model (Claude-Opus in blue, Gemini-3.1 in orange, GPT-5.3 in teal, and Qwen3.5-397b in red), while shade encodes the evaluation metric, with the lightest ribbon denoting Visual Appeal (VA), the medium shade Usability (Usab), and the darkest shade Prompt Adherence (PA); these abbreviations (PA, Usab, VA) label the ribbons in the legend. Ribbon width is uniform and does not encode uncertainty. All twelve ribbons exhibit an upward trend from Ideation to Refinement. The wide dispersion observed at Ideation, spanning approximately 3.0 to 4.2, converges into a substantially narrower band of
The Human Creativity Benchmark
Figure 18: Head-to-head pairwise win rates for ad-image generation in the Refinement stage. Each cell reports the row model’s win rate against the column model; diagonal self-matches are excluded.
23
Figure 34: Model win rates across the three phases for landing-page generation. Claude-Opus-4.6 leads Ideation at 80%, followed by Gemini-3.1-Pro-Preview at 58%, Qwen3.5-397b-a17b at 38%, and finally GPT-5.3-Codex at 25%. In Mockup, Gemini-3.1-ProPreview leads at 69%, followed by Claude-Opus-4.6 at 50%, Qwen3.5-397b-a17b at 44%, and GPT-5.3-Codex at 37%. By Refinement, Claude-Opus-4.6 leads again at 60%, followed by Gemini-3.1-Pro-Preview at 52%, Qwen3.5-397b-a17b at 48%, and GPT-5.3-Codex at 40%.
The Human Creativity Benchmark
Figure 19: Overall pairwise win rates across the three pipeline stages for product-video generation. Veo3.1 leads Ideation at 61%, followed by Kling-v3.0-pro (51%), Grok-Imagine-Video (46%), and Seedance-v1.5-pro (42%). In Mockup, Kling-v3.0-pro rises to the top at 61%, ahead of Veo3.1 (56%), Grok-Imagine-Video (44%), and Seedance-v1.5-pro (39%). By Refinement the order reshuffles again: Grok-Imagine-Video leads at 56%, with Kling-v3.0-pro and Seedance-v1.5-pro nearly tied just above break-even (52% and 53%), and Veo3.1 falling to last at 39%. The dashed line marks the 50% break-even point.
Figure 35: Mean scalar ratings (1–5) for landing-page deliverables across the three pipeline stages, broken out by evaluation question. The dashed line marks the neutral midpoint (3.0). Prompt Adherence. Claude-Opus-4.6 ties Gemini-3.1-Pro-Preview for the Ideation lead at 3.9 (ahead of GPT-5.3-Codex and Qwen3.5-397b-a17b, both 3.3), dips to last in Mockup at 3.6 while the others sit at 3.9–4.0, then jumps to the top in Refinement at 4.4, followed by GPT-5.3-Codex (4.2), and Gemini-3.1-Pro-Preview and Qwen3.5-397b-a17b (both 4.1). Gemini-3.1-Pro-Preview is the most stable, holding 3.9 → 4.0 → 4.1 across the three stages. Usability. Claude-Opus-4.6 starts strongest in Ideation at 4.3, ahead of Gemini-3.1-Pro-Preview (3.8), Qwen3.5-397b-a17b (3.5), and GPT-5.3-Codex (3.3). The field tightens in Mockup, where Gemini-3.1-Pro-Preview and Qwen3.5-397b-a17b lead at 4.0, followed by Claude-Opus-4.6 (3.9) and GPT-5.3-Codex (3.7). By Refinement, Claude-Opus-4.6 and GPT-5.3-Codex tie at 4.1, ahead of Qwen3.5-397b-a17b (4.0) and Gemini-3.1-Pro-Preview (3.9). Visual Appeal. Claude-Opus-4.6 leads Ideation at 4.0, followed by Gemini-3.1-Pro-Preview (3.6), Qwen3.5-397b-a17b (3.2), and GPT-5.3-Codex (3.1). In Mockup the models bunch together, with Gemini-3.1-Pro-Preview just ahead at 3.7 and Claude-Opus-4.6, GPT-5.3-Codex, and Qwen3.5-397b-a17b all at 3.6. In Refinement, Claude-Opus-4.6 leads at 4.3, followed by GPT-5.3-Codex (4.2), Gemini-3.1-Pro-Preview (4.1), and Qwen3.5-397b-a17b (4.0).
24
The Human Creativity Benchmark
Figure 20: Head-to-head pairwise win rates for product-video generation in the Ideation stage. Each cell reports the row model’s win rate against the column model; diagonal self-matches are excluded.
25
The Human Creativity Benchmark
Figure 21: Head-to-head pairwise win rates for product-video generation in the Mockup stage. Each cell reports the row model’s win rate against the column model; diagonal self-matches are excluded.
26
The Human Creativity Benchmark
Figure 22: Head-to-head pairwise win rates for product-video generation in the Refinement stage. Each cell reports the row model’s win rate against the column model; diagonal self-matches are excluded.
27
The Human Creativity Benchmark
Figure 25: Overall pairwise win rates across the three pipeline stages for brand-design generation. GPT-Image-1.5 leads Ideation at 63%, followed by Flux-2-max (50%), Gemini-3-Pro-Image-Preview (49%), and Seedream-4.5 (38%). In Mockup, Gemini-3-ProImage-Preview rises to the top at 63%, ahead of GPT-Image-1.5 (53%), Seedream-4.5 (42%), and Flux-2-max (41%). By Refinement, Gemini-3-Pro-Image-Preview extends its lead to 66%, followed by GPT-Image-1.5 (57%) and Seedream-4.5 (48%), while Flux-2-max falls to 29%. The dashed line marks the 50% break-even point.
28
The Human Creativity Benchmark
Figure 26: Head-to-head pairwise win rates for brand-design generation in the Ideation stage. Each cell reports the row model’s win rate against the column model; diagonal self-matches are excluded.
29
The Human Creativity Benchmark
Figure 27: Head-to-head pairwise win rates for brand-design generation in the Mockup stage. Each cell reports the row model’s win rate against the column model; diagonal self-matches are excluded.
30
The Human Creativity Benchmark
Figure 28: Head-to-head pairwise win rates for brand-design generation in the Refinement stage. Each cell reports the row model’s win rate against the column model; diagonal self-matches are excluded.
31
The Human Creativity Benchmark
Figure 29: Scalar performance across the three phases for brand-design generation, comprising one ribbon for each model–metric combination (twelve in total). Color encodes the model (Gemini-3 in purple, GPT-Image in orange, Seedream-4.5 in brown, and Flux-2 in green), while shade encodes the evaluation metric, with the lightest ribbon denoting Visual Appeal (VA), the medium shade Usability (Usab), and the darkest shade Prompt Adherence (PA); these abbreviations (PA, Usab, VA) label the ribbons in the legend. Ribbon width is uniform.GPT-Image records the highest scores at Ideation, led by Prompt Adherence at 4.26, but its three metrics decline over the following phases. Gemini-3 moves in the opposite direction, rising to the top by Refinement with Prompt Adherence at 4.08 and Usability at 4.06. Flux-2 falls steadily across phases to the lowest scores in the final phase, reaching 2.92 on Prompt Adherence, 2.81 on Usability, and 2.61 on Visual Appeal. Seedream-4.5 remains in the middle of the range throughout. The dashed line denotes the neutral midpoint of 3.0.
32
The Human Creativity Benchmark
Figure 30: Head-to-head pairwise win rates for desktop-app generation in the Ideation stage. Each cell reports the row model’s win rate against the column model; diagonal self-matches are excluded.
33
The Human Creativity Benchmark
Figure 31: Head-to-head pairwise win rates for desktop-app generation in the Mockup stage. Each cell reports the row model’s win rate against the column model; diagonal self-matches are excluded.
34
The Human Creativity Benchmark
Figure 32: Head-to-head pairwise win rates for desktop-app generation in the Refinement stage. Each cell reports the row model’s win rate against the column model; diagonal self-matches are excluded.
35
The Human Creativity Benchmark
Figure 36: Head-to-head pairwise win rates for landing-page generation in the Ideation stage. Each cell reports the row model’s win rate against the column model; diagonal self-matches are excluded.
36
The Human Creativity Benchmark
Figure 37: Head-to-head pairwise win rates for landing-page generation in the Mockup stage. Each cell reports the row model’s win rate against the column model; diagonal self-matches are excluded.
37
The Human Creativity Benchmark
Figure 38: Head-to-head pairwise win rates for landing-page generation in the Refinement stage. Each cell reports the row model’s win rate against the column model; diagonal self-matches are excluded.
38
The Human Creativity Benchmark
Figure 39: Scalar performance across the three phases for landing-page generation, comprising one ribbon for each model–metric combination. Color encodes the model (Claude-Opus in blue, Gemini-3.1 in orange, GPT-5.3 in teal, and Qwen3.5-397b in red), while shade encodes the evaluation metric, with the lightest ribbon denoting Visual Appeal (VA), the medium shade Usability (Usab), and the darkest shade Prompt Adherence (PA); these abbreviations (PA, Usab, VA) label the ribbons in the legend. Ribbon width is uniform. The wide dispersion observed at Ideation, spanning approximately 3.0 to 4.2, narrows into a tighter band of roughly 3.9 to 4.4 by Refinement. Claude-Opus on Prompt Adherence records the highest score at both Ideation (4.23) and Refinement (4.43), though it falls toward the middle of the range at Mockup before climbing back. The lowest-scoring models at Ideation, GPT-5.3 and Qwen3.5-397b at approximately 3.0 to 3.2, show the largest gains across phases, closing much of the gap by Refinement. The dashed line denotes the neutral midpoint of 3.0.
Figure 40: Spread of axes agreement across domains.
39