AtelierEval: Agentic Evaluation of Humans & LLMs as Text-to-Image Prompters
Hanjun Luo 1 * Zhimu Huang 1 * Sylvia Chung 2 Yiran Wang 3 Yingbin Jin 4 Jialin Li 1 Jiang Li 1 Xinfeng Li 5 Hanan Salam 1
arXiv:2605.22645v1 [cs.AI] 21 May 2026
Abstract
to academic diagrams (Ko et al., 2023; Esser et al., 2024; Jaiprakash & Prakash, 2025; Yang et al., 2025b; Sordo et al., 2025). As generation quality and controllability improve, T2I systems increasingly depend on the Prompting Proficiency of upstream prompters, i.e., the ability to transform user intent into executable prompts that reliably produce desired outputs (Liu & Chilton, 2022; Canossa et al., 2025).
Text-to-image (T2I) systems increasingly rely on upstream prompters, either humans or multimodal large language models (MLLMs), to translate user intent into detailed prompts. Yet current benchmarks fix the prompt and only evaluate T2I models, leaving the prompting proficiency of this upstream component entirely unmeasured. We introduce AtelierEval, the first unified benchmark that quantifies prompting proficiency across 360 expert-crafted tasks. Grounded in a cognitive view, it spans three task categories and instantiates tasks using a taxonomy of real-world challenges, with interfaces for both humans and MLLMs. To enable scalable and reliable evaluation, we propose AtelierJudge, a skill-based, memory-augmented agentic evaluator. It produces subjective and objective scores for prompt–image pairs, achieving a Spearman correlation of 0.81 with human experts, approaching human performance. Extensive experiments benchmark 8 MLLMs against 48 human users across 4 T2I backends, validate AtelierEval as a robust diagnostic tool, and reveal the superiority of mimicry over planning, advocating for an image-augmented direction for future prompters. Our work is released to support future research.
Figure 1. MLLMs act as prompters in diverse T2I workflows, translating user intent into effective prompts.
This proficiency remains an important bottleneck in practice, as effective prompting requires substantial expertise to encode semantics, constraints, and stylistic intent in a single prompt (Cao et al., 2023). In response, contemporary T2I workflows rely on multimodal large language models (MLLMs) to support human prompters. These workflows typically follow two integration patterns as illustrated in Figure 1: (i) Leading platforms (e.g., ChatGPT (OpenAI, 2025b), Gemini (Citron, 2024), and Doubao (Doubao, 2025)) adopt an implicit integration pattern, embedding an internal MLLM middleware to automatically translate user inputs into model-friendly prompts. (ii) Advanced creators employ explicit reasoning workflows, where state-of-the-art (SOTA) MLLMs are utilized as external prompting assistants (Gani et al., 2023; Xiang et al., 2025). By employing a general-purpose yet carefully designed system prompt, they leverage the visual grounding and reasoning capabilities of MLLMs to decompose complex visual intent into detailed, parameterized prompts that meet production-level needs.
1. Introduction The rapid advancement of text-to-image (T2I) models is reshaping creative workflows, enabling the efficient production of complex visual content from commercial illustrations *
Equal contribution 1 New York University Abu Dhabi Zhejiang University 3 University of Electronic Science and Technology of China 4 The Hong Kong Polytechnic University 5 Nanyang Technological University. Correspondence to: Hanjun Luo <[email protected]>. 2
Proceedings of the 43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026. Copyright 2026 by the author(s).
1
AtelierEval
a robust 95.5% accuracy on objective constraints.
Prompting-proficiency evaluation targets the input end of T2I pipelines and is complementary to conventional benchmarks that assess generative models given fixed prompts. Nonetheless, existing evaluation practices have not kept pace with this evolution: there is still no standardized way to quantify the prompting proficiency of humans and MLLMs (Gu et al., 2023; Schulhoff et al., 2025). Research on human prompting remains confined to fragmented qualitative user studies (Sanchez, 2023; Holzner et al., 2025; Heigl, 2025), while standardized benchmarks focus solely on T2I models (Bakr et al., 2023; Zhang & Tang, 2024), systematically overlooking upstream instruction builders. This evaluation gap can significantly mislead research on prompt optimizers. Typically validated on T2I benchmarks that presuppose executable inputs, these methods prioritize modelspecific prompt polishing for visual quality or alignment (Pavlichenko & Ustalov, 2023; Yeh et al., 2024; You et al., 2025), leaving prompting proficiency—the model-agnostic translation from intent to executable prompts—largely unexplored. This gap represents a key bottleneck for the democratization of T2I systems (Mahdavi Goloujeh et al., 2024) and for recognizing the contributions of prompters (Chong et al., 2025), yet remains unmeasured, limiting the advancement of prompting, both as a human skill and as a scalable technology (Hartwig et al., 2025).
Leveraging AtelierEval, we benchmark a stratified spectrum of 8 MLLMs, ranging from lightweight opensource models to SOTA models, against 48 users (24 novices, 24 skilled users). Prompters are evaluated across 4 T2I models, representing varying levels of LLM middleware intervention. The results show AtelierEval effectively quantifies the model-agnostic prompting proficiency gaps between different tiers of humans and MLLMs, as evidenced by consistent prompter rankings across all evaluated T2I backends. Moreover, our experiments reveal several key insights. In particular, while advanced T2I backends with MLLM middleware tend to homogenize subjective quality across prompters, they also induce logical conflicts with external MLLM reasoning in several settings, leading to objective performance degradation. In contrast, imitation prompting consistently avoids these conflicts, motivating a shift from symbolic planning to image-augmented prompting for the development of future agents. Our core contributions are summarized as follows: ❶ Unified Benchmark. We introduce AtelierEval, the first systematic benchmark to quantify T2I prompting proficiency. Grounded in cognitive science, it comprises 360 expert-crafted tasks spanning modalities and features a unified interaction protocol for both humans and MLLMs. ❷ Agentic Evaluator. We propose AtelierJudge, a cognitivemimetic and skill-based agent with memory for retrievalaugmented evaluation. It achieves superior alignment with human experts compared to baselines, offering a principled blueprint for future T2I evaluators. ❸ Validation & Insights. We validate our framework’s effectiveness in evaluating prompting proficiency. We also reveal key insights and provide open-source infrastructure, facilitating both prompt engineering education and the development of prompting agents.
In light of this gap, we introduce AtelierEval, the first unified benchmark to evaluate both humans and LLMs quantitatively in their role as prompters. Inspired by Structure of Intellect (SI) theory (Guilford, 1967), we operationalize 360 expert-crafted tasks into three cognitive categories across two modalities: Open-ended, Constrained, and Imitation. Drawing on prior analyses (Huang et al., 2023; Chen et al., 2025b), we taxonomize common T2I generation failures into semantic interpretation and constraint realization, yielding ecologically grounded, execution-challenging tasks to expose prompting proficiency. Furthermore, AtelierEval offers a unified protocol with an intuitive user interface (UI) for humans and a standardized toolkit for MLLMs, enabling rigorous comparisons between humans and MLLMs in the T2I domain. To empower AtelierEval with scalable precision, we propose AtelierJudge, a cognitive-mimetic Agent-as-a-Judge system orchestrating modular skills and 5 task-specific memories. Grounded in Dual-Process Theory (Kahneman, 2011), the agent routes skills to decouple evaluation into two parallel processes: ❶ subjective scoring via retrieval-augmented generation (RAG) (Lewis et al., 2020), and ❷ objective constraint checks. This synergy of objective logic and calibration via human-curated exemplars not only mitigates selfpreference bias but also architecturally mirrors human cognition. Empirically, AtelierJudge achieves a Spearman correlation of 0.81, nearing the human agreement (0.83) while surpassing baselines (0.55) in subjective scoring, along with
2. Related Work Prompting Proficiency Benchmarking. Current T2I benchmarks offer comprehensive coverage, ranging from general quality and alignment (Saharia et al., 2022; Wu et al., 2023) through complex metrics like text rendering and spatial relationships (Bosheah & Bilicki, 2025) to specific applications (Chang et al., 2025; Tao et al., 2025) and safety metrics (Luo et al., 2024a;c; 2026; Chinchure et al., 2024), yet they primarily focus on T2I models themselves. By treating the inputs as static prompts, they systematically overlook the role of prompters. Although numerous prompt optimizers have emerged, their validation often directly applies the aforementioned T2I benchmarks (Hao et al., 2023; Mañas et al., 2024; Zhang et al., 2024; Luo et al., 2024b). This evaluation paradigm confines prompting proficiency to model-specific, metric-specific, and fragmented paramet-
2
AtelierEval
ric adaptation (Mo et al., 2024; Li et al., 2024c), rather than a generalized request-to-prompt translation capability. While research on human prompting has evolved from simple prompt refinement (Feng et al., 2024; Hei et al., 2024; Wang et al., 2024b) to analyzing users’ request-to-prompt performance (Liu et al., 2025; Huang & Xie, 2025), this body of work remains confined to small-scale and qualitative studies (Tsao et al., 2025; Gulzar et al., 2025). Consequently, this gap urgently calls for a standardized benchmark that can isolate prompting proficiency from model execution and support systematic comparison across prompters.
formally distinguish prompting proficiency from existing evaluation paradigms by modeling the generation process x = M(p) with two distinct alignment targets: the Intent Representation I (the user’s ground truth goal in a pre-execution form) and the Literal Intent Ip (the explicit semantics and implied aesthetics of prompt p). Notably, I may be abstract, structured, or perceptual, provided that it is not directly executable by T2I models. Let π : I → p be the prompter’s policy and M be the T2I model.
Agentic Evaluation Paradigms. T2I evaluation has long relied on automated metrics such as CLIP Score, with corresponding evaluators (Heusel et al., 2017; Hessel et al., 2021). However, recent studies indicate that these coarse-grained metrics exhibit poor correlation with human perception, particularly for complex spatial structures, logical consistency, and subtle aesthetic nuances (Ghosh et al., 2023; Li et al., 2024a). This discrepancy has compelled the community to resort to unscalable and unreproducible manual evaluation (Lee et al., 2023; Wiles et al., 2025). In response, the LLM/MLLM-as-a-Judge paradigm—leveraging LLMs to simulate human judgment—has been widely adopted (Zheng et al., 2023; Li et al., 2024b). Recent advancements have introduced the Agent-as-a-Judge, which significantly enhances evaluation performance (Zhuge et al., 2024). Nevertheless, this shift remains underexplored in the T2I domain. Current MLLM-based scorers predominantly perform static visual question answering (VQA) (Hu et al., 2023; Sun et al., 2025). Not only do they struggle with the dual challenge of “objective logical constraints” and “subjective aesthetic appreciation” when evaluating prompts (Fu et al., 2024; Jin & Chua, 2025), but they are also prone to severe model biases and self-preferences (Chen et al., 2024).
Figure 2. Three cognitively grounded categories form a complete task partition across 4 application context dimensions and 24 tags. Details of the application contexts are presented in Appendix B.
Paradigm 1: Model Benchmarking. Mainstream benchmarks fix p to evaluate the generative capabilities of a set of candidate models M. They assume p is the ground truth, defining the optimal model M∗ ∈ M that maximizes the expected score S over a fixed prompt distribution D: M∗ = argmax Ep∈D [S(M(p), Ip )] ,
3. AtelierEval: Benchmark Design
(1)
M
For systematic evaluation, AtelierEval operationalizes prompting proficiency as a structured cognitive process. Section 3.1 formalizes the prompter-model interaction as an optimization problem, defining prompting proficiency rigorously. Drawing on Guilford’s SI theory, Section 3.2 dissects this process into three distinct cognitive categories to ensure comprehensive capability coverage. Section 3.3 details the construction of our 360-task dataset based on taxonomic challenge primitives. Section 3.4 establishes a unified protocol to benchmark humans and MLLMs.
Paradigm 2: Prompt Optimization. Optimizers (e.g., token searchers, LLM-based rewriters) refine an initial prompt pinit into an optimized p∗ within a semantic neighborhood N (pinit ) to boost generation quality or alignment: p∗ = argmax S(M(p), I).
(2)
p∈N (pinit )
This utility-oriented paradigm prioritizes instance-specific maximizing S, reduces prompting to polishing the explicit pinit , and neglects the fundamental translation of I. Paradigm 3: Prompting Proficiency. Unlike Paradigm 2, the paradigm of prompting proficiency is capabilityoriented. We evaluate the prompter’s policy π itself. We fix the distribution of intents D and evaluate π’s ability to structurally decode I across varying models:
3.1. Problem Formulation Based on established research on prompt engineering (Zamfirescu et al., 2023), we conceptualize prompting as the process of bridging the gap between user intent and executable system actions (Gulf of Execution) (Norman, 2013). We
π ∗ = argmax EI∈D [EM [S(M(π(I)), I)]] . π
3
(3)
AtelierEval
3.3. Dataset Construction
Prompting proficiency aims to quantify the prompter’s ability to generalize, translate, and adhere to intents, establishing a metric for the prompter’s intrinsic proficiency.
To focus our evaluation on prompting proficiency, tasks are designed to involve intents difficult for the model to satisfy directly, but whose satisfaction probability can be improved through effective translation. Drawing on prior studies (Ribeiro et al., 2020; Oppenlaender, 2024), we abstract failure modes from such intents into two dimensions, semantic interpretation {Si } and constraint realization {Cj }, and define two corresponding sets of challenge primitives for task construction (Table 1). Tasks are instantiated by experts via cross-combining subsets of these primitives and grounding them in common T2I application contexts, with task complexity emerging from their composition. The details of task construction process and expert involvement are provided in Appendix C & N. The three task categories are instantiated as follows, each with 120 tasks: ➠ OE Tasks instantiate {Si } under high narrative noise. These tasks focus on semantic translation, without introducing explicit {Cj }. ➠ CO Tasks emphasize {Cj } under low-noise, structured specifications. {Si } are still present but expressed explicitly with limited ambiguity to provide constraint context. ➠ IM Tasks instantiate visual counterparts of challenge primitives through image inputs. Primitives that cannot be reliably inferred from images, including S2 , S4 , and C5 , are excluded as stated in Appendix C.3.
3.2. Task Categorization & Cognitive Mapping Equation (3) provides a principled way to evaluate prompters. However, our goal is a systematic and diagnostic evaluation that avoids reliance on fragmented, ad-hoc task collections and reveals the specific capability dimensions that distinguish different prompters, thereby informing both human prompting education and prompting agent development. Currently, prompt engineering is largely treated as a trial-and-error process, with evaluation confined to overall outcome quality (Strobelt et al., 2022; Knoth et al., 2024). To move beyond such black-box evaluation, a structured decomposition of the policy π is necessary. To prevent adhoc or arbitrary decompositions, we adopt the Operations dimension of SI theory and map all five operations to our framework. Specifically, the analysis-oriented memory and evaluation operations are delegated to AtelierJudge (Section 4), while task design focuses on the three constructive operations—divergent production, convergent production, and cognition. Accordingly, as shown in Figure 2, we partition the task space into the following three categories: ➠ Open-ended Creation (OE) tasks correspond to divergent production and focus on translating abstract, scenario-driven natural language requests—typically information-sparse and containing narrative noise—into executable prompts. These tasks emphasize thematic, atmospheric, and stylistic intent without providing explicit execution constraints. Prompters must extract key information, complete and organize elements, and translate implicit intent into a complete prompt. ➠ Constrained Creation (CO) tasks correspond to convergent production and focus on constructing prompts under explicit rules. Such tasks provide structured constraints across multiple dimensions, which are semantically clear yet difficult to execute directly. Prompters must jointly integrate complex constraints into a single prompt that is more likely to satisfy all conditions during generation. ➠ Imitation (IM) tasks correspond to cognition. Given a target image, prompters identify key features and translate perceptual information into prompts to reproduce similar visual results. These tasks reflect the information encoding and representation processes involved in cognition. Notably, AtelierEval is restricted to single-turn, pure text-to-image settings. From an information-theoretic perspective (Cover, 1999), OE, CO, and IM correspond to expansion, convergence, and fidelity-preserving encoding, respectively, forming a complete task partition for this setting. Formal justification is provided in Appendix A.
Table 1. Challenge Primitives of AtelierEval. Primitive decomposition and task instances are provided in Appendix E. Primitive
Definition
Semantic Challenge Primitives S1 Abstract Intent Abstract or affective intent requiring visualization. S2 Audience Intent Intent specified via target audience preferences. S3 Implicit Style Style or medium implied but not explicitly stated. S4 Semantic Negation What should not be generated at the semantic level. Constraint Challenge Primitives C1 Attribute Binding Binding attributes to the correct entities. C2 Spatial Relation Explicit spatial or layout relationships. C3 Quantity Exact numerical constraints on object counts. C4 Text Exact text content and spelling. C5 Hard Constraint Global, non-relaxable constraints on generation.
3.4. Unified Interaction Protocol We design a unified and comparable interaction protocol for humans and MLLMs to evaluate their prompting policy π : I → p under controlled conditions. The protocol adopts a single-turn, text-only input paradigm (stated in Section 3.2) without immediate generation feedback to isolate and measure the prompter’s ability to translate I into p in one decision, excluding feedback-driven iterative optimization. By uniformly constraining both humans and MLLMs, the protocol enables a direct and comparable assessment of prompting proficiency. At the implementation level, we provide dual interfaces for humans and MLLMs 4
AtelierEval
Figure 3. Illustration of how AtelierJudge works in AtelierEval. It decouples evaluation into two parallel processes independently applied to prompt and image: a subjective branch that performs memory-augmented quality evaluation, and an objective branch to verify constraint adherence. Both processes are executed through a skill library, from which the evaluator selects a task-conditioned sequence.
with identical semantics. The UI follows the interaction pattern of SD-WebUI (AUTOMATIC1111, 2022), a widely adopted interface in current T2I tools, and is implemented as a web interface built with Gradio (Abid et al., 2019) and deployed on HF Spaces (Wolf et al., 2020). The interface is intentionally kept minimal, retaining only the task description and text input areas to avoid introducing extraneous variables. Correspondingly, MLLMs receive the same task specification through a standardized API and output prompts. Prompt–image pairs generated from both interfaces are processed through the same pipeline to ensure fairness and reproducibility. Complete workflow and examples are provided in Appendix U.
individual constraint violations from being influenced by overall judgments, consistent with established dual-process accounts of evaluation in cognitive science (Evans, 2008; Evans & Stanovich, 2013). Based on this abstraction, as illustrated in Figure 3, AtelierJudge adopts a modular skill library that operationalizes the separation between subjective and objective evaluation. Each skill encapsulates a focused evaluation dimension, operates over either the prompt or the generated image, and is flexibly composed to support different task categories. This structured library enables an agentic evaluation process that systematically composes subjective and objective skills across the prompt and image modalities, resulting in comprehensive and interpretable evaluation of prompting proficiency.
4. AtelierJudge: Agentic Evaluation
4.1. System 1: Subjective Evaluation
Dual-Process Theory. Evaluating prompting proficiency inherently involves both assessing subjective aspects (e.g., aesthetics) and verifying explicit task requirement satisfaction. These two aspects are complementary but require fundamentally different evaluation mechanisms. To coherently integrate these distinct aspects into a unified evaluation framework, we draw inspiration from Dual-Process Theory, which characterizes reasoning as comprising two complementary systems, commonly referred to as System 1 (S1) and System 2 (S2). Conceptually, S1 is associated with holistic, experience-driven judgment, whereas S2 emphasizes analytic reasoning over explicit structures and conditions.
AtelierJudge executes subjective evaluation skills for degreebased assessment to characterize the perceptual quality of submitted results. The design draws from the memoryaugmented evaluation paradigm, shown to improve LLM judgment consistency and human alignment in complex subjective tasks (Luo et al., 2025a; Huang et al., 2026). The subjective evaluation is implemented as five modular skills, each paired with an expert-curated exemplar memory for calibration: three prompt-level skills across all task categories, and two image-level skills for OE and CO tasks. Subjective evaluation for images is not applied for IM tasks which aims to reproduce a target image. Each memory contains 120 representative exemplars spanning multiple tasks and quality levels. All exemplars are annotated with a 1–5 Likert score (Likert, 1932) for each subjective dimension, together with a brief rationale justifying the assigned score. Prompt dimensions include Instructional Clarity, Creative Elaboration, Terminology Proficiency, and Intent Formalization. Image dimensions include Mood & Atmosphere,
Our Design. In our setting, these systems are adopted as a functional abstraction. S1 supports subjective assessment by grounding judgments in gold exemplar memory. S2 supports objective verification by decomposing task alignment into explicit, independent checkpoints. This formulation provides a principled explanation for our design: perceptual impressions about quality are stabilized through memorybased calibration, while analytic decomposition prevents 5
AtelierEval
Visual Composition, Color & Lighting, and Technical Flawlessness. These dimensions are derived from prior studies on text-to-image generation and human perceptual evaluation (Otani et al., 2023; Liang et al., 2024; Chen et al., 2025a), and definitions are detailed in Appendix F. Details of memory construction are described in Appendix G.
without shared intermediate state or ordering dependencies. After skill execution, the system aggregates outputs from different skills. Multi-dimensional scores from subjective evaluation skills are directly retained, while binary outputs from objective evaluation skills are accumulated to compute a task-level constraint satisfaction rate in [0, 1]. Subjective and objective metrics remain strictly decoupled, providing interpretable and orthogonal evaluation signals for further analysis.
During evaluation, AtelierJudge first routes each submission to the corresponding skills based on the task category. The prompt and the generated image are then decoupled and evaluated independently. For each modality, the selected skill retrieves the Top-K exemplars from its associated memory based on cosine similarity in modality-specific embedding spaces, with each exemplar annotated with human-curated scores and rationales. Conditioned on the predefined scoring criteria and retrieved exemplars, AtelierJudge then produces scores for each subjective dimension.
5. Experiments 5.1. General Experimental Settings MLLM Selection. We evaluate the prompting proficiency of 8 MLLMs, selected to form a structured capability spectrum. Models are grouped into three capability tiers (T0 –T2 ), namely SOTA, mid-tier, lightweight, with comparable benchmark performance within each tier to control for capacity effects in analyses (Hendrycks et al., 2020; Yue et al., 2024; Rein et al., 2024). Specifically, T0 includes GPT-5.2 (OpenAI, 2025c), Claude-4.5-Sonnet (Cl4.5) (Anthropic, 2025), and Gemini-3-Pro-preview (Gem-3) (Google, 2025c); T1 includes GPT-4.1 (OpenAI, 2025a), Qwen-3-VL-235B-A22B (Qwen-L) (Yang et al., 2025a), and Gemini-2.0-Flash (Gem-2) (Google, 2025b); T2 includes GPT-4.1-Nano (GPT-4n) and Qwen-3-VL-8B (QwenS). GPT models enable controlled analysis along a single model lineage. Qwen models represent widely adopted open-source MLLMs (Guo et al., 2025). GPT-5.4 (OpenAI, 2026), GPT-5.2, Cl-4.5, and Gem-3 are used for evaluator selection.
4.2. System 2: Objective Check AtelierJudge executes objective evaluation skills for binary verification, determining whether results satisfy task constraints that are inherently T/F in semantics. For each task, evaluation is formulated as an expert-defined, fixed checklist of constraints. The objective skills apply zero-shot QA/VQA to perform constraint checks separately on the prompt and the generated image. Each constraint corresponds to a pair of checkpoints: a prompt checkpoint for explicit specification of constraints, and an image checkpoint for its actual realization. This paired design decouples the prompter’s intent specification from the model’s execution, enabling fine-grained analysis of prompting proficiency. The sources of task constraints vary across task types. CO task constraints are given as explicit, structured specifications. OE task constraints are derived from descriptions with clear T/F requirements. IM task constraints are derived from perceptible attributes of the target image that admit binary judgment. These checklists evaluate attribute-level semantic consistency rather than pixel-level reconstruction, allowing multiple valid realizations that preserve the target image’s key objects, relations, and style cues. Appendix E provides representative checklist examples.
T2I Backends. We evaluate prompting proficiency on 4 widely applied T2I backends, representing typical patterns in architectures and MLLM involvement: Gemini-3-ProImage-Preview (nBanana) (Google, 2025a), GPT-Image1-All (GI-1) (OpenAI, 2024), and Flux.1 Pro (Labs et al., 2025) as commercial models; SDXL (Podell et al., 2023) as the open-source model. Specifically, GI-1 incorporates an MLLM middleware to rewrite user inputs into structured internal prompts before image generation, while nBanana is built on Gem-3 and supports internal reasoning via undisclosed mechanisms. Flux Pro and SDXL do not include MLLM middleware. Flux Pro adopts a Transformer-based architecture with joint modeling of text conditions, whereas SDXL employs a diffusion-based architecture in which text conditions are injected via a relatively independent encoder (Radford et al., 2021). As Flux Pro and GI-1 achieve comparable performance on existing benchmarks (Huang et al., 2025), we use them as a roughly controlled pair when analyzing the impact of middleware. Collectively, they span the spectrum of text processing paradigms in T2I, enabling holistic analysis of prompting dynamics across diverse systems. Detailed settings for all MLLMs and T2I backends
4.3. Skill Routing AtelierJudge decomposes evaluation into five composable skills (Zhang et al., 2026), and each skill encapsulates a single scoring or verification logic. In a complete run, AtelierJudge first executes safety filter skills (see Appendix H for details) to filter out submissions with severe safety risks. For passed submissions, the system concurrently schedules skills along the subjective and objective branches, and within each branch further executes the corresponding skills on the prompt and image modalities, forming a 2 × 2 parallel evaluation process. All skills execute independently, 6
AtelierEval
are provided in Appendix K.
up-shift a large number of samples, resulting in an inflated distribution of scores 4 and 5. This makes it difficult for the models to distinguish between “good” and “perfect.” Regarding the models, Gem-3 tends to assign more balanced scores, GPT-5.2 is slightly closer to humans in relative ranking, and GPT-5.4 shows higher coarse-grained agreement but also a more lenient judging tendency. AtelierJudge addresses this failure by utilizing memory-augmented evaluation, recovering the critical gradients between scores 4/5 or 3/4. This calibration capability enables MLLMs to adjust their judgments according to human experts, establishing a decisive advantage in discriminative capability.
Human Study Design. We recruit 48 human participants as prompters into two predefined groups: 24 novices and 24 skilled users (see Appendix M for criteria and statistics). Each participant completes 30 tasks, with 10 tasks per category, randomly assigned via a balanced Latin square design. Each task is completed twice per group, and scores are averaged at the group level. To reduce ordering effects and fatigue, tasks are organized into 6 randomized rounds, each containing 5 tasks of the same category. Participants interact through our UI under the same protocol as MLLMs (Section 3.4), submit only natural-language prompts (structured formats such as lists are allowed), and are not informed of the underlying T2I backend. The use of LLM-based tools is prohibited. Informed consent and participant instructions are provided in Appendix T.1.
Objective. Table 3 demonstrates AtelierJudge’s exceptional reliability. GPT-5.4 achieves the best overall performance (95.5% Acc and 93.9% F1), combining strong promptside verification with the highest image accuracy in the most challenging visual dimension. Meanwhile, Gem-3 retains the highest prompt-side accuracy, and Cl-4.5 achieves the strongest prompt-side F1, indicating complementary strengths across textual constraint parsing and binary decision calibration. Overall, the high-precision performance, with prompt accuracy consistently exceeding 95% and image accuracy surpassing 90%, confirms that our design meets the strict requirements for automated evaluation.
5.2. Meta-Evaluation of AtelierJudge Setup. We validate AtelierJudge on an expert-annotated set of 360 stratified prompt–image pairs (one per task), detailed in Appendix P. Text and image retrieval are implemented with Nomic-Embed-Text-V1.5 (dim=512) (Nussbaum et al., 2024) and DINO-V2-Giant (Oquab et al., 2023), respectively, with a retrieval count of K=3 for both modalities. For subjective metrics, we calculate results per subjective dimension and modality, and report the macro-averaged mean absolute error (MAE), Within-1 accuracy (W1-A), and Spearman correlation (ρ) (Spearman, 1961). For objective metrics, we report the accuracy (Acc) and F1-score (F1), micro-averaged globally across all checkpoints. Design justifications, including ablations, are detailed in Appendix Q.
Table 3. Objective verification results by modality. Model
Prompt
Image
Overall
Acc (%) F1 (%) Acc (%) F1 (%) Acc (%) F1 (%) GPT-5.4 GPT-5.2 Gem-3 Cl-4.5
97.6 96.6 97.8 97.4
97.2 95.8 97.0 98.0
93.4 91.8 90.7 86.1
90.6 90.6 89.6 86.6
95.5 94.2 94.3 91.7
93.9 93.1 93.3 92.3
Table 2. Subjective verification results. Human performance is estimated via leave-one-out cross-validation among three experts
5.3. Benchmarking Prompting Proficiency Model
MAE ↓
W1-A ↑
ρ↑
We design two MLLM prompting strategies to simulate ”Novice” behavior via direct natural language prompting and ”Skilled” behavior via structured reasoning, with specific prompts provided in Appendix L. We employ GPT-5.4 as the core evaluator, selected for its superior human alignment in subjective evaluation and robustness in visual evaluation as verified above. Regarding metric calculation, we first aggregate modality-specific subjective and objective scores at the task instance level, which are subsequently analyzed or aggregated for specific comparisons. Case studies and detailed results are provided in Appendix O and R. Our key Observations are presented below.
Base Ours∆(%) Base Ours∆(%) Base Ours∆(%) GPT-5.4 GPT-5.2 Gem-3 Cl-4.5 Human
0.74 0.72 0.65 0.78
0.33↓55.4 0.34↓52.8 0.35↓46.2 0.37↓52.6 0.29
0.67 0.64 0.68 0.61
0.95↑41.8 0.93↑45.3 0.93↑36.8 0.91↑49.2 0.97
0.55 0.56 0.51 0.48
0.81↑47.3 0.79↑41.1 0.77↑51.0 0.73↑52.1 0.83
Subjective. As shown in Table 2, integrating AtelierJudge yields significant improvements across all baselines, with GPT-5.4 achieving the strongest alignment. Its MAE decreases to 0.33, W1-A rises to 0.95, demonstrating exceptional absolute calibration precision. Crucially, on ρ, which measures ranking correctness, it reaches 0.81, narrowing the gap to human performance to just 0.02. This breakthrough stems from correcting the core defects of the baselines. Our qualitative analysis reveals that all MLLM baselines tend to
Obs.1 Effective decoupling of prompting proficiency. As presented in Table 4, across all T2I backends, prompter rankings remain highly consistent, despite localized rank inversions. T0 MLLMs consistently outperform T2 models on average across all backends, and we observe no significant homophily bias (e.g., GPT models do not exhibit 7
AtelierEval Table 4. Experimental results averaged by task categories. For both humans and MLLMs, each prompt is executed to produce 4 images using each T2I backend, and the highest-scored image is retained. This setting approximates limited retries in practical workflows. In Appendix S, our empirical analysis shows that 4 samples are sufficient for stable evaluation with negligible variance at large scales. Novice and Skilled denote prompting strategies for MLLMs and user groups for humans. Novice
Skilled
GPT-5.2 Gem-3 Cl-4.5 GPT-4.1 Gem-2 Qwen-L GPT-4n Qwen-S Human GPT-5.2 Gem-3 Cl-4.5 GPT-4.1 Gem-2 Qwen-L GPT-4n Qwen-S Human Prompt
Obj. Subj.
63.3 4.05
59.2 3.81
54.5 3.45
50.7 3.46
43.5 2.96
54.3 3.65
47.8 3.27
50.6 3.15
56.5 2.90
63.7 4.10
65.5 4.01
60.1 3.77
56.1 3.72
48.2 3.23
58.8 3.92
53.2 3.48
52.7 3.38
80.6 3.88
Image
Obj. Subj.
65.7 3.96
64.7 4.01
64.1 3.88
63.0 3.92
60.9 3.82
61.3 3.99
60.8 3.89
57.8 3.90
56.9 3.01
65.4 3.96
65.4 4.00
65.1 3.92
63.9 3.99
62.9 3.89
63.1 3.98
62.8 3.92
57.2 3.87
76.7 3.97
nBanana
Obj. Subj.
73.2 4.14
70.6 4.10
70.0 4.02
67.6 4.07
66.0 3.95
67.1 4.13
64.7 4.08
64.1 4.05
58.2 3.11
73.5 4.17
73.9 4.15
72.5 4.08
70.3 4.17
68.7 4.06
70.9 4.12
67.7 4.08
65.4 4.01
84.9 4.11
GI-1
Obj. Subj.
71.9 4.21
69.8 4.22
70.3 4.13
67.6 4.18
66.2 4.12
66.7 4.20
65.7 4.17
64.9 4.14
66.4 3.25
73.4 4.22
72.5 4.24
72.0 4.17
69.4 4.25
68.8 4.13
69.8 4.19
68.2 4.12
64.2 4.10
83.5 4.01
Flux Pro
Obj. Subj.
64.1 3.92
63.8 3.98
61.1 3.83
62.3 3.88
57.7 3.68
59.5 3.92
59.3 3.78
54.4 3.81
56.8 2.94
62.2 3.90
63.2 4.00
62.1 3.84
63.6 3.93
60.7 3.79
60.5 3.94
60.7 3.85
54.0 3.82
73.1 3.93
SDXL
Obj. Subj.
53.6 3.56
54.5 3.75
54.8 3.56
54.6 3.58
53.8 3.55
51.8 3.71
53.6 3.53
48.1 3.60
46.3 2.74
52.6 3.54
52.1 3.62
53.8 3.60
52.5 3.62
53.2 3.56
50.9 3.69
54.6 3.60
45.4 3.57
65.4 3.84
disproportionate gains on GI-1), validating that our metric captures a transferable, intrinsic capability. Moreover, the ranking of subjective scores largely aligns with objective scores, indicating that performance improvements are holistic rather than narrowly optimized. Notably, while T0 MLLMs perform on par with skilled users, even T2 MLLMs significantly outperform novice users, underscoring the potential of MLLMs in prompting. The gap is mainly reflected in objective metrics, while subjective quality is much closer.
negative returns in generated images. This trend is consistent across backends and MLLMs, suggesting a systematic over-structuring effect: under loosely constrained settings, rigid structures interfere with T2I models’ own creativity. Obs.3 The paradox of constraints. A deeper decomposition reveals a counterintuitive phenomenon in CO tasks, highlighting the non-monotonic nature of performance gains. On T2I systems with strong middleware (e.g., GI-1), as presented in Table 5, we observe severe logical interference, as external MLLM reasoning reduces objective scores from a direct-input baseline of 69.6% to 47.1%. We attribute this to conflicts between external reasoning and internal rewriting. In contrast, for models without MLLM middleware such as Flux Pro, external reasoning remains positive (32.2% → 37.5%). Crucially, skilled users are immune to this interference, achieving the highest accuracy of 81.5%. This demonstrates that high-level capability extends beyond reasoning alone and encompasses adaptability. Table 5. Performance comparison on objective scores of CO and IM tasks. Direct denotes using the task descriptions as input to T2I models. G, F, and P denoting GI-1, Flux Pro, and Prompt.
Figure 4. Performance on OE tasks across the proficiency spectrum. MLLM scores are averaged within tiers under novice prompting.
Novice Prompter
Skilled Prompter
Direct GPT-5.2 Gem-3 Cl-4.5 Human GPT-5.2 Gem-3 Cl-4.5 Human
Obs.2 Homogenization and over-structuring in creation. In OE tasks, the middleware plays a decisive role, acting as a performance baseline that significantly compresses variance among prompters. As shown in Figure 4, while the gap between T2 and T1 remains observable, the marginal gains from T1 to T0 diminish due to homogenization. This effect is also visible in the aggregated subjective scores over all tasks (Table 4). The results also suggest that T0 MLLMs outperform skilled users, possibly due to larger vocabularies and faster generation, primarily reflected in OE subjective image quality. Notably, the Skilled prompting strategy also faces homogenization, consistently yielding negligible or
CO G CO F CO P
69.6 32.2 /
47.2 38.6 37.5
45.5 39.4 36.2
48.6 34.4 27.8
62.8 50.1 48.2
48.1 38.5 35.6
46.3 39.2 37.4
46.7 38.1 33.2
81.5 72.4 78.1
IM G IM F IM P
/ / /
71.6 59.8 56.7
70.5 60.5 53.5
65.6 54.8 41.2
43.8 40.9 40.4
76.5 57.9 64.6
77.5 57.3 66.6
74.3 57.1 59.0
70.4 53.3 71.3
Obs.4 Future of Image-Augmented Prompting. IM tasks exhibit distinct behaviors. Crucially, for instance, even when facing constraint complexity comparable to CO settings (up to 20 checkpoints per task), T0 MLLMs with skilled prompts not only excel on the GI-1 backend (76.1%) but demonstrate a clear advantage over human experts (70.4%). 8
AtelierEval
We attribute it not merely to the MLLM’s powerful visual encoder, whose precision in extracting and verbalizing image features surpasses human linguistic articulation, but fundamentally to the prevalence of mimicry over planning. Under equally complex constraints, imitation-based prompting significantly outperforms pure planning. This aligns with the cognitive concept of recognition over recall (Srivastava & Vul, 2017), where access to perceptual exemplars substantially reduces the burden of reconstructing complex structures. This finding also supports an emerging workflow that utilizes MLLMs to reverse-engineer prompts from reference images for subsequent intent injection, also validated by our small-scale qualitative studies. Consequently, we empirically suggest that visual exemplar guidance, namely ImageAugmented Prompting, constitutes a promising paradigm.
ture systems may retrieve high-quality visual exemplars aligned with user intent and generate prompts by adapting these references to task-specific requirements. This approach aligns with the cognitive ease of recognition over recall, mitigating the limitations of purely linguistic reasoning in current text-to-image models. ❷ Human–LLM collaboration in prompting. While AtelierEval intentionally isolates prompters to measure intrinsic prompting proficiency, an important extension is to study collaborative workflows between humans and MLLMs (Li et al., 2026; Luo et al., 2025b; Zou et al., 2026a). Human intervention may partially mitigate constraint failures of MLLMs in constrained creation tasks, while MLLMs may in turn support humans as drafting agents with strong lexical coverage and imitation ability. Systematically evaluating such collaborative settings would require new protocols beyond single-agent prompting, which we leave for future work.
6. Limitations We note several limitations of our work: ❶ Demographic Bias in Human Studies. Participants in our study are demographically concentrated (Appendix M.2). While this distribution is intentionally aligned with the current active user base of T2I systems, human-related experimental results may carry sampling bias. AtelierJudge may inherit these biases.
❸ Beyond single-turn prompting. This benchmark adopts a single-turn interaction paradigm to enable controlled and diagnostic evaluation. Extending it to interactive, multi-turn prompting with visual feedback, tool usage, or search-based optimization (Du et al., 2026; Zou et al., 2026b) remains an open challenge and may require new task abstractions and evaluation methodologies beyond isolated prompting policies.
❷ Interaction Assumptions and Scope. By restricting AtelierEval to a single-turn, pure text-to-image prompting setting, we construct a complete and minimal task partition under this interaction assumption. This abstraction enables controlled, consistent, and diagnostic analysis of prompting proficiency, but does not cover real-world workflows involving multi-turn interaction, iterative refinement with visual feedback, multimodal input, or search-based prompt optimization algorithms.
❹ Unified Multimodal Models. An important future direction is to extend AtelierEval to unified multimodal models (UMMs) that can act both as prompters and as image generators (Deng et al., 2025; Chen et al., 2025c; Xie et al., 2025). Unlike conventional evaluations that separately assess language-side reasoning or image-generation quality, AtelierEval provides a unified protocol to disentangle intent-to-prompt translation from prompt-toimage execution within the same system. This makes it particularly suitable for studying self-prompting, promptmodel co-adaptation, and the interaction between internal reasoning and visual generation in unified models.
❸ Task Difficulty Modeling. In T2I prompting, task difficulty lacks a widely accepted objective metric. Existing work typically characterizes tasks by challenge types or counts, but cannot reliably distinguish which tasks are more difficult for humans or models. Accordingly, AtelierEval balances and reports challenge primitives without explicitly modeling task difficulty.
8. Conclusion In this paper, we extend the evaluation paradigm of T2I systems beyond model-centric, prompt-fixed benchmarking to explicitly assess prompting proficiency. To operationalize this paradigm, we introduce AtelierEval, a unified benchmark that evaluates both humans and MLLMs as prompters. We further propose AtelierJudge to enable scalable and precise evaluation with AtelierEval. By formulating prompting proficiency as a cognitive competence, AtelierEval helps transform prompting from a heuristic artifact to a principled and measurable capability, supporting the democratization of generative AI.
7. Future Work We identify four promising avenues for future research: ❶ Image-augmented and imitation-based prompting. A central insight of this work is that IM tasks achieve higher objective constraint satisfaction than symbolic prompting in CO tasks, even with comparable numbers of constraints. This suggests a promising direction for future prompting agents that leverage image-augmented or retrieval-based mechanisms, analogous to RAG (Li et al., 2025b). Instead of encoding complex constraints purely symbolically, fu9
AtelierEval
Acknowledgements
the IEEE/CVF International Conference on Computer Vision, pp. 20041–20053, 2023.
This work is supported in part by the NYUAD Center for Interdisciplinary Data Science & AI (CIDSAI), funded by Tamkeen under the NYUAD Research Institute Award CG016.
Biega, A. J., Potash, P., Daumé, H., Diaz, F., and Finck, M. Operationalizing the legal principle of data minimization for personalization. In Proceedings of the 43rd international ACM SIGIR conference on research and development in information retrieval, pp. 399–408, 2020.
Impact Statement
Bosheah, Z. and Bilicki, V. Challenges in generating accurate text in images: A benchmark for text-to-image models on specialized content. Applied Sciences, 15(5): 2274, 2025.
This work introduces AtelierEval, an evaluation framework for T2I prompting proficiency, together with its agentic evaluator, AtelierJudge. They provide infrastructure for measuring and analyzing how humans and MLLMs perform as T2I prompters across a wide range of tasks. Standardized benchmarking for T2I prompting significantly enhances the reproducibility and comparability of research on T2I prompting, reducing fragmented and repeat experimentation. By systematically evaluating the capability of MLLMs versus human prompters, our work serves as a diagnostic tool that democratizes access to high-quality creation. As open-source tools equipped with intuitive interfaces, they also provide novices with a standardized environment for self-assessment and practice.
Canossa, A., Berger, L. T., Fellner, L., van der Maden, W., Juul, J., and Zhu, J. Algorithmic creativity: How visual and ai literacy impact the use of text-to-image tools in design tasks. In International Conference on HumanComputer Interaction, pp. 3–13. Springer, 2025. Cao, T., Wang, C., Liu, B., Wu, Z., Zhu, J., and Huang, J. Beautifulprompt: Towards automatic prompt engineering for text-to-image synthesis. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: Industry Track, pp. 1–11, 2023.
However, powerful prompting tools carry risks, as they may be exploited to optimize manipulative visual content or lower the threshold for generating unsafe material. Furthermore, our evaluators inevitably inherit biases, risking the reinforcement of aesthetic and representational inequalities. To mitigate these risks, we restrict tasks to non-sensitive scenarios and employ an integrated SafetyFilter. Crucially, we promote transparency by open-sourcing our toolkit with explicit documentation of scope and limitations, while cautioning against sole reliance on automated scoring in highstakes contexts involving human creators. Detailed ethical protocols are provided in Appendix W.
Chang, Y., Feng, Y., Sun, J., Ai, J., Li, C., Zhou, S. K., and Zhang, K. Sridbench: Benchmark of scientific research illustration drawing of image generation model. arXiv preprint arXiv:2505.22126, 2025. URL https://ar xiv.org/abs/2505.22126. Chen, B., Zhang, Z., Langrené, N., and Zhu, S. Unleashing the potential of prompt engineering for large language models. Patterns, 2025a. Chen, D., Chen, R., Zhang, S., Liu, Y., Wang, Y., Zhou, H., Zhang, Q., Wan, Y., Zhou, P., and Sun, L. MLLM-as-a-judge: Assessing multimodal LLM-as-ajudge with vision-language benchmark. arXiv preprint arXiv:2402.04788, 2024. ICML 2024 (Oral).
References Abid, A., Abdalla, A., Abid, A., Khan, D., Alfozan, A., and Zou, J. Gradio: Hassle-free sharing and testing of ml models in the wild. arXiv preprint arXiv:1906.02569, 2019.
Chen, K., Lin, Z., Xu, Z., Shen, Y., Yao, Y., Rimchala, J., Zhang, J., and Huang, L. R2i-bench: Benchmarking reasoning-driven text-to-image generation. arXiv preprint arXiv:2505.23493, 2025b.
Anthropic. Claude Sonnet 4.5 System Card. Technical report, Anthropic, September 2025. URL https://ww w.anthropic.com/news/claude-sonnet-4 -5. Model: claude-sonnet-4-5-20250929.
Chen, X., Wu, Z., Liu, X., Pan, Z., Liu, W., Xie, Z., Yu, X., and Ruan, C. Janus-pro: Unified multimodal understanding and generation with data and model scaling, 2025c. URL https://arxiv.org/abs/2501.17811.
AUTOMATIC1111. Stable diffusion webui. https:// github.com/AUTOMATIC1111/stable-diffu sion-webui, 2022. Accessed: 2026-01-23.
Chen, Y. T., Smith, A. D., Reinecke, K., and To, A. Why, when, and from whom: considerations for collecting and reporting race and ethnicity data in hci. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pp. 1–15, 2023.
Bakr, E. M., Sun, P., Shen, X., Khan, F. F., Li, L. E., and Elhoseiny, M. Hrs-bench: Holistic, reliable and scalable benchmark for text-to-image models. In Proceedings of 10
AtelierEval
Chinchure, A., Shukla, P., Bhatt, G., Salij, K., Hosanagar, K., Sigal, L., and Turk, M. Tibet: Identifying and evaluating biases in text-to-image generative models. In European Conference on Computer Vision, pp. 429–446. Springer, 2024.
Evans, J. S. B. and Stanovich, K. E. Dual-process theories of higher cognition: Advancing the debate. Perspectives on psychological science, 8(3):223–241, 2013. doi: 10.1 177/1745691612460685. Feng, Y., Wang, X., Wong, K. K., Wang, S., Lu, Y., Zhu, M., Wang, B., and Chen, W. Promptmagician: Interactive prompt engineering for text-to-image creation. IEEE Transactions on Visualization and Computer Graphics, 30 (1):295–305, 2024. doi: 10.1109/TVCG.2023.3327168.
Chong, L., Lo, I.-P., Rayan, J., Dow, S., Ahmed, F., and Lykourentzou, I. Prompting for products: investigating design space exploration strategies for text-to-image generative models. Design Science, 11:e2, 2025. Citron, D. New in gemini: Custom gems and improved image generation with imagen 3. https://blog.g oogle/products-and-platforms/product s/gemini/google-gemini-update-augus t-2024/, August 2024. The Keyword (Google Blog), Accessed: 2026-01-13.
Fu, X., He, M., Lu, Y., Wang, W. Y., and Roth, D. Commonsense-t2i challenge: Can text-to-image generation models understand commonsense? arXiv preprint arXiv:2406.07546, 2024. Gani, H., Bhat, S. F., Naseer, M., Khan, S., and Wonka, P. Llm blueprint: Enabling text-to-image generation with complex and detailed prompts. arXiv preprint arXiv:2310.10640, 2023.
Cover, T. M. Elements of information theory. John Wiley & Sons, 1999. Deng, C., Zhu, D., Li, K., Gou, C., Li, F., Wang, Z., Zhong, S., Yu, W., Nie, X., Song, Z., Shi, G., and Fan, H. Emerging properties in unified multimodal pretraining, 2025. URL https://arxiv.org/abs/2505.14683.
Gao, B., Xie, H., Yu, S., Wang, Y., Zuo, W., and Zeng, W. Exploring user acceptance of al image generator: unveiling influential factors in embracing an artistic aigc software. In International Conference on AI-generated Content, pp. 205–215. Springer, 2023. doi: 10.1007/97 8-981-99-7587-7 17.
Dong, W., Xue, S., Duan, X., and Han, S. Prompt tuning inversion for text-driven image editing using diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 7430–7440, 2023.
Ghosh, D., Hajishirzi, H., and Schmidt, L. Geneval: An object-focused framework for evaluating text-to-image alignment. arXiv preprint arXiv:2310.11513, 2023.
Doubao. Doubao text-to-image technical report released! full disclosure of data processing, pre-training, and rlhf workflow. https://seed.bytedance.com/e n/blog/doubao-text-to-image-technic al-report-released-full-disclosure-o f-data-processing-pre-training-and-r lhf-workflow, March 2025. ByteDance Seed Blog, Accessed: 2026-01-13.
Google. Gemini 3 Pro Image External Model Card. https: //storage.googleapis.com/deepmind-med ia/Model-Cards/Gemini-3-Pro-Image-Mod el-Card.pdf, November 2025a. Version 2, dated 20 November 2025. Google. Gemini 2.0 flash model card. Model card (official), 2025b. URL https://modelcards.withgoogl e.com/assets/documents/gemini-2-flash .pdf.
Du, S., Liu, J., Du, W., Huang, Y., Li, J., Luo, Y., Zhang, X., Conitzer, V., and Kingsford, C. Why search when you can transfer? amortized agentic workflow design from structural priors, 2026. URL https://arxiv.org/ abs/2604.25012.
Google. Gemini 3 pro model card. Model card (preview), 2025c. URL https://storage.googleapis .com/deepmind-media/Model-Cards/Gemin i-3-Pro-Model-Card.pdf. Published November 2025. Includes details about Gemini 3 Pro capabilities and evaluation.
Esser, P., Kulal, S., Blattmann, A., Entezari, R., Müller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boesel, F., Podell, D., Dockhorn, T., English, Z., and Rombach, R. Scaling rectified flow transformers for high-resolution image synthesis. In Forty-first International Conference on Machine Learning, 2024. URL https://openre view.net/forum?id=FPnUhsQJ5B.
Gu, J., Han, Z., Chen, S., Beirami, A., He, B., Zhang, G., Liao, R., Qin, Y., Tresp, V., and Torr, P. A systematic survey of prompt engineering on vision-language foundation models. arXiv preprint arXiv:2307.12980, 2023.
Evans, J. S. B. Dual-processing accounts of reasoning, judgment, and social cognition. Annu. Rev. Psychol., 59 (1):255–278, 2008.
Guilford, J. P. The Nature of Human Intelligence. McGrawHill, New York, 1967. 11
AtelierEval
Gulzar, S., Harun, J., and Yahaya, N. Integration of generative artificial intelligence in design education: Evidence review (2020–2024). International Journal of Academic Research in Business and Social Sciences, 15(11):284– 317, 2025. doi: 10.6007/IJARBSS/v15-i11/26903.
Huang, B. and Xie, H. Promptnavi: Text-to-image generation through interactive prompt visual exploration. Computers & Graphics, 132:104417, 2025. doi: 10.1016/j.ca g.2025.104417. URL https://www.sciencedir ect.com/science/article/pii/S0097849 325002584.
Guo, J., Luo, X., Zheng, J., Wang, Y., Chang, K.-W., Wang, W., and Liu, J. Quantized-tinyllava: A new multimodal foundation model enables efficient split learning, 2025.
Huang, K., Sun, K., Xie, E., Li, Z., and Liu, X. T2icompbench: A comprehensive benchmark for open-world compositional text-to-image generation. Advances in Neural Information Processing Systems, 36:78723–78747, 2023.
Hao, Y., Chi, Z., Dong, L., and Wei, F. Optimizing prompts for text-to-image generation. Advances in Neural Information Processing Systems, 36:66923–66939, 2023.
Huang, K., Duan, C., Sun, K., Xie, E., Li, Z., and Liu, X. T2i-compbench++: An enhanced and comprehensive benchmark for compositional text-to-image generation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025.
Hartwig, S., Engel, D., Sick, L., Kniesel, H., Payer, T., Poonam, P., Glockler, M., Bauerle, A., and Ropinski, T. A survey on quality metrics for text-to-image generation. IEEE Transactions on Visualization and Computer Graphics, 2025.
Huang, W.-C., Zhang, W., Liang, Y., Bei, Y., Chen, Y., Feng, T., Pan, X., Tan, Z., Wang, Y., Wei, T., Wu, S., Xu, R., Yang, L., Yang, R., Yang, W., Yeh, C.-Y., Zhang, H., Zhang, H., Zhu, S., Zou, H. P., et al. Rethinking memory mechanisms of foundation agents in the second half: A survey. arXiv preprint arXiv:2602.06052, 2026.
Hei, N., Guo, Q., Wang, Z., Wang, Y., Wang, H., and Zhang, W. A user-friendly framework for generating modelpreferred prompts in text-to-image synthesis. In Proceedings of the AAAI Conference on Artificial Intelligence, pp. 2139–2147, 2024. Heigl, R. Generative artificial intelligence in creative contexts: a systematic review and future research agenda. Management Review Quarterly, pp. 1–38, 2025.
Jaiprakash, S. P. and Prakash, C. S. Exploring text-to-image generation models: Applications and cloud resource utilization. Computers and Electrical Engineering, 123: 110194, 2025.
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020.
Jin, Z. and Chua, T.-S. Compose your aesthetics: Empowering text-to-image models with the principles of art. arXiv preprint arXiv:2503.12018, 2025. URL https://arxiv.org/abs/2503.12018.
Hessel, J., Holtzman, A., Forbes, M., Le Bras, R., and Choi, Y. Clipscore: A reference-free evaluation metric for image captioning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 7514–7528, 2021.
Jose, C., Moutakanni, T., Kang, D., Baldassarre, F., Darcet, T., Xu, H., Li, D., Szafraniec, M., Ramamonjisoa, M., Oquab, M., et al. Dinov2 meets text: A unified framework for image-and pixel-level vision-language alignment. In Proceedings of the Computer Vision and Pattern Recognition Conference, pp. 24905–24916, 2025.
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Advances in Neural Information Processing Systems, volume 30, 2017.
Kahneman, D. Thinking, fast and slow. macmillan, 2011.
Holzner, N., Maier, S., and Feuerriegel, S. Generative ai and creativity: A systematic literature review and metaanalysis. arXiv preprint arXiv:2505.17241, 2025.
Knoth, N., Tolzin, A., Janson, A., and Leimeister, J. M. Ai literacy and its implications for prompt engineering strategies. Computers and Education: Artificial Intelligence, 6:100225, 2024.
Hu, Y., Liu, B., Kasai, J., Wang, Y., Ostendorf, M., Krishna, R., and Smith, N. A. Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 20406–20417, 2023.
Ko, H.-K., Park, G., Jeon, H., Jo, J., Kim, J., and Seo, J. Large-scale text-to-image generation models for visual artists’ creative works. In Proceedings of the 28th international conference on intelligent user interfaces, pp. 919–933, 2023. 12
AtelierEval
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y. Large language models are zero-shot reasoners. Advances in neural information processing systems, 35: 22199–22213, 2022.
Li, H., Dong, Q., Chen, J., Su, H., Zhou, Y., Ai, Q., Ye, Z., and Liu, Y. Llms-as-judges: a comprehensive survey on llm-based evaluation methods. arXiv preprint arXiv:2412.05579, 2024b.
Krippendorff, K. Content analysis: An introduction to its methodology. Sage publications, 2018.
Li, J., Li, Y., Han, L., Tang, R., and Wang, W. Towards generalizable implicit in-context learning with attention routing. arXiv preprint arXiv:2509.22854, 2025a.
Kumar, N. Midjourney statistics 2026 (active users & revenue). https://www.demandsage.com/midjo urney-statistics/, 2026. DemandSage, accessed 2026-01-27.
Li, J., Chen, Z., Luo, H., and Salam, H. Prefix: Understand and adapt to user preference in human-agent interaction. arXiv preprint arXiv:2602.06714, 2026.
Labs, B. F., Batifol, S., Blattmann, A., Boesel, F., Consul, S., Diagne, C., Dockhorn, T., English, J., English, Z., Esser, P., Kulal, S., Lacey, K., Levi, Y., Li, C., Lorenz, D., Müller, J., Podell, D., Rombach, R., Saini, H., Sauer, A., and Smith, L. Flux.1 kontext: Flow matching for incontext image generation and editing in latent space, 2025. URL https://arxiv.org/abs/2506.15742.
Li, W., Wang, J., and Zhang, X. Promptist: Automated prompt optimization for text-to-image synthesis. In CCF international conference on natural language processing and Chinese computing, pp. 295–306. Springer, 2024c. Li, Y., Cao, Y., He, H., Cheng, Q., Fu, X., Xiao, X., Wang, T., and Tang, R. M²IV: Towards efficient and fine-grained multimodal in-context learning via representation engineering. In Second Conference on Language Modeling, 2025b. URL https://openreview.net/forum ?id=9ffYcEiNw9.
Lee, C., Roy, R., Xu, M., Raiman, J., Shoeybi, M., Catanzaro, B., and Ping, W. Nv-embed: Improved techniques for training llms as generalist embedding models. arXiv preprint arXiv:2405.17428, 2024.
Liang, Y., He, J., Li, G., Li, P., Klimovskiy, A., Carolan, N., Sun, J., Pont-Tuset, J., Young, S., Yang, F., et al. Rich human feedback for text-to-image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 19401–19411, 2024.
Lee, T., Yasunaga, M., Meng, C., Mai, Y., Park, J. S., Gupta, A., Zhang, Y., Narayanan, D., Teufel, H., Bellagente, M., Kang, M., Park, T., Leskovec, J., Zhu, J.-Y., Li, F.-F., Wu, J., Ermon, S., and Liang, P. S. Holistic evaluation of textto-image models. In Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., and Levine, S. (eds.), Advances in Neural Information Processing Systems, volume 36, pp. 69981–70011. Curran Associates, Inc., 2023. URL https://proceedings.neurips.cc/paper _files/paper/2023/file/dd83eada2c3c7 4db3c7fe1c087513756-Paper-Datasets_an d_Benchmarks.pdf.
Likert, R. A technique for the measurement of attitudes. Archives of psychology, 1932. Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L. Microsoft coco: Common objects in context. In European conference on computer vision, pp. 740–755. Springer, 2014.
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., Riedel, S., and Kiela, D. Retrieval-augmented generation for knowledge-intensive nlp tasks. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Advances in Neural Information Processing Systems, volume 33, pp. 9459–9474. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/p aper_files/paper/2020/file/6b4932302 05f780e1bc26945df7481e5-Paper.pdf.
Liu, V. and Chilton, L. B. Design guidelines for prompt engineering text-to-image generative models. In Proceedings of the 2022 CHI conference on human factors in computing systems, pp. 1–23, 2022. Liu, Y., He, M., Yao, F., Ji, Y., Tao, S., Du, J., Li, J., Gao, J., Li, Z., Yang, H., et al. Taming text-to-image synthesis for novices: User-centric prompt generation via multiturn guidance. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025.
Li, B., Lin, Z., Pathak, D., Li, J., Fei, Y., Wu, K., Ling, T., Xia, X., Zhang, P., Neubig, G., and Ramanan, D. Genaibench: Evaluating and improving compositional textto-visual generation. arXiv preprint arXiv:2406.13743, 2024a.
Luo, H., Deng, Z., Chen, R., and Liu, Z. Faintbench: A holistic and precise benchmark for bias evaluation in text-to-image models. arXiv preprint arXiv:2405.17814, 2024a. 13
AtelierEval
Luo, H., Deng, Z., Huang, H., Liu, X., Chen, R., and Liu, Z. Versusdebias: Universal zero-shot debiasing for textto-image models via slm-based prompt engineering and generative adversary. arXiv preprint arXiv:2407.19524, 2024b.
OpenAI. Gpt image 1: Natively multimodal image generation. OpenAI API Documentation, 2024. URL https://platform.openai.com/docs/mod els/gpt-image-1. First generation of native GPTbased image synthesis models, distinct from DALL-E 3 diffusion models.
Luo, H., Huang, H., Deng, Z., Li, X., Wang, H., Jin, Y., Liu, Y., Xu, W., and Liu, Z. Bigbench: A unified benchmark for evaluating multi-dimensional social biases in text-toimage models. arXiv preprint arXiv:2407.15240, 2024c.
OpenAI. Introducing gpt-4.1 in the api, April 2025a. URL https://openai.com/index/gpt-4-1/. Official announcement blog post.
Luo, H., Dai, S., Ni, C., Li, X., Zhang, G., Wang, K., Liu, T., and Salam, H. Agentauditor: Human-level safety and security evaluation for llm agents. arXiv preprint arXiv:2506.00641, 2025a.
OpenAI. The new chatgpt images with gpt-image-1.5. ht tps://openai.com/index/new-chatgpt-i mages-is-here/, December 2025b. Introduces gptimage-1.5 with improved editing, text rendering, and 4x faster generation.
Luo, H., Ni, C., Wen, J., Huang, Z., Wang, Y., Liao, B., Chung, S., Jin, Y., Li, X., Xu, W., Wang, X., and Salam, H. Hai-eval: Measuring human-ai synergy in collaborative coding, 2025b. URL https://arxiv.org/abs/ 2512.04111.
OpenAI. Update to gpt-5 system card: Gpt-5.2, 2025c. URL https://openai.com/index/gpt-5-syste m-card-update-gpt-5-2/. Accessed: 2026-0125.
Luo, H., Huang, Z., Huang, H., Deng, Z., Chen, R., Li, X., Liu, Z., and Salam, H. Biasig: Benchmarking multidimensional social biases in text-to-image models. arXiv preprint arXiv:2604.11934, 2026.
OpenAI. Gpt-5.4 thinking system card, Mar 2026. URL https://openai.com/index/gpt-5-4-thi nking-system-card/. Accessed: 2026-05-17. Oppenlaender, J. A taxonomy of prompt modifiers for text-to-image generation. Behaviour & Information Technology, 43(15):3763–3776, 2024.
Mahajan, S., Rahman, T., Yi, K. M., and Sigal, L. Prompting hard or hardly prompting: Prompt inversion for text-toimage diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 6808–6817, 2024.
Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., ElNouby, A., et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023.
Mahdavi Goloujeh, A., Sullivan, A., and Magerko, B. Is it ai or is it me? understanding users’ prompt journey with text-to-image generative ai tools. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, pp. 1–13, 2024.
Otani, M., Togashi, R., Sawai, Y., Ishigami, R., Nakashima, Y., Rahtu, E., Heikkilä, J., and Satoh, S. Toward verifiable and reproducible human evaluation for text-to-image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 14277–14286, 2023.
Mañas, O., Astolfi, P., Hall, M., Ross, C., Urbanek, J., Williams, A., Agrawal, A., Romero-Soriano, A., and Drozdzal, M. Improving text-to-image consistency via automatic prompt optimization. Transactions on Machine Learning Research, 2024. URL https://arxiv.or g/abs/2403.17804.
Pavlichenko, N. and Ustalov, D. Best prompts for text-toimage models and how to find them. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 2067– 2071, 2023.
Mo, W., Zhang, T., Bai, Y., Su, B., Wen, J.-R., and Yang, Q. Dynamic prompt optimizing for text-to-image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 26627–26636, 2024.
Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., and Rombach, R. Sdxl: Improving latent diffusion models for high-resolution image synthesis. arXiv preprint arXiv:2307.01952, 2023.
Norman, D. The design of everyday things: Revised and expanded edition. Basic books, 2013.
Qin, X., Li, S., Cai, Y., and Wang, L. Enhancing counterfactual explanations with feasibility and diversity. In 2025 IEEE International Conference on Data Mining Workshops (ICDMW), pp. 2310–2319. IEEE, 2025a.
Nussbaum, Z., Morris, J. X., Duderstadt, B., and Mulyar, A. Nomic embed: Training a reproducible long context text embedder. arXiv preprint arXiv:2402.01613, 2024. 14
AtelierEval
Qin, X., Yu, R., Khayati, A., Qiu, Z., Zou, G., Li, Y., and Wang, L. Interpretable and interactive deep survival analysis with time-dependent extreme gradient integration. In 2025 IEEE International Conference on Data Mining (ICDM), pp. 673–682. IEEE, 2025b.
Shannon, C. E. A mathematical theory of communication. The Bell system technical journal, 27(3):379–423, 1948.
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. Learning transferable visual models from natural language supervision. In Meila, M. and Zhang, T. (eds.), Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pp. 8748–8763. PMLR, 18–24 Jul 2021. URL https:// proceedings.mlr.press/v139/radford21a. html. Rein, D., Hou, B. L., Stickland, A. C., Petty, J., Pang, R. Y., Dirani, J., Michael, J., and Bowman, S. R. Gpqa: A graduate-level google-proof q&a benchmark. In First Conference on Language Modeling, 2024.
Siméoni, O., Vo, H. V., Seitzer, M., Baldassarre, F., Oquab, M., Jose, C., Khalidov, V., Szafraniec, M., Yi, S., Ramamonjisoa, M., et al. Dinov3. arXiv preprint arXiv:2508.10104, 2025. Sordo, Z., Chagnon, E., and Ushizima, D. A review on generative ai for text-to-image and image-to-image generation and implications to scientific images. arXiv preprint arXiv:2502.21151, 2025. Spearman, C. The proof and measurement of association between two things., 1961. Srivastava, N. and Vul, E. A simple model of recognition and recall memory. Advances in Neural Information Processing Systems, 30, 2017. Strobelt, H., Webson, A., Sanh, V., Hoover, B., Beyer, J., Pfister, H., and Rush, A. M. Interactive and visual prompt engineering for ad-hoc task adaptation with large language models. IEEE transactions on visualization and computer graphics, 29(1):1146–1156, 2022.
Ribeiro, M. T., Wu, T., Guestrin, C., and Singh, S. Beyond accuracy: Behavioral testing of NLP models with checklist. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 4902– 4912. Association for Computational Linguistics, 2020.
Sun, K., Fang, R., Duan, C., Liu, X., and Liu, X. T2ireasonbench: Benchmarking reasoning-informed textto-image generation. arXiv preprint arXiv:2508.17472, 2025.
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E. L., Ghasemipour, K., Gontijo Lopes, R., Karagol Ayan, B., Salimans, T., Ho, J., Fleet, D. J., and Norouzi, M. Photorealistic text-to-image diffusion models with deep language understanding. In Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., and Oh, A. (eds.), Advances in Neural Information Processing Systems, volume 35, pp. 36479–36494. Curran Associates, Inc., 2022. URL https://proceedings.neurips.cc/p aper_files/paper/2022/file/ec795aead ae0b7d230fa35cbaf04c041-Paper-Confere nce.pdf.
Tao, S., Liang, I., Peng, C., Wang, Z., Palani, S., and Dow, S. P. Designweaver: Dimensional scaffolding for text-toimage product design. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, pp. 1–26, 2025. Tsao, J., Liang, C. X., Nogues, C., and Wong, A. Perceptions and integration of generative artificial intelligence in creative practices and industries: A scoping review and conceptual model. AI & Society, 2025. doi: 10.1007/s00146-025-02667-2. Wang, L., Qin, X., Jiang, J., Li, Y., and Liaw, W. Interactive explainable deep survival analysis. In 2024 46th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC), pp. 1–4. IEEE, 2024a.
Sanchez, T. Examining the text-to-image community of practice: Why and how do people prompt generative ais? In Proceedings of the 15th Conference on Creativity and Cognition, pp. 43–61, 2023. Schulhoff, S., Ilie, M., Balepur, N., Kahadze, K., Liu, A., Si, C., Li, Y., Gupta, A., Han, H., Schulhoff, S., Dulepet, P. S., Vidyadhara, S., Ki, D., Agrawal, S., Pham, C., Kroiz, G., Li, F., Tao, H., Srivastava, A., Costa, H. D., Gupta, S., Rogers, M. L., Goncearenco, I., Sarli, G., Galynker, I., Peskoff, D., Carpuat, M., White, J., Anadkat, S., Hoyle, A., and Resnik, P. The prompt report: A systematic survey of prompt engineering techniques, 2025. URL https://arxiv.org/abs/2406.06608.
Wang, Z., Huang, Y., Song, D., Ma, L., and Zhang, T. Promptcharm: Text-to-image generation through multimodal prompting and refinement. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, pp. 1–21, 2024b. Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in 15
AtelierEval
neural information processing systems, 35:24824–24837, 2022.
Yeh, S.-Y., Li, Y., Park, S.-H., Oh, G., Wang, X., Song, M., Yu, Y., and Lai, S.-H. Tipo: Text to image with text presampling for prompt optimization. arXiv preprint arXiv:2411.08127, 2024.
Wiles, O., Zhang, C., Albuquerque, I., Kajić, I., Wang, S., Bugliarello, E., Onoe, Y., Papalampidi, P., Ktena, I., Knutsen, C., Rashtchian, C., Nawalgaria, A., PontTuset, J., and Nematzadeh, A. Revisiting text-to-image evaluation with gecko: On metrics, prompts, and human ratings, 2025. URL https://arxiv.org/abs/24 04.16820.
You, J., Lin, Y., and Hu, B. Enhancing aesthetic image generation with reinforcement learning guided prompt optimization in stable diffusion. Journal of Visual Communication and Image Representation, pp. 104641, 2025. Yuan, S., Sun, X., Kim, H., Yu, S., and Tomasi, C. Optical flow training under limited label budget via active learning. In European Conference on Computer Vision (ECCV), 2022.
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 conference on empirical methods in natural language processing: system demonstrations, pp. 38–45, 2020.
Yue, X., Ni, Y., Zhang, K., Zheng, T., Liu, R., Zhang, G., Stevens, S., Jiang, D., Ren, W., Sun, Y., et al. Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9556–9567, 2024.
Wu, X., Hao, Y., Sun, K., Chen, Y., Zhu, F., Zhao, R., and Li, H. Human preference score v2: A solid benchmark for evaluating human preferences of text-to-image synthesis. arXiv preprint arXiv:2306.09341, 2023.
Zamfirescu, P. J. D., Wong, R. Y., Hartmann, B., and Yang, Q. Why johnny can’t prompt: how non-ai experts try (and fail) to design llm prompts. In Proceedings of the 2023 CHI conference on human factors in computing systems, pp. 1–21, 2023.
Wu, Z., Zhang, H., Lin, F., Xu, W., Xu, X., Chen, Y., Zou, H. P., Chen, S., Zhang, W., Liu, X., Yu, P. S., and Wang, H. Gam: Hierarchical graph-based agentic memory for llm agents, 2026. URL https://arxiv.org/abs/ 2604.12285.
Zhang, H., Wang, M., He, S., and Ming, A. Ak4prompts: aesthetics-driven automatically keywords-ranking for prompts in text-to-image models. In Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, pp. 1661–1669, 2024.
Xiang, D., Xu, W., Chu, K., Ding, T., Shen, Z., Zeng, Y., Su, J., and Zhang, W. Promptsculptor: Multi-agent based text-to-image prompt optimization. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pp. 774– 786, 2025.
Zhang, H., Fan, S., Zou, H. P., Chen, Y., Wang, Z., Zhou, J., Li, C., Huang, W.-C., Yao, Y., Zheng, K., Liu, X., Li, X., and Yu, P. S. Coevoskills: Self-evolving agent skills via co-evolutionary verification, 2026. URL https: //arxiv.org/abs/2604.01687.
Xie, J., Mao, W., Bai, Z., Zhang, D. J., Wang, W., Lin, K. Q., Gu, Y., Chen, Z., Yang, Z., and Shou, M. Z. Show-o: One single transformer to unify multimodal understanding and generation, 2025. URL https://arxiv.org/abs/ 2408.12528.
Zhang, N. and Tang, H. Text-to-image synthesis: A decade survey. arXiv preprint arXiv:2411.16164, 2024.
Yang, A., Li, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Gao, C., Huang, C., Lv, C., Zheng, C., Liu, D., Zhou, F., Huang, F., Hu, F., Ge, H., Wei, H., Lin, H., Tang, J., Yang, J., Tu, J., Zhang, J., Yang, J., Yang, J., Zhou, J., Zhou, J., Lin, J., Dang, K., Bao, K., Yang, K., Yu, L., Deng, L., Li, M., Xue, M., Li, M., Zhang, P., Wang, P., Zhu, Q., Men, R., Gao, R., Liu, S., Luo, S., Li, T., Tang, T., Yin, W., Ren, X., Wang, X., Zhang, X., Ren, X., Fan, Y., Su, Y., Zhang, Y., Zhang, Y., Wan, Y., Liu, Y., Wang, Z., Cui, Z., Zhang, Z., Zhou, Z., and Qiu, Z. Qwen3 technical report, 2025a. URL https: //arxiv.org/abs/2505.09388.
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., and Stoica, I. Judging LLM-as-a-judge with MT-bench and chatbot arena. In Advances in Neural Information Processing Systems, volume 36, 2023. Zhuge, M., Zhao, C., Ashley, D., Wang, W., Khizbullin, D., Xiong, Y., Liu, Z., Chang, E., Krishnamoorthi, R., Tian, Y., et al. Agent-as-a-judge: Evaluate agents with agents. arXiv preprint arXiv:2410.10934, 2024. Zou, H. P., Huang, W.-C., Wu, Y., Guo, J., Chen, Y., Miao, C., Nguyen, H., Zhou, Y., Zhang, W., Fang, L., Zhang, H., Wang, F., Zhang, P., Wang, H., He, L., Li, Y., Li, D., Jiang, R., Liu, X., and Yu, P. S. Llm-based human-agent
Yang, P., Cheung, N.-M., and Ma, X. Text to image generation and editing: A survey. arXiv preprint arXiv:2505.02527, 2025b. 16
AtelierEval
collaboration and interaction systems: A survey, 2026a. URL https://arxiv.org/abs/2505.00753. Zou, H. P., Miao, C., Huang, W.-C., Chen, Y., Zhou, Y., Zhang, H., Wu, Y., Fang, L., Gu, Z., Zhang, Z., Zheng, K., Wang, F., Nian, Y., Li, S., Fan, W., He, L., Zhang, W., Liu, X., and Yu, P. S. When users change their mind: Evaluating interruptible agents in long-horizon web navigation, 2026b. URL https://arxiv.org/ abs/2604.00892.
17
AtelierEval
APPENDIX A
B
C
Formal Justification of Task Completeness
20
A.1
Information Sources and Primitive Operations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20
A.2
Text-only Tasks: OE and CO . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20
A.3
Image-conditioned Tasks: IM . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21
A.4
Coverage and Minimality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21
A.5
Scope . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21
Application Context Categorization
22
B.1
Tagging Scheme . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22
B.2
Distribution of Tags Across Task Categories . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22
Dataset Construction Details
23
C.1
OE Task Instantiation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23
C.2
CO Task Instantiation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23
C.3
IM Task Instantiation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24
D
Model/Seed-Agnostic Design of IM Tasks
25
E
Examples of Task Instances
25
F
E.1
OE Task Example . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25
E.2
CO Task Example . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27
E.3
IM Task Example . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29
Subjective Evaluation Dimensions
31
F.1
Image-Level Dimensions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31
F.2
Prompt-Level Dimensionse . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 32
G
Evaluation Exemplar Memory Construction
32
H
Details of Safety Filter Skill
33
I
Scoring Guidelines for Subjective Evaluation
34
J
Prompts for AtelierJudge
35
J.1
Subjective Skill Prompts . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 35
J.2
Objective Skill Prompts . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 37
K
Model Hyperparameters
39
L
Prompts for MLLM Novice and Skilled Conditions
39
M
N
L.1
Novice MLLM Prompts . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 39
L.2
Skilled MLLM Prompts . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 40
Participant Selection & Statistics
42
M.1
Participant Selection Criteria . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42
M.2
Participant Demographic Statistics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 43
Human Expert Involvement
44
18
AtelierEval
O
Subjective Evaluation Case Studies
44
O.1
Subjective Prompt Evaluation (Open-Ended Task) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 44
O.2
Subjective Image Evaluation (Constrained Task) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 44
P
Details of Validation Set
45
Q
Design Validation & Ablation Study of AtelierJudge
45
R
Detailed Benchmarking Results
47
S
Stability Analysis of Evaluation Scale
48
T
User Study Materials
49
U
T.1
Informed Consent Form . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 49
T.2
Pre-Test Questionnaire . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 51
Participant Workflow and Interface Design
53
U.1
Welcome and Authentication . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 53
U.2
Assessment Structure and Randomization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 53
U.3
Task Type Interfaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 54
U.4
Coverage and Minimality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 57
V
Computational Resource Consumption
58
W
Detailed Ethical Considerations and Procedures
58
W.1
Ethics Approval and Consent . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 58
W.2
Participant Anonymity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 58
19
AtelierEval
A. Formal Justification of Task Completeness This appendix justifies the completeness claim in Section 3.2. Under a restricted interaction setting, we argue that the three task categories (OE, CO, IM) form a complete and minimal set of information-processing primitives for prompting. Assumption 1 (Single-turn, Pure Text-to-Image).
We work under the following setting:
The prompter interacts with the text-to-image model in a single turn and produces exactly one textual prompt. The generation model conditions solely on this text prompt, without access to any image-based inputs, intermediate feedback, or external control signals. This is exactly the interface instantiated by AtelierEval. Under Assumption 1, prompting can be viewed as an information-processing problem at the model interface, in the sense of a source whose information must be transmitted through a constrained channel (Shannon, 1948). Let I denote the latent abstract intent, O the observable task input, p the textual prompt, π the prompting policy, M the text-to-image model, and S(·, ·) an intent-consistency score. The prompter selects p = π(O) to maximize expected intent alignment: π ∗ = arg max EI EM S(M(π(O)), I) . (4) π
A.1. Information Sources and Primitive Operations We decompose the observable input as O = (Otext , Oimg ), where Otext is the textual description and Oimg an optional target image. Under Assumption 1, any prompting policy takes the form p = π(I, Otext , Oimg ),
(5)
and the model M only observes p. Thus all task-relevant information that can influence M must flow from the three sources (I, Otext , Oimg ) into the single textual channel. We distinguish three primitive information-processing operations at this interface: 1. Textual convergence: transforming, compressing, and reconciling constraints that are already present in Otext into a single executable prompt (reformatting, reordering, aggregating, resolving conflicts). 2. Intent-driven expansion: injecting additional structure, detail, and style into p that is not specified in Otext or Oimg but is consistent with the abstract intent I. 3. Visual-to-textual encoding: describing aspects of Oimg (composition, layout, style, key objects) in text so that they can be realized by M. By construction, every non-trivial contribution to p must belong to exactly one of these three cases: it either reorganizes information already in Otext , introduces new information from I, or encodes information from Oimg . There is no additional side channel into M. A.2. Text-only Tasks: OE and CO When the input is purely textual, we write Otext for clarity. Let H(I | Otext )
(6)
denote the conditional entropy of the abstract intent I given the textual input Otext , using standard information-theoretic notation (Cover, 1999). We use H(I | Otext ) qualitatively: it measures the residual uncertainty about I after reading the task description, allowing us to distinguish regimes of relatively higher or lower uncertainty. Real task descriptions often mix abstract, underspecified cues (themes, moods, high-level scenarios) with explicit, low-level constraints (counts, layouts, text rendering). Accordingly, both convergence and expansion can appear in a single text-only task. We therefore define OE and CO by their dominant operation: 20
AtelierEval
• Open-ended creation (OE): H(I | Otext ) is relatively high, and the dominant burden on the prompter is intent-driven expansion. Most of the non-trivial content in p originates from I rather than Otext . • Constrained creation (CO): H(I | Otext ) is relatively low, but Otext contains many explicit constraints not yet compressed into a prompt. The dominant burden is textual convergence: most of the non-trivial content in p reorganizes and integrates information already present in Otext . Our OE and CO tasks are constructed so that one of these operations is clearly dominant while the other is relatively simple or trivial. A.3. Image-conditioned Tasks: IM When Oimg is present, the prompter cannot pass it directly to M and must instead encode its content into text. Let Iimg denote the semantic intent induced by the target image, and let D(·, ·) be a semantic distortion measure. The core information-processing problem in IM tasks can be written as p∗ = arg min D Iimg , M(p) , (7) p
which is analogous to a fidelity-constrained encoding problem in rate–distortion theory (Cover, 1999). Since the only path from Oimg to M is through p, the prompter must perform visual-to-textual encoding. In our benchmark, IM tasks are imitation-oriented, so this encoding is explicitly fidelity-preserving. More creative image-conditioned tasks (e.g., “keep the composition but change the mood”) can be decomposed into the same primitives: encode the preserved aspects of Oimg (IM), expand I with new elements or styles (OE), and converge all constraints into a single prompt (CO). A.4. Coverage and Minimality Coverage. Under Assumption 1, any prompting policy π can only draw information from (I, Otext , Oimg ). Every nontrivial contribution to p is therefore: (i) convergence on Otext , (ii) expansion from I, or (iii) encoding from Oimg . Text-only tasks (Oimg absent) necessarily combine (i) and (ii), with OE and CO distinguished by which operation is dominant. Image-conditioned tasks necessarily invoke (iii), because information about Oimg has no other path to the model. Mixed cases are compositions of these operations. Thus, OE, CO, and IM together cover all possible information-processing patterns available at the prompt interface. Minimality.
Each primitive is necessary. There exist tasks whose successful solution requires:
• non-trivial intent-driven expansion (OE-style underspecified scenarios); • non-trivial textual convergence (CO-style densely constrained specifications); • non-trivial visual-to-text encoding (IM-style imitation from a target image). Removing any primitive would make the corresponding family of tasks inexpressible as a single prompt. Conversely, any hypothetical new primitive must either: (a) reorganize information already present in Otext or Oimg (a special case of convergence or encoding), or (b) inject information consistent with I that is not specified in Otext or Oimg (a special case of expansion). It does not introduce a new information-flow direction, and can be expressed as a tactic within OE, CO, or IM, or as their composition. In this sense, OE, CO, and IM form a complete and minimal task partition for prompting under Assumption 1, from an information-theoretic perspective. A.5. Scope If Assumption 1 is relaxed (e.g., multi-round interaction, direct image conditioning at the model interface, external tools), additional information channels become available and further operations—such as iterative refinement or tool-augmented retrieval—may be required. These settings fall outside the scope of AtelierEval and are left for future work. 21
AtelierEval
B. Application Context Categorization This section introduces our application context categorization and tagging scheme, which is used to characterize and summarize how our tasks map onto real-world T2I creative workflows, ensuring broad coverage during task construction and supporting qualitative analysis. This multi-dimensional, non-mutually exclusive design reflects that real-world T2I briefs often span multiple overlapping concerns (e.g., characters and objects both being central) (Lin et al., 2014), and thus better captures how entities, structure, style, and thematic context jointly shape prompting requirements. B.1. Tagging Scheme While task construction is driven by challenge primitives and cognitive categories, we explicitly considered common T2I application contexts throughout the design process. We treat the application context of each task as a multi-label annotation vector
C = {E, S, V, T },
(8)
where E denotes Entities, S denotes Structure, V denotes Visual Style, and T denotes Theme & Context. The tagging scheme is not intended as a strict ontology. Instead, it is a lightweight, non-exclusive label set designed to reflect how tasks map onto common T2I application scenarios. Concretely: ➠ Entities (E) describe what is depicted. ➠ Structure (S) describes how content is arranged in space or sequence. ➠ Visual Style (V ) describes how the image is rendered. ➠ Theme & Context (T ) describes the high-level narrative or aesthetic context. Each task may carry multiple tags both across and within dimensions, e.g., a fantasy character sheet rendered in cel-shaded style with a knolling layout. Tags are only used for coverage analysis and qualitative discussion. B.2. Distribution of Tags Across Task Categories We summarize the application-context coverage considered during task construction over all 360 tasks using the tagging scheme. Table 6 reports the occurrence counts of all used tags.
Table 6. Tags and their occurrence counts in the three task categories. The table is intended as a coverage summary rather than a balanced distribution, illustrating the broad coverage of real-world T2I application scenarios considered in AtelierEval.
Tag
OE
CO
IM
Entities #Object #Character #Environment #Typography #Data Element
OE
CO
IM
43 50 7 15 5
37 12 52 26 8
48 6 18 27 19
21 24 26 25 15 27 16
56 14 9 2 27 15 3
35 17 13 11 13 4 1
Visual Style 56 49 45 6 17
96 37 34 76 32
39 58 54 18 14
Structure #Portrait CloseUp #Full Body Shot #Schematic Diagram #Knolling Layout #Sequential Panel #UI Interface #Isometric View
Tag #Photorealistic #Traditional Media #Vector Flat #3D Render #Cel Shaded Theme & Context
27 18 18 10 7 8 0
25 21 25 15 6 6 6
22 20 26 6 15 12 4
#Corporate Clean #SciFi Cyberpunk #Fantasy Mythic #Abstract Conceptual #Cute Pop #Retro Vintage #Horror Dark
22
AtelierEval
C. Dataset Construction Details Table 7. Structural challenge load statistics aggregated across the 360 tasks. For each task category, we report summary statistics of the number of semantic and constraint challenge primitive instances per task.
# semantic primitive instances / task
# constraint primitive instances / task
Category
Min
Max
Mean
Std
Min
Max
Mean
Std
OE CO IM
6 3 3
10 4 7
8.2 3.7 4.3
0.9 0.4 1.0
0 10 7
0 18 16
0.0 12.7 12.4
0.0 3.1 2.8
This section supplements Section 3.3 by detailing the task instantiation procedures. We organize the discussion by the three task categories and describe the corresponding instantiation procedures for each. Illustrative examples are provided in Appendix E. The summarized statistics of challenge primitives are provided in Table 7. C.1. OE Task Instantiation Open-ended Creation (OE) tasks are designed to evaluate a prompter’s ability to translate abstract, fuzzy, and unstructured requests into executable prompts. They focus on upstream semantic analysis and creative translation. OE tasks are presented as full natural-language paragraphs that deliberately introduce narrative noise to simulate real creative briefs. The description contains a substantial amount of contextually meaningful information that is common in real-world workflows but has no direct visual counterpart, and prompters must filter such narrative context, identify visually relevant intent, and translate it into an executable specification. OE tasks only use semantic challenge primitives {Si } and exclude explicit {Cj }. They are instantiated across diverse T2I application contexts to avoid tying the evaluation to any particular subject matter or style. Formally, we write an OE task as:
tOE =
Ntext , n1 S1 + n2 S2 + n3 S3 + n4 S4 , At ,
6 ≤
4 X
ni ≤ 10,
(9)
i=1
where Ntext is the high-noise natural-language description, At is the multi-label application-context tag vector defined in Appendix B, and ni ∈ Z≥0 is the number of instances of semantic primitive Si embedded in the task. P The challenge load of an OE task is thus determined jointly by the total number of embedded semantic cues i ni and the translation difficulty of their concrete instances. Even within the same primitive type, instances can vary in how hard they are to render into visual terms (e.g., a simple emotional descriptor versus a highly abstract aesthetic intent, or a broad audience description versus a highly specific cultural group). During task design, we explicitly control both factors: every OE task contains between 6 and 10 key semantic cues that require translation, and we avoid stacking multiple high-difficulty instances within a single task, thereby constraining the challenge load to a reasonable and comparable range across OE tasks. Before inclusion in the benchmark, all OE tasks undergo manual validation to ensure that tasks are understandable and executable by two independent domain experts. Each expert independently translates the task into prompts; a task is accepted only if both agree that the intent is clear and unambiguous, and that all four text-to-image models used in our experiments can produce images within multiple generations that broadly satisfy the intended requirements. C.2. CO Task Instantiation Constrained Creation (CO) tasks are designed to evaluate a prompter’s ability to integrate multiple explicit constraints into an executable prompt under clearly specified requirements. In task construction, each CO task starts with a concise natural-language description that provides basic context and semantic direction. This description explicitly contains a small number of semantic challenge primitives {Si }, specifying affect, audience, or stylistic intent with relatively low ambiguity. For each task, 3–4 types of semantic primitives from S1 –S4 are involved, and each type typically appears only once. In contrast, constraint challenge primitives {Cj } are provided through structured fields, explicitly separating semantic background from executable constraints. Each CO task includes 3–5 different types of constraint primitives, and each type may contain multiple concrete constraint instances. They are also instantiated across diverse T2I application contexts to 23
AtelierEval
avoid tying the evaluation to any particular subject matter or style. Formally, we write a CO task as tCO =
Dtext , m1 S1 + m2 S2 + m3 S3 + m4 S4 , n1 C1 + n2 C2 + n3 C3 + n4 C4 + n5 C5 , At ,
(10)
with the following constraints: 3≤
4 X
mi ≤ 4,
3≤
i=1
5 X
I(nj > 0) ≤ 5,
(11)
j=1
where Dtext denotes the clear natural-language task description, mi ∈ {0, 1} indicates whether semantic primitive Si is present, and nj ∈ Z≥0 denotes the number of concrete constraint instances associated with primitive Cj . The challenge load of a CO task is determined jointly by the number of active constraint primitives and the total number of executable constraint instances. During task design, we explicitly control these factors by constraining the total number of key constraint checks to the range above, while avoiding stacking too many high-strictness constraint instances within a single task. Before inclusion in the benchmark, all CO tasks are manually validated by two independent domain experts. The experts independently write prompts for each task and examine whether, using the four text-to-image models employed in our experiments, the generated images are logically interpretable. A task is included only if the experts agree that the task description is unambiguous, the structured constraints contain no systematic conflicts, and that across repeated generations or different prompt attempts, each category of specified constraints admits at least one satisfiable generation instance. C.3. IM Task Instantiation Imitation (IM) tasks are designed to evaluate a prompter’s ability to translate visual observations into executable prompts (Yuan et al., 2022). Each IM task is instantiated from a target image, which serves as the ground truth specification and is firstly obtained through mixed sources, including real photographs sourced from CC0 or CC0-equivalent open-license repositories, AI-generated images, and manual drafts. Then we manually refine all collected images to ensure full compliance with a custom checklist for evaluation purposes and prevent potential data leakage. Such refinement includes adding or removing text or objects, adjusting color, scaling or cropping, cloning or removing stamps, and AI-based editing. IM tasks instantiate visual counterparts of challenge primitives. We restrict IM to primitives with sufficient visual observability. Specifically, IM tasks exclude S2 (Audience Intent), S4 (Semantic Negation), and C5 (Hard Constraint). Audience intent reflects creator-side goals that are not uniquely encoded in visual appearance; semantic negation specifies what should not be generated, which is underdetermined from a single image; and global hard constraints cannot be inferred from an image without additional contextual assumptions. Including these primitives would introduce systematic ambiguity. Formally, we write an IM task as tIM = Itarget , k1 VS1 + k2 VS3 + n1 VC1 + n2 VC2 + n3 VC3 + n4 VC4 , At , (12) with the following bound on checklist size: 10 ≤
X
km +
m
X
nj ≤ 20.
(13)
j
Here Itarget denotes the finalized target image, VS1 and VS3 are the visual counterparts of S1 and S3 , and VC1 –VC4 are the visual counterparts of constraint primitives C1 –C4 . The challenge load of an IM task is determined by the number and diversity of checklist items extracted from the target image. During task design, we explicitly control this load by constraining the total number of checklist items to a moderate range, as shown in Equation 13, and by avoiding excessive stacking of highly ambiguous or fine-grained visual properties. Before inclusion in the benchmark, all IM tasks undergo manual validation by two independent domain experts. Each expert independently attempts to reproduce the target image by writing prompts and generating images; a task is included only if both experts can, through multiple attempts, produce at least one image that fully satisfies all checklist items. This validation ensures that IM tasks are practically reproducible and that evaluation reflects the prompter’s ability to encode visual semantics, rather than artifacts of target construction. 24
AtelierEval
Ultimately, we note that fixed-seed reconstruction, i.e., evaluation strategies that fix the random seed of a text-to-image model and assess a prompter’s ability to reproduce a target image under the same model configuration, is a deliberately excluded design choice for IM tasks. A detailed discussion of this design decision is provided in Appendix D.
D. Model/Seed-Agnostic Design of IM Tasks This appendix clarifies the motivation behind the model- and seed-agnostic design of IM tasks. We argue that such a formulation is necessary to ensure that IM evaluation reflects prompting proficiency itself. Conventional Fixed-Seed Design. A natural and widely adopted evaluation method for imitation ability is to fix both the generation model and random seed, and evaluate a prompter by reproducing a target image under identical settings with a reconstruction-based score. This setup offers strong determinism: when the prompt exactly matches the original, the target image can in principle be reproduced without ambiguity. While such strategies have been explored in prior work (Dong et al., 2023; Mahajan et al., 2024), they do not meet our evaluation requirements in several important respects: ➠ Model Coverage. Most commercial text-to-image models, including widely used systems such as the DALL·E and NanoBanana series, do not expose random seed control, and restricting evaluation to seed-controllable models would therefore substantially narrow benchmark coverage. Relying exclusively on open-source models does not resolve this issue, as current open-source systems lack sufficient robustness to support the full range of complex IM tasks, limiting prompter evaluation. ➠ Evaluation Stability. Although the random seed controls stochastic initialization, text-to-image models remain highly sensitive to prompt-level variations through their conditional representations. Even minor linguistic changes that preserve semantic equivalence can therefore induce substantial differences in composition and object configuration under the same seed. This sensitivity introduces large variance into reconstruction-based scores, making evaluation outcomes depend more on randomness than on a prompter’s underlying ability. ➠ Evaluation Objective. Fixed-model, fixed-seed reconstruction evaluates whether a prompter can reproduce a specific realization of a generative model, effectively measuring prompt inversion or model-specific overfitting. In contrast, IM tasks aim to assess a prompter’s ability to analyze visual content and encode its underlying semantics in a transferable, model-agnostic form. Consequently, success under reconstruction-based scores does not necessarily reflect stronger prompting ability. Our Design. We deliberately avoid reconstruction-based evaluation under fixed model–seed configurations for IM tasks. Instead, target images are treated as semantic specifications and are collected from multiple sources, as detailed in Appendix C.3. This design supports multiple text-to-image backends without privileging any particular model, and aligns the evaluation objective with prompting proficiency rather than reconstruction fidelity. Empirically, this design choice is supported by our experimental results. As shown in Table 15, prompter performance on IM tasks exhibits consistent rankings across all evaluated T2I backends, indicating that the evaluation captures the model-agnostic notion of prompting proficiency.
E. Examples of Task Instances To make the design principles and construction procedures more concrete, we also present one representative task from each category. For each instance, we (i) show the original task as presented to prompters and (ii) analyze how its construction arises from composing the challenge primitives in Table 1. We additionally visualize the full prompting pipeline for each task instance (Figures 5, Figures 6 and Figures 8). Notably, these examples are selected to illustrate how variations in challenge load naturally emerge from different combinations of primitives, without relying on ad-hoc complexity labels. E.1. OE Task Example Task Description. The OE track probes divergent production under high narrative noise and without explicit constraint primitives {Cj }. Task oe 29 below is instantiated in a commercial illustration context and targets the Environment and Photorealistic application categories. Task ID: oe 29
Title: Sunset Mountain Landscape
Task: The Aether Hotel chain is commissioning art prints for their new Alpine location, and our studio is bidding for
25
AtelierEval
the project. The theme is Majesty of Nature. The art consultants brief asks for a landscape photography-style scene of mountains at sunset. Key elements are warm golden light, mist-filled valleys, and dramatic depth. This is for their high-end guests, so it must look premium and peaceful. It must look hyper-realistic, like a high-end photo. It cannot be a simple drawing or painting, and it absolutely must not contain any people, roads, or buildings. Challenge Primitive Decomposition. This instance deliberately concentrates semantic challenge primitives while avoiding explicit structural constraints: • Abstract Intent (S1 ). The theme “Majesty of Nature” and the requirement that the scene feel “premium and peaceful” encode affective goals that must be visualized rather than directly executed.
• Audience Intent (S2 ). The mention of “high-end guests” and the hotel commissioning context implicitly targets a luxury audience, steering style and polish without prescribing concrete attributes.
• Implicit Style (S3 ). Phrases such as “landscape photography-style” and “hyper-realistic, like a high-end photo” specify a photographic look and camera-like realism without enumerating technical parameters (lens, focal length, etc.), requiring the prompter to complete these details.
• Semantic Negation (S4 ). The brief bans “simple drawing or painting” and “any people, roads, or buildings.” These exclusions operate at the semantic level (which concepts must be absent), not as hard procedural constraints C5 . No explicit constraint primitives C1 –C4 are introduced: there is no fixed object count, layout template, or textual content to render. The task thus isolates the semantic interpretation and translation component of prompting proficiency. To operationalize this decomposition in our objective evaluation, we summarize the image-level and prompt-level criteria in Table 8. Table 8. Objective image- and prompt-based checklists for OE task oe 29. Image-based checklist I1. Mountains are visible in the image. I2. The scene is set at sunset. I3. Mist or fog is visible in valleys. I4. No human figures are visible in the image. I5. No roads or buildings are visible in the image. I6. The visual style is photorealistic
Prompt-based checklist P1. The prompt mentions mountains or related terms. P2. The prompt mentions sunset or related terms. P3. The prompt mentions mist, fog, or related terms. P4. The prompt excludes people, human figures, or related terms. P5. The prompt excludes roads, buildings, or man-made structures. P6. The prompt specifies photographic or photorealistic as the style.
Design Rationale. This task’s challenge load lies primarily in filtering narrative noise and aggregating a small, coherent set of semantic requirements. A competent prompter can translate the brief into a concise prompt (e.g., by foregrounding the Alpine setting, golden-hour lighting, and absence of man-made elements) without managing complex combinatorial constraints. Typical failure modes include (i) under-specifying the atmosphere, yielding generic mountain photos that miss the “Majesty of Nature” mood (S1 ), or (ii) neglecting semantic negation, leading to cabins, hiking trails, or human figures in the generations (S4 ). These errors directly diagnose weaknesses in semantic decoding rather than structural prompt construction. Qualitative generations across models. To see how they manifest in practice across our model zoo, we also show qualitative generations for this instance in Figure 5. A notable failure mode arises from the hotel-related narrative context: several prompters overemphasize the commissioning background, leading SDXL to generate interior hotel scenes rather than outdoor mountain landscapes. 26
AtelierEval
Figure 5. Qualitative generations for OE task oe 29 from the skilled condition.
E.2. CO Task Example Task Description. CO tasks emphasize convergent production under low narrative noise, concentrating multiple explicit constraints {Cj } that must be jointly realized in a single prompt. Task co 106 below is instantiated in a sequential-panel comic design setting. Task ID: co 106
Title: Morning Routine Four-Panel Comic
Brief: A lifestyle blog needs a simple comic strip about daily routines for their wellness section. The comic must be relatable and clean, for general readers. This is a slice-of-life comic strip. The comic must not include any speech bubbles inside panels. Character and Object Attributes: The main character must have black hair in all four panels. The character must wear a blue shirt in all panels. The alarm clock must be red. Layout: Standard horizontal four-panel layout. Panel 1: character asleep with alarm ringing. Panel 2: character groggily sitting up. Panel 3: character stretching. Panel 4: character smiling while holding coffee mug. Quantity: Exactly four panels of equal size arranged horizontally. The same character must appear in all four panels. Text Rendering: Below each panel, captions must read: “Panel 1: 6:00 AM”, “Panel 2: Wake up”, “Panel 3: Stretch”, “Panel 4: Ready for the day”. Global Constraints: The comic must use only black (hair and outlines), blue (shirt), red (alarm clock), and white background. All panels must have identical dimensions. Art style must be simple line art with flat colors only. Challenge Primitive Decomposition. Compared to OE, this instance carries a much denser mix of constraint primitives, while semantic primitives are expressed explicitly with little ambiguity: • Semantic primitives. The slice-of-life setting and requirement that the strip be “relatable and clean” instantiate S1 and S2 , but here they function mainly as context that bounds aesthetics rather than as implicit cues. • Attribute Binding (C1 ). Hair color, shirt color, and alarm clock color must be consistently bound to the correct entities across all panels (e.g., no panel where the shirt changes color). • Spatial Relation (C2 ). The standard horizontal four-panel layout, panel-wise narrative progression, and the positioning of captions beneath each panel all impose explicit spatial structure. 27
AtelierEval
• Quantity (C3 ). The requirement of “exactly four panels of equal size arranged horizontally” and “the same character must appear in all four panels” imposes strict cardinality constraints on both panels and character instances.
• Text (C4 ). The four captions require precise text rendering and spelling, with panel indices and phrases that must appear verbatim under the corresponding panels.
• Hard Constraint (C5 ). The restricted color palette and prohibition of speech bubbles act as global, non-relaxable constraints that any valid solution must respect.
The resulting challenge load is thus dominated by {C1 , . . . , C5 }, with semantics largely transparent. This matches the CO design goal of making the main difficulty the integration of many interacting constraints into one executable prompt. We summarize the image-level and prompt-level criteria in Table 9.
Table 9. Objective image- and prompt-based checklists for CO task co 106. Image-based checklist I1. Four panels are present in the image. I2. A character is present in all panels. I3. An alarm clock is present in the image. I4. A coffee mug is present in the image. I5. The character has black hair in all four panels. I6. The character wears a blue shirt in all panels. I7. The alarm clock is red. I8. The layout is standard horizontal four-panel. I9. Panel 1 shows character asleep with alarm ringing. I10. Panel 2 shows character groggily sitting up. I11. Panel 3 shows character stretching. I12. Panel 4 shows character smiling while holding coffee mug. I13. Exactly four panels of equal size are arranged horizontally (not more, not fewer). I14. Text below panel 1 reads ‘Panel 1: 6:00 AM’ or caption reads ‘6:00 AM’. I15. Text below panel 2 reads ‘Panel 2: Wake up’ or caption reads ‘Wake up’. I16. Text below panel 3 reads ‘Panel 3: Stretch’ or caption reads ‘Stretch’. I17. Text below panel 4 reads ‘Panel 4: Ready for the day’ or caption reads ‘Ready for the day’. I18. No speech bubbles are visible inside panels. I19. The comic uses only four colors: black (hair and outlines), blue (shirt), red (alarm clock), and white background. I20. The art style is simple line art with flat colors only. I21. The visual style is flat vector or simple comic strip illustration.
Prompt-based checklist P1. The prompt mentions four panels or panel sequence. P2. The prompt mentions a character, person, or related terms. P3. The prompt mentions an alarm clock or related terms. P4. The prompt mentions a coffee mug, mug, or related terms. P5. The prompt describes character with black hair in all panels. P6. The prompt describes character wearing blue shirt in all panels. P7. The prompt specifies alarm clock as red. P8. The prompt describes horizontal four-panel layout or four panels in a row. P9. The prompt describes panel 1 with character asleep and alarm ringing. P10. The prompt describes panel 2 with character groggily sitting up. P11. The prompt describes panel 3 with character stretching. P12. The prompt describes panel 4 with character smiling holding coffee mug. P13. The prompt specifies exactly four panels of equal size arranged horizontally. P14. The prompt specifies caption ‘Panel 1: 6:00 AM’ or ‘6:00 AM’ below panel 1. P15. The prompt specifies caption ‘Panel 2: Wake up’ or ‘Wake up’ below panel 2. P16. The prompt specifies caption ‘Panel 3: Stretch’ or ‘Stretch’ below panel 3. P17. The prompt specifies caption ‘Panel 4: Ready for the day’ or ‘Ready for the day’ below panel 4. P18. The prompt excludes speech bubbles inside panels or mentions no speech bubbles/captions only below. P19. The prompt specifies four colors: black (hair/outlines), blue (shirt), red (alarm clock), white background. P20. The prompt describes simple line art with flat colors or minimal style. P21. The prompt specifies flat vector, comic strip, slice-of-life comic, or related flat comic style terms.
28
AtelierEval
Figure 6. Qualitative generations for CO task co 106 from the skilled condition.
Design Rationale. This task’s challenge lies in many ways. Prompters must simultaneously (i) encode a multi-panel narrative, (ii) maintain consistent character identity and attributes across panels, (iii) control color usage globally, and (iv) specify exact captions. Typical failure modes include attribute leakage (e.g., shirt or alarm clock colors drifting in one panel; violating C1 ), incorrect panel count or arrangement (C2 –C3 ), missing or misspelled captions (C4 ), or the model introducing extra colors or speech bubbles (breaching C5 ). These errors directly reveal whether the prompter can systematically translate a structured, low-noise spec into a compact yet sufficiently over-specified prompt. Qualitative generations across models. Beyond the objective checklist in Table 9, it is also informative to inspect how different prompter–backend combinations actually instantiate this four-panel comic; Figure 6 visualizes these generations for task co 106. E.3. IM Task Example
Figure 7. Target image for IM task im 43. Prompters only see the image, not the internal seed prompt.
29
AtelierEval
Task Description. IM tasks assess cognition-oriented prompting: given only a target image, the prompter must encode perceptual information into text so that a T2I model can reproduce it. Figure 7 shows the target image for task im 43, which belongs to the Environment, Object, 3D Render, and Abstract Conceptual categories. Task ID: im 43
Instruction to prompters:
You are given the target image in Figure 7. Write a single English prompt that would enable a text-to-image model to reproduce an image as close as possible to this target, including its composition, geometry, materials, lighting, and overall mood. Table 10. Objective image- and prompt-based checklists for IM task im 43. Image-based checklist
Prompt-based checklist
I1. A floating platform or circular stage is present in the center of the image I2. The platform has multiple vertical cylindrical pillars or columns standing on it I3. The pillars have pink or magenta colored sections
P1. The prompt mentions a floating platform or circular stage or related terms P2. The prompt describes multiple vertical cylindrical pillars or columns or related terms P3. The prompt specifies pink or magenta colored sections on pillars or related terms P4. The prompt mentions teal or turquoise colored sections on pillars or related terms P5. The prompt describes curved or arched tops connecting pillars or arch structures or related terms P6. The prompt mentions a large white cloud above the platform or in the center or related terms P7. The prompt describes a thin line or string or connection between cloud and platform or related terms P8. The prompt specifies floating above water or reflective surface or related terms P9. The prompt mentions multiple horizontal circular tiers or layers on platform or related terms P10. The prompt describes blue and white colored tiers or layers or related terms P11. The prompt mentions reflections in water or mirrored elements below or related terms P12. The prompt describes small disc-shaped objects or floating discs in water or related terms P13. The prompt mentions 3D render or surreal digital art or related terms P14. The prompt describes pastel color palette or pink blue teal white colors or related terms P15. The prompt mentions symmetrical composition or centered platform or related terms
I4. The pillars have teal or turquoise colored sections I5. Some pillars have curved or arched tops connecting two vertical sections I6. A large white cloud formation is positioned above the platform in the center I7. A thin vertical line or string appears to connect the cloud to the platform I8. The platform appears to be floating above water or a reflective surface I9. The platform has multiple horizontal circular tiers or layers I10. The tiers are colored in blue and white tones I11. Reflections of the pillars are visible in the water below I12. Multiple small disc-shaped objects are floating in the water beneath the platform I13. The visual style is 3D rendered or surreal digital art I14. The color palette includes pastel tones of pink, blue, teal, and white I15. The composition is symmetrical with the platform centered
Challenge Primitive Decomposition. Unlike OE and CO, the challenge primitives here are instantiated visually rather than textually. Consistent with the IM design, audience intent S2 , semantic negation S4 , and hard constraints C5 are absent. • Abstract Intent (S1 ). The image conveys a calm, surreal dreamscape with a “cloud tree” floating above a mirror-like sea. Prompters must infer and verbalize this atmosphere (e.g., serenity, dreaminess, futurism) from purely visual cues. • Implicit Style (S3 ). The pastel palette, smooth gradients, soft focus, and 3D-rendered look imply a minimalistic surreal style and high-resolution rendering, which are not named explicitly but must be captured (e.g., “dreamy surrealism, pastel 3D render”). • Attribute Binding (C1 ). Several distinctive object–attribute bindings must be preserved: the cloud shaped like a tree sitting on a transparent glass pedestal, pastel pink and aqua arches arranged around a white platform, and smooth stepping stones leading toward the center. • Spatial Relation (C2 ). The composition is strongly center-focused and symmetrical, with circular, tiered platforms and 30
AtelierEval
clear foreground–midground–background separation, plus horizontal reflections in the water. Capturing these relationships is crucial for reproducing the scene. • Quantity (C3 ). While exact counts (e.g., number of arches or stepping stones) are not emphasized, the prompt must encode the presence of multiple repeated elements in a ring-like arrangement and a path of stepping stones, which serve as soft quantity cues. These primitives together require the prompter to perform fine-grained visual analysis and encode it into text without any scaffolding from a textual brief. We summarize the image-level and prompt-level criteria in Table 10. Design Rationale and Difficulty. This task combines a relatively simple, symmetric global layout with several non-trivial perceptual details. Novice prompters often under-specify either the abstract mood (omitting the surreal, dreamlike quality; weakening S1 and S3 ) or the structural composition (forgetting the glass pedestal, circular tiers, or stepping-stone path; harming C1 –C2 ). Skilled prompters, in contrast, tend to systematically decompose the image into objects, materials, spatial relations, and rendering style, then recombine them into a compact prompt that preserves both aesthetics and geometry. This behavior directly reflects the cognition-oriented encoding ability that IM tasks are designed to probe. Qualitative generations across models. Finally, to complement the analysis of the IM instance, Figure 8 shows how our full set of prompter–backend combinations attempt to recreate the target scene from Figure 7.
Figure 8. Qualitative generations for IM task im 43 from the novice condition.
F. Subjective Evaluation Dimensions This section details the subjective evaluation dimensions used by ATELIER J UDGE to assess the quality of prompt–image pairs, 4 dimensions for prompts and 4 dimensions for images. These dimensions are designed to capture perceptual, aesthetic, and semantic qualities that are difficult to verify through binary constraints, while remaining grounded in observable and model-agnostic criteria. All 8 dimensions are rated on a 1–5 Likert scale. Higher scores indicate stronger and more consistent satisfaction of the listed criteria, rather than merely increased verbosity or detail. The same scale and criteria are used for both human experts and AtelierJudge. Full scoring rubrics are provided in Appendix I. F.1. Image-Level Dimensions Mood & Atmosphere. This dimension evaluates whether the image conveys an emotional tone and overall atmosphere that aligns with the intent expressed or implied in the prompt. It focuses on the coherence between affective intent and visual 31
AtelierEval
realization, including whether the emotional quality (e.g., joyful, melancholic, tense, serene) is appropriate in strength and consistency. Visual elements such as color palette, lighting, composition, and texture are considered insofar as they jointly support the intended mood and communicate both explicit and implicit emotional cues. Visual Composition. This dimension assesses the structural quality of the image and the unity among its visual elements. Key considerations include the presence of a clear visual focus, effective layering and depth, and harmonious relationships between subject and background. The evaluator examines whether spatial arrangement, balance, and visual flow guide attention naturally, and whether multiple elements are integrated into a coherent whole through compositional principles such as alignment, repetition, contrast, and spatial relationships. In design-oriented tasks (e.g., logos or posters), emphasis is placed on meaningful integration rather than simple aggregation. Color & Lighting. This dimension evaluates the harmony and appropriateness of color usage and lighting. Assessment includes whether the color scheme supports the intended atmosphere, whether saturation and brightness are well balanced, and whether lighting direction and intensity are visually and physically plausible. Shadows and highlights are examined for consistency and their contribution to depth, form, and mood, as well as the overall coordination between color and lighting. Technical Flawlessness. This dimension focuses on identifying common technical artifacts and visual defects characteristic of generative models. Evaluators look for issues such as anatomical distortions (especially of hands and faces), implausible physical structures, incorrect perspective, unnatural textures or repetitive patterns, blurred or fused object boundaries, and other noticeable generation artifacts. The absence of such flaws indicates higher technical soundness, independent of stylistic or aesthetic preference. F.2. Prompt-Level Dimensions Instructional Clarity. This dimension evaluates the linguistic clarity and executability of the prompt. It considers whether the prompt is grammatically complete, logically coherent, and free from ambiguity that could lead to misinterpretation. Clear identification of the main subject, stable descriptions without internal contradictions, and well-structured sentences that convey intent unambiguously are key factors. The goal is to assess whether the prompt can be reliably understood and acted upon by a text-to-image system without requiring guesswork. Creative Elaboration. This dimension assesses the richness and specificity of visual information provided in the prompt. Evaluation focuses on whether the prompt goes beyond minimal descriptions to include concrete details about subjects, environments, materials, lighting, atmosphere, or composition. Prompts that demonstrate imagination through distinctive, sensory-rich descriptions are rated higher than those that remain generic or underspecified. Terminology Proficiency. This dimension evaluates the prompter’s ability to appropriately use domain-relevant visual terminology in a model-agnostic manner. This includes accurate and consistent use of concepts from photography, visual arts, and rendering (e.g., depth of field, impressionist style, volumetric lighting), without internal contradictions. Terminology should be precise, non-conflicting, and well matched to the task intent. The evaluation explicitly discourages reliance on model-specific syntax or engine-dependent keywords, emphasizing transferable visual literacy over system-specific tricks. Intent Formalization. This dimension assesses the prompter’s ability to translate abstract, implicit, or high-level intent into concrete visual specifications. Evaluators examine whether abstract emotions, concepts, audience requirements, or implicit stylistic expectations are effectively grounded in describable visual elements, styles, or compositional choices. High-quality prompts avoid passing vague abstractions directly to the model and instead operationalize them into visual cues that faithfully serve the original intent.
G. Evaluation Exemplar Memory Construction This appendix details the construction and quality control of the exemplar memories used by the subjective evaluation skills of AtelierJudge. All exemplars are constructed and fixed prior to the main experiments to ensure consistency and reproducibility. An overview of the exemplar memory construction and quality-control pipeline is shown in Figure 9.
32
AtelierEval
Figure 9. Overview of the memory construction pipeline.
For each task in AtelierEval, we construct two evaluation exemplars, one for the prompt and one for the image (except for IM tasks). Each exemplar consists of 4 components: ❶ the task directly from AtelierEval, ❷ the prompt or generated image used for evaluation, ❸ human-annotated scores across subjective dimensions, and ❹ brief rationales explaining the assigned scores. Prompts are collaboratively authored by two experts to ensure broad coverage of different subjective quality patterns and score ranges. In total, 360 prompts are created. For each prompt, one of the four evaluated T2I backends is used to generate a corresponding image, yielding 360 images in total. For IM tasks, only prompts are retained, while generated images are used solely to verify prompt feasibility. After image generation, we construct subjective human-labeled scores for each exemplar. Three experts independently assign scores to each prompt and generated image as two mutually independent evaluation targets, following the same subjective scoring guidelines as AtelierJudge (Appendix I). To ensure balanced coverage, the exemplar set is curated such that, for each task category, every subjective dimension contains representative exemplars at all score levels (Qin et al., 2025a). We assess inter-annotator agreement (IAA) using Krippendorff’s α (Krippendorff, 2018), which is suitable for ordinal data. Only exemplars meeting the predefined agreement threshold (α ≥ 0.75) are retained. This threshold typically allows one expert’s score to differ from the other two by at most one scale point. Exemplars failing the criterion are discarded, and the corresponding prompts are rewritten or images regenerated. For retained exemplars, the three experts further discuss them to reach a final consensus score. Exemplars for which consensus cannot be reached are also discarded. Final scores stored in the exemplar memories are determined by expert consensus rather than numerical averaging. After the consensus scores are finalized, we attach rationales to each exemplar, with one rationale corresponding to each subjective dimension. For scalability, rationales are generated post-hoc by GPT-5.2. For each dimension, the model is provided with a structured chain-of-thought prompt containing (i) the task, (ii) the prompt or generated image, (iii) the scoring guideline for the dimension, and (iv) the finalized expert score. The model is instructed to produce a concise, criterion-grounded explanation that justifies the given score (“knowing the answer and explaining why”). Although the rationale drafter is implemented using GPT-5.2 and may introduce stylistic self-consistency, all rationales are generated post-hoc given fixed expert consensus scores. To further mitigate such bias, before being admitted to the exemplar memory, each generated rationale is reviewed and approved by three human experts through a joint deliberative process. During this review, experts verify that the rationale faithfully reflects the scoring rubric, accurately explains the assigned score, and does not introduce extraneous or misleading information. Only rationales that pass expert review are stored together with the finalized scores and subsequently used for memory-augmented evaluation. Through this process, we obtain 600 high-quality exemplars across five memories to support the subjective evaluation skills of AtelierJudge.
H. Details of Safety Filter Skill AtelierJudge includes a safety filtering design, which serves as non-scoring gates to ensure ethical compliance. Safety filtering consists of two skills operating on different modalities: a QA-based prompt safety filter skill and a VQA-based image safety filter skill, which are executed sequentially in the skill routing process but remain logically decoupled. ➠ Prompt Safety Filter Skill. Before image generation, all prompts are checked by this skill, which is implemented using GPT-5.2. This skill operates solely on the prompt and detects standard categories of unsafe content (e.g., violence, sexual content, hate, child safety violations, and illegal activities). Across all experiments, only 2 prompts out of 7,200 total prompts were initially flagged and later manually confirmed to be safe. 33
AtelierEval
➠ Image safety filter skill. As stated in Section 5.1, we evaluate four T2I backends, and three of them are accessed via provider APIs that already include safety mechanisms. Among all submissions, only Flux-Pro refused image generation for 5 prompts out of the 7,200 prompts. All five prompts were subsequently confirmed to be safe by human experts. To enforce a unified standard across backends, an image safety filter skill—also implemented via GPT-5.2—is applied to all generated images, regardless of the backend. This skill uses the same set of safety categories as the prompt safety filter skill. No generated image was flagged as unsafe by the unified image safety filter skill. Submissions that cannot produce an image are excluded from all aggregate metrics. The extremely low trigger rate of the safety filter skills indicates that the expert-designed tasks are well aligned with content safety requirements and that these skills function as non-intrusive gating components rather than confounding factors in our experiments. For reproducibility, the full prompts and code used to implement the safety filter skills are open-sourced in our repository.
I. Scoring Guidelines for Subjective Evaluation This appendix presents the detailed scoring rubrics used to assess subjective quality in AtelierEval. To ensure a unified evaluation standard, these rubrics are employed identically by both human experts during the construction of the exemplar memory, and AtelierJudge during the automated agentic evaluation process. Prompt Evaluation Guidelines 1. Instructional Clarity (Grammar, logic, lack of ambiguity, structure) • 1 (Failure): Incoherent, contradictory, or grammatically broken. Impossible to execute reliably. • 2 (Poor): Significant confusion or conflicting instructions. Logic is hard to follow. • 3 (Acceptable): Understandable but contains minor ambiguities or loose sentence structure. Requires some model guesswork. • 4 (Good): Generally clear and well-structured. Only very minor phrasing issues. • 5 (Excellent): Perfectly structured, logical in flow, and unambiguous. Explicitly identifies subjects and relations. Zero chance of misinterpretation. 2. Creative Elaboration (Richness, detail, sensory specificity) • 1 (Empty): Bare minimum description. Lacks any detail beyond the core subject. • 2 (Generic): Uses clichéd descriptions. Details are vague (e.g., ”nice background”). • 3 (Basic): Provides standard details (color, size) but lacks imagination or sensory depth. • 4 (Detailed): Good use of adjectives and specific descriptions. Sets a clear scene. • 5 (Rich): Highly evocative. Describes textures, atmosphere, materials, and specific nuances. Demonstrates strong imagination. 3. Terminology Proficiency (Use of visual/artistic vocabulary, model-agnosticism) • 1 (Poor): Uses wrong terms, relies on “magic words” (e.g., “4k”, “trending”), or non-visual text. • 2 (Naive): Uses artistic terms incorrectly or relies heavily on engine-specific syntax (e.g., –v 6.0) inside the text. • 3 (Average): Uses basic visual terms correctly (e.g., “oil painting”, “sunset”). • 4 (Advanced): Uses more complex terms (e.g., “macro lens”, “impasto”) correctly. • 5 (Expert): Precise use of domain-specific vocabulary (e.g., “volumetric lighting”, “depth of field”, “chiaroscuro”) correctly and effectively. 4. Intent Formalization (Translating abstract goals into concrete visual specs) • 1 (Abstract): Pastes abstract concepts (e.g., “a sad vibe”) directly without visual translation. • 2 (Mostly Abstract): Slight attempt at visual description, but mostly relies on the model to interpret feelings. • 3 (Mixed): Partially translates intent but relies on some abstract descriptions. • 4 (Mostly Concrete): Most abstract concepts are converted to visual cues, with minor gaps. • 5 (Concrete): Fully operationalizes abstract intent into observable visual elements (e.g., translates “sadness” into “muted blue tones, rain-streaked windows, slumped posture”).
Image Evaluation Guidelines 1. Mood & Atmosphere (Emotional tone, consistency with intent) • 1 (Mismatch): The image conveys the completely wrong emotion or has no discernible atmosphere.
34
AtelierEval
• 2 (Weak): The mood is barely present or confusing. • 3 (Generic): The mood is somewhat aligned but weak or inconsistent. Lacks strong emotional impact. • 4 (Strong): The atmosphere is clear and mostly consistent with the intent. • 5 (Evocative): The atmosphere is palpable, consistent, and perfectly matches the intended emotional tone (e.g., tension, serenity). 2. Visual Composition (Structure, balance, focus, depth) • 1 (Chaotic): Cluttered, lacks a focal point, or poor spatial arrangement. Hard to parse. • 2 (Unbalanced): Elements feel randomly placed. Poor use of space. • 3 (Standard): Functional composition. Center-focused or basic rule-of-thirds, but lacks depth or dynamic flow. • 4 (Good): Clear focal point and good balance. Uses space well. • 5 (Masterful): Excellent use of depth, layering, and guiding lines. Visual elements are integrated harmoniously; the eye is led naturally. 3. Color & Lighting (Harmony, direction, saturation, physics) • 1 (Bad): Clashing colors, flat lighting, or physically impossible shadows. Looks washed out or oversaturated. • 2 (Dull): Colors are muddy or lighting makes the subject hard to see. • 3 (Passable): Lighting is logical but flat. Colors are acceptable but not distinct or strictly harmonized. • 4 (Cohesive): Good color harmony and clear light source. • 5 (Cinematic): Lighting creates volume and mood. Color palette is sophisticated and cohesive. Shadows and highlights are physically accurate and aesthetic. 4. Technical Flawlessness (Artifacts, distortions, anatomy, rendering) • 1 (Broken): Severe artifacts (mangled hands, extra limbs), blurred boundaries, or distinct digital noise. Unusable. • 2 (Obvious Flaws): Distracting distortions or mutations are immediately visible. • 3 (Minor Flaws): Generally good, but contains noticeable small artifacts, slight perspective issues, or unnatural textures on close inspection. • 4 (Clean): High quality rendering with only negligible, hard-to-spot imperfections. • 5 (Flawless): Clean, crisp rendering. Anatomy, perspective, and textures appear natural and intentional. No visible generative artifacts.
J. Prompts for AtelierJudge J.1. Subjective Skill Prompts Prompt Subjective Skill Template System Prompt You are an expert Text-to-Image Prompt Engineer and Evaluator. Your goal is to assess the quality of a text prompt based on how effectively it translates a user’s intent into an executable description. You MUST follow these rules: • Act as a calibrated judge using the provided Reference Exemplars. • Evaluate strictly based on the provided 1-5 scale definitions. • Output ONLY a valid JSON object containing scores and rationales. User Prompt: [Start of Exemplar 1] Task: <<EXEMPLAR 1 TASK>> Prompt to Evaluate: <<EXEMPLAR 1 PROMPT>> <<PROMPT EVALUATION GUIDELINES>> Step-by-Step Reasoning Instructions: • Analyze the Task Description to understand the goal. • Read the Candidate Prompt. • For each dimension, assign a score (integer 1-5) and write a brief rationale explaining why. Output JSON: <<EXEMPLAR 1 JSON>>
35
AtelierEval
[End of Exemplar 1] ... [Start of Exemplar k] Task: <<EXEMPLAR k TASK>> Prompt to Evaluate: <<EXEMPLAR k PROMPT>> Evaluation Guidelines: Evaluate the Candidate Prompt on the following 1-5 scale. Read the definitions carefully. <<PROMPT EVALUATION GUIDELINES>> Step-by-Step Reasoning Instructions: • Analyze the Task Description to understand the goal. • Read the Candidate Prompt. • For each dimension, assign a score (integer 1-5) and write a brief rationale explaining why. Output JSON: <<EXEMPLAR k JSON>> [End of Exemplar k] Task: <<TASK>> Prompt to Evaluate: <<PROMPT>> Evaluation Guidelines: Evaluate the Candidate Prompt on the following 1-5 scale. Read the definitions carefully. <<PROMPT EVALUATION GUIDELINES>> Step-by-Step Reasoning Instructions: • Analyze the Task to understand the goal. • Read the Prompt. • For each dimension, assign a score (integer 1-5) and write a brief rationale explaining why. Output JSON:
Image Subjective Skill Template System Prompt You are an expert Art Director and Visual Critic. Your goal is to assess the quality of a generated image based on its aesthetic value, technical execution, and alignment with visual intent. You MUST follow these rules: • Act as a calibrated judge using the provided Reference Exemplars. • Evaluate strictly based on the provided 1-5 scale definitions. • Output ONLY a valid JSON object containing scores and rationales. User Prompt: [Start of Exemplar 1] Task: <<EXEMPLAR 1 TASK>> Image to Evaluate: <<EXEMPLAR 1 IMAGE INPUT>> Evaluation Guidelines: Evaluate the Candidate Image on the following 1-5 scale. Read the definitions carefully. <<IMAGE EVALUATION GUIDELINES>> Step-by-Step Reasoning Instructions: • Analyze the Task Description to understand the goal. • Observe the Candidate Image. • For each dimension, assign a score (integer 1-5) and write a brief rationale explaining why. Output JSON: <<EXEMPLAR 1 JSON>> [End of Exemplar 1] ... [Start of Exemplar k] Task: <<EXEMPLAR k TASK>>
36
AtelierEval
Image to Evaluate: <<EXEMPLAR k IMAGE INPUT>> Evaluation Guidelines: Evaluate the Candidate Image on the following 1-5 scale. Read the definitions carefully. <<IMAGE EVALUATION GUIDELINES>> Step-by-Step Reasoning Instructions: • Analyze the Task to understand the goal. • Observe the Image. • For each dimension, assign a score (integer 1-5) and write a brief rationale explaining why. Output JSON: <<EXEMPLAR k JSON>> [End of Exemplar k] Task: <<TASK>> Image to Evaluate: <<IMAGE INPUT>> Evaluation Guidelines: Evaluate the Candidate Image on the following 1-5 scale. Read the definitions carefully. <<IMAGE EVALUATION GUIDELINES>> Step-by-Step Reasoning Instructions: • Analyze the Task Description to understand the goal. • Observe the Candidate Image. • For each dimension, assign a score (integer 1-5) and write a brief rationale explaining why. Output JSON:
J.2. Objective Skill Prompts Prompt Objective Skill Template System Prompt: You are a strict prompt checklist evaluator. You MUST follow these rules: • You MUST evaluate EACH checklist item independently. • You MUST base your judgment ONLY on the given prompt text. • You MUST output ONLY a valid JSON object. • JSON keys MUST EXACTLY match the checklist text (character-for-character). • JSON values MUST be ONLY 0 or 1. • You MUST NOT include explanations, comments, or any text outside the JSON. • If a requirement is NOT clearly and explicitly specified in the prompt, you MUST output 0. User Prompt: You are given a text prompt for image generation and a checklist of requirements. Your task is to determine whether the prompt clearly specifies EVERY requirement. Prompt: <<PROMPT>> Checklist (evaluate each item one by one): <<CHECKLIST>> Instructions: 1. Go through EVERY checklist item. Do not skip any item.
37
AtelierEval
2. For each item, output 1 if the requirement is clearly and explicitly specified in the prompt; otherwise output 0. 3. Your final answer MUST be ONLY a single valid JSON object. 4. Do NOT add, remove, or modify any checklist item text. Required Output Format (example): { “Checklist item text 1”: 1, “Checklist item text 2”: 0 } Now read the prompt carefully and output ONLY the JSON object.
Image Objective Skill Template System Prompt: You are a strict visual checklist evaluator. You MUST follow these rules: • You MUST evaluate EACH checklist item independently. • You MUST base your judgment ONLY on the given image. • You MUST output ONLY a valid JSON object. • JSON keys MUST EXACTLY match the checklist text (character-for-character). • JSON values MUST be ONLY 0 or 1. • You MUST NOT include explanations, comments, or any text outside the JSON. • If a requirement is NOT clearly and fully satisfied, you MUST output 0. User Prompt: You are given an image generated by a text-to-image model and a checklist of visual requirements. Your task is to determine whether the image satisfies EVERY checklist item. Image: <<IMAGE>> Checklist (evaluate each item one by one): <<CHECKLIST>> Instructions: 1. Go through EVERY checklist item. Do not skip any item. 2. For each item, output 1 if the requirement is clearly satisfied by the image; otherwise output 0. 3. Your final answer MUST be ONLY a single valid JSON object. 4. Do NOT add, remove, or modify any checklist item text. Required Output Format (example): { “Checklist item text 1”: 1, “Checklist item text 2”: 0 } Now examine the image and output ONLY the JSON object.
38
AtelierEval
K. Model Hyperparameters MLLMs. The following hyperparameters are shared by all MLLMs used in our experiments, both when acting as prompters and as the evaluator in AtelierJudge. All models are called by API, utilizing the newest version on Dec. 29, 2025. Multimodal LLM Hyperparameters (Prompter & Evaluator) • Context window cap: 16,384 tokens; • Temperature: 0.0; • Top-p: 0.7; • Maximum output tokens: 4,096.
T2I Models. We provide the key configuration details for the locally deployed SDXL. Notably, due to the CLIP-based text encoder, SDXL operates with a maximum context length of 77 tokens, which may result in truncation for some prompts. For the remaining T2I models which are all accessed through DALLE-style APIs, no hyperparameters are exposed to the user except the image resolution (1024×1024). Text-to-Image Hyperparameters (SDXL) • Model: SDXL-base-1.0; • Image size: 1024×1024; • Sampler / scheduler: Euler; • Inference steps: 30; • Guidance scale (CFG): 7.5; • Precision: FP16; • Refiner: None; • Batch Size: 4; • Random seed: independently sampled per image.
L. Prompts for MLLM Novice and Skilled Conditions This appendix documents the prompts used to instantiate the novice and skilled MLLM prompter conditions in our experiments. The novice condition uses minimal task instructions, whereas the skilled condition employs structured system prompts that explicitly encode high-level cognitive strategies, chain-of-thought (Wei et al., 2022; Kojima et al., 2022) style reasoning and self-verification steps. We present the exact prompts below for reproducibility. L.1. Novice MLLM Prompts Novice Group Prompt Template (OE / CO) Convert the following request into an image generation prompt. Use natural language only. Request: <<TASK>>
Novice Group Prompt Template (IM) Write an image generation prompt that would produce an image similar to the one shown. Use natural language only. Image: <<IMAGE>>
39
AtelierEval
L.2. Skilled MLLM Prompts Skilled Group Prompt Template (OE) System Prompt: You are an expert at translating creative concepts into vivid, detailed prompts for text-to-image generation systems. User Prompt: Your Task Convert the user’s request into an evocative, well-structured image generation prompt using natural language only. Structured formats such as lists are allowed. Professional Approach 1. Extract Core Intent from Narrative • The request may contain contextual noise (e.g., “for next season,” “the boss mentioned”). • Identify what truly matters for visual output vs. background context. • Look for implicit requirements: target audience, emotional tone, abstract concepts. 2. Translate Abstract to Concrete • Abstract emotions (e.g., “loneliness,” “eco-friendly”) → specific visual metaphors. • Audience-based descriptions (e.g., “for veterans,” “professional”) → appropriate style choices. • Implicit styles (e.g., “logo,” “poster”) → concrete artistic techniques. • Negative constraints (e.g., “no modern elements”) → explicit exclusions. 3. Build Rich Visual Descriptions • Describe the scene or subject with sensory detail. • Establish clear focal points and visual hierarchy. • Specify artistic style, medium, or aesthetic approach. • Include atmospheric elements: lighting, color palette, mood. • Add quality and technical specifications as appropriate. 4. Self-Verification • Check: Have I captured all implicit requirements? • Check: Does this prompt create a compelling, coherent vision? • Check: Is there enough detail to guide generation? Output Format Provide ONLY the image generation prompt. Do not include explanations or meta-commentary. Now, convert the following request into an image generation prompt: <<TASK>>
Skilled Group Prompt Template (CO) System Prompt: You are an expert at translating precise technical requirements into accurate, executable prompts for text-to-image generation systems. User Prompt: Your Task Convert the user’s request into a clear, detailed image generation prompt using natural language only. Structured formats such as lists are allowed. Professional Approach 1. Analyze Requirements Thoroughly • Identify the core subject and purpose. • Note ALL specified constraints and requirements. • Recognize any negative constraints (what must NOT appear). • Understand the visual style or context implied.
40
AtelierEval
2. Prioritize Precision and Completeness • Every stated constraint is a REQUIREMENT, not a suggestion. • Pay special attention to exact colors, positions, and quantities. • Ensure specific text content and placement are correctly specified. • Maintain correct attribute bindings and spatial relationships. • Explicitly state exclusions (what should NOT be included). 3. Structure Your Prompt Logically • Start with the main subject and setting. • Describe the visual style and medium. • Detail spatial arrangements and compositions. • Specify exact colors, text, quantities, and attributes. • Explicitly state any exclusions or negative constraints. • Include quality and technical specifications. 4. Self-Verification Checklist • Check: Have I addressed EVERY specified constraint? • Check: Are colors, numbers, and positions exact? • Check: Is all required text content included correctly? • Check: Have I stated what should NOT appear? • Check: Is the language clear and unambiguous? Output Format Provide ONLY the image generation prompt. Do not include explanations or meta-commentary. Now, convert the following request into an image generation prompt: <<TASK>> Skilled Group Prompt Template (IM) System Prompt: You are an expert at analyzing images and generating precise descriptive prompts for text-to-image generation systems. User Prompt: Your Task Carefully analyze the provided image and convert it into a detailed, accurate image generation prompt using natural language only. Structured formats such as lists are allowed. Professional Approach 1. Systematic Visual Analysis • Overall composition: What is the main subject? How is the scene organized? • Style and medium: What artistic style, rendering technique, or photographic approach? • Atmosphere and mood: What emotional tone or aesthetic does it convey? • Technical aspects: Lighting, perspective, color palette, depth of field. 2. Precise Element Identification • Count carefully: How many of each object? (Avoid counting errors.) • Attributes matter: Which object has which color/property? (Avoid attribute confusion.) • Spatial relationships: What is where? (Left/right, foreground/background, on/under.) • Text content: If text appears, what does it say exactly? (OCR carefully.) • Fine details: Materials, textures, expressions, accessories. 3. Describe Systematically • Start with subject and main action (if any). • Describe artistic style, medium, and rendering approach. • Detail specific elements with their attributes (colors, materials, states). • Specify spatial arrangement and composition.
41
AtelierEval
• Capture atmosphere, lighting, and mood. • Note any text, symbols, or graphic elements. • Include quality and technical characteristics. 4. Self-Verification Against Image • Check: Does my description match what I actually see? • Check: Have I counted objects correctly? • Check: Are attributes bound to the right objects? • Check: Have I captured the style and mood accurately? • Check: Is any text transcribed correctly? Output Format Provide ONLY the image generation prompt. Do not include explanations or meta-commentary. Now, analyze the provided image and convert it into an image generation prompt: <<IMAGE>>
M. Participant Selection & Statistics This section describes how human participants were selected for the study, and reports summary statistics of the final participant pool. We detail the recruitment channels, inclusion criteria, group assignment procedure, and aggregated demographic characteristics of the participants. The detailed informed consent form and screening questions for participant selection are provided in Appendix T. M.1. Participant Selection Criteria Human participants were recruited on a voluntary basis through multiple channels, including personal contacts, online advertisements on social media, and announcements on university forums. To ensure competency with experimental tools and minimize learning effects related to tool unfamiliarity, all participants were required to meet the following eligibility criteria: Eligibility Criteria • At least 18 years of age; • Proficiency in reading and writing English sufficient to understand task instructions; • Ability to complete the study using a desktop or laptop environment with stable internet access. From the initial pool of applicants, we selected 24 participants for the novice group and 24 participants for the skilled group. We applied group-specific screening criteria to ensure a clear separation between novice and skilled prompters. These criteria are intended to operationalize working definitions of novice and skilled prompters solely for the purpose of this benchmark. Both groups can reliably complete the tasks without additional training, while differing in the extent of their prompting experience and systematic control strategies. Novice Group Screening Criteria • Limited prior exposure to text-to-image generation tools, typically involving exploratory use; • Ability to perform basic prompt-based image generation (e.g., entering a textual description and triggering generation), without systematic experience in prompt engineering; • No regular use of advanced controls, model parameters, or structured prompting workflows.
42
AtelierEval
Skilled Group Screening Criteria • Sustained and regular use of text-to-image generation tools over an extended period; • Familiarity with common prompting strategies and techniques for controlling visual attributes (e.g., composition, lighting, style, or subject binding); • Demonstrated awareness of technical or artistic control dimensions in text-to-image generation, as assessed by the pre-test questionnaire.
Candidates who applied for the skilled group were required to provide evidence of prior work to support their self-reported expertise. This typically consisted of example prompts and corresponding generated images or links to publicly available portfolios voluntarily shared by the candidate. In some borderline cases, additional clarification was obtained through a brief online discussion with cameras turned off to better understand the candidate’s experience. Only candidates who satisfied the corresponding group-specific screening criteria were included in the final sample. Candidates who did not meet the requirements for the skilled group were either considered for the novice group (if consistent with the novice criteria) or not enrolled in the study. At present, T2I prompting lacks standardized skill metrics or objective proficiency tests. Skilled user identification thus relies on behavioral and experiential signals. Indeed, this lack of objective criteria directly motivates the development of AtelierEval, which aims to quantify prompting proficiency beyond coarse self-reported categories. M.2. Participant Demographic Statistics We report aggregated demographic statistics of the final participant pool in Figure 10. Participants in our study are mainly young adults with regions of residence primarily concentrated in North America and East Asia. This demographic profile reflects a substantial subset of the current active user population of T2I systems (Gao et al., 2023; Kumar, 2026). However, these attributes are reported to provide contextual information about the study population rather than to serve as explanatory variables for performance differences. Following prior work that cautions against collecting and interpreting race or ethnicity without a clearly articulated analytical purpose (Biega et al., 2020; Chen et al., 2023), we report region of residence as a less sensitive more appropriate descriptor of participant background given the research scope.
Figure 10. Gender, age range, and region of residence distributions are shown separately for the novice and skilled groups.
43
AtelierEval
N. Human Expert Involvement We recruited a total of six human experts to facilitate the development and validation of our framework. Each expert possesses at least three years of research experience specifically within the T2I domain. These experts were organized into task-specific groups to handle dataset construction, exemplar memory construction, and evaluator verification. To prevent data contamination and ensure rigorous validation, the three experts assigned to exemplar memory construction were strictly distinct from the three experts responsible for evaluator verification. Dataset construction was performed by a subset of three experts selected from this total pool of six. Specific methodologies regarding dataset design and manual validation are detailed in Appendix C. The construction of the exemplar memories used by AtelierJudge is described in Appendix G, while the meta-evaluation protocol and human expert baseline are detailed in Section 5.2.
O. Subjective Evaluation Case Studies To demonstrate the rigor of our assessment logic and the necessity of decoupling subjective quality from objective constraints, we provide two qualitative case studies from the AtelierEval dataset below. O.1. Subjective Prompt Evaluation (Open-Ended Task) Task (OE 45): ”I’m the art director for Metropolis Weekly. Our next cover story is a long-form investigative piece on The Air We Breathe. The deadline is tight. We need a cover image that shows the impact of air pollution on city life. The mood must be oppressive but also realistic. Our readers are educated urban-dwellers, so it needs to be subtle but powerful. It must not look like a futuristic, Blade Runner city. It needs to look like a real, modern city. And no cartoons.”
Prompt (From Qwen-3-VL-235B-A22B, Expert): ”A realistic, modern cityscape at dusk, with dense smog hanging low over the skyline, casting a hazy, gray pall over the streets. The buildings are contemporary but not futuristic—glass towers and concrete structures with subtle signs of wear. People move through the streets, some wearing masks, others coughing or looking away from the polluted air. The atmosphere feels heavy and oppressive, with muted colors and soft shadows emphasizing the weight of the pollution. The image is grounded in reality, avoiding any sci-fi or cartoonish elements, and conveys a quiet, powerful message about the unseen dangers of air quality in urban life.”
AtelierJudge Evaluation: ➠ Instructional Clarity (Score: 5): The prompt is syntactically flawless and logically structured. It explicitly addresses negative constraints (”not futuristic”) and defines the subject clearly, leaving no ambiguity for the model. ➠ Creative Elaboration (Score: 3): The imagery relies on standard, first-order associations with pollution (masks, coughing, gray smog). While descriptive, it lacks the unique, subtle narrative framing requested (”subtle but powerful”). ➠ Terminology Proficiency (Score: 2): The prompt relies entirely on layperson’s adjectives (”heavy”, ”muted”, ”soft shadows”) rather than professional photographic or rendering terminology (e.g., ”volumetric fog”, ”desaturated grading”, ”telephoto compression”). ➠ Intent Formalization (Score: 3): A partial failure in translation. While some mood is visualized, the prompt resorts to pasting the abstract goal directly (”conveys a quiet, powerful message...”) instead of operationalizing that message into observable visual cues. O.2. Subjective Image Evaluation (Constrained Task) This example illustrates a case of high-quality hallucination, where the model ignores explicit constraints (Objective Failure) but produces an aesthetically coherent image. This highlights the necessity of decoupling subjective quality from objective constraint verification. Task (CO 65): ”A cafe needs a promotional banner for its website. The design should feel nostalgic and warm, for a general audience. The style is retro 1950s advertisement. The banner must not contain any modern devices... The coffee cup must be red... ONLY red and cream colors. Text ’Fresh Coffee’ must be in the center.” Prompt (From Gemini-3-Pro-Preview and Flux.1 Pro, Novice): 44
AtelierEval
Figure 11. The generated image evaluated by AtelierJudge.
AtelierJudge Evaluation: ➠ Mood & Atmosphere (Score: 3): While the image captures the ”warm and nostalgic” keyword, it completely misses the ”Retro Advertisement” genre, presenting a cinematic still instead of a graphical banner. The mood is pleasant but fails to convey the intended commercial tone. ➠ Visual Composition (Score: 4): (Note: Specific layout constraints are penalized in the Objective Check.) As a standalone image, the composition is well-balanced and strictly framed. The visual weight is distributed effectively between the two subjects, creating a functional, albeit standard, photographic composition. ➠ Color & Lighting (Score: 4): (Note: Palette restrictions are penalized in the Objective Check.) Subjectively, the warm lighting and retro color grading are harmonious and aesthetically pleasing, despite fulfilling the wrong color palette constraints (full color vs. red/cream). ➠ Technical Flawlessness (Score: 3): The image exhibits distinct ”AI artifacts”: the skin textures appear overly smooth and plastic (”waxy”), and the steam from the cups has an unnatural, illustrative look that clashes with the realistic style. These recognizable generative traces prevent a higher score.
P. Details of Validation Set This appendix details the construction of the expert-annotated validation set used to meta-evaluate AtelierJudge. This set consists of 360 prompt–image pairs, one per task in AtelierEval, and is used to assess both the subjective scoring skills and the objective checklist-based skills of AtelierJudge. Sampling and composition. For each of the 360 tasks, we select exactly one prompt–image pair from the full pool of collected submissions. Each pair contains the original prompt provided to a T2I backend and the corresponding generated image, so that the two modalities always describe the same submission. Pairs are drawn via stratified random sampling over task type, so that the marginal distributions of tasks, systems, and difficulty closely match those of the overall benchmark. Each pair is associated with the task-specific checklist used for objective constraint verification. Annotation. For objective validation, one domain expert first labels whether the constraint is explicitly specified in the prompt and whether it is satisfied in the image. A second expert then independently reviews all labels and corrects any mistakes; residual disagreements are resolved through discussion, and the agreed labels are treated as the final objective annotations. For subjective validation, three independent experts rate each prompt and each image separately on all subjective dimensions. We take the arithmetic mean of the three scores as the subjective ground truth.
Q. Design Validation & Ablation Study of AtelierJudge In this appendix, we validate the design choices of the subjective skills in AtelierJudge. We investigate the impact of retrieval strategies (Li et al., 2025a; Wu et al., 2026), the number of retrieved exemplars (K), and the choice of embedding models. 45
AtelierEval
All experiments are conducted using GPT-5.2. We report the MAE, W1-A, and ρ against human expert ratings on the validation set (Appendix P). Table 11. Ablation study on memory retrieval strategies. Strategy
MAE ↓
W1-A ↑
ρ↑
Zero-shot Fixed Few-shot Random Retrieval Similarity Retrieval (Ours)
0.72 0.55 0.61 0.34
0.64 0.81 0.75 0.93
0.56 0.68 0.62 0.79
Impact of Retrieval Strategy. To verify the necessity of the similarity-based retrieval mechanism, we compare our method against three baselines, including the zero-shot baseline, a fixed few-shot baseline that utilizes a fixed set of 3 diverse exemplars (one per task category and covering 1-5 scores) for all queries, and a random baseline that randomly selects 3 exemplars from the corresponding memory for each query. As shown in Table 11, while providing fixed or random examples improves performance over the zero-shot baseline by establishing a general scoring scale, AtelierJudge significantly outperforms all other strategies. This confirms that retrieving semantically relevant exemplars is crucial for aligning the evaluator with the specific nuances of the evaluated task. Table 12. Impact of the number of retrieved exemplars (K). Setting
MAE ↓
W1-A ↑
ρ↑
Baseline (No Memory)
0.72
0.64
0.56
K=1 K=2 K = 3 (Ours) K=4
0.56 0.39 0.34 0.35
0.83 0.89 0.93 0.91
0.63 0.76 0.79 0.78
Sensitivity to Exemplar Count (K). We investigate the optimal number of retrieved exemplars K by varying it from 1 to 4. As presented in Table 12, the performance of AtelierJudge achieves peak when K = 3, as multiple exemplars provide a more robust triangulation of the scoring criteria. However, increasing K to 4 yields diminished performance. We attribute it to the context window with less relevant information, especially when processing images. Consequently, we adopt K = 3 as the optimal value. Table 13. Comparison of different embedding models for retrieval. Text Encoder
Image Encoder
Baseline (No Retrieval) NV-Embed-v2 (7B) Nomic-Embed-Text Nomic-Embed-Text (Ours)
DINO-V2-Giant DINO-V3 DINO-V2-Giant
MAE ↓
W1-A ↑
ρ↑
0.72
0.64
0.56
0.41 0.35 0.34
0.88 0.92 0.93
0.74 0.78 0.79
Impact of Embedding Models. We justify our choice of embedding models by comparing them against large-scale models. Following the selection strategies of recent studies (Luo et al., 2025a; Jose et al., 2025), we select our embedding models as Nomic-Embed-Text-V1.5 (dim=512) for text and DINO-V2-Giant for images. We replace the text encoder with NV-Embed-v2 (Lee et al., 2024) and the image encoder with DINO-V3 (Siméoni et al., 2025), two widely-recognized 7B models for comparison. The results in Table 13 reveal two insights. First, scaling up the image encoder to DINO-V3 brings no observable improvement, suggesting that DINO-V2 already captures sufficient perceptual features for this task. Second, surprisingly, NV-Embed-v2 leads to a performance degradation. We hypothesize that while NV-Embed-v2 excels in general benchmarks, its high-dimensional embedding space (dim=4096) may over-emphasize semantic minutiae rather than the instructional alignment features required for our specific retrieval task. These findings validate our choice. 46
AtelierEval
R. Detailed Benchmarking Results
Table 14. Detailed experimental results for novice and skilled prompters across different T2I backends on CO tasks. Subjective (Subj.) scores are reported on a 5-point scale, and Objective (Obj.) scores are reported as percentages Prompt
Model
GPT-5.2 Gem-3 Cl-4.5 GPT-4.1 Gem-2 Qwen-L GPT-4n Qwen-S Human
no. sk. no. sk. no. sk. no. sk. no. sk. no. sk. no. sk. no. sk. no. sk.
Image Avg.
nBanana
GI-1
Flux Pro
SDXL
Subj.
Obj.
Subj.
Obj.
Subj.
Obj.
Subj.
Obj.
Subj.
Obj.
Subj.
Obj.
2.99 2.93 2.81 2.81 2.57 2.69 2.66 2.74 2.27 2.32 2.73 2.82 2.50 2.61 2.49 2.47 3.10 4.05
37.5 35.6 36.2 37.4 27.8 33.2 32.9 36.1 22.8 27.1 31.9 35.3 28.4 32.4 26.5 26.9 48.2 78.1
3.65 3.60 3.68 3.65 3.50 3.54 3.57 3.65 3.41 3.45 3.64 3.58 3.49 3.52 3.50 3.43 2.77 3.80
41.8 41.9 41.6 41.5 40.7 41.4 42.4 42.0 38.8 39.4 41.3 40.0 42.1 43.2 40.1 39.2 51.4 75.7
3.83 3.86 3.76 3.81 3.65 3.73 3.67 3.86 3.54 3.71 3.79 3.75 3.73 3.75 3.71 3.59 2.98 3.92
47.3 46.8 45.5 46.0 45.0 45.7 47.9 47.5 44.9 46.1 43.9 44.0 45.7 46.6 46.5 46.1 55.5 82.8
3.88 3.90 3.88 3.91 3.77 3.83 3.86 3.94 3.77 3.77 3.88 3.79 3.82 3.73 3.78 3.71 2.54 3.87
47.2 48.1 45.5 46.3 48.6 46.7 46.2 46.5 45.7 45.8 45.5 44.6 47.4 48.7 46.7 45.8 62.8 81.5
3.60 3.48 3.63 3.61 3.43 3.43 3.51 3.56 3.24 3.28 3.54 3.46 3.32 3.42 3.36 3.32 2.98 3.82
38.6 38.5 39.4 39.2 34.4 38.1 40.4 41.0 32.7 34.5 40.1 38.2 38.3 38.4 34.8 34.6 50.1 72.4
3.27 3.16 3.46 3.29 3.17 3.19 3.26 3.23 3.10 3.04 3.36 3.30 3.07 3.19 3.16 3.10 2.57 3.61
34.2 34.1 35.8 34.4 34.9 34.9 35.0 33.1 31.9 31.3 35.7 33.4 37.0 39.3 32.3 30.5 37.2 66.0
Table 15. Detailed experimental results for novice and skilled prompters across different T2I backends on IM tasks. Subjective (Subj.) scores are reported on a 5-point scale, and Objective (Obj.) scores are reported as percentages Prompt
Model
GPT-5.2 Gem-3 Cl-4.5 GPT-4.1 Gem-2 Qwen-L GPT-4n Qwen-S Human
no. sk. no. sk. no. sk. no. sk. no. sk. no. sk. no. sk. no. sk. no. sk.
Image Avg.
nBanana
GI-1
Flux Pro
SDXL
Subj.
Obj.
Subj.
Obj.
Subj.
Obj.
Subj.
Obj.
Subj.
Obj.
Subj.
Obj.
4.46 4.66 4.25 4.62 3.87 4.32 3.63 4.07 3.23 3.50 3.79 4.39 3.55 3.84 3.26 3.64 3.02 3.94
56.7 64.6 53.5 66.6 41.2 59.0 32.9 45.1 24.8 34.4 38.5 52.8 30.6 43.4 42.6 46.3 40.4 71.3
/ / / / / / / / / / / / / / / / / /
60.9 62.9 60.4 63.1 57.1 61.7 55.4 58.4 52.6 57.4 48.9 57.9 49.9 54.5 41.7 40.9 40.0 58.9
/ / / / / / / / / / / / / / / / / /
74.8 79.2 72.4 81.9 68.4 77.3 62.1 69.4 60.1 67.0 60.5 74.1 55.6 64.7 53.7 56.4 37.2 73.0
/ / / / / / / / / / / / / / / / / /
71.6 76.5 70.5 77.5 65.6 74.3 62.0 68.1 58.9 66.8 58.4 70.9 56.4 62.3 54.0 50.8 43.8 70.4
/ / / / / / / / / / / / / / / / / /
59.8 57.9 60.5 57.3 54.8 57.1 55.7 57.5 50.0 55.2 46.0 53.5 49.7 53.0 37.2 36.3 40.9 53.3
/ / / / / / / / / / / / / / / / / /
37.5 37.9 38.4 35.7 39.5 38.3 41.7 38.7 41.2 40.4 30.7 32.9 37.7 38.0 22.7 20.5 38.0 38.9
47
AtelierEval Table 16. Detailed experimental results for novice and skilled prompters across different T2I backends on OE tasks. Subjective (Subj.) scores are reported on a 5-point scale, and Objective (Obj.) scores are reported as percentages Prompt
Model
GPT-5.2 Gem-3 Cl-4.5 GPT-4.1 Gem-2 Qwen-L GPT-4n Qwen-S Human
no. sk. no. sk. no. sk. no. sk. no. sk. no. sk. no. sk. no. sk. no. sk.
Image Avg.
nBanana
GI-1
Flux Pro
SDXL
Subj.
Obj.
Subj.
Obj.
Subj.
Obj.
Subj.
Obj.
Subj.
Obj.
Subj.
Obj.
4.70 4.72 4.37 4.60 3.90 4.30 4.10 4.34 3.38 3.88 4.43 4.56 3.76 4.00 3.71 4.02 2.58 3.65
95.8 90.8 87.8 92.5 94.4 88.2 86.3 87.0 82.8 83.1 92.4 88.4 84.3 83.8 82.8 84.9 80.9 92.3
4.27 4.32 4.34 4.35 4.26 4.30 4.28 4.33 4.23 4.32 4.33 4.39 4.29 4.31 4.30 4.31 3.25 4.14
94.3 91.4 92.0 91.7 94.4 92.2 91.3 91.4 91.4 91.8 93.6 91.3 90.5 90.7 91.7 91.5 79.4 95.5
4.45 4.48 4.44 4.49 4.39 4.44 4.47 4.47 4.36 4.41 4.47 4.49 4.44 4.42 4.40 4.43 3.24 4.30
97.5 94.4 93.8 93.7 96.7 94.6 92.7 93.9 92.9 93.1 96.8 94.6 92.7 91.7 92.2 93.6 81.9 98.9
4.53 4.55 4.56 4.57 4.50 4.51 4.51 4.55 4.46 4.49 4.51 4.59 4.52 4.51 4.50 4.49 3.95 4.15
96.9 95.5 93.4 93.8 96.8 94.9 94.7 93.7 94.1 93.8 96.2 94.0 93.3 93.7 94.1 96.1 92.6 98.6
4.25 4.32 4.33 4.39 4.22 4.25 4.26 4.31 4.11 4.29 4.29 4.41 4.23 4.28 4.26 4.32 2.90 4.05
94.0 90.1 91.4 93.2 94.2 91.0 90.8 92.4 90.5 92.5 92.5 89.9 89.8 90.7 91.2 91.2 79.4 93.6
3.85 3.93 4.03 3.94 3.95 4.01 3.89 4.00 4.00 4.09 4.07 4.07 3.99 4.02 4.04 4.03 2.91 4.07
89.0 85.7 89.4 86.2 90.0 88.2 87.2 85.6 88.4 87.8 88.9 86.5 86.1 86.5 89.2 85.3 63.7 91.2
S. Stability Analysis of Evaluation Scale In the main experiments summarized in Table 4, each prompter–task pair is evaluated under a fixed sampling scheme: for every task, the prompter produces a single natural-language prompt, each prompt is executed by every T2I backend to generate four images, and AtelierJudge retains the highest-scored image (top-1) to compute all metrics. This setting involves two potential sources of sampling variability: (i) the MLLM prompter when generating prompts and (ii) the T2I backend when generating images. Although all MLLMs are decoded with temperature 0.0 and shared hyperparameters (Appendix M), making prompt sampling nearly deterministic, it might be questioned whether using only one prompt and four images per task is sufficient for stable evaluation. This appendix therefore studies the sampling stability of our evaluation scale and examines how scores change when increasing the number of prompts and images. Experimental setup. We conduct a controlled stability study under a single representative configuration. The prompter is GPT-5.2 and the T2I backend is GI-1, a strong pair that is also used extensively in the main results. We evaluate on a stratified subset of 120 tasks from AtelierEval, balanced across the three task categories (OE, CO, IM) introduced in Section 3.2. Decoding and evaluator hyperparameters follow Appendix K: GPT-5.2 runs with temperature 0.0, and all prompts and images are scored by AtelierJudge as described in Section 4 using the subjective dimensions and memories from Appendix H and Appendix J. For all conditions we reuse the full evaluation pipeline of Section 3.4 and Section 5.1 and vary only the numbers of prompts and images. For the prompt-scale analysis we vary the number of independently sampled prompts per task,Nprompt ∈ {1, 2, 3, 4}, while keeping the image sampling scheme fixed to four images per prompt. For each Nprompt we run the complete pipeline and report two aggregated metrics over tasks: the mean subjective prompt score (1–5) and the mean prompt-side checklist satisfaction rate in percent, referred to as prompt accuracy. Even though decoding uses temperature 0.0, we still treat repeated generations as independent runs and check whether these aggregate metrics drift as Nprompt increases. For the image-scale analysis we fix a single prompt per task and vary the number of images sampled from GI-1, Nimg ∈ {1, 2, 4, 8, 16}. For each task and each Nimg , GI-1 generates Nimg images under the same configuration as in the main experiments. AtelierJudge scores all images independently, and we retain the highest-scored image (top-1) when aggregating metrics, exactly as in the main experiments. For each Nimg we again report two metrics averaged over tasks: the mean subjective score of the top-1 image and the mean image-side checklist satisfaction rate in percent (top-1 image accuracy). In both analyses, checklist satisfaction is computed following the objective evaluation protocol of Section 4.2. 48
AtelierEval
Results and discussion. Figure 12 shows the stability results. Across all values of Nprompt and Nimg , the curves for subjective scores and objective checklist-based accuracies remain nearly flat, and the small fluctuations are within random variation, indicating that neither metric is materially affected by the number of sampled prompts or images. Taken together, these observations justify the default design used throughout the main experiments: evaluating each prompter–task pair with one prompt and four images per T2I backend, retaining the top-1 image for scoring. This configuration lies on the plateau of both stability curves and therefore provides a computationally efficient yet converged operating point. The small variations observed across different Nprompt and Nimg further indicate that our evaluation scale is robust to these sampling hyperparameters for both subjective scores and objective accuracy, and that the main experimental conclusions are stable.
Figure 12. Prompt-scale and image-scale stability for GPT-5.2 as prompter and GI-1 as the T2I backend. Left: effect of the number of prompt samples per task on mean prompt score (left axis) and prompt accuracy (right axis). Right: effect of the number of image samples per prompt on mean top-1 image score (left axis) and top-1 image accuracy (right axis). Accuracy is the checklist-based constraint satisfaction rate defined in Section 4.2.
T. User Study Materials This section presents all materials provided to the user study participants, including the informed consent form and the pre-test screening questionnaire. To preserve anonymity during the review process, certain experiment-irrelevant details have been omitted. Detailed ethical considerations about our human study are provided in Appendix W. T.1. Informed Consent Form The following is a static version of the informed consent form for inclusion. In the actual study, participants were required to check a box and sign, indicating they had read, understood, and voluntarily agreed to participate. Research Project Title: AtelierEval: Agentic Evaluation of Text-to-Image Prompters Principal Investigator: xxxxxxx Institution: xxxxxxxxxxxxxx Contact Information: [email protected] 1. I NTRODUCTION You are invited to participate in an academic study aimed at establishing a standardized benchmark for evaluating Prompting Proficiency in the era of Generative AI. Your participation will provide critical data to understand how humans translate visual intent into textual instructions compared to AI. This study is conducted entirely in English and requires proficiency in reading complex instructions and writing descriptive text. To ensure the ethical conduct of this research and the protection of all participants, this study has been reviewed and approved by the Institutional Review Board (IRB). If you have any questions, please contact the Principal Investigator by email. 2. P ROCEDURES If you agree to participate in this study, you will be asked to complete the following steps: • Pre-Test Screening: You will first complete a brief questionnaire (approx. 5–10 minutes) regarding your experience with Text-to-Image tools and background. This ensures you meet the study’s criteria and allows us to categorize participants into “Novice” or “Skilled” groups. • Mandatory Technical Setup: 49
AtelierEval
– Screen Recording: To ensure integrity and verify that no unauthorized AI tools are used, you are required to record your entire screen during the session. Before starting the recording, please close all personal messaging apps and irrelevant browser tabs to protect your privacy. You must record the full screen. To ensure manageable file sizes and clarity, please restrict the recording resolution from 1920×1200 to 1280×720. Please note that failure to provide a valid screen recording, evidence of prohibited tool usage, or obvious lack of effort (e.g., irrelevant inputs) will result in disqualification and forfeiture of all compensation. – Network Access: You are kindly required to access HuggingFace and DeepL in the test. If you experience difficulties accessing them due to network restrictions, we will provide a secure VPN service for the duration of the study. • Prompting Tasks: If selected, you will access a web-based interface (Gradio-based, similar to Stable Diffusion WebUI) to write text prompts to generate images based on specific requirements. You will complete a total of 10 prompting tasks for each of the below three task types: – Open-ended Creation: Writing prompts based on abstract requests to test creativity. – Constrained Creation: Writing prompts based on requests that strictly adhere to logical or technical constraints. – Imitation: Writing prompts to reverse-engineer and reproduce a reference image. To encourage careful thinking and prevent low-quality submissions, a mandatory minimum timer for 1 min is required for each task. The submission button will remain disabled until the timer expires. You are required to complete 5 tasks each round, 6 rounds in total. Between rounds, you may take a short break to rest (typically up to 10 minutes). Please note that if there is no activity for more than 10 minutes, the session may be automatically ended, and you may need to restart the study. • Tool Usage Policy: – Allowed: You are permitted to use standard mode of DeepL Translator strictly for language translation. – Prohibited: You may NOT use AI chatbots (e.g., ChatGPT, Claude, Gemini). You may NOT use any other functions of DeepL, such as DeepL Write for AI Polish/Rewrite. The content of the prompt must originate from you. • Time Commitment: – Novice Group: Estimated total time is approximately 45 to 60 minutes. – Skilled Group: Estimated total time is approximately 1.5 to 3 hours, reflecting the expectation of high-fidelity, professional-grade inputs. • Data Collection: The system will record your submitted text prompts, timestamps, generated images, and operational logs. Upon completion, you must manually upload the screen recording file to the provided secure Google Drive or Baidu Netdisk link. 3. R ISKS • Risks: As approved by the IRB, the risks associated with this study are minimal. You may experience some stress or frustration if the generated images do not meet your expectations, which is a common occurrence in generative AI. All provided images and task descriptions are screened to eliminate any harmful contents. You may contact us if you find anything that makes you uncomfortable. All your personal identity data will be strictly anonymized. 4. C OMPENSATION Upon completion of all 30 tasks and the questionnaires, novice participants will receive a compensation of 12 USD (or the equivalent in local currency), while skilled participants will receive 50 USD. Payments can be processed via Amazon Gift Card, Zelle, Alipay, or WeChat Pay, with the specific method to be coordinated with each participant following the study’s conclusion. 5. C ONFIDENTIALITY We will take strict measures to protect your privacy. • Anonymous Access: To ensure your anonymity, you will not use your personal HuggingFace account. You will be provided with a uniformly assigned, anonymous HuggingFace account to access HuggingFace Space for the Gradio-based user interface for the tasks. The credentials for this account will be sent to your registered email address. 50
AtelierEval
• Data Usage: The personal information you provide will be used for specific, distinct purposes. Your contact information, typically email, will be used strictly for study-related communication and compensation. Your demographic and background information will be used for anonymized statistical analysis in our research. Your screen recording is strictly for verification. It will be accessible only to the research team and will be permanently deleted after the verification process is complete. • Publication: All published research findings will use fully anonymized, aggregated data. No information that could personally identify you will be disclosed. 6. VOLUNTARY PARTICIPATION Your participation is voluntary. You may withdraw at any time without penalty. T.2. Pre-Test Questionnaire The following questionnaire was administered to screen and assign participants. This questionnaire is designed to understand your background to ensure you meet the participation criteria for this study and to assign you to the most suitable task group. The information you provide will be kept strictly confidential. Please ensure that all information you provide is truthful. We reserve the right to withhold compensation if any of the information is found to be false. We will try our best to assign you to the group you applied for, but we do not guarantee it. PART 1: BASIC I NFORMATION 1. Name: 2. Email Address: 3. Which group are you applying for? (Select one) • Skilled • Novice PART 2: D EMOGRAPHIC & E DUCATIONAL BACKGROUND 4. Age range: • 18–22 • 22-25 • 26–30 • 30–40 • 40+ • Prefer not to say 5. Gender: • Male • Female • Non-binary • Prefer not to say 6. Current primary region of residence or work (for the past 3+ years): • East Asia (e.g., Mainland China, Hong Kong SAR, Taiwan, Japan, Korea) • Southeast Asia (e.g., Singapore, Malaysia, Thailand) • South Asia • Middle East • North America 51
AtelierEval
• Europe • Latin America • Sub-Saharan Africa • Oceania • Other • Prefer not to say 7. What is your current or highest level of education? • Secondary Education or Technical College • Undergraduate • Master • PhD • Prefer not to say 8. Major(s) (Use “None” if Secondary Education): PART 3: E NGLISH P ROFICIENCY 9. Is English your native language? (Yes / No) 10. (If No) Standardized English test scores (if applicable): • TOEFL • IELTS • Duolingo • CET-4 • CET-6 • Other • I have not taken any PART 4: T EXT- TO -I MAGE E XPERIENCE & P ROFICIENCY This section determines your group. Please answer truthfully. 11. How long have you been using T2I tools? • I have never used them. • Less than 1 month. • 1–6 months. • 6 months – 1 year. • More than 1 year. 12. Which types of tools do you use regularly? (Select all that apply) • Conversational Interfaces: Interfaces where you describe the image in natural language, and an AI assistant rewrites/handles the prompt for you (e.g., the image generation function of ChatGPT, Bing Image Creator, Doubao, Gemini). • Professional/Direct Web Tools: Web interfaces where your exact text is sent to the model without hidden rewriting (e.g., Google AI Studio, Midjourney, Ideogram). • Local Interfaces: Advanced local interfaces allowing model switching, ControlNet, or node-based editing (e.g., Stable Diffusion WebUI, ComfyUI). • Design Software Integration: Image generation tools integrated in professional design software (e.g., Photoshop Generative Fill). 52
AtelierEval
13. Knowledge Check: Which of the following concepts can you confidently explain or use? (Check all that apply) This helps us gauge your technical and artistic depth. • Technical: Seed / Randomness • Technical: CFG Scale (Classifier-Free Guidance) • Technical: Checkpoints / Base Models / LoRA / Embeddings / Textual Inversion • Artistic: Composition Rules (e.g., Rule of Thirds, Golden Ratio) • Artistic: Lighting Styles (e.g., Volumetric, Chiaroscuro, Rim Light) • Artistic: Camera Angles/Lens (e.g., Isometric, Macro, Wide-angle) • None of the above 14. Self-Assessment: How do you typically construct a prompt? • Natural Language: “A cat sitting on a bench, sunny day.” • Keyword Stacking: “Cat, bench, park, sun, 4k, high quality, masterpiece.” • Structured Engineering: I systematically organize subject, medium, style, artist references, and technical parameters. PART 5: P ORTFOLIO V ERIFICATION (R EQUIRED FOR S KILLED U SERS ) To verify your expertise level, please share 1–3 recent examples of your work. Both the prompt text and the generated image are required. Alternatively, if you have a portfolio on a social media or AI art platform (e.g., Civitai, ArtStation), please provide the link. 15. Portfolio Link or Upload Description:
U. Participant Workflow and Interface Design This section provides a comprehensive walkthrough of the AtelierEval assessment platform, illustrating both the participant experience and the underlying workflow design (Wang et al., 2024a; Qin et al., 2025b). We present the complete procedural flow that participants encounter during evaluation, showcasing the user interface design and highlighting the framework’s ecological validity. This walkthrough serves dual purposes: (1) demonstrating the practical implementation of our benchmark, and (2) providing transparency for reproducibility. The entire protocol is designed as a self-contained experience, with all tasks performed within a user-friendly, Gradio-based web interface hosted on Hugging Face Spaces. U.1. Welcome and Authentication Participants begin by accessing the assessment platform via a provided Hugging Face Space link using their assigned anonymous account credentials. Upon loading, they are greeted with a welcome screen that provides a concise overview of the assessment’s structure and objectives, as shown in Figure 13. The interface explains that the assessment evaluates three distinct skill dimensions: (1) Open-Ended Creation for testing creative interpretation and problem-solving, (2) Constrained Creation for evaluating logical precision under multiple technical requirements, and (3) Imitation for measuring visual analysis and reverse-engineering capabilities. Participants are instructed to enter their assigned anonymous email address into the text field and click the “Login and Start Assessment” button to proceed. This authentication step ensures proper data association while maintaining participant anonymity throughout the study. The clean, welcoming design with clear instructions minimizes cognitive load and technical barriers, allowing participants to focus on the prompting tasks themselves. U.2. Assessment Structure and Randomization After successful authentication, participants are informed that they will complete two independent assessment rounds. Each round consists of 15 prompting tasks distributed across the three task types (5 tasks per type). Participants may take short breaks between rounds. To ensure robust evaluation and prevent order effects, the system implements two levels of randomization: 53
AtelierEval
Figure 13. The welcome and login interface. Participants enter their anonymous email address and receive a clear overview of the three assessment modules before beginning.
• Task Type Order Randomization: Within each round, the three task types (Open-Ended, Constrained, Imitation) are presented in a randomized sequence. This means a participant might encounter Constrained tasks first in Round 1 but Imitation tasks first in Round 2. • Task Instance Randomization: For each task type, the 5 specific task instances are randomly sampled from a curated pool of 360 validated problems (120 per task type). This ensures that no two participants receive identical task sets, while maintaining consistent difficulty and construct validity across all instances. This randomization protocol serves dual purposes: 1. Methodological rigor: Controls for learning effects and position bias in our experimental design. 2. Ecological realism: Simulates the unpredictable variety of real-world creative demands that professional prompters encounter in practice. U.3. Task Type Interfaces Having authenticated and understood the overall assessment structure, participants proceed to the core evaluation component. The following subsections describe the three task types that participants encounter in randomized order. Each task type begins with a dedicated introduction screen explaining its specific objectives and evaluation criteria, followed by the task interface where participants construct their prompts. Navigation buttons (“Previous” and “Next”) allow participants to review and modify their responses or adjust task order if needed. U.3.1. O PEN -E NDED C REATION M ODULE When participants first encounter an Open-Ended Creation block, they are presented with an introductory screen (Figure 14) that explains the purpose and nature of these tasks. The interface informs participants that they will receive natural, conversational project briefs similar to real-world creative requests from clients, editors, or teammates. Each scenario includes context, audience, and implicit expectations, but deliberately avoids rigid checklists to test the participant’s ability to independently interpret requirements and translate abstract concepts into concrete visual specifications. The screen emphasizes that this task type evaluates creative interpretation and professional communication skills—the ability to decode intent, balance competing priorities, and craft prompts that capture the envisioned aesthetic without explicit guidance. For each task within this module, participants are presented with a detailed creative scenario. Figure 15 shows a representative example where a participant needs to create cover art for a children’s book. The task description provides rich context 54
AtelierEval
Figure 14. Introduction screen for the Open-Ended Creation module, explaining the conversational nature of creative briefs and the assessment objectives.
including the target audience (ages 4–7), thematic requirements (whimsical, magical, peaceful), stylistic constraints (soft digital painting, not photorealistic), and technical specifications (no dark or scary clouds). Participants need to synthesize this multifaceted information into a single, coherent prompt entered in the text box labeled “Enter your prompt here.” Once satisfied with their response, participants click “Next” to proceed to the subsequent task.
Figure 15. Example task interface for Open-Ended Creation. The scenario provides rich contextual information that needs to be translated into an effective prompt.
U.3.2. C ONSTRAINED C REATION M ODULE When participants enter a Constrained Creation block, the introductory screen (Figure 16) explains that this module evaluates their ability to precisely control AI output under multiple simultaneous constraints. Unlike the open-ended tasks, success here requires accurately encoding all explicit requirements—including color restrictions, spatial layouts, exact quantities, and logical bindings—without introducing conflicts or omissions. The interface emphasizes that this module tests logical reasoning and systematic control skills, challenging participants to orchestrate competing technical demands into a single, compliant prompt. Each Constrained Creation task presents a detailed list of explicit requirements in a structured, bulleted format. Figure 17 illustrates a representative example: creating an Instagram advertisement for iced tea. The task specifies: • Title: Product name must appear as text overlay • Pairing: Specific hand positions for two objects • Constraints: Restricted color palette (peach, mint green, beige only) • Layout: Product photography style illustration • Quantity: Exact number of visible hands (two) 55
AtelierEval
Figure 16. Introduction screen for the Constrained Creation module, highlighting the focus on technical precision and constraint satisfaction.
• Text: Brand name must be clearly printed on the can • Prohibitions: No plastic items allowed Participants need to construct a prompt that satisfies all constraints simultaneously without internal contradictions. The task description is displayed at the top, followed by the structured constraint list, with the prompt input area below.
Figure 17. Example task interface for Constrained Creation. The bulleted list explicitly enumerates all requirements, testing the participant’s ability to encode multiple constraints into a single compliant prompt.
U.3.3. I MITATION AND R EPRODUCTION M ODULE When participants encounter the Imitation module, the introductory screen (Figure 18) explains that they will be shown target images and need to deconstruct their visual components—including subject matter, composition, lighting, style, and technical details—into descriptive prompts that can reproduce the images as closely as possible. This module tests visual analysis and technical vocabulary skills, requiring participants to translate what they observe into precise, structured language that captures both obvious elements and subtle stylistic nuances. For each Imitation task, the interface adopts a side-by-side layout (Figure 19). The target image is displayed prominently on the left side of the screen, while the prompt input area occupies the right side. This parallel presentation allows participants to continuously reference the target image while constructing their descriptive prompt, facilitating iterative refinement of their textual description. The example shown in Figure 19 presents a cross-sectional diagram of Earth’s interior structure. Participants need to identify and describe not only the obvious elements (planetary sphere, labeled layers) but also technical details such as the cutaway 56
AtelierEval
Figure 18. Introduction screen for the Imitation and Reproduction module, emphasizing the reverse-engineering challenge.
visualization style, color-coding scheme, text annotations, and scientific illustration aesthetic. The task description simply states: “Your goal is to write a prompt that replicates the target image on the left as closely as possible,” providing no additional hints about which features to prioritize.
Figure 19. Example task interface for Imitation. The side-by-side layout presents the target image (left) alongside the prompt input area (right), enabling continuous visual reference during prompt construction.
U.4. Workflow Integration and Data Collection Throughout the entire assessment, the system automatically records all participant interactions, including: • Submitted prompt text for each task • Task type and instance identifier for each response • Task progression order (reflecting the randomized sequence) • Round completion status (Round 1 vs. Round 2) • Operational logs including interface interactions The two-round structure serves multiple purposes: it increases the sample size per participant for more robust statistical analysis, allows examination of within-subject consistency, and provides sufficient task diversity through the randomization mechanism. Each participant ultimately completes 30 tasks in total (15 tasks × 2 rounds), with 10 responses per task type collected across both rounds. Upon completing all tasks, participants upload their mandatory screen recording to a secure cloud storage location (Google Drive or Baidu Netdisk). The screen recording is strictly for verification purposes and will be permanently deleted after the verification process is complete. This self-contained, browser-based workflow ensures a standardized experience across all participants while maintaining high ecological validity through multiple design principles: 57
AtelierEval
1. Realistic task scenarios: All tasks are grounded in authentic professional use cases, from creative briefs to technical specifications to visual reproduction challenges. 2. Professional-grade interface: The Gradio-based UI mirrors industry-standard text-to-image platforms, ensuring participants interact with familiar design patterns and workflows. 3. Authentic cognitive demands: Time constraints, task complexity, and the need for multifaceted decision-making reflect real-world prompting scenarios. 4. Naturalistic interaction patterns: Participants construct prompts without artificial restrictions, using their own vocabulary, style, and problem-solving approaches. The combination of randomized task presentation and comprehensive data collection enables rigorous assessment of prompting proficiency across multiple skill dimensions while controlling for potential confounds such as learning effects and order bias. This design allows our findings to generalize effectively to real-world text-to-image prompting contexts.
V. Computational Resource Consumption Benchmarking Consumption. The primary resource consumption stems from API invocations for three commercial backends (Gemini-3-Pro-Image-Preview, GPT-Image-1-All, and Flux.1 Pro), generating 7,200 images each, alongside MLLM overheads for prompting and evaluation. Given shared institutional access and pricing variability, we report usage via query volume rather than monetary expenditure. Locally, SDXL also generated 7,200 images on a workstation with NVIDIA RTX 5080 (16GB). Precise GPU hours are not isolated due to the mixed-workload nature of the deployment environment. Evaluator Consumption. To assess evaluator efficiency, we conducted a controlled experiment using GPT-5.2 on 30 sampled tasks involving full objective and subjective cycles. Based on official pricing, the zero-shot QA/VQA baseline incurred $0.35 versus $0.89 for AtelierJudge—a margin we consider justifiable given the substantial reliability gains. Additionally, owing to engineering optimizations, the embedding calculation for each T2I-prompter pair requires less than 5 minutes on our local workstation, rendering this overhead negligible compared to the T2I generation costs.
W. Detailed Ethical Considerations and Procedures This section summarizes the ethical considerations and related procedures of the user study of our work. The full informed consent form and all questionnaires are provided in Appendix T. W.1. Ethics Approval and Consent This study was reviewed and approved by the Institutional Review Board (IRB) at the first author’s institution. Prior to any experimental tasks, all participants were provided with a comprehensive Informed Consent Form (Appendix T.1). This document detailed the research objectives, experimental procedures, potential risks, and compensation details. Participants were required to explicitly indicate their agreement to proceed. Participants were informed of their right to withdraw from the experiment at any stage without penalty. To ensure the capacity for informed consent, all participants were required to be at least 18 years of age. W.2. Participant Anonymity This subsection outlines the data protection protocols enforced to guarantee participant confidentiality and data integrity throughout the user study component of our work. Anonymity in Tests. The core strategy for maintaining anonymity was the utilization of pre-configured, anonymous Hugging Face accounts. The research team assigned a unique, anonymous account credential to each participant. These credentials were securely transmitted to participants via email before the session began, and users were required to use these credentials to access the tests. By adhering to this protocol, we ensured that all interactions within the Gradio-based UI in Hugging Face Space were completely decoupled from the participants’ real-world identities. 58
AtelierEval
Anonymity in Tool Usage. To further safeguard participant privacy, we implemented strict protocols for tool deployment. For VPN, we used Clash, an open-source VPN tool. Only generic config files were distributed without any user accounts or credentials. For DeepL, participants were required to utilize the non-registered version. For screen recording, participants were instructed to avoid displaying Personally Identifiable Information (PII) during the session. All recordings are subject to unified post-hoc deletion immediately after the validity verification of the submissions. Participant Information Verification. To maintain the consistency of the skilled group and to validate self-assessments provided in the pre-test questionnaire (Appendix T.2), we conducted a limited eligibility verification for participants who self-identified as skilled users. This verification primarily involved reviewing publicly available portfolios or example works voluntarily provided by participants (e.g., prompts and generated images shared on content platforms). In a small number of cases, participants were invited to participate in a brief online discussion (with cameras turned off). All verification steps were performed solely to confirm eligibility for the skilled group. Any personal information accessed during this process was handled confidentially and was not retained beyond the verification stage. De-identification and Aggregation. Following the conclusion of data collection, a comprehensive de-identification procedure was executed. For analysis purposes, only the anonymous Hugging Face IDs were retained. Identifying information was stored separately and never linked in the released dataset. Sensitive contact information required for compensation was isolated in encrypted storage with restricted access. The final research data and questionnaire responses were linked exclusively via anonymous IDs. All findings presented in this paper and any open-source data are fully anonymized and aggregated, ensuring that no individual participant can be re-identified.
59