ConceptioArchivearXiv CS
arXiv CSopen access

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
artificialintelligenceknowledgerepresentationreasoning
artificial intelligence, reasoning, knowledge representation

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation Ziwei Zhou * 1 Zeyuan Lai * 2 Rui Wang 1 Yifan Yang 3 Zhen Xing 1 Yuqing Yang 3 Qi Dai 3 Lili Qiu 3 Chong Luo 3

arXiv:2604.08540v1 [cs.CV] 9 Apr 2026

Abstract

Others

Text-to-Audio-Video (T2AV) generation is rapidly becoming a core interface for media creation, yet its evaluation remains fragmented. Existing benchmarks largely assess audio and video in isolation or rely on coarse embedding similarity, failing to capture fine-grained joint correctness required by realistic prompts. We introduce AVGen-Bench, a task-driven benchmark for T2AV generation, featuring high-quality prompts across 11 real-world categories. To support comprehensive assessment, we propose a multi-granular evaluation framework that combines lightweight specialist models with Multimodal Large Language Models (MLLMs), enabling evaluation from perceptual quality to fine-grained semantic controllability. Our evaluation reveals a pronounced gap between strong audio-visual aesthetics and weak semantic reliability, including persistent failures in text rendering, speech coherence, physical reasoning, and universal breakdown in musical pitch control. Code and benchmark resources are available at http://aka.ms/avgenbench.

AVGen-Bench (Ours)

Seperate AV Evaluation

Joint AV Evaluation

General & Coarsegrained Evaluation Simple & Unclassified Prompts

Rich & Fine-grained Evaluation Complicated Prompts

Figure 1. Comparison between AVGen-Bench and existing benchmarks. Unlike prior works that rely on separate audio/visual evaluations and simple prompts, AVGen-Bench introduces (1) joint audio-visual evaluation, (2) fine-grained metrics across 10 dimensions, and (3) rich, complex prompts with high token counts to ensure a rigorous assessment.

chronized and semantically correct audio can dramatically enhance immersion—for example, the crisp cutting sound in a fruit-slicing clip or intelligible dialogue in a conversational scene. As frontier systems such as Sora 2 (OpenAI, 2025), Veo 3.1 (DeepMind, 2026b), and Kling 2.6 (KuaishouTechnology, 2026) emerge, T2AV generation is quickly becoming the default interface for user-centric creation. Despite rapid architectural progress, the field faces a critical bottleneck: the lack of a rigorous and holistic evaluation framework for T2AV. Most existing benchmarks for generative models were designed for uni-modal settings. Visual benchmarks such as VBench (Huang et al., 2024) and VBench++ (Huang et al., 2025) focus exclusively on video quality, while audio benchmarks typically evaluate sound in isolation. More recent efforts attempt to combine audio and video evaluation (Wang et al., 2025a; Zhang et al., 2025; Liu et al., 2025; Hu et al., 2025), but they still fall short in two key aspects. First, they often rely on coarse-grained metrics that score overall audio, video, or audio-visual quality, without distinguishing specific capabilities or failure modes. Second, joint evaluation is commonly reduced to embedding similarity using models such as CLIP (Radford et al., 2021) or CLAP (Wu et al., 2023), which is insufficient for verifying fine-grained semantic alignment required by realistic prompts.

1. Introduction The landscape of generative video is undergoing a fundamental shift from silent Text-to-Video (T2V) synthesis (OpenAI, 2024; Wan et al., 2025; Wu et al., 2025) to multimodal Text-to-Audio-Video (T2AV) generation (Low et al., 2025; HaCohen et al., 2026; AI, 2026). This transition is not merely an incremental feature upgrade. In many real-world AIGC scenarios, audio is essential for conveying information, realism, and engagement. A visually plausible video without sound is often flat and uninformative, while syn1

Fudan University 2 University of Science and Technology of China 3 Microsoft Research Asia. Correspondence to: Yifan Yang <[email protected]>.

This limitation becomes particularly evident in real T2AV usage. Users typically provide a single textual prompt that 1

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation (a1) Prompted Text Rendering Failures

Veo3.1 Fast

Ovi

(a2) Incidental Text Rendering Failures

LTX-2 Seedance 1.5 Pro

Veo3.1 Fast

LTX-2

Veo 3.1 Fast

Prompt: "... The player performs four block chords in sequence, one per bar: C major, then G major, then A minor, then F major...

All Models: Random Notes

Seedance 1.5 Pro

Prompt: "...the title CITY OF GLASS appears..."

(c) Speech Generation Errors Incidental

(b) Pitch Inaccuracy

(d) Semantic Misalignment

Explicitly Prompted

Prompt: "A teammate voice in Ovi: "Thank you sciencecomms says, 'Reloading—cover tavator wassupkitsignsolivo, me!'" to look up tonight." Kling 2.6: "Cover me!"

Prompt: "Four-shot teaser with comedic timing and punchy sound cues. Shot 1: A quiet office kitchen; a kettle begins to whistle as a deadpan voice says, ... Shot 2: Close-up of a mug slipping; it hits the floor and ... Shot 3: Smash cut to a formal meeting room; polite applause abruptly stops as someone coughs loudly into a microphone, ..." Identity Drift: Significant identity loss occurs during shot transitions or large pose changes

Prompt: "A small piece of sodium metal is dropped onto the surface with a tiny plunk ..." Expectation1: A visible trail of white smoke and gas bubbles follows the reaction, though it originates from the bottom. Expectation 2: The metal piece sinks directly to the floor of the tank. It does not stay on the surface or (e) Violation of Physical Laws skate/dart across it as required.

Crowd Degradation: In multi-face scenarios, rendering quality and stability collapse.

(f) Face Rendering Failures

Figure 2. Qualitative examples of failure modes across different fine-grained dimensions. (a1) Explicitly prompted text rendering. (a2) Incidental text rendering in background elements. (b) Fine-grained musical control (Pitch Accuracy). (c) Speech generation regarding incidental coherence and explicit instruction following. (d) Holistic semantic alignment in complex multi-shot narratives. (e) High-level physical plausibility and dynamic constraints. (f) Facial consistency failures, illustrating identity drift across shot transitions and degradation in multi-face crowd scenes. Red markings and crosses (✗) indicate generated errors.

interleaves visual and acoustic requirements—often implicitly—rather than specifying audio and video separately. Under such settings, current models exhibit recurring yet undermeasured failure modes: speech content that is unintelligible or incorrect, environmental sounds that do not align with visual events, mismatched lip movements, incorrect musical notes despite realistic playing motions, and violations of basic physical or causal logic. Figure 2 illustrates representative examples of these phenomena. Without a benchmark that explicitly targets these joint, fine-grained behaviors, it is difficult to diagnose model weaknesses or guide future progress.

enables meaningful evaluation of not only perceptual quality, but also whether a model can accomplish what the user intends in a given scenario. Furthermore, we propose a comprehensive, multi-granular evaluation suite for T2AV. Beyond basic uni-modal aesthetics and audio-visual synchronization, our framework introduces targeted metrics for fine-grained controllability and semantic correctness, including scene text legibility, facial identity consistency, pitch accuracy in music generation, speech intelligibility, and physical plausibility. Methodologically, we adopt a hybrid evaluation strategy that integrates lightweight specialist models with Multimodal Large Language Models (MLLMs). This design leverages the complementary strengths of both paradigms: specialist models provide precise signal-level measurements, while MLLMs enable high-level semantic reasoning and holistic intent verification.

To bridge this gap, we introduce AVGen-Bench, a taskdriven benchmark dedicated to Text-to-Audio-Video generation. Instead of tailoring prompts to fit available metrics, AVGen-Bench is grounded in realistic user intents and application scenarios. Our prompt suite spans 11 daily-life categories, covering professional media production (e.g., movie trailers and advertisements), creator economy applications (e.g., music tutorials and gameplay), and physically grounded world simulation tasks. This task-centric design

In summary, our contributions are threefold: (1) A TaskDriven T2AV Benchmark. We present AVGen-Bench, a curated benchmark with high-quality prompts across 11 real-world categories, shifting evaluation from metric-driven 2

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

Task-Driven Prompt

Multi-Granular Evaluation Suite Mono-modal Quality

Professional Media News

Ads Movie Creator Economy

ASMR

Music

Cooking

Gaming

Visual Quality

Audio Quality

Q-Align

AudioboxAesthetic

T2AV

Cross Modal Alignment AV Sync

Lip Sync

Syncformer

Syncnet

Fine-grained Evaluation Modules

Models

Text Rendering

Facial Consistency

Pitch Accuracy

Paddle-OCR

InsightFace

Basic-Pitch

Speech Intelligibility & Coherence

Physical Plausibility

Holistic Semantic Alignment

Reasoner World Simulator Physics

Chemistry

Animals

Sports

Gemini

Whisper

VideoPhy-2AutoEval

Gemini

Figure 3. Overview of the AVGen-Bench framework. The benchmark features a Task-Driven Prompt Set (left) categorized into three real-world application domains: Professional Media, Creator Economy, and World Simulation. The generated content is evaluated via our Multi-Granular Evaluation Suite (right), which employs a hybrid strategy combining lightweight specialist models (orange) for signal-level precision and MLLMs (purple) for high-level semantic reasoning and physical plausibility analysis. Table 1. Comparison with existing benchmarks. AVGen-Bench features the highest average prompt complexity (Avg. Tokens) and a comprehensive set of evaluation metrics covering all audio modalities.

design to user-centric task understanding. (2) A MultiGranular, Hybrid Evaluation Framework. We introduce a unified evaluation suite that jointly assesses uni-modal quality, audio-visual consistency, and fine-grained semantic alignment by combining specialist models with MLLMs. (3) A Systematic Diagnosis of T2AV Failure Modes. Through extensive evaluation, we reveal a sharp gap between strong audio-visual aesthetics and weak fine-grained semantic control, highlighting critical challenges in text, speech, and physical reasoning.

2. Related Works

Benchmark

Task

#Metrics

Avg. Tokens

VBench 2.0 TTA-Bench JavisBench Harmony-Bench VerseBench UniAVGen

T2V T2A T2AV TI2AV TI2AV TI2AV

18 10 4 6 4 3

26.56 20.00 65.00 68.00 -

Audio Types SFX, Music, Speech SFX, SFX, Music, Speech SFX, Speech Speech

AVGen-Bench (Ours)

T2AV

10

88.54

SFX, Music, Speech

nAI, 2025), Veo 3.1 (DeepMind, 2026b), Wan 2.6 (AI, 2026), and Kling 2.6 (KuaishouTechnology, 2026), demonstrate high-fidelity synchronized synthesis. In the open domain, models like Ovi (Low et al., 2025) and JavisDiT (Liu et al., 2025) explore dual-stream Diffusion Transformers, while LTX-2 (HaCohen et al., 2026) employs flow matching. Recently, hybrid architectures such as MAViD (Pang et al., 2025) have emerged, combining Autoregressive (AR) modeling with diffusion to enhance cross-modal consistency. Complementary to these unified approaches are conditional pipelines, including Video-to-Audio (e.g., MMAudio (Cheng et al., 2025), Kling-Foley (Wang et al., 2025c)) and Audio-to-Video (e.g., MTVCraft (Weng et al., 2025), Wan-S2V (Gao et al., 2025)) systems. While useful, these cascaded methods differ fundamentally from the holistic world modeling aim of simultaneous generation.

2.1. Audio-Video Generation Models Text-to-Video (T2V) Synthesis. The advent of Sora (OpenAI, 2024) marked a paradigm shift, demonstrating the scalability of Diffusion Transformers (DiT) (Peebles & Xie, 2023) for video synthesis. This catalyzed a rapid transition from earlier U-Net architectures(Blattmann et al., 2023; Guo et al., 2024) to DiT and Flow Matching (Lipman et al., 2023) paradigms. Consequently, a wave of high-fidelity T2V models has emerged, ranging from proprietary systems (KuaishouTechnology, 2026; Runway, 2024) to powerful open-weight ones such as HunyuanVideo (Wu et al., 2025), LTX-Video (HaCohen et al., 2024) and Wan (Wan et al., 2025). Despite achieving cinema-grade visual quality, these “silent” models lack the acoustic dimension essential for immersive world modeling.

2.2. Text-to-Audio-Video Benchmarks

Joint Audio-Video Generation. To bridge the modality gap, research has pivoted towards unified T2AV architectures. Leading proprietary systems, including Sora 2 (Ope-

Prior protocols typically isolate modalities. In the visual domain, the VBench series (Huang et al., 2024; 2025; Zheng 3

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

(a) Distribution of the 235 curated prompts across 3 main domains and 11 sub-categories.

(b) Distribution of audio types, audio source relation, and shot counts.

Figure 4. Dataset-level statistics of the prompts used in AVGen-Bench.

3.1. Task-Driven Prompt Curation

et al., 2025) sets the standard for video quality but inherently neglects the acoustic dimension. Conversely, audio benchmarks like TTA-Bench (Wang et al., 2025b) focus on text-to-audio generation but often face scalability bottlenecks, relying heavily on subjective human evaluation to compensate for the poor perceptual correlation of traditional automated metrics. Recent studies have attempted to combine audio and video evaluation into unified benchmarks. Early works such as HarmonyBench (Hu et al., 2025), UniAVGen (Zhang et al., 2025), and VerseBench (Wang et al., 2025a) assess the capability to generate both modalities together. However, these benchmarks are often too general (coarse-grained). They typically score overall audio-visual quality but fail to distinguish specific errors, such as incorrect pitch or rhythm.

To ensure our benchmark reflects realistic usage rather than merely categorizing static visual concepts, we adopt a topdown, intent-first curation strategy. We first defined a comprehensive taxonomy of user scenarios for AI video generation, and then implemented a “Human-in-the-Loop” generation pipeline. Specifically, we utilized GPT-5.2 (OpenAI, 2026) to generate candidate prompts based on our scenario definitions, followed by a rigorous manual review process to filter for complexity, clarity, and diversity. As illustrated in Figure 4a, the resulting dataset consists of 235 highly curated tasks, systematically distributed across 3 main domains and 11 real-world sub-categories. Notably, to simulate professional editing workflows, the dataset maintains an average of 1.6 shots per prompt, with 44% of samples involving speech and 88% containing environmental sound effects as demonstrated in Figure 4b.

Similarly, benchmarks like JavisBench(Liu et al., 2025) rely on embedding models such as CLIP(Radford et al., 2021) and CLAP (Wu et al., 2023). While useful for general matching, these “black box” metrics cannot verify fine-grained details like specific musical notes or precise synchronization. Consequently, they often fail to detect hallucinations, highlighting the need for the interpretable evaluation suite we propose. We provide a comparison between our benchmark and existing benchmarks in Table 1.

Crucially, a distinct feature of our framework is that the prompt curation is entirely decoupled from the evaluation metrics. Unlike prior benchmarks that often reverseengineer prompts to fit specific available detectors (e.g., curating speech prompts solely because a TTS metric is available), our prompts are derived strictly from genuine user needs. This design choice ensures that AVGen-Bench is both scalable and customizable—users can easily extend the prompt set to new domains. The resulting prompt set is organized into three task domains:

3. AVGen-Bench This section details the architecture of AVGen-Bench. We begin by outlining our task-driven prompt construction strategy, structured around diverse daily-life categories to probe model capabilities boundaries. Following this, we introduce our evaluation protocol, discussing the rationale behind our hybrid design and specifying the implementation of individual metrics for uni-modal quality, cross-modal alignment, and fine-grained semantic control.

Professional Media Production. This domain assesses the model’s capacity to synthesize cinema-grade content suitable for professional workflows. For Commercial Ads, we curated a dataset of classic Bumper Ads from YouTube and employed Gemini 3 Pro (DeepMind, 2026a) to reversecaption these videos into anonymized textual descriptions, ensuring the prompts describe visual styles without relying on specific brand logos. For Movie Trailers, we in4

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

structed GPT-5.2 to construct multi-shot scripts, requiring the model to maintain visual consistency and narrative continuity across varying camera angles and scene transitions.

a robust proxy for subjective listening tests. Cross-Modal Alignment. Proper timing between video and audio is crucial for creating realistic content. We evaluate this using two specific methods. For general synchronization (e.g., impact sounds), we use Syncformer (Iashin et al., 2024). It calculates the time difference between visual motion and the start of the sound. Additionally, since humans are very sensitive to mismatched speech, we use the standard SyncNet (Chung & Zisserman, 2016) model for Lip Synchronization. This measures the error (in frames) between lip movements and speech, ensuring that characters appear to speak naturally.

Creator Economy. Geared towards the booming sector of user-generated content, this domain covers ASMR, Cooking Tutorials, Gameplays, and Musical Instrument Tutorials. A critical innovation in the Musical Instrument Tutorial category is the injection of fine-grained acoustic constraints. We explicitly included requirements for specific musical scales (e.g., “C Major scale”) or chords in the prompts. This design rigorously tests whether the model can perform precise audio-visual alignment—generating the correct audio frequencies corresponding to the visual finger positions—rather than merely producing generic music.

3.2.2. F INE - GRAINED E VALUATION M ODULES

World Simulator. This domain probes the model’s understanding of fundamental laws governing the physical world, spanning Physics, Chemistry, Sports, and Animals. Notably, for Physics and Chemistry, we employed an “Underspecified Prompting” strategy. In these prompts, we intentionally omit explicit descriptions of the physical outcome. For example, in a prompt describing a Newton’s Cradle experiment, we describe the setup but do not specify how many balls should recoil. This forces the model to rely on its “world knowledge” to simulate the correct physical dynamics, rather than simply following a textual instruction.

General aesthetic metrics often gloss over specific semantic failures. To address this, we introduce a suite of hybrid evaluation pipelines. By chaining specialist models (as feature extractors) with Gemini 3 Flash (as the reasoning engine), we can rigorously audit the model’s adherence to fine-grained constraints. Scene Text Rendering. To evaluate the accuracy and contextual validity of generated text, we implement a “detectaggregate-verify” pipeline. First, we utilize PaddleOCR (Cui et al., 2025a) to extract text content and bounding boxes from each video frame. Addressing temporal redundancy, we apply a spatiotemporal clustering algorithm to aggregate spatially proximal text instances across adjacent frames into consolidated sequences. Finally, these parsed sequences are fed into the MLLM for a dual-objective assessment: (1) verifying strict adherence to any text explicitly specified in the prompt, and (2) evaluating the semantic coherence of incidental text (e.g., scrolling tickers in news broadcasts). This ensures that even unprompted text elements are legible and contextually appropriate, rather than manifesting as gibberish or visual artifacts.

3.2. Evaluation Suite To provide a holistic assessment of generative quality, we construct a comprehensive evaluation suite for AVGenBench that utilizes a hybrid methodology, integrating lightweight specialist models with Multimodal Large Language Models (MLLMs). This architecture allows us to bridge the gap between low-level signal fidelity and highlevel semantic reasoning, covering three critical dimensions: uni-modal aesthetics, cross-modal alignment, and text-tomedia consistency. Furthermore, we introduce a set of targeted evaluation modules specifically designed to probe capabilities where current models empirically struggle, such as scene text rendering and fine-grained audio control.

Facial Consistency. To quantify identity preservation and stability without referencing external character facial features, we implement a reference-free “Detect-Track-Cluster” pipeline augmented by MLLM-derived constraints. We first employ InsightFace (Buffalo-L) (Deng et al., 2019) to extract facial embeddings and bounding boxes frameby-frame. To handle occlusion and temporal discontinuity, we construct “tracklets” using a hybrid heuristic combining IoU overlap and cosine similarity. Subsequently, we apply DBSCAN clustering on these tracklets to discover distinct identities (clusters), filtering for “primary characters” based on temporal occupancy ratios. The final consistency score is a weighted aggregate of two dimensions: (1) Identity Count Accuracy (40%): We compare the number of discovered primary clusters against the ground-truth character count predicted by Gemini based on the prompt, penalizing hallucinations or erasure. (2) Identity Stability (60%):

3.2.1. BASIC E VALUATION M ODULES Uni-modal Quality. We begin by assessing the perceptual quality of the visual and acoustic modalities independently. For the visual domain, we leverage Q-Align (Wu et al., 2024), a state-of-the-art MLLM-based evaluator fine-tuned to correlate closely with human aesthetic judgments. Unlike distribution-based metrics (e.g., FVD (Unterthiner et al., 2019)), Q-Align provides a direct score reflecting visual fidelity and technical quality. For the audio domain, we utilize the aesthetic assessment module from Audiobox (Tjandra et al., 2025) (Audiobox-Aesthetic). This model evaluates acoustic clarity and production quality, serving as 5

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation (a) Scene Text Render (Detect-Aggregate-Verify) FrameWise Text & BBox PaddleOCR & Conf. (Text Extraction & Bounding Box)

Video

MLLM

Spatiotemporal Clustering

Prompt-Specific Text Check

(Sequence Aggregation)

Incidental Text Check

(b) Facial Consistency (Detect-Track-Cluster)

Text Rendering Score

MLLM

("Play a C Major

Constraint Extraction

Chord")

Basic-Pitch Audio Waveform

(AMT)

{'type':'chord', 'value':['C', 'E', 'G']} MIDI Events & Chord Frames

MLLM Symbolic Logic Verification

Text Prompt ("A Newton's cradle...")

Feature Extraction)

(IoU & Consine Sim)

(Identity Discovery)

MLLM Identity Count (vs Prompt) Identity Stability

Facial Consistency Score

(d) Speech Intelligibility & Coherence (ASR-Reasoning) Verbatim Mode

MLLM

Text Prompt

Pitch Accuracy Score

("The anchor says Breaking News")

Audio Waveform

(Low-Level Kinematic Expectations: Check) ['one ball MLLM recoil',...] MLLM High-Level Causal Expectation Extraction Reasoning

DBSCAN Clustering

Video

(e) Physical Plausibility (Dual-Stream Evaluation) VideoPhy2AutoEval

Tracklet Generation

InsightFace

(c) Pitch Accuracy (Symbolic-Neural Verification)

Text Prompt

(Face Detection &

Adaptive Mode Selection

Contextual Mode

WhisperLarge-V3

Transcribed Text: "Breaking News..."

(Transciption)

MLLM Adaptive Verification

Speech Score

(f) Holistic Semantic Alignment (Decompose-and-Verify)

Low-Level and HighLevel Physical Plausibility Score

MLLM

Text Prompt ("Four-shot mockumentary teaser with on-set chatter and abrupt cuts....")

MLLM Constraint Decomposition

Narrative Beats

Audio Events

Visual Attributes Camera Control

(Gemini 3 Flash, Evidence-based Verification)

Video

Holistic Alignment Score

Audio

Figure 5. Detailed workflows of the six Fine-grained Evaluation Modules in AVGen-Bench. The suite employs hybrid strategies combining specialist models (blue nodes) and MLLMs (purple nodes) to evaluate: (a) Scene Text Rendering (OCR + Verification); (b) Facial Consistency (InsightFace + DBSCAN); (c) Pitch Accuracy (Audio-to-MIDI + Theory Check); (d) Speech Intelligibility (ASR + Contextual Logic); (e) Physical Plausibility (Kinematics + Causal Reasoning); and (f) Holistic Semantic Alignment (Constraint Decomposition).

For each primary cluster, we measure the 50th percentile (P50 ) internal cosine similarity of its tracklets to assess the robustness of identity preservation over time.

without specifying content), the system evaluates Semantic Coherence, detecting whether the generated speech aligns with the visual context and narrative intent or degenerates into unintelligible gibberish.

Pitch Accuracy. General audio encoders fail to verify finegrained music theory constraints. We address this via a Symbolic-Neural Verification pipeline. First, we feed the text prompt into Gemini to perform Constraint Extraction & Gating, extracting explicit musical constraints (e.g., “C Major chord”) into a structured JSON format while filtering out abstract prompts (e.g., “jazzy vibe”) to avoid invalid penalization. For applicable prompts, we then employ BasicPitch (Bittner et al., 2022) for Automatic Music Transcription (AMT), converting the audio waveform into symbolic MIDI events and aggregating note onsets within an 80ms window into “chord frames.” Finally, the extracted MIDI events are fed back to Gemini for Symbolic Logic Verification, where the MLLM verifies whether the generated note sequences strictly adhere to the music theory requirements defined in the prompt.

Physical Plausibility. We evaluate physical realism through two decoupled modules targeting different levels of abstraction. For Low-Level Kinematic Plausibility, we employ VideoPhy2-AutoEval (Bansal et al., 2025). This specialist model acts as a “physics engine checker,” scoring the video based on motion smoothness and trajectory realism to detect basic artifacts like jittery motion independent of semantic context. In parallel, for High-Level Causal Reasoning (e.g., “Sodium dropped into water”), we implement a Two-Stage Semantic Verification pipeline using Gemini inspired by PhyT2V (Xue et al., 2025). This involves first extracting a list of Observable Expectations (e.g., “violent bubbling”) from the prompt, followed by Semantic Adjudication, where the MLLM logs observable events in the video to calculate a Semantic Physics Score based purely on the alignment between expected physical outcomes and visual evidence.

Speech Intelligibility & Coherence. Unlike general audio metrics, we aim to verify the semantic content of speech using a cascade ASR-Reasoning pipeline. We utilize FasterWhisper (Radford et al., 2022), which integrates Voice Activity Detection (VAD) to effectively filter non-speech noise and accelerate inference, for robust transcription. We then employ Gemini for semantic auditing, introducing an Adaptive Compliance Mechanism. Specifically, in Verbatim Mode (triggered when the prompt explicitly prescribes dialogue), the pipeline enforces strict lexical matching. Conversely, in Contextual Mode (for prompts implying speech

Holistic Semantic Alignment. While embedding-based metrics capture high-level relevance, they often fail to penalize subtle contradictions. To address this, we implement a “Decompose-and-Verify” pipeline using Gemini as a multimodal auditor. The MLLM first performs Constraint Decomposition, parsing the prompt into checkable constraints across four dimensions: (1) Narrative Beats, (2) Visual Attributes (object counts, colors), (3) Audio Events, and (4) Cinematography. Subsequently, it per-

6

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

forms Evidence-based Scoring by scanning the video to verify each constraint against visual/audio evidence. The final score provides a nuanced assessment of how well the generated content fulfills the user’s intent beyond simple semantic similarity.

compact model-level comparison, we also report a Total score in Table 2: Total = 0.2Sbasic + 0.2Scross + 0.6Sfine , where Sbasic = mean(Vis × 100, Aud(PQ) × 10), Scross = mean(100 · max(0, 1 − AV/0.5), 100 · max(0, 1 − Lip/8)), and Sfine = mean(Text, Face, Music, Speech, Lo-Phy × 20, Hi-Phy, Holistic).

4. Experiment

Basic Uni-modal Quality. As presented in Table 2, the evaluated models demonstrate exceptional performance in the visual domain. The consistently high Visual Quality scores (e.g., Seedance-1.5 Pro reaching 0.970 and Veo 3.1 reaching 0.960) indicate that current T2AV systems have largely mastered the synthesis of high-fidelity imagery. Qualitative inspection confirms that this metric aligns strongly with subjective perception: models with top-tier scores consistently produce videos with professional lighting, composition, and ”cinematic” aesthetics.

4.1. Experimental Setup Models Evaluated. To ensure a comprehensive assessment of the current T2AV landscape, we select a diverse set of state-of-the-art models spanning both commercial services and research frameworks. For proprietary systems, we evaluate market-leading models accessed via their official APIs, including Sora 2 (OpenAI, 2025), Kling 2.6 (KuaishouTechnology, 2026), Wan 2.6 (AI, 2026), and Seedance-1.5 Pro (Seedance et al., 2025). Additionally, we include Google’s Veo 3.1 (DeepMind, 2026b), testing both its Fast and Quality variants to analyze the trade-off between inference speed and generation fidelity. In the open-source domain, we evaluate representative unified models, specifically LTX-2.3, LTX-2 (HaCohen et al., 2026), and Ovi (Low et al., 2025). Furthermore, to benchmark modular cascaded approaches, we include a standard T2V+V2A pipeline combining Wan 2.2 (Wan et al., 2025) with HunyuanVideoFoley (Shan et al., 2025). We also incorporate Text-toImage-to-Audio-Video (T2Image+TI2AV) pipelines by pairing both the open-source Emu3.5 (Cui et al., 2025b) and the proprietary NanoBanana2 (Raisinghani, 2026) with the open-source MOVA (SII-OpenMOSS Team et al., 2026) model.

In contrast, Audio Quality scores—specifically measured by the Production Quality (PQ) sub-metric of AudioboxAesthetic—are relatively lower, suggesting that acoustic synthesis still trails behind visual generation. We observe a clear correlation between PQ and auditory clarity: highscoring models (e.g., Seedance-1.5 Pro at 7.48) generate crisp, studio-like sound, whereas lower scores typically correspond to audible background noise or signal artifacts. Basic Cross-modal Alignment. Regarding temporal synchronization, results indicate that current models have not yet achieved frame-perfect alignment. For general AV Sync, the mean absolute offset ranges from 0.2s to 0.44s, while Lip Sync errors span from 2.0 to over 5 frames. These figures reveal a tangible gap from ideal performance, particularly in speech scenarios where even minor offsets (e.g., > 2 frames) can disrupt the perceptual illusion of a talking head.

Implementation Details. To maintain a fair comparison, we standardize the output resolution for the majority of models to 720p (1280×720), with the exception of pipelines utilizing MOVA, which are evaluated using its 360p version. Regarding temporal duration, we target a length of 10 seconds for most models (e.g., Kling 2.6, Wan 2.6, LTX-2.3, LTX-2, Ovi). Exceptions are dictated by specific architectural or API constraints: Veo 3.1 is evaluated at 8 seconds (its maximum supported duration), Sora 2 at 12 seconds due to fixed duration quantization, and the Wan 2.2 pipeline at 5 seconds (16 fps) reflecting the T2V model’s native generation limit. For open-source models, inference is performed using official checkpoints with default sampling parameters recommended by their respective authors.

Fine-grained Visual: Text Rendering Quality. As indicated in Table 2, text rendering remains a significant bottleneck. Our analysis reveals a distinct performance dichotomy governed by text prominence and explicitness. Models generally succeed in rendering explicitly prompted text when the target string is short and occupies a dominant spatial region (e.g., a large movie title). However, performance degrades rapidly as text length increases or spatial resolution decreases, frequently resulting in ”glyph collapse” or unintelligible gibberish. More critically, regarding incidental text—contextual writing not explicitly defined in the prompt (e.g., small print on a clapperboard)—we observe a universal failure mode across all evaluated models. Instead of generating coherent context-appropriate characters, models consistently hallucinate messy, graffiti-like scribbles.

4.2. Experimental Results We provide detailed analysis of each evaluation module. To complement the quantitative metrics, we visually analyze representative failure cases across different finegrained dimensions in Figure 2. Additional examples and extended analysis are provided in Appendix A. For

Fine-grained Visual: Facial Consistency. Maintaining character identity across time remains a persistent challenge for all T2AV models. As shown in Table 2, even the top-

7

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation Table 2. Quantitative comparison on AVGen-Bench. We evaluate models across three granularities: Basic Uni-modal (Visual/Audio Aesthetic), Basic Cross-modal (Sync), and our proposed Fine-grained Modules. Best scores are highlighted in bold, and second-best are underlined. Note that for AV-Sync and Lip-Sync, lower (↓) is better; for others, higher (↑) is better. We also report an aggregate Total score (Scheme-2). Wan2.2+HunyuanVideo-Foley denotes a cascaded pipeline of T2V followed by V2A. Emu3.5+MOVA and NanoBanana2+MOVA are both T2Image+TI2AV cascaded pipelines. Proprietary components are marked with orange background, while open-source components are marked with blue background. Models are sorted by Overall score in descending order. Basic Uni-modal

Basic Cross-modal

Fine-grained Visual

Fine-grained Audio

Vis ↑

Aud (PQ) ↑

AV ↓

Lip ↓

Text ↑

Face ↑

Music ↑

Speech ↑

Lo-Phy ↑

Hi-Phy ↑

Holistic ↑

Veo 3.1-fast

0.960

6.64

0.21

2.39

75.10

52.77

3.13

94.53

3.68

67.43

86.27

67.87

Veo 3.1-quality

0.954

6.77

0.24

3.59

76.53

52.90

5.00

96.09

3.74

68.53

84.10

66.28

Model

Fine-grained Macro

Overall Total ↑

Sora-2

0.848

5.91

0.25

4.50

74.84

51.17

7.81

88.63

4.05

78.95

88.89

64.16

Wan2.6

0.959

7.15

0.30

4.32

76.95

49.27

1.75

89.33

3.69

66.92

80.98

62.97

Seedance-1.5 Pro

0.970

7.48

0.26

3.43

38.28

54.42

1.88

93.45

3.72

66.88

77.38

62.55

Kling-V2.6

0.906

6.93

0.21

2.30

14.52

57.33

5.00

89.62

3.84

63.92

76.74

61.82

LTX-2.3

0.858

7.11

0.36

2.00

54.17

45.06

1.38

86.66

3.99

64.31

65.22

59.97

NanoBanana2 + MOVA

0.890

6.71

0.44

2.70

68.26

41.33

0.59

82.45

3.91

60.95

72.48

58.10

LTX-2

0.828

6.84

0.23

4.76

24.76

48.53

5.75

87.07

4.05

60.20

66.59

56.62

Emu3.5 + MOVA

0.911

6.80

0.38

4.83

64.72

48.44

0.62

81.74

3.89

55.85

66.55

56.12

Wan2.2 + HunyuanVideo-Foley

0.936

6.60

0.23

5.38

48.46

36.23

3.44

53.40

3.90

54.11

60.63

53.29

Ovi

0.839

6.31

0.37

5.40

41.36

49.05

11.25

76.49

3.93

52.92

57.45

52.02

performing model (Kling-V2.6) only achieves a consistency score of 57.33, while others hover around 48-54. We identify two primary degradation patterns: (1) Temporal Identity Drift: Identity features are highly unstable during discontinuities. When a character reappears after a shot transition, or undergoes large pose changes (e.g., turning their head), models often fail to recall the original face embeddings, effectively generating a new person. (2) Crowd Degradation: We observe a distinct ”inverse scaling” law regarding the number of faces. In multi-face scenarios (e.g., a cheering crowd), the rendering quality and stability of individual faces collapse significantly compared to single-portrait shots, resulting in distorted features and severe flickering.

When prompts imply speech without dictating a script (Incidental Mode), open-source models like Ovi (76.49) and LTX-2 frequently generate unintelligible gibberish or “alien languages.” (2) Partial Instruction Dropping: In Verbatim Mode, even capable models often omit specific words or truncate sentences when long or complex dialogue is explicitly required. Physical Plausibility. The evaluation results highlight significant deficits in how models model the physical world. First, in Low-Level Kinematic Plausibility, most models fail to surpass the passing threshold (a score of 4.0 in VideoPhy2). This indicates that the underlying physics of generated videos are often flawed, frequently exhibiting unnatural motion or object instability. Second, regarding High-Level Causal Reasoning, models demonstrate a lack of precise “world knowledge,” leading to incorrect physical phenomena. For instance, in the prompt describing “sodium dropped into water,” almost all models fail to correctly simulate the sodium floating on the water surface (due to density differences); instead, they often depict it sinking or simply changing color without the correct physical dynamics.

Fine-grained Audio: Pitch Accuracy. A critical finding in our benchmark is that current T2AV models completely fail to understand musical notes. As shown in Table 2, all models achieve extremely low scores (< 12/100), indicating a lack of basic music theory knowledge. While models can correctly generate the timbre of an instrument, they cannot follow instructions regarding specific notes or pitch. When prompted to play a specific scale (e.g., “C Major”) or chord sequence, models simply generate random notes that have no connection to the prompt.

Holistic Semantic Alignment. Finally, when evaluating overall alignment, we observe that models frequently ignore specific visual and audio controls as the prompt becomes more complex. This issue is particularly severe in opensource models, which often fail to capture multiple constraints simultaneously. While proprietary models demonstrate a significant advantage (likely due to richer training data), they still struggle with complex audio layering. For instance, when a prompt requires multiple overlapping sounds—such as background music, footsteps, and speech occurring at the same time—even top-tier models tend to

Fine-grained Audio: Speech Intelligibility & Coherence. As reported in Table 2, Google’s Veo 3.1 series demonstrates dominant performance in speech generation, with the Quality variant achieving a remarkable score of 96.09 and Fast at 94.53. This suggests that Veo has largely bridged the gap between video generation and TTS, maintaining high clarity even in complex scenes. However, significant limitations persist in other systems. We identify two primary failure modes: (1) Hallucination in Contextual Speech: 8

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation Table 3. Human validation of fine-grained evaluation metrics. We report the Pearson correlation between our automated scores and expert human judgments across six fine-grained dimensions. All results are computed on a shared subset of 85 tasks annotated by 10 expert raters. Higher is better. Dimension

Protocol

Pearson ↑

Text Rendering Pitch Accuracy Facial Consistency Speech Intelligibility & Coherence Physical Plausibility Holistic Semantic Alignment

Pointwise Pointwise Pairwise Pairwise Pairwise Pairwise

0.9657 0.5544 0.8270 0.8300 0.8290 0.8402

Table 5. Repeated-run stability of the MLLM-assisted evaluation. We repeat the full evaluation pipeline 3 times on the same generated outputs for two representative models, Veo 3.1 Fast and LTX-2. We report the mean, standard deviation, and value range of each fine-grained metric across runs. Lower standard deviation indicates better stability.

Model

Metric

Mean

Std. ↓

Range

Veo 3.1 Fast

Text Face Music Speech Lo-Phy Hi-Phy Holistic

74.75 52.97 2.76 94.48 74.79 70.73 85.47

0.83 0.00 0.08 0.02 0.00 1.70 0.28

73.95–75.90 52.97–52.97 2.65–2.81 94.46–94.51 74.79–74.79 69.16–73.09 85.06–85.68

LTX-2

Text Face Music Speech Lo-Phy Hi-Phy Holistic

26.91 45.54 7.57 86.96 81.45 63.11 66.64

0.59 0.00 1.24 0.12 0.00 0.54 0.60

26.29–27.71 45.54–45.54 5.88–8.82 86.82–87.11 81.45–81.45 62.51–63.82 65.84–67.30

Table 4. Inter-rater agreement on the shared user-study subset. We report inter-rater reliability across 10 expert raters on the same 85 tasks. For pointwise dimensions, we use weighted Cohen’s κ; for pairwise dimensions, we report Cohen’s κ. Higher is better. Dimension

Agreement Metric

Score ↑

Text Rendering Pitch Accuracy Facial Consistency Speech Intelligibility & Coherence Physical Plausibility Holistic Semantic Alignment

Weighted Cohen’s κ Weighted Cohen’s κ Cohen’s κ Cohen’s κ Cohen’s κ Cohen’s κ

0.9116 0.3156 0.8511 0.9272 0.8455 0.8909

“drop” some audio elements, failing to generate a complete acoustic scene.

0.9657 for Text Rendering, 0.8270 for Facial Consistency, 0.8300 for Speech Intelligibility & Coherence, 0.8290 for Physical Plausibility, and 0.8402 for Holistic Semantic Alignment. These results indicate that our specialist-model + MLLM evaluation pipeline is well aligned with expert perception on a broad range of fine-grained T2AV capabilities.

4.3. User Study To validate the reliability of our fine-grained evaluation framework, we conducted a larger-scale human study covering all six fine-grained dimensions in AVGen-Bench: Text Rendering, Pitch Accuracy, Facial Consistency, Speech Intelligibility & Coherence, Physical Plausibility, and Holistic Semantic Alignment. We recruited 10 expert raters and asked them to annotate a shared subset of 85 tasks. This subset was used both for evaluating the correlation between our automated metrics and human judgments, and for measuring inter-rater agreement.

The only relatively weaker dimension is Pitch Accuracy, where the Pearson correlation is 0.5544. We attribute this mainly to a floor effect: current T2AV systems perform extremely poorly on explicit pitch control, causing human ratings to cluster within a narrow low-score range and making correlation estimates less stable. In other words, this lower correlation reflects the immaturity of current models on pitch-controllable generation, rather than the absence of meaningful signal in the evaluation itself.

Following the nature of each evaluation target, we adopted two annotation protocols. For dimensions that require absolute judgment of a single output—namely Text Rendering and Pitch Accuracy—we used pointwise scoring. For dimensions that are more naturally assessed in relative terms— Facial Consistency, Speech Intelligibility & Coherence, Physical Plausibility, and Holistic Semantic Alignment— we used pairwise comparison. This hybrid design mirrors the structure of our automatic evaluation pipeline and allows us to assess both metric validity and annotation consistency under realistic conditions.

To further assess the reliability of the annotations, we also compute inter-rater agreement on the same shared subset, as shown in Table 4. Agreement is high on most dimensions, with weighted Cohen’s κ of 0.9116 for Text Rendering and Cohen’s κ of 0.8511, 0.9272, 0.8455, and 0.8909 for Facial Consistency, Speech Intelligibility & Coherence, Physical Plausibility, and Holistic Semantic Alignment, respectively. Pitch Accuracy again shows lower agreement (0.3156), consistent with the same floor-effect phenomenon.

The human–metric correlation results are summarized in Table 3. Overall, our automated metrics show strong agreement with expert judgment on five out of six fine-grained dimensions. In particular, the Pearson correlation reaches

Overall, these results provide strong evidence that our finegrained evaluation is both human-aligned and annotationstable on the dimensions most relevant to realistic T2AV generation. 9

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

level comparison under prompt subsampling, and that the full benchmark scale is sufficient for statistically meaningful evaluation. Taken together, these results show that AVGen-Bench is not only human-aligned, but also stable with respect to repeated evaluation and prompt subsampling, supporting its use as a reliable benchmark for T2AV generation.

5. Conclusion

Figure 6. Benchmark-scale robustness under prompt subset resampling. We repeatedly sample prompt subsets at different ratios and recompute the overall normalized score 200 times. The solid lines denote the mean score over random subsets, the error bars indicate one standard deviation, and the dashed lines mark the corresponding full-benchmark score. Results for both Veo 3.1 Fast and LTX-2 remain close to the full-score baseline, with smaller variance at larger subset ratios, indicating that AVGenBench yields stable model comparison under prompt subsampling.

In this paper, we introduced AVGen-Bench, a task-driven framework for T2AV evaluation. Our results reveal a sharp dichotomy: while state-of-the-art models excel at general audio-visual aesthetics, creating cinematic content, they fail significantly at fine-grained semantic control. This is evidenced by low scores in tasks requiring precise pitch, text rendering, and physical logic. These findings suggest that current training paradigms based on coarse alignment are insufficient. Future research must prioritize finer-grained supervision to transition from probabilistic texture generators to physically grounded world models.

4.4. Stability of the Evaluation

References

Beyond human alignment, we further assess the stability of our MLLM-assisted evaluation from two complementary perspectives: run-to-run consistency and benchmarkscale robustness.

AI, W. Wan 2.6: Ai video generation model, 2026. URL https://www.wan-ai.co/wan-2-6. Accessed: 2026-01-22. Bansal, H., Peng, C., Bitton, Y., Goldenberg, R., Grover, A., and Chang, K.-W. Videophy-2: A challenging actioncentric physical commonsense evaluation in video generation, 2025. URL https://arxiv.org/abs/2503. 06800.

First, we measure run-to-run consistency by repeating the full evaluation pipeline 3 times on the same generated outputs. Table 5 reports the mean, standard deviation, and value range of each fine-grained metric for two representative models, Veo 3.1 Fast and LTX-2. The observed fluctuations are generally small across runs. For example, on Veo 3.1 Fast, the standard deviation is only 0.83 for Text, 0.08 for Music, 0.02 for Speech, and 0.28 for Holistic evaluation. LTX-2 shows similarly stable behavior across most dimensions. These results indicate that, although our framework includes an MLLM-based reasoning component, the resulting scores are stable in practice under repeated evaluation.

Bittner, R. M., Bosch, J. J., Rubinstein, D., MeseguerBrocal, G., and Ewert, S. A lightweight instrumentagnostic model for polyphonic note transcription and multipitch estimation. In Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), Singapore, 2022. Blattmann, A., Rombach, R., Ling, H., Dockhorn, T., Kim, S. W., Fidler, S., and Kreis, K. Align your latents: Highresolution video synthesis with latent diffusion models. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 22563–22575, 2023. doi: 10.1109/CVPR52729.2023.02161.

Second, we test whether the benchmark scale is sufficient for stable model comparison. Specifically, we repeatedly sample random prompt subsets at different ratios (20%, 40%, 60%, and 80%) and recompute the overall normalized score. For each ratio, we repeat the sampling procedure 200 times. Figure 6 shows the mean subsampled score together with one standard deviation for two representative models, Veo 3.1 Fast and LTX-2, and compares them against the corresponding full-benchmark score. In both cases, the subset-based estimates remain close to the full score, while the variance decreases steadily as the subset ratio increases. This indicates that AVGen-Bench provides stable model-

Cheng, H. K., Ishii, M., Hayakawa, A., Shibuya, T., Schwing, A., and Mitsufuji, Y. MMAudio: Taming multimodal joint training for high-quality video-to-audio synthesis. In CVPR, 2025. Chung, J. S. and Zisserman, A. Out of time: automated lip sync in the wild. In Workshop on Multi-view Lip-reading, ACCV, 2016. 10

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

Cui, C., Sun, T., Lin, M., Gao, T., Zhang, Y., Liu, J., Wang, X., Zhang, Z., Zhou, C., Liu, H., Zhang, Y., Lv, W., Huang, K., Zhang, Y., Zhang, J., Zhang, J., Liu, Y., Yu, D., and Ma, Y. Paddleocr 3.0 technical report, 2025a. URL https://arxiv.org/abs/2507.05595.

Hu, T., Yu, Z., Zhang, G., Su, Z., Zhou, Z., Zhang, Y., Zhou, Y., Lu, Q., and Yi, R. Harmony: Harmonizing audio and video generation through cross-task synergy, 2025. URL https://arxiv.org/abs/2511.21579.

Cui, Y., Chen, H., Deng, H., Huang, X., Li, X., Liu, J., Liu, Y., Luo, Z., Wang, J., Wang, W., Wang, Y., Wang, C., Zhang, F., Zhao, Y., Pan, T., Li, X., Hao, Z., Ma, W., Chen, Z., Ao, Y., Huang, T., Wang, Z., and Wang, X. Emu3.5: Native multimodal models are world learners, 2025b. URL https://arxiv.org/abs/2510. 26583. DeepMind, G. Gemini 3.0 pro, 2026a. URL https:// deepmind.google/models/gemini/pro/. Accessed: 2026-01-23. DeepMind, G. Veo 3.1: Video, meet audio, 2026b. URL https://deepmind.google/technologies/ veo/. Accessed: 2026-01-22.

Huang, Z., He, Y., Yu, J., Zhang, F., Si, C., Jiang, Y., Zhang, Y., Wu, T., Jin, Q., Chanpaisit, N., Wang, Y., Chen, X., Wang, L., Lin, D., Qiao, Y., and Liu, Z. VBench: Comprehensive benchmark suite for video generative models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024. Huang, Z., Zhang, F., Xu, X., He, Y., Yu, J., Dong, Z., Ma, Q., Chanpaisit, N., Si, C., Jiang, Y., Wang, Y., Chen, X., Chen, Y.-C., Wang, L., Lin, D., Qiao, Y., and Liu, Z. VBench++: Comprehensive and versatile benchmark suite for video generative models. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025. doi: 10.1109/TPAMI.2025.3633890. Iashin, V., Xie, W., Rahtu, E., and Zisserman, A. Synchformer: Efficient synchronization from sparse cues. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2024.

Deng, J., Guo, J., Niannan, X., and Zafeiriou, S. Arcface: Additive angular margin loss for deep face recognition. In CVPR, 2019.

KuaishouTechnology. Kling ai global, 2026. URL https: //klingai.com/global/. Accessed: 2026-01-22.

Gao, X., Hu, L., Hu, S., Huang, M., Ji, C., Meng, D., Qi, J., Qiao, P., Shen, Z., Song, Y., Sun, K., Tian, L., Wang, G., Wang, Q., Wang, Z., Xiao, J., Xu, S., Zhang, B., Zhang, P., Zhang, X., Zhang, Z., Zhou, J., and Zhuo, L. Wan-s2v: Audio-driven cinematic video generation, 2025. URL https://arxiv.org/abs/2508.18621.

Lipman, Y., Chen, R. T. Q., Ben-Hamu, H., Nickel, M., and Le, M. Flow matching for generative modeling. In International Conference on Learning Representations (ICLR), 2023. URL https://openreview.net/ forum?id=PqvMRDCJT9t.

Guo, Y., Yang, C., Rao, A., Liang, Z., Wang, Y., Qiao, Y., Agrawala, M., Lin, D., and Dai, B. Animatediff: Animate your personalized text-to-image diffusion models without specific tuning. International Conference on Learning Representations, 2024.

Liu, K., Li, W., Chen, L., Wu, S., Zheng, Y., Ji, J., Zhou, F., Jiang, R., Luo, J., Fei, H., and Chua, T.-S. Javisdit: Joint audio-video diffusion transformer with hierarchical spatio-temporal prior synchronization. In arxiv, 2025. Low, C., Wang, W., and Katyal, C. Ovi: Twin backbone cross-modal fusion for audio-video generation, 2025. URL https://arxiv.org/abs/2510.01284.

HaCohen, Y., Chiprut, N., Brazowski, B., Shalem, D., Moshe, D., Richardson, E., Levin, E., Shiran, G., Zabari, N., Gordon, O., Panet, P., Weissbuch, S., Kulikov, V., Bitterman, Y., Melumian, Z., and Bibi, O. Ltx-video: Realtime video latent diffusion, 2024. URL https: //arxiv.org/abs/2501.00103.

OpenAI. Video generation models as world simulators, 2024. URL https://openai.com/research/ video-generation-models-as-world-simulators. Accessed: 2024-02-15.

HaCohen, Y., Brazowski, B., Chiprut, N., Bitterman, Y., Kvochko, A., Berkowitz, A., Shalem, D., Lifschitz, D., Moshe, D., Porat, E., Richardson, E., Shiran, G., Chachy, I., Chetboun, J., Finkelson, M., Kupchick, M., Zabari, N., Guetta, N., Kotler, N., Bibi, O., Gordon, O., Panet, P., Benita, R., Armon, S., Kulikov, V., Inger, Y., Shiftan, Y., Melumian, Z., and Farbman, Z. Ltx-2: Efficient joint audio-visual foundation model, 2026. URL https: //arxiv.org/abs/2601.03233.

OpenAI. Sora 2 System Card, 2025. URL https://cdn.openai.com/pdf/ 50d5973c-c4ff-4c2d-986f-c72b5d0ff069/ sora_2_system_card.pdf. Accessed: 2026-0122. OpenAI. Introducing gpt-5.2, 2026. URL https:// openai.com/index/introducing-gpt-5-2/. Accessed: 2026-01-23. 11

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

Pang, Y., Liu, J., Tan, L., Zhang, Y., Gao, F., Deng, X., Kang, Z., Wei, X., and Liu, Y. Mavid: A multimodal framework for audio-visual dialogue understanding and generation, 2025. URL https://arxiv.org/abs/ 2512.03034.

Tjandra, A., Wu, Y.-C., Guo, B., Hoffman, J., Ellis, B., Vyas, A., Shi, B., Chen, S., Le, M., Zacharov, N., Wood, C., Lee, A., and Hsu, W.-N. Meta audiobox aesthetics: Unified automatic quality assessment for speech, music, and sound. 2025. URL https://arxiv.org/abs/ 2502.05139.

Peebles, W. and Xie, S. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 4195–4205, October 2023.

Unterthiner, T., van Steenkiste, S., Kurach, K., Marinier, R., Michalski, M., and Gelly, S. Towards accurate generative models of video: A new metric & challenges, 2019. URL https://arxiv.org/abs/1812.01717.

Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning (ICML), pp. 8748–8763. PMLR, 2021.

Wan, T., Wang, A., Ai, B., Wen, B., Mao, C., Xie, C.-W., Chen, D., Yu, F., Zhao, H., Yang, J., et al. Wan: Open and advanced large-scale video generative models, 2025. URL https://arxiv.org/abs/2503.20314.

Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I. Robust speech recognition via largescale weak supervision, 2022. URL https://arxiv. org/abs/2212.04356.

Wang, D., Zuo, W., Li, A., Chen, L.-H., Liao, X., Zhou, D., Yin, Z., Dai, X., Jiang, D., and Yu, G. Universe-1: Unified audio-video generation via stitching of experts. arXiv preprint arXiv:2509.06155, 2025a.

Raisinghani, N. Nano banana 2: Combining pro capabilities with lightning-fast speed. https: //blog.google/innovation-and-ai/ technology/ai/nano-banana-2/, February 2026. Google Blog. Accessed: 2026-03-16.

Wang, H., Liu, C., Chen, J., Liu, H., Jia, Y., Zhao, S., Zhou, J., Sun, H., Bu, H., and Qin, Y. Tta-bench: A comprehensive benchmark for evaluating text-to-audio models. arXiv preprint arXiv:2509.02398, 2025b.

Runway. Introducing gen-3 alpha, 2024. URL https://runwayml.com/research/ introducing-gen-3-alpha. Accessed: 2026-0123.

Wang, J., Zeng, X., Qiang, C., Chen, R., Wang, S., Wang, L., Zhou, W., Cai, P., Zhao, J., Li, N., et al. Kling-foley: Multimodal diffusion transformer for high-quality videoto-audio generation. arXiv preprint arXiv:2506.19774, 2025c.

Seedance, T., Chen, H., Chen, S., Chen, X., Chen, Y., Chen, Y., Chen, Z., Cheng, F., Cheng, T., Cheng, X., Chi, X., et al. Seedance 1.5 pro: A native audiovisual joint generation foundation model, 2025. URL https://arxiv.org/abs/2512.13507.

Weng, S., Zheng, H., Chang, Z., Li, S., Shi, B., and Wang, X. Audio-sync video generation with multi-stream temporal control. NeurIPS, 2025. Wu, B., Zou, C., Li, C., Huang, D., Yang, F., Tan, H., Peng, J., Wu, J., Xiong, J., Jiang, J., et al. Hunyuanvideo 1.5 technical report, 2025. URL https://arxiv.org/ abs/2511.18870.

Shan, S., Li, Q., Cui, Y., Yang, M., Wang, Y., Yang, Q., Zhou, J., and Zhong, Z. Hunyuanvideo-foley: Multimodal diffusion with representation alignment for highfidelity foley audio generation, 2025. URL https: //arxiv.org/abs/2508.16930. SII-OpenMOSS Team, Yu, D., Chen, M., Chen, Q., Luo, Q., Wu, Q., Cheng, Q., Li, R., Liang, T., Zhang, W., Tu, W., Peng, X., Gao, Y., Huo, Y., Zhu, Y., Luo, Y., Zhang, Y., Song, Y., Xu, Z., Zhang, Z., Yang, C., Chang, C., Zhou, C., Chen, H., Ma, H., Li, J., Tong, J., Liu, J., Chen, K., Li, S., Wang, S., Jiang, W., Fei, Z., Ning, Z., Li, C., Li, C., He, Z., Huang, Z., Chen, X., and Qiu, X. Mova: Towards scalable and synchronized video-audio generation, February 2026. URL https://arxiv.org/abs/2602. 08794. Technical report. Corresponding authors: Xie Chen and Xipeng Qiu. Project leaders: Qinyuan Cheng and Tianyi Liang. 12

Wu, H., Zhang, Z., Zhang, W., Chen, C., Liao, L., Li, C., Gao, Y., Wang, A., Zhang, E., Sun, W., Yan, Q., Min, X., Zhai, G., and Lin, W. Q-align: Teaching LMMs for visual scoring via discrete text-defined levels. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pp. 54015–54029. PMLR, 21–27 Jul 2024. Wu, Y., Chen, K., Zhang, T., Hui, Y., Berg-Kirkpatrick, T., and Dubnov, S. Large-scale contrastive languageaudio pretraining with feature fusion and keyword-tocaption augmentation. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 1–5. IEEE, 2023.

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

Xue, Q., Yin, X., Yang, B., and Gao, W. Phyt2v: Llmguided iterative self-refinement for physics-grounded textto-video generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 18826–18836, June 2025. Zhang, G., Zhou, Z., Hu, T., Peng, Z., Zhang, Y., Chen, Y., Zhou, Y., Lu, Q., and Wang, L. Uniavgen: Unified audio and video generation with asymmetric cross-modal interactions, 2025. URL https://arxiv.org/abs/ 2511.03334. Zheng, D., Huang, Z., Liu, H., Zou, K., He, Y., Zhang, F., Gu, L., Zhang, Y., He, J., Zheng, W.-S., et al. Vbench-2.0: Advancing video generation benchmark suite for intrinsic faithfulness. arXiv preprint arXiv:2503.21755, 2025.

13

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

A. Additional Qualitative Results In this section, we provide extended qualitative examples to further illustrate the failure modes discussed in the main paper. We categorize these failures into three groups: (1) Text Rendering Failures (Figure 7), (2) Consistency and Speech Failures (Figure 8), and (3) Physical and Semantic Logic Failures (Figure 9).

Figure 7. Extended Examples of Text Rendering Failures. Top (Prompted Text): Models struggle with ”glyph collapse” and layout errors when prompted with specific strings like ”Your customers are talking” or ”EIGHTY-SEVEN SECONDS”. Even high-performing models like Veo 3.1 and Wan 2.6 often fail to render the text perfectly legible or place it on the correct object. Bottom (Incidental Text): A pervasive failure mode where models hallucinate gibberish for background text that was not explicitly prompted, such as website content, car license plates, or studio backdrops. This highlights a lack of ”world knowledge” regarding how text naturally appears in real-world scenes.

B. Human Evaluation Protocols and Interfaces To ensure the reproducibility and rigorousness of our meta-evaluation, we developed a unified annotation platform using Gradio. We employed a hybrid annotation strategy, selecting the most appropriate protocol (Pairwise vs. Pointwise) based on the nature of the specific evaluation dimension. 14

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

B.1. Hybrid Annotation Strategy 1. Pairwise Comparison for Subjective Quality (Speech & Semantic). For dimensions where quality is often relative or nuanced—such as Speech Quality and Holistic Semantic Alignment—we utilized a Blind A/B Testing protocol (Figure 11a). • Rationale: Determining ”which voice sounds more natural” is cognitively easier and more consistent via side-by-side comparison than absolute scoring. • Mechanism: Annotators are presented with two anonymized videos (randomized Left/Right order) and the strict prompt constraints. They must vote for the superior model or select ”Tie”. Notably, the interface explicitly displays required speech lines to force verification of verbatim adherence. 2. Pointwise Scoring for Objective Correctness (Text Rendering). Conversely, text rendering requires an absolute assessment of legibility and spelling correctness. A pairwise comparison might result in a ”Tie” if both models produce gibberish, failing to capture the absolute failure. Therefore, we adopted a Pointwise Protocol (Figure 11b). • Rationale: Text quality is objective (e.g., a typo is a typo). Absolute scoring allows us to quantify the exact success rate of each model. • Rubric: We used a 3-point scale: Good (Fully legible and correct), OK (Minor artifacts but legible), and Poor (Illegible/Hallucinated/Missing).

15

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

Figure 8. Extended Examples of Face Inconsistency and Speech Generation Errors. Top (Face Inconsistency): We observe two distinct patterns of identity loss: (1) Identity Drift across shot transitions, where a character’s appearance changes significantly after a cut; and (2) Crowd Degradation, where faces in multi-person scenes (e.g., boxing audience) become distorted. Bottom (Speech Generation): Models frequently fail to adhere to linguistic or speaker constraints. Failures include generating the wrong language (e.g., Spanish instead of English), producing rhythmic noise instead of dialogue, or assigning dialogue to the wrong speaker count.

16

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

Figure 9. Extended Examples of Physical Violations and Semantic Misalignment. Top (Violation of Physical Laws): Models fail to simulate complex physical phenomena driven by sound. Left (Chladni Plate): Models fail to generate the correct geometric sand patterns corresponding to resonant frequencies. Right (Chemical Reaction): Models fail to depict the correct color oscillations or liquid dynamics in a Briggs-Rauscher reaction setup. Bottom (Semantic Misalignment): In complex multi-shot narratives (e.g., a vacation ad), models often miss key semantic constraints, such as specific actions (”hitting a beach ball”) or correct text sequencing (”Book your family home now”).

17

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

Figure 10. Deep Dive into Pitch Accuracy Failures via Symbolic-Neural Verification. We illustrate the disconnect between visual realism and acoustic logic in music generation. Top (Piano): The prompt strictly requests a ”C-G-Am-F” chord progression. While models generate convincing visuals of hands on keys, the extracted MIDI data reveals that the audio contains wrong chords, random melodic noise, or chaotic note clusters, failing to follow basic music theory constraints. Bottom (Guitar): The prompt requests a specific single note (A4) plucked four times. Models fail to isolate the pitch, instead generating complex, unprompted chords or multi-string noise. This confirms that current T2AV models function as ”texture generators” rather than grounded simulators of physical acoustic events.

18

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

(a) Pairwise Interface (Speech/Semantic): Used for relative quality assessment. Features blind A/B testing with explicit constraint display.

(b) Pointwise Interface (Text Rendering): Used for absolute quality assessment. Features a 3-point rubric (Good/OK/Poor) to judge objective legibility. Figure 11. Overview of the Custom Gradio Annotation Suite. We tailored the annotation interface to the specific nature of the task. (a) For subjective dimensions, we enforce strict side-by-side comparison to reduce inter-rater variance. (b) For objective dimensions like text, we use absolute scoring to capture specific failure modes.

19

Record · ID 2632 · SHA-256 27f5de7c1c3e0d9a
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.