Accepted for publication at IEEE VIS 2026 Workshop on GenAI, Agents, and the Future of VIS.
Semantic Action Graph: A Shared Representation for Agent Grounding and Human Interpretation of Sports Highlights Tica Lin
*, Deepak Chandran
, Gauri Jagatap , Chen Chen David Gunawan , and Josh Kimball
†, Andrea Fanelli
,
arXiv:2609.20768v1 [cs.HC] 17 Sep 2026
Dolby Laboratories
Figure 1: We design the semantic action graph as a shared representation for sports highlights, expressing a match as performer, action, recipient, moment, and state nodes joined by semantic edges. SportSAGE instantiates it with a highlight agent (bottom) to compute statistics, compose narratives, and return clips as node sequences with frame boundaries read from the graph. The same structure drives the interface (top) for filtering (a) and navigating (b) the graph. Viewers can specify preferred narrative focus (c) and view the resulting personalized highlights (d).
A BSTRACT
The schema demonstrates three key properties: 1) connected event sequences, 2) a shared, closed vocabulary, and 3) frame-addressable moments, making it suitable to serve two consumers at once: an agentic pipeline that composes narrated highlights, and a visual interface through which viewers query and inspect the same structure. We instantiate it in SportSAGE, a design probe pairing a four-module highlight pipeline with a graph interface, and report feedback from 12 soccer fans. Participants were satisfied with the quality of the generated highlights and narratives, and used the graph interface to search, navigate, and interpret the match highlights. These results provide early evidence that one small, human-readable schema can ground agent generation and support human interpretation at the same time.
Generative agents are increasingly used to select and narrate video highlights, but they typically operate over unstructured or framelevel representations. Their output is consequently difficult for a viewer to verify and steer toward individual preferences. We present the semantic action graph, a lightweight domain schema that represents a sports match as performer, action, recipient, moment, and state nodes connected by role, temporal, and outcome edges. * T. Lin, D. Chandran, G. Jagatap, A. Fanelli, D. Gunawan, and J. Kimball are with Dolby Laboratories. E-mail: [email protected] † C. Chen is now with XPENG.
Keywords: Graph-Based Video Representation, Highlight Generation, Agentic Systems, Human-AI Interaction, Sports Visualization.
1
Accepted for publication at IEEE VIS 2026 Workshop on GenAI, Agents, and the Future of VIS. 1 I NTRODUCTION Sports highlights are central to how fans engage with a game beyond the live match, condensing key moments into a reel with narrative that explains their significance. Yet fans differ in what they want to see and how they want it framed—a single manually edited reel can’t serve everyone. As a result, current highlights often miss the moments individual fans care about, offer little narrative context, and provide no effective way to find a specific play. Recent systems use large language and vision-language models to automatically generate highlights and narratives. One line of work selects or ranks the moments worth showing, by scoring frame importance [17], combining domain metrics with contextual reasoning [15], or assisting an editor with language-based retrieval [29]. Another line of work generates language to accompany footage already chosen, such as live commentary [2] or answers to questions about a play [16, 24]. These two halves are usually developed in isolation, and systems that produce both together are limited [4]. Because the generated prose isn’t linked to the underlying events, fans can’t verify its accuracy or customize results meaningfully. If the narrative mentions a defense-splitting pass, nothing in the interface confirms it happened, who made it, or when. A fan wanting an earlier starting point or the passes leading to shots rather than the shots themselves, has no way to specify this, since the system doesn’t expose those events. Existing pipelines work over frame embeddings and text tokens, which enable generation but aren’t units a person can inspect or reference. We attribute both limitations to the lack of a shared intermediate representation. Visualization tools have produced structured video indices that support browsing [8, 19], but generative models don’t consume them. Agentic pipelines that build explicit structures keep them internal, discarding them before output reaches the viewer [13, 30]. A representation that simultaneously grounds the AI pipeline and supports human interpretation remains absent. To bridge the gap, we present the semantic action graph, a lightweight schema that represents a sports match as performer, action, recipient, moment, and state nodes connected by role, temporal, and outcome relations. Designed for both agent and human consumers, the graph captures connected event sequences using a closed vocabulary and frame-addressable moments. An agentic pipeline can use it to compute match statistics, compose narratives, and retrieve clips, while a visual interface renders the identical structure for viewers to filter, inspect, and navigate. We demonstrate the schema through SportSAGE, a design probe for soccer highlights that pairs a four-module agentic pipeline with a graph-based interface. Grounded in three design goals derived from a formative study with six fans, SportSAGE takes structured match event data from an official league feed [9] and produces personalized narrated highlights that viewers can verify and navigate. We collected feedback from 12 soccer fans using SportSAGE on a professional match. Participants valued receiving customized highlights with varying narrative focuses of the same match, and used the graph layer to inspect generated clips and check the narratives against the video. These results serve as early evidence that the schema can ground agentic highlight creation and guide human interaction with its output. This work contributes the following: (1) the semantic action graph, a domain schema that serve both an agentic pipeline and a viewer-facing interface; (2) SportSAGE, a design probe demonstrating the schema in soccer highlight generation; and (3) user feedback narrative personalization and the value of a graph layer for interpreting and interacting with generated video content.
as commentator speech and crowd noise [25], modeled viewer arousal [12], and multimodal excitement features [20]. More recently, perception models have been trained to perform action spotting over large annotated datasets [7, 11] to identify key moments in a match. Apostolidis et al. [3] survey the wider summarization literature. However, the highlight these approaches return is a ranked set of moments rather than a composed sequence, limiting viewer interpretability and control. Language models have recently been applied to the same task. Structured representations such as scene graphs improve video understanding [13, 30] and reduce hallucination [31], though in these systems the structure remains internal to the model. Our work extends this line by defining an action-centric schema consumed by both the generative pipeline and the viewer. 2.2
Interfaces for casual game exploration
Visualization research makes video navigable by pairing it with structured metadata. Matejka et al. [19] presented Video Lens, which indexes baseball footage by attributes such as pitch type and player so fans can filter moments and play them back. Pavel et al. [21, 22] structure lecture and film video through aligned transcripts, and Deng et al. [8] reduce the effort needed to annotate racket sports events for such structured event indices. Sports visualization has produced interfaces for both analysts and fans. SoccerStories [23] supported visual analysis of soccer phases, and Stein et al. [28] combined movement data with video. Closer to casual viewing, Lin et al. [18] augmented sports video with embedded visualizations, and Lee et al. [16] answered tactical questions with narratives and embedded visualization. These systems show that structured metadata makes video navigable and supports fans engage with game context interactively. The structures they expose, however, are built for human retrieval alone. Our work explores a shared structure for both generative pipeline and human interfaces. 3
F ORMATIVE S TUDY
To understand how fans experience current highlights, we conducted 30-minute semi-structured interviews with six soccer fans (ages 1855; 1 female, 5 male), including one casual viewer (F2), two regular fans (F1, F5), three die-hard fans (F3, F4, F6). Interviews covered viewing habits, highlight consumption, difficulties in finding specific content, and desired features. Recordings were transcribed and analyzed using reflexive thematic analysis [6], with two researchers iteratively refining themes to consensus. Several gaps in current highlights impede fans’ engagement, which we categorized below. Content Gap. Participants reported that highlights omit surrounding events and the build-up to key moments. Four of six preferred seeing the actions leading to a goal over the finish alone: F3 wanted a clip to play “from when the action starts, the attack phase or previous defense phase, not just the actual moment of the goal.” Narrative Gap. Participants also wanted to understand why a moment mattered, not just what happened. There is often an issue with missing broader context and storytelling components: “Sometimes the highlight is just the moment of a goal, but I didn’t know there was a red card before that” (F1). Personalization Gap. Current highlights use a one-size-fits-all approach that ignores individual preferences. Most fans we interviewed have clear preferences in highlight narrative types: “There is no ability to walk in onto a particular player. I’d love to be able to be more personalized” (F4); “Everyone’s different. People care about the big moments but I kind of want to see the build-up” (F5). Discovery Gap. Locating a specific moment was reported as difficult, particularly for less prominent events: “if it’s a small moment, I usually have a hard time, like a penalty or a smaller moment that’s not normally captured” (F1). These gaps translate into three requirements on the underlying representation, which we use as design goals below. G1: Play
2 R ELATED W ORK 2.1 Sports highlight generation Automatically extracting highlights from sports video has been studied for over two decades, through proxies for excitement such
2
Accepted for publication at IEEE VIS 2026 Workshop on GenAI, Agents, and the Future of VIS.
Figure 2: (a) Semantic Action Graph schema. Five node types record who acted, what occurred, who was affected, when it occurred, and what game state resulted; five edge types carry role, sequence, and outcome. (b) A goal sequence expressed in the schema. The then edge chains the action sequence, including two passes, the tackle, and the shot that produced the score.
should be represented as connected sequences rather than isolated moments, so that build-up and consequence are available to both the system and the viewer. G2: The units the system reasons over should form a vocabulary viewers can also use to express preferences. G3: Individual moments should be addressable, so that a viewer can reach a specific event directly rather than by scrubbing. 4 4.1
the schema. This property makes the boundary of the system’s capability explicit while allowing structured extension, such as adding new action types (e.g., different foul types) or descriptors for players (e.g., appearance, speed). Frame-addressable moments (G3). Each Moment node carries a frame index of the event from the source feed, providing clear clip boundary when composing a highlight. A highlight generation agent can directly reference the moment nodes from the desired action sequence for accurate timestamps, avoiding incorrect generated time codes. For the viewer, the same property lets the viewer navigate the video through meaningful structure,(Fig. 1b), linking generated units to the footage to verify the narrative or localize specific moment effectively. Overall, the semantic action graph schema builds upon the common metadata that professional event feed carries (e.g., IPTC [14]) and adds two structural relations that metadata alone does not express. First, the then and produce relations make explicit the temporal and causal relationships between distinct action records, i.e., one action follows from another, and that an action changes the game state. Second, the sport-specific vocabulary is fixed in advance and can be retargeted — participants, actions, sequence, and outcome are shared across most sports, while event granularity and narrative style are not, so a long build-up chain in soccer and a single last-minute shot in basketball are both expressible. Our contribution is not a new data model but a minimal relational layer that connects atomic event records to narrative structure. Additionally, the schema is intentionally small. It does not represent spatial position, formation, or tactical intent, which matter for analysis but are not required to cover much of the descriptive content of sports commentary. The schema therefore serves as a base for agentic highlight generation from metadata that professional feeds already provide, and it can be extended when spatial data such as player location, speed, or formation are available.
S EMANTIC ACTION G RAPH Schema
A match is represented as a graph over five node types, as shown in Fig. 2a. The Performer nodes record who carried out an action and the Recipient nodes who was affected by it, both drawn from the match roster. The Action nodes record what occurred, drawn from a sport-specific catalog (e.g., for soccer Tackle, Foul, ShotOnTarget, Save, Score, YellowCard, and others). The Moment nodes carry temporal context in the form of game clock, video time, and frame index. The State nodes carry the game state an action produced, such as the score. Five edge types connect these nodes with semantic relationships. Three encode the roles within a single event: perform ( → ) links a performer to an action, affect ( → ) links an action to its recipient, and happen ( → ) links a moment to the action occurring at it. The remaining two carry structure across events: produce ( → ) links an action to the state it results in, and then ( → ) chains actions that are consecutive within a possession. Fig. 2b shows a goal expressed in the schema, in which a pass, a second pass, a tackle, and a shot are linked by then edges, ending in a produce edge to the resulting score. 4.2
Design Properties
The schema demonstrates three key properties informed by the design goals in Sec. 3. Connected event sequences (G1). With explicit edge definition in the graph, i.e. then edges connecting discrete actions, a highlight can be defined as a traversal of the graph rather than as a fixed time window around a detected event. For example, a highlight generation pipeline can add build-ups to the goal event by chaining back to a possession boundary rather than determining how many seconds before a goal as a parameter. The same edge structure also gives the viewer a compact representation of what a clip contains and in what order, which aligns with how a human viewer describes the event sequences. A shared, closed vocabulary (G2). In our schema, action nodes come from a fixed catalog of types, and performer and recipient nodes from the match roster. Because the entities and events are defined in advance, the pipeline’s generation freedom lies in phrasing rather than in facts. The same catalog is also presented to the viewer as the filter panel (Fig. 1a), so the system and the user shared same descriptive units. This closure also bounds what a viewer can ask for. A request outside the catalog, such as one about crowd celebrations, has no corresponding node and cannot be satisfied without extending
5
S PORT SAGE: A H IGHLIGHT G ENERATION P IPELINE
We built SportSAGE as a design probe to explore how a semantic action graph can support both agentic highlight generation and viewer-facing interpretation. 5.1
Graph Construction and Agentic Pipeline
We constructed the semantic action graph for a selected soccer match (i.e., 2024-2025 Pokal Cup Finale match between VfB Stuttgart and Arminia Bielefeld [1]) from official league match event data. The raw event logs contain timestamped, frame-aligned records of actions, performers, and recipients at roughly 2-5 second granularity. Such format is routinely produced for professional soccer [5]. We mapped the record onto the closed action catalog, including 13 actions covering possess, pass, tackle, foul types, shot types, kick types, and others. A 96-minute match yields 1,330 action nodes. The SportSAGE highlight pipeline is composed of four modules that operate on these records, as shown in Fig. 3.
3
Accepted for publication at IEEE VIS 2026 Workshop on GenAI, Agents, and the Future of VIS.
Figure 3: SportSAGE Agentic Pipeline. All four modules operate on semantic action graph data; none reads video. 1. Game Analyzer computes match statistics. 2. Narrative Creator generates narratives tailored to user preferences. 3. Highlight Generator assembles 10 candidate clips bounded by Moment frame indices. 4. Ranking Agent selects the top 5 and provides feedback to (3). Modules 2-4 are LLM agents.
AI highlight and narrative quality. Participants were satisfied with the highlight sets produced by the agentic pipeline. On a 7-point scale (1 = poor, 7 = excellent), they rated overall highlight quality positively for both the default set (M = 5.25) and the set matched to their chosen narrative focus (M = 5.25). Perceived quality was therefore unchanged by personalization. The difference appeared in preference, where participants favored the customized narrative over the default (M = 5.92). This is consistent with the narrative and personalization needs found in the formative study. They praised clip coverage and accuracy (“it showed enough of the buildup to be engaging, the pivotal moment and some of the celebration” (P4)) and valued the accompanying narrative for letting them “follow the progression of events easily” (P7). Distinct narrative focus mattered for both die-hard and casual fan: “This [tactical] type of description is more favorable for me when I want to search and watch highlights or game re-cap” (P8) and “It makes it easy to watch for a casual viewer like me who isn’t too invested in soccer” (P2). Graph layer usefulness. Participants interpreted the semantic action graph correctly and used it to check narratives against the video and to recover game context. They rated the interface useful (M = 5.96) and completed the three tasks reliably, including filtering shot events (12/12), locating a specific player action (11/12), and interpreting video events from the graph (11/12). In particular, participants valued the customizable nature and the granularity of the graph, which support fans “breaking down highlights to understand the plays better” (P2). These preliminary insights suggest that AI highlights grounded in the semantic action graph yield satisfying quality and have the potential to address the content, narrative, personalization, and discovery gaps found in our formative study.
1) Game Analyzer is a deterministic component. It takes the graph as a CSV conforming to the schema, and computes match statistics and per-player activity profiles, removing arithmetic from the language model [26]. 2) Narrative Creator takes the computed statistics with a chosen narrative preference (action-focused, tactical analysis, playercentric, casual-viewing), and generates a highlight narrative using templates built around common soccer archetypes, such as Momentum Shifts, Tactical Masterclass, and Individual Brilliance. 3) Highlight Generator assembles ten candidate clips, each a set of nodes whose frame indices give the clip boundaries. 4) Ranking Agent returns a top-five set and, through an optional iteration-limited reflection step [27], feeds corrective feedback back. The three generative modules use Claude 3.5 Sonnet. Each receives the Game Analyzer’s summaries and a pool of candidate clips identified by their constituent records, without accessing the original video. The schema thus serves as the sole interface between the agent and the match. Statistics are computed deterministically, and clip boundaries derived from Moment frame indices rather than generated timecodes, ensuring the highlight output is factual. 5.2 Interface The interface (Fig. 1) exposes the same schema through four components. A search panel filters by action type, player, and period, using the catalog of Sec. 4. A node-link view renders the events in the current filter result or generated clip along a timeline, encoding team membership in node borders and player identity in thumbnails; selecting a node moves the video to that frame. A personalization panel offers four narrative focuses (Action, Tactical, Casual, Player) and team/player preferences. A highlight list presents each generated clip, and selecting a clip updates the node-link view to its constituent events. The frontend uses React with Cytoscape.js [10].
7
6 U SER F EEDBACK We collected feedback from 12 soccer fans (2 female, 10 male; ages 18-54; 6 casual, 3 regular, 3 die-hard) in a 40-minute online session on SportSAGE highlights and the graph utility. Participants first viewed both a default highlight set of the soccer match (five clips) and a customized highlight set with their preferred narrative focus (Tactical, Casual, or Player). They provided feedback on their preference and perceived quality of the highlights. They then used the SportSAGE interface to examine the AI highlight clips with semantic action graph and perform reasoning tasks on the interface, including locating specific game moments from graph, and linking the highlight events to its semantic structure. We summarized user feedback into two aspects: 1) AI highlight and narrative quality, and 2) graph layer usefulness.
C ONCLUSION & F UTURE W ORK
We introduced the semantic action graph, a lightweight schema that represents matches as categorized, temporally ordered event sequences shared by both the agentic pipeline and the viewer-facing interface, making generative output in a narrative-rich domain more steerable and navigable. SportSAGE applies this schema to soccer highlight generation, enabling personalized narratives that viewers can verify and navigate through the graph interface. As a design probe, SportSAGE was evaluated on a single match with one model and no unstructured baseline, so we cannot yet attribute participants’ satisfaction to the shared representation itself. Future work should compare against an unstructured baseline, extend the schema with spatial and tactical attributes, and explore constructing the graph directly from video rather than official feeds.
4
Accepted for publication at IEEE VIS 2026 Workshop on GenAI, Agents, and the Future of VIS. ACKNOWLEDGMENTS
[18] T. Lin, C. Zhu-Tian, Y. Yang, D. Chiappalupi, J. Beyer, and H. Pfister. The quest for omnioculars: Embedded visualization for augmenting basketball game viewing experiences. IEEE transactions on visualization and computer graphics, 29(1):962–972, 2022. 2 [19] J. Matejka, T. Grossman, and G. Fitzmaurice. Video lens: Rapid playback and exploration of large video collections and associated metadata. In Proc. ACM Symp. on User Interface Software and Technology (UIST), pp. 541–550. ACM, 2014. doi: 10.1145/2642918.2647366 2 [20] M. Merler, K.-N. C. Mac, D. Joshi, Q.-B. Nguyen, S. Hammer, J. Kent et al. Automatic curation of sports highlights using multimodal excitement features. IEEE Trans. Multimedia, 21(5):1147–1160, 2019. doi: 10.1109/TMM.2018.2876046 2 [21] A. Pavel, D. B. Goldman, B. Hartmann, and M. Agrawala. SceneSkim: Searching and browsing movies using synchronized captions, scripts and plot summaries. In Proc. ACM Symp. on User Interface Software and Technology (UIST), pp. 181–190. ACM, 2015. doi: 10.1145/ 2807442.2807502 2 [22] A. Pavel, C. Reed, B. Hartmann, and M. Agrawala. Video digests: A browsable, skimmable format for informational lecture videos. In Proc. ACM Symp. on User Interface Software and Technology (UIST), pp. 573–582. ACM, 2014. doi: 10.1145/2642918.2647400 2 [23] C. Perin, R. Vuillemot, and J.-D. Fekete. SoccerStories: A kick-off for visual soccer analysis. IEEE Trans. Vis. Comput. Graph., 19(12):2506– 2515, 2013. doi: 10.1109/TVCG.2013.192 2 [24] J. Rao, Z. Li, H. Wu, Y. Zhang, Y. Wang, and W. Xie. Multi-agent system for comprehensive soccer understanding. In Proc. ACM Int. Conf. on Multimedia, pp. 3654–3663, 2025. doi: 10.1145/3746027. 3755144 2 [25] Y. Rui, A. Gupta, and A. Acero. Automatically extracting highlights for TV baseball programs. In Proc. ACM Int. Conf. on Multimedia, pp. 105–115. ACM, 2000. doi: 10.1145/354384.354443 2 [26] T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, E. Hambro et al. Toolformer: Language models can teach themselves to use tools. In Advances in Neural Information Processing Systems, vol. 36, pp. 68539–68551, 2023. 4 [27] N. Shinn, F. Cassano, A. Gopinath, K. Narasimhan, and S. Yao. Reflexion: Language agents with verbal reinforcement learning. In Advances in Neural Information Processing Systems, vol. 36, pp. 8634–8652, 2023. 4 [28] M. Stein, H. Janetzko, A. Lamprecht, T. Breitkreutz, P. Zimmermann, B. Goldlücke et al. Bring it to the pitch: Combining video and movement data to enhance team sport analysis. IEEE Trans. Vis. Comput. Graph., 24(1):13–22, 2018. doi: 10.1109/TVCG.2017.2745181 2 [29] B. Wang, Y. Li, Z. Lv, H. Xia, Y. Xu, and R. Sodhi. LAVE: LLMpowered agent assistance and language augmentation for video editing. In Proc. Int. Conf. on Intelligent User Interfaces (IUI), pp. 699–714, 2024. doi: 10.1145/3640543.3645147 2 [30] J. Yang, W. Peng, X. Li, Z. Guo, L. Chen, B. Li et al. Panoptic video scene graph generation. In Proc. IEEE/CVF CVPR, 2023. doi: 10. 1109/CVPR52729.2023.00634 2 [31] S. Yin, C. Fu, S. Zhao, T. Xu, H. Wang, D. Sui et al. Woodpecker: Hallucination correction for multimodal large language models. Science China Information Sciences, 67(12), 2024. doi: 10.1007/s11432-024 -4251-x 2
We acknowledge the Deutsche Fußball Liga (DFL) for granting access to official match data and video footage used in this research. We also thank the interviewees in our formative study for sharing their experiences. R EFERENCES [1] 2024-25 German Cup, Final. ESPN German Cup Final, 2025. [Online; accessed 13-November-2025]. 3 [2] P. Andrews, O. E. Nordberg, S. Zubicueta Portales, N. Borch, F. Guribye, and K. Fujita. AiCommentator: A multimodal conversational agent for embedded visualization in football viewing. In Proc. Int. Conf. on Intelligent User Interfaces (IUI), pp. 14–34, 2024. doi: 10 .1145/3640543.3645148 2 [3] E. Apostolidis, E. Adamantidou, A. I. Metsai, V. Mezaris, and I. Patras. Video summarization using deep neural networks: A survey. Proc. IEEE, 109(11):1838–1863, 2021. doi: 10.1109/JPROC.2021.3117472 2 [4] A. Barua, K. Benharrak, M. Chen, M. Huh, and A. Pavel. Lotus: Creating short videos from long videos with abstractive and extractive summarization. In Proc. Int. Conf. on Intelligent User Interfaces (IUI), pp. 967–981, 2025. 2 [5] M. Bassek, R. Rein, H. Weber, and D. Memmert. An integrated dataset of spatiotemporal and event data in elite soccer. Scientific Data, 12(1):195, 2025. doi: 10.1038/s41597-025-04505-y 3 [6] V. Braun and V. Clarke. Using thematic analysis in psychology. Qualitative Research in Psychology, 3(2):77–101, 2006. doi: 10.1191/ 1478088706qp063oa 2 [7] A. Deliège, A. Cioppa, S. Giancola, M. J. Seikavandi, J. V. Dueholm, K. Nasrollahi et al. SoccerNet-v2: A dataset and benchmarks for holistic understanding of broadcast soccer videos. In Proc. IEEE/CVF CVPR Workshops, pp. 4503–4514, 2021. doi: 10.1109/CVPRW53098. 2021.00508 2 [8] D. Deng, J. Wu, J. Wang, Y. Wu, X. Xie, Z. Zhou et al. EventAnchor: Reducing human interactions in event annotation of racket sports videos. In Proc. CHI Conf. on Human Factors in Computing Systems, pp. 1–13, 2021. doi: 10.1145/3411764.3445431 2 [9] Deutsche Fußball Liga (DFL). Official match data. https:// www.dfl.de/en/topics/match-data/official-match-data/, 2024. Accessed: 2025-11-13. 2 [10] M. Franz, C. T. Lopes, G. Huck, Y. Dong, O. Sumer, and G. D. Bader. Cytoscape.js: A graph theory library for visualisation and analysis. Bioinformatics, 32(2):309–311, 2016. doi: 10.1093/bioinformatics/ btv557 4 [11] S. Giancola, M. Amine, T. Dghaily, and B. Ghanem. SoccerNet: A scalable dataset for action spotting in soccer videos. In Proc. IEEE/CVF CVPR Workshops, pp. 1792–1810, 2018. doi: 10.1109/CVPRW.2018. 00223 2 [12] A. Hanjalic. Adaptive extraction of highlights from a sport video based on excitement modeling. IEEE Trans. Multimedia, 7(6):1114–1122, 2005. doi: 10.1109/TMM.2005.858397 2 [13] Z. Huang, Y. Ji, X. Wang, N. Mehta, T. Xiao, D. Lee et al. Building a mind palace: Structuring environment-grounded semantic graphs for effective long video analysis with LLMs. In Proc. IEEE/CVF CVPR, 2025. doi: 10.48550/arXiv.2501.04336 2 [14] International Press Telecommunications Council. IPTC sport schema. https://sportschema.org/, 2024. Version 1.1, approved October 2024. Licensed CC-BY 4.0. 3 [15] J. Kang, S. Kwon, J. Lee, and B.-H. Kim. DIAMOND: An LLMdriven agent for context-aware baseball highlight summarization. arXiv preprint arXiv:2506.02351, 2025. doi: 10.48550/arXiv.2506.02351 2 [16] C. Lee, T. Lin, H. Pfister, and Z. Chen. Sportify: Question answering with embedded visualizations and personified narratives for sports video. IEEE Trans. Vis. Comput. Graph., 31(1):12–22, 2025. doi: 10. 1109/TVCG.2024.3456332 2 [17] M. Lee, D. Gong, and M. Cho. Video summarization with large language models. In Proc. IEEE/CVF CVPR, pp. 18981–18991, 2025. 2
5