ConceptioArchivearXiv CS
arXiv CSopen access

Diffusion TV: Experiencing Diffusion Models through Tangible, Embodied Interaction

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

Diffusion TV: Experiencing Diffusion Models through Tangible, Embodied Interaction SIHWA PARK, York University, Canada Diffusion TV is an interactive AI art installation that offers a tangible and embodied experience of diffusion models through a modified CRT TV. By physically manipulating the TV’s antenna, audiences control the clarity of AI-generated images and sounds, metaphorically enacting the denoising process that underlies diffusion-based generation. Using the tuning knob, participants switch between three channels featuring AI-generated animals from the Past (extinct species), Present (endangered species), and Future (speculative

arXiv:2609.05404v1 [cs.HC] 4 Sep 2026

creatures), situating the interaction within a temporal and ecological narrative. Through continuous audiovisual feedback and physical interaction, Diffusion TV foregrounds the generative process over final outputs, allowing audiences to explore intermediate states as experiential material. Rather than providing explicit technical explanation, the work presents an alternative, embodied mode of explainable AI that invites exploratory engagement with and reflection on generative technologies. Additional Key Words and Phrases: interactive installation, AI art, generative AI, diffusion models, explainable AI (XAI), tangible interaction, CRT television ACM Reference Format: Sihwa Park. 2026. Diffusion TV: Experiencing Diffusion Models through Tangible, Embodied Interaction. In Proceedings of Explainable AI for the Arts Workshop 2026 (XAIxArts 2026). ACM, New York, NY, USA, 5 pages.

1

Introduction and Background

Since the introduction of Denoising Diffusion Probabilistic Models (DDPMs) by Ho et al. [10], diffusion-based approaches have become a dominant paradigm for generative modeling across image, audio, and video. These models frame generation as an iterative transformation from noise to structured data, making the process temporally observable. In explainable AI (XAI) [1], intermediate outputs produced during this denoising process are generally used to analyze model behavior and interpret how structure emerges over time [14]. Outside research contexts, however, such intermediate states are rarely exposed to broader audiences and are typically presented as pre-rendered or static visualizations, limiting opportunities for interactive or embodied exploration. While XAI has been widely studied in functional and task-oriented domains, its application in creative contexts remains comparatively underexplored [2]. As generative AI increasingly shapes cultural production, questions arise as to how its processes can be made legible and meaningful to non-experts beyond conventional technical explanation, which commonly relies on 2D, screen-based graphical user interfaces (GUIs) [4, 6]. In this context, alternative approaches have emerged. The XAIxArts workshop series, for example, emphasizes artistic practice as a means of explainability, advancing alternatives to technocentric XAI within the arts [3]. Hemment et al. [9] introduce experiential AI as an approach that leverages artistic practices and tangible experiences to mediate between opaque computational systems and human understanding. Tangible explainable AI [4] and graspable AI [6] demonstrate how physical interaction, Author’s Contact Information: Sihwa Park, [email protected], York University, Toronto, ON, Canada. Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from [email protected]. © 2026 Copyright held by the owner/author(s). Publication rights licensed to ACM. Manuscript submitted to ACM Manuscript submitted to ACM

1

2

Sihwa Park

material metaphors, and aesthetic qualities can support experiential understanding, allowing users to engage with AI systems not only intellectually but also perceptually and emotionally. This paper presents Diffusion TV, an interactive AI art installation that explores alternative ways of engaging with complex AI mechanisms through embodied interaction. Using a modified cathode-ray tube (CRT) television (TV), the work enables audiences to interact with the denoising process of diffusion models by mapping antenna and knob interactions to intermediate stages of image and sound generation. In doing so, Diffusion TV reframes denoising as a performative and exploratory act, treating intermediate states as primary experiential material. Rather than prioritizing precise technical explanation, the work emphasizes intuitive, sensory engagement, inviting audiences to explore generative processes and reflect on broader themes of disappearance, preservation, and the relationship between humanity, technology, and the environment. 2

Diffusion TV Design

(a)

(b)

(c)

(d)

Fig. 1. Content display sequence with antenna interaction: (a) the noisiest image of an animal, (b) an intermediate denoising state as the antenna rotates, (c) the final denoised image with additional information, and (d) when no rotation input is detected, the noisiest image of the next animal appears after a few seconds.

Situating the work within traditions of media archaeology and artistic reuse of outdated technologies [13], Diffusion TV employs the format of a nostalgic CRT TV as both interface and conceptual framework. The physical affordances of CRT interaction, such as tuning knobs and antennas, are mapped to the denoising process of diffusion models, allowing audiences to engage with generative transformation through familiar broadcast metaphors. By rotating the antenna, viewers control the visual and auditory clarity of the AI-generated images and sounds of an animal, moving across intermediate stages from noise to structured output. Once a denoising sequence reaches clarity, the system resets after a short interval, introducing a new animal and maintaining a continuous, evolving interaction (see Figure 1). The content is organized through a multichannel structure accessible via the very-high frequency (VHF) tuner knob, consisting of three channels: Past, Present, and Future. The Past channel presents extinct animals, evoking a sense of loss and remembrance. The Present focuses on endangered animals, offering a moment of reflection on the current environmental crisis. The Future channel features speculative creatures imagined with AI, encouraging contemplation about where humanity and technology are headed. This structure situates the generative process within temporal and ecological narratives. At the point of clarity, brief contextual information appears in a lower-third format reminiscent of broadcast TV. Varying by channel, this information provides minimal cues, such as identifiers, temporal references, or speculative descriptors. Moreover, analogous to conventional television programming, Diffusion TV continuously regenerates the Manuscript submitted to ACM

Diffusion TV: Experiencing Diffusion Models through Tangible, Embodied Interaction

3

images and sounds of the animals at configurable intervals, such as daily updates. This content update ensures repeated visits yield different audiovisual experiences. 3

Animal Data

Several sources, including the International Union for Conservation of Nature’s Red List of Threatened Species [11] and the World Wildlife Fund website [18], were used to identify sixteen extinct and fourteen endangered species, along with contextual information such as last-seen year, habitat, and estimated population size. Image and sound generation prompts were manually designed and iteratively refined for each species. For the Future channel, twelve speculative species were generated using ChatGPT o3 [12], including each species’ name, emergence year, environmental trigger, short bio, and prompts for image and sound generation. 4

Technical Overview

Fig. 2. Schematic diagram of Diffusion TV

As Figure 2 illustrates, Diffusion TV consists of a modified CRT TV (GoldStar CMX-4200 13", 1987), a Raspberry Pi 5 (RPi) running a custom client program, and a sensor system for capturing user interaction. The RPi outputs the generated audiovisual content through a modified HDMI-to-radio frequency (RF) modulator for display on the TV. Images and sounds for all animals were pre-generated with Stable Diffusion XL [15] and Stable Audio Open [5], using the animal dataset and custom Python code executed in Google Colab [7]. Model inference was performed using the Diffusers [17] library, with custom callback functions used to extract intermediate outputs at each denoising step. These outputs were stored and made accessible to the client via Google Drive [8]. User interaction is captured through rotary and magnetic encoders integrated into the TV. A telescopic antenna is mechanically linked to a rotary encoder via a custom 3D-printed connector, mapping horizontal rotation to a continuous control signal. This signal is mapped to an index that selects corresponding image and sound pairs from different denoising stages. Channel selection is detected via a magnetic encoder attached to the VHF tuner, enabling switching between content sets associated with the Past, Present, and Future channels. The client, developed in Processing [16], manages content playback and interaction mapping. It dynamically updates audiovisual output in response to user input, selecting pre-generated media based on antenna position and channel state. The client also periodically checks for newly generated content and updates the dataset asynchronously without interrupting interaction. Manuscript submitted to ACM

4 5

Sihwa Park Exhibition and Discussion

Fig. 3. Diffusion TV installation

As shown in Figure 3, Diffusion TV was first exhibited at the International Symposium on Electronic/Emerging Art 2025, held from May 23 to 29, 2025, at the Hangaram Design Museum in Seoul, South Korea, where it was open to the public. The following reflections are based on the author’s exhibition observations and video recordings rather than a structured empirical evaluation. Audiences typically learned the interaction through exploration, with prior familiarity with CRT TVs shaping engagement: older participants navigated intuitively, while younger viewers sometimes struggled with the interface. The antenna-based denoising interaction was generally perceived as intuitive and engaging; however, some participants did not fully explore sequential content or expected additional system behaviors. Overall, the tuning metaphor supported an experiential engagement with the generative process and encouraged thematic interpretation without explicit instruction. Video documentation of the Diffusion TV installation and audience interactions is available at https://sihwapark. com/Diffusion-TV. 6

Conclusion and Future Work

Diffusion TV demonstrates the potential of embodied, metaphorical engagement as an alternative mode of XAI that moves beyond explicit technical explanation in conventional 2D GUI-based approaches. While the tangible interaction with the denoising process supports intuitive and experiential engagement, it does not guarantee precise comprehension of diffusion models, and audience responses varied depending on prior familiarity with AI. Observations were based on informal feedback and qualitative reflection rather than systematic evaluation. Future work will involve structured audience studies to examine how embodied interaction shapes public interpretation of generative AI systems. In addition, the system will be extended toward real-time or hybrid inference pipelines to maintain responsive interaction while enabling more flexible generative behavior, alongside further refinement of tangible controls to support richer exploration of the generative process. Manuscript submitted to ACM

Diffusion TV: Experiencing Diffusion Models through Tangible, Embodied Interaction

5

Acknowledgments This project was undertaken thanks in part to funding from the Connected Minds Program, supported by Canada First Research Excellence Fund, Grant #CFREF-2022-00010. References [1] Alejandro Barredo Arrieta, Natalia Díaz-Rodríguez, Javier Del Ser, Adrien Bennetot, Siham Tabik, Alberto Barbado, Salvador Garcia, Sergio Gil-Lopez, Daniel Molina, Richard Benjamins, Raja Chatila, and Francisco Herrera. 2020. Explainable Artificial Intelligence (XAI): Concepts, Taxonomies, Opportunities and Challenges toward Responsible AI. 58 (2020), 82–115. doi:10.1016/j.inffus.2019.12.012 [2] Nick Bryan-Kinns, Corey Ford, Alan Chamberlain, Steven David Benford, Helen Kennedy, Zijin Li, Wu Qiong, Gus G. Xia, and Jeba Rezwana. 2023. Explainable AI for the Arts: XAIxArts. In Proceedings of the 15th Conference on Creativity and Cognition (New York, NY, USA, 2023-06-19) (C&C ’23). Association for Computing Machinery, 1–7. doi:10.1145/3591196.3593517 [3] Nick Bryan-Kinns, Shuoyang Zheng, Francisco Castro, Makayla Lewis, Jia-Rey Chang, Gabriel Vigliensoni, Terence Broad, Michael Paul Clemens, and Elizabeth Wilson. 2025. XAIxArts Manifesto: Explainable AI for the Arts. In Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems (New York, NY, USA, 2025-04-25) (CHI EA ’25). Association for Computing Machinery, 1–8. doi:10.1145/3706599.3716227 [4] Ashley Colley, Kaisa Väänänen, and Jonna Häkkilä. 2022. Tangible Explainable AI - an Initial Conceptual Framework. In Proceedings of the 21st International Conference on Mobile and Ubiquitous Multimedia (New York, NY, USA, 2022-12-29) (MUM ’22). Association for Computing Machinery, 22–27. doi:10.1145/3568444.3568456 [5] Zach Evans, Julian D. Parker, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons. 2025. Stable Audio Open. In ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 1–5. doi:10.1109/ICASSP49660.2025.10888461 [6] Maliheh Ghajargar, Jeffrey Bardzell, Alison Marie Smith-Renner, Kristina Höök, and Peter Gall Krogh. 2022. Graspable AI: Physical Forms as Explanation Modality for Explainable AI. In Proceedings of the Sixteenth International Conference on Tangible, Embedded, and Embodied Interaction (New York, NY, USA, 2022-02-13) (TEI ’22). Association for Computing Machinery, 1–4. doi:10.1145/3490149.3503666 [7] Google. 2026. Google Colaboratory. https://colab.research.google.com/. Accessed: 2026-04-24. [8] Google. 2026. Google Drive. https://drive.google.com/. Accessed: 2026-04-24. [9] Drew Hemment, Dave Murray-Rust, Vaishak Belle, Ruth Aylett, Matjaz Vidmar, and Frank Broz. 2024. Experiential AI: Between Arts and Explainable AI. 57, 3 (2024), 298–306. doi:10.1162/leon_a_02524 [10] Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising Diffusion Probabilistic Models. In Proceedings of the 34th International Conference on Neural Information Processing Systems (Red Hook, NY, USA, 2020-12-06) (NIPS ’20). Curran Associates Inc., 6840–6851. [11] International Union for Conservation of Nature. 2026. The IUCN Red List of Threatened Species. https://www.iucnredlist.org/. Accessed: 2026-04-24. [12] OpenAI. 2026. ChatGPT. https://chatgpt.com/. Accessed: 2026-04-24. [13] Jussi Parikka. 2012. What Is Media Archaeology? Polity Press, Cambridge, UK. [14] Ji-Hoon Park, Yeong-Joon Ju, and Seong-Whan Lee. 2024. Explaining Generative Diffusion Models via Visual Analysis for Interpretable DecisionMaking Process. 248 (2024), 123231. doi:10.1016/j.eswa.2024.123231 [15] Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. 2024. SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis. In The Twelfth International Conference on Learning Representations. https: //openreview.net/forum?id=di52zR8xgf [16] Processing Foundation. 2026. Processing. https://processing.org/. Accessed: 2026-04-24. [17] Patrick von Platen, Suraj Patil, Anton Lozhkov, Pedro Cuenca, Nathan Lambert, Kashif Rasul, Mishig Davaadorj, Dhruv Nair, Sayak Paul, William Berman, Yiyi Xu, Steven Liu, and Thomas Wolf. 2022. Diffusers: State-of-the-art diffusion models. https://github.com/huggingface/diffusers. [18] World Wildlife Fund. 2026. World Wildlife Fund. https://www.worldwildlife.org/. Accessed: 2026-04-24.

Manuscript submitted to ACM

Record · ID 660839 · SHA-256 825ba0b3d1f2f448
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.