Sound generator tools powered by artificial intelligence (AI) create sound based on a given source image. It made us wonder — how would AI logics represent what a museum sounds like? 

With generous support from CLCF, we have been investigating the intersection of museums, sound studies, and artificial intelligence. In July we had the pleasure of presenting our findings at the Digital Humanities 2025 Conference in Lisbon, Portugal. In this post, we outline a new interdisciplinary methodology for sampling sound in museums and the results of our autoethnographic experiments with AI.

🎶 Sound in Museums

Snippets of song trace the path through the Museum of Broadway, from Oklahoma to the Phantom of the Opera. Heady, pounding music reverberates through the Museum of Sex, while the sounds of carnival-style games drift from the floor above. At the Transit Museum, the creaks of old subway turnstiles punctuate the air. 

Susan Stewart argued in 1999 that museums have historically served as “empires of sight.” However, now in 2025, specialized museums increasingly incorporate audio technologies into their exhibits. We argue that these technologies increase accessibility, dismantle colonial empires of sight, and broaden the possibilities of accessing cultural heritage through other sensory modalities. These audio technologies thus foster unique possibilities for affective experiential processing.

Through our work we ask: In what ways does the use of sound embolden immersive and affective relationships to cultural heritage in museums? How does the addition of sonic media in these spaces subvert or challenge the dominance of visual material in museums?

📝 Our Methodology

We use listening as an autoethnographic methodology to investigate the social, personal, relational, and experiential qualities of sound in museums. 

We conducted ethnographic field work in three specialized museums – the New York Transit Museum, the Museum of Sex, and the Broadway Museum – and two universal museums – The Brooklyn Museum and The Royal Ontario Museum.

We repurposed tools traditionally used in soundscape studies for environmental analysis and deep listening practice, and amended them to align with museum listening. We began each visit with a brief soundwalk through the museums (a deep listening practice with special attention to what we are hearing). We then recorded various locations throughout the museums through a combination of field notes, photos, and videos. We concluded our visits by identifying three sonically diverse areas for deepened listening practice using the soundcount methodology. 

The soundcount sheet is a methodology developed by R. Murray Schafer which is used in soundscape studies to examine the sonic balance or imbalance of environmental spaces. We amended the categories to better reflect museum spaces and conducted sonic analyses with the soundcount sheet in 3 locations per museum, resulting in 30 soundcount sheets across 5 museums.

This methodology was remarkably effective for sampling the varieties of sound within the museum space. It drew our attention to visitor demographics, helping us to better identify how they were moving through the space and what features of the exhibit they were interacting with. It also provided a useful comparator between museums. By effectively sampling the types of sounds present in the museum, we were able to reflect on the ways in which sound is curated and unplanned, constant or ephemeral, and negotiated between the venue and the visitor.

✋ Informative, Immersive, Interpretive… and Interactive

Our research expands on the work of scholars Foteini Salmouka and Andromache Gazi. They  classify sound in museums into three categories: informative, immersive, and interpretive. Informative sound functions like audio descriptions: it uses a narrator and is meant to support, enhance, or replace written text. A classic example of this would be an audio box, like the box pictured below from the Royal Ontario Museum.

They further separate more “thematic” sound into “interpretive sound.” Examples of interpretive sound would include archival recordings, like having a recording of a Winston Churchill speech playing in an exhibit about him, or — from our field research — how the Museum of Broadway plays different musical soundtracks throughout the museum. 

Their third and final category is immersive: sound that deliberately creates a mood and an affective connection to the exhibit. So for example, the use of soundscapes to create immersive, affective experiences in otherwise dimly lit rooms at the Museum of Sex. 

Based on our autoethnographic study, we identified a fourth category: interactive. Through our soundcounts, many of the sounds that we identified were generated by visitors interacting with the exhibits — pulling levers, walking through old subway turnstiles, playing games, and more. 

As argued by Stewart, museums have historically served as colonial “empires of sight” that privilege sight over any other senses, reifying Western ways of knowing. We argue that sonic technologies like the ones we recorded subvert and challenge the dominance of visual material in museums. Moreover, the addition of tactile, interactive elements in an exhibit break down this empire of sight on two fronts: through tactile sensation and the resulting creation of interactive sound.

💻 Generating Sound with AI

After collecting data on how sound currently exists in museums, we began our investigation into AI sound generation. Using photographs that we had taken during our field visits, we generated   30 AI soundscapes of the ROM and Brooklyn Museum to examine the AI’s biases and expectations of museum soundscapes. We then compared the AI output to our field recordings.

Let’s do an example together. Here is a recording from the Royal Ontario Museum. You’ll hear us move from one area which features the previously mentioned audio box to a period-specific room, audibly separated by alcoves of music:

ROM Field Recording

Next, we generated the AI-soundscape. We used Krotos Studios, a sound editing software and newly integrated AI tool. Importantly, this program solely uses their “team uploaded” audios as source material for their soundscape generation — i.e. the material has been collected with consent. The software allows a user to upload an image into Krotos; their AI system then offers a description of what might be audible in the image (up to 250 words). Clicking generate, the program then chooses 4 audio presents to generate a mix.

Uploading an image to the software, four sound sources are situated in a mixing field, enabling a sound editor to adjust levels, source, and location, among other traits.

We gave Krotos Studio this photograph of where we recorded the audio (above). Krotos Studio described this image vaguely as “museum ambience, antique furniture creaks, occasional footsteps, distant voices.” Here is the AI generated soundscape:

ROM AI Soundscape

Krotos’s soundscape reflects generalized perceptions of museums as silent spaces transformed by chatter and movement. What Krotos doesn’t account for are the purposeful instances of sound design. On one hand this is likely due to the limitations of the technology; while there are some musical components from which it can generate a suggestion, the primary purpose of the tool is to be an ambient generator. It similarly doesn’t account for the “leaky” qualities of sound. Sound is reverberant and can be heard in the movement of a field recordist through the hallways between exhibitions, sounds fade into the background but they are still audible.

One of the things that interested us the most about this project was seeing the limitations of how AI ‘thinks’ museums sound. At best, the AI is able to suggest footsteps and voices. The lack of any informative, interpretive, immersive, or interactive sound is understandable, but it also reifies this idea of museums as empires of sight. Generative AI is only as good as its training data — and this reflects a specific Western conception of what museums are and how they should sound.

✍️Next Steps

Having finished our data collection this spring, we are moving into data analysis. In this next phase of research we will code the characteristics of our AI generated soundscapes to try and reverse engineer the logics of the AI soundscape generator. This artistic ethnographic experimentation will enable us to reflect on the relationship between visuality and sonic media within museums, while simultaneously engaging in an analysis of the possibilities, limitations, and ethics of AI generated material or technologies within cultural heritage sites. 

Stay tuned for our academic article! We will also be generating an accompanying listening experience about the process and outcomes of our research on sound design in museums.

Cate Cleo Alexander

CLCF Postdoctoral Fellow

Dr. Cate Cleo Alexander is a graduate of the Faculty of Information at the University of Toronto. Cate employs a wide variety of methodologies in her research, including autoethnography, digital ethnography, media historiographies, sampling/scraping, qualitative coding,& artistic autoethnographic experiments with genAI

Lauren Knight

Research Assistant

Lauren Knight (she/her) is a sound artist and Ph.D. candidate (SSHRC CGS-D) in the Faculty of Information at the University of Toronto. Her research interests include acoustic ecology, cultural sound studies, media history, and research creation.