Contributor Cataloging in the American Archive of Public Broadcasting: Documenting the Past, Thinking Forward

The following was submitted by AV Cataloging and Metadata Intern, Avery Schanbacher.

In preserving and providing access to historic public television and radio from across the country, the American Archive of Public Broadcasting (AAPB) provides a window into both national news and local stories. Within a single collection of materials from a station or program, who and what is represented in these broadcasts over time creates an illustrative cross-section of life’s throughlines and changes. When the presence and contribution of those featured in archival public media is brought out of the time-based media that it’s embedded in through searchable metadata, these stories then become more navigable, allowing users to discover new content, search for contributors, and trace themes. At both GBH Archives and the AAPB, increasing discoverability by documenting individuals represented in public media is an ongoing mission, and one I was eager to support as an Audiovisual and Metadata Training Data Intern. 

Working towards these goals of greater accessibility, my fellow intern and I used specialized cataloging tools incorporating elements of machine learning to catalog collections and create metadata documenting contributors at scale, working within what could otherwise be an impenetrable volume of material. Part of GBH’s collaboration with the CLAMS project at Brandeis University, these tools allow catalogers to utilize information-rich frames already present in audiovisual and public media materials to describe programs in greater depth and help users locate materials of interest.1 Common scenes with text like slates, bumpers, copyright frames, and most importantly, chyrons (text in the lower third of the screen that identifies an individual), hold valuable information about a program and its contributors that can be incredibly useful to viewers and researchers, including when a program aired, who is portrayed onscreen, and who aided in the program’s production. Through the use of a vision language model, which extracts text from scenes in a program, and a large language model that formats that text into preliminary catalog data that can then be selected and edited by the cataloger, the CLAMS app allows scenes with text to become a source for locating and learning about individuals appearing within a program. 

For me, seeing this work in action was a valuable way to demystify applications of machine learning in LIS, while remaining conscious of larger issues surrounding AI and its impacts. In approaching this work, I was grateful for the opportunity to engage in open discussions about responsible AI use and the importance of prioritizing trustworthiness and equitability in machine-assisted metadata creation while working towards practical and creative solutions for increasing discoverability.

Over the course of the summer, I worked alongside fellow Audiovisual Metadata and Training Data Intern Jenn Leishman and Metadata Operations Manager Owen King to continue applying and adjusting these workflows to increase documentation of contributors within the AAPB’s collections, and explore what that discoverability means. At the beginning of the summer, we focused on episodes of the New Jersey Nightly News (NJN), streamlining cataloging workflows and building off the work of previous GBH Archives interns who documented major contributors from the producing team (“usual suspects”) and identified guidelines for cataloging within the NJN collection. In total, we cataloged six months of NJN episodes, following news coverage, new members of the staff, and changes in the community from June to December 1981, from labor movements and political shifts to cultural and environmental phenomena. 

For the remainder of the summer, I focused on cataloging episodes of GBH’s locally focused news production Greater Boston. Cataloging Greater Boston episodes from 2012-2018 was a fascinating way to revisit recent events from a new lens. While the show focuses heavily on national politics, especially during a period of recent history that saw immense shifts at the national level, it also traces key figures in Boston’s city leadership, citizens and community organizations making an impact on the city, and important stories affecting life in the Greater Boston area. 

Using the unique format and visual style of the show to inform specialized guidance, I drafted a set of program-specific guidelines for cataloging the Greater Boston collection, including cataloging goals, a list of information to record, instructions for recording information from specific styles of scenes with text, and a list of regular commentators and members of the production team. This allowed our general instructions for our cataloging workflows to be easily applied and adapted to a new collection with different goals and characteristics. Once I identified a list of GBH contributors who appeared regularly on Greater Boston, I was able to include their names in a list of “registered people” in our prompt for normalizing extracted text and formatting catalog data, with instructions to correct spelling errors if text similar to one of these names appeared, and to automatically flag their roles as members of the production team. This enabled us to increase the initial accuracy of the catalog data as much as possible through specific guidelines for automatically formatting data. 

At the same time, our workflows evolved as we incorporated a new vision language model, Qwen3.5, for extracting data from scenes with text, which increased accuracy and reduced the amount of corrections needed for the final catalog data extracted from these frames. Although spelling errors are much less common and extracted text is more complete with this change to the vision language model we utilize, errors can still occur and models can still perform inequitably for names of different linguistic origins. This means that human review during the cataloging process remains vital in ensuring names are recorded and formatted correctly. Having eyes on the materials and being able to watch a part of a program while cataloging also allows us to observe context and make corrections, apply content warnings if necessary, and even document contributors who are heard and not seen, and thus do not appear in scenes with text pulled by the CLAMS app. While extracted catalog data is becoming more and more accurate, working responsibly with that output and knowing where bias still may be present requires critical thinking and human perspectives, demonstrating the continued benefit and necessity of human involvement during this process.

These accelerated contributor cataloging workflows allow us to vastly expand archivists’ exposure to the content of materials in the AAPB, and consider issues like privacy and content sensitivity in new ways. While evaluating privacy concerns and sensitive information in archival collections can be challenging due to the practical barriers of surfacing sensitive information from large quantities of material that may not be described on a granular level,2 cataloging scenes with text enabled us to identify areas within a program that might need attention and flag them for review, or indicate that a content warning should be added to the corresponding frame or scene. 

Being able to make onscreen text associated with contributors searchable also forced us to consider the impacts of greater discoverability beyond just increasing searchability and access to materials. Chyron text, especially attributes listed below a name, contain immense amounts of detail that can be valuable to researchers and bring light to contributors’ stories, like their role or positions, affiliations and associations, places they’re connected to, or what they’re being featured for. But these attributes can also include information that may be sensitive, contain harmful language, or pose privacy concerns. While attributes add valuable dimension to contributors’ stories, what are the ethics and implications of bringing this information, likely shared willingly but recorded in the temporally confined setting of a television broadcast, into a new digital setting where it will become searchable in perpetuity? How does this impact our responsibilities as catalogers of these collections?

To begin to address these questions, we turned to Helen Nissenbaum’s theory of contextual integrity, which posits that information’s original context and flow of distribution determine whether or not privacy has been breached in its disclosure.3 When evaluating privacy concerns within contributor metadata in the AAPB under this framework, we considered the impact of greater discoverability beyond the context of the original program, including the extent to which metadata may be indexed in search engine results or included in AI-generated summaries. Many individuals documented in AAPB collections have no current presence online, meaning that any new information made available about them could make up a significant portion of the entire digital record of that individual, and therefore have a disproportionate effect on how they may be perceived and understood in the future. If that new information being added to a minimal digital presence linked someone’s name to a word like “inmate” or “victim,” or used harmful language to describe an identity or disability without the source or context immediately apparent, what could that mean for the individual being described? For chyrons that include harmful or reductive language, language that carries strong negative stigma, or directly reference a traumatic event, we felt that these implications were especially important to consider.4 

Motivated by these potential impacts of decontextualization, and by concern for the replication of harmful language in catalog data, we drafted guidelines for a new flag to be applied during cataloging that would allow catalogers to ensure sensitive chyron attributes are viewed alongside their original context, allowing another option for handling attributes containing sensitive information or harmful language with care and without limiting contributor visibility. While assessing sensitivity remains subjective and will continue to require judgment and careful consideration from archivists, we created standardized criteria for applying this option that utilize understandings of historical context surrounding harmful language and societal stigma to help define what information in a chyron may be sensitive, as well as ideas drawn from contextual integrity to inform decision making about privacy.

When embarking on this project, it was important to me to think not only about documenting contributors to public media, but to consider what our obligations to those contributors are. Working with the GBH Archives and AAPB in support of greater discoverability was an incredible chance to explore these ideas, and how we can use the opportunities presented by scaled cataloging workflows to interact with materials on a deeper level and address issues like these in ways that do contributors and the stories they’re a part of justice. 

Footnotes

  1. Kyeongmin Rim, Owen C. King, Kelley Lynch, Marc Verhagen, and James Pustejovsky, “A Platform for AI-Assisted Archival Metadata Generation,” In Culture and Computing, HCII 2025, edited by Matthias Rauterberg 2025, and Lecture Notes in Computer Science vol. 15800, Springer, 2025, https://doi.org/10.1007/978-3-031-93160-4_12↩︎
  2. Steven Bingo, “Of Provenance and Privacy: Using Contextual Integrity to Define Third-Party Privacy,” The American Archivist vol. 74, no. 2 (2011): 506–21, https://doi-org.ezproxysim.flo.org/10.17723/aarc.74.2.55132839256116n4. ↩︎
  3. Helen Nissenbaum, “Privacy as Contextual Integrity,” Washington Law Review vol. 79, no. 1 (2004): 119-158, https://digitalcommons.law.uw.edu/wlr/vol79/iss1/10 ↩︎
  4. Tulane University Libraries, “A Guide to Inclusive and Reparative Archival Description at Tulane University Libraries and Newcomb Archives and Vorhoff Collection,” Version 1, April 2024, https://www2.archivists.org/sites/all/files/TUSC%20Guidelines%20for%20Inclusive%20and%20Conscious%20Description_0.pdf ↩︎

Leave a Reply