Provenance Metadata for Slop Control, with Transcripts and other Digital Objects

The following post was written by Metadata Operations Manager, Owen King, and first published as a blog post on the AI4LAM website.  GBH Archives became a founding member of AI4LAM last year, and collaborates within that organization to develop and communicate novel and responsible approaches to the use of AI within libraries, archives, and museums.

For present purposes, think of AI slop as the low-quality generative AI outputs that are unreliable as sources of information.

Here are two facts about slop that are of special relevance to libraries, archives, and museums:

  1. Not all derivative digital objects created with AI are slop.
  2. It is not always easy to discern what is slop, or the degree of sloppiness.

Consider two exemplary kinds of derivative objects: transcripts of radio shows and summaries of personal letters. Collecting institutions constantly create these sorts of derivatives within our processing workflows. To create them by hand is relatively straightforward, but also time-consuming. Fortunately, advances in AI since the introduction of the transformer architecture have yielded artificial systems that can produce high quality transcripts and summaries. These days, many of these AI-generated derivatives are quite usable, to the point that having them is far better than nothing, especially for collection management and access. In other words, the AI-derivatives are typically much more valuable than mere slop.

But, as with most AI products, your mileage may vary. At worst, a generated transcript or summary might be full of nonsense, stringing together words that are evidently meaningless even at a glance. That is the sloppiest slop, and the easiest to spot. But other problems are less visible and more insidious. A transcript might omit phrases, systematically mis-hear speech in certain accents, or replace distinctive terms of art with more common phrases that sound similar. Along the same lines, a summary might omit central points, misrepresent relations among ideas, or entirely fabricate details. With these problems the sloppiness is less obvious, but, for that reason, potentially more problematic. Furthermore, uncovering such faults may be as tedious as creating the derivatives by hand without any AI assistance at all.

So, to reiterate: Not all AI-generated derivatives are slop, but it is often hard to tell. In the face of this kind of unreliability and uncertainty, we need to know more about how the derivative objects were created. For the case of transcripts, a partial solution is provided by Transcript Provenance Metadata Elements (TPME). Designed and refined over the last two years by the AI4LAM Speech-to-Text Working Group, TPME specifies a content standard for documentation of the AI processes by which a transcript was produced. 

A TPME record tells you where a transcript came from. For example, it might tell you that a transcript or captions file was generated 

  • from the MP3 access copy of a recording
  • using OpenAI’s Python wrapper
  • using the “tiny” size of the model
  • with the language set to English
  • enforcing a maximum line length of 40 characters. 

A lack of further provenance data would imply that no subsequent post-processing or correction had been performed on the transcript. 

Such a TPME record is quite useful. To continue with the example, although the “tiny” model in the Whisper family is considerably better than speech-to-text models from 5 or 6 years ago, it has well-known shortcomings. Not only does it have a relatively high word error rate across different types of content; it has particular problems: It often transcribes instrumental music or background noise as common words or phrases, and it (like all models in the Whisper family) tends to produce nonsense when speakers change languages. Depending on the content, the transcript with this provenance is likely to have significant problems and would be a good candidate to be corrected or redone. However, without a TPME record, a glance at the transcript would not necessarily reveal its problems. In short, although provenance data is not a direct measure of quality, this data can serve as a proxy metric, indicating known patterns of accuracy and defects.

Beyond its value for managing transcripts, provenance metadata offers benefits for end users. In particular, it supports transparency about how AI has been used in collection processing: TPME records themselves could be made available, or they could be used to conditionally display user-facing indicators of the use of AI. These measures align with the transparency obligations of the EU AI Act and the principles for the use of generative AI recommended by the Association for Computing Machinery’s Technology Policy Council. TPME may serve other communicative purposes as well, for example, providing a framework for clarity with vendors about the procurement and acceptance criteria for transcription services.

The Speech-to-Text Working Group released version 1.0 of TPME on Zenodo in July. While AI continues its rapid evolution, and record-keeping practices evolve in tandem, the current version of TPME has proved suitable for use in audiovisual archives. GBH Archives creates TPME records for all transcripts of items in the American Archive of Public Broadcasting. Elements of TPME are used along with the FADGI Guidelines for Embedded Metadata in WebVTT Files for transcripts in the audiovisual special collections at Emory University Libraries. The National Library of Norway has also developed a pilot implementation of TPME that would inform the upcoming transcription of their 1M-hour collection of radio material. Meanwhile, other organizations, such as the Ukrainian History and Education Center, are currently exploring implementation of TPME.

With the core schema in place, work continues. The Speech-to-Text Working Group will make updates and extensions as needed by community members. Anyone interested in this project is invited to make comments or suggestions in the current working draft of the TPME schema documentation. Suggested clarifications and proposals for new data elements are especially welcome. Another possible direction of future development would be creation of controlled vocabularies for element values. 

On the implementation side, we recognize that provenance data is most reliable when captured as the work happens. For example, a transcript editing application can report whether a transcript has actually been corrected by a human, rather than relying on that being asserted after the fact. So, we are exploring TPME integration into tooling for generating and editing transcripts, particularly in the Hyperaudio Lite Editor.

Zooming back out to the larger domain of AI-generated derivative objects, this blog post began by discussing not just transcripts but also summaries. As it turns out, many of the same needs arise across both transcripts and summaries—as well as translations, audio descriptions, thumbnails, and other derivative representations of primary digital objects. For this reason, TPME has been noted by the Trust in Archives Initiative as a source of guidance for taxonomies of AI-generated or AI-edited media. Along these lines, another possible path forward would be to treat TPME as a special case of a more general schema for recording digital provenance across AI-generated alterations. That, of course, would require a coordinated effort across the world of libraries, archives, and museums.

Leave a Reply