The fine print, printed large

Methodology

The Yap Index runs public podcast episodes through the Speko speech-to-text API (/v1/transcribe) and counts what comes back. As of the last run that is 10 shows, 50 episodes, and 85 hours of audio, fetched from each show's own public RSS feed. We never host or redistribute the audio.

What we measure

  • Words per minute: total transcribed words divided by total audio minutes. The steadiest number here.
  • Words per hour (yap density): the same math on an hourly scale.
  • Longest run: an estimate of the longest stretch of continuous speech, derived from text cadence. The transcripts carry no word timestamps, so this is inference, not measurement. Treat it as directional.
  • Fillers per minute: counts of um, uh, like, you know, sort of, kind of. Uncalibrated: transcription models disagree about how many disfluencies survive into text, so cross-show comparisons are entertainment, not evidence.

What we deliberately do not publish

  • Per-host stats. The pipeline runs without speaker diarization, so we cannot honestly attribute words to individual people. Show-level only.
  • Interruption counts and airtime splits. Same reason. When the measurement is not defensible, it does not ship.

Error expectations

Long episodes are transcribed in chunks with overlap, and chunk joins occasionally duplicate or drop a few words. WPM figures are stable to within a few percent; longest-run and filler figures carry wider error bars. A calibration harness with hand-labeled reference audio gates which metrics are allowed on the board; fillers have not passed it yet, which is why they are labeled uncalibrated wherever they appear.

Removal

If it is your show and you want off the board, delist it. The show is hidden immediately and the removal is logged. No argument.

Generated July 13, 2026. Providers and models are recorded per episode in the public data file.