Models read text
An hour of video containing your best explanation of a topic contributes almost nothing to how an engine describes you unless that explanation exists as text somewhere it can read. The medium is not the problem; the absence of a transcript is.
Transcripts alone aren’t enough
Raw transcripts are unstructured, repetitive and full of speech artefacts, so they rarely contain a cleanly quotable passage. Editing one into a structured article is where the value is — the same requirement as any page that wants to be quoted.
Your best explanation trapped in an audio file is invisible to the systems now answering questions about you.
One recording, several artefacts
A single conversation can yield an article, several answers to specific questions, and the raw media. That is a far better return than publishing the recording alone, and it costs an afternoon of editing.
Interviews are unusually good source material
Someone asking you questions produces exactly the question-and-answer structure that retrieval favours, in the phrasing real people use — the same reason real support questions make good FAQ material.
The corroboration still counts
Appearing on someone else’s podcast produces a description of you at an independent source, which is the corroboration that carries weight — provided their show notes contain enough text to be read.
Host the text yourself
Publish the edited article on your own site and link the media. Otherwise the platform hosting the video becomes the source that gets cited, the same trap as syndicating your best pages elsewhere.