# Text Version of Podcasts

**URL:** <https://forums.thimbleweedpark.com/t/text-version-of-podcasts/221>\
**Category:** Podcasts\
**Created:** [May 6, 2017, 7:45pm UTC](https://forums.thimbleweedpark.com/t/text-version-of-podcasts/221 "2017-05-06T19:45:09Z")\
**Posts on this page:** 20\
**Page:** 4

<div class="post-metadata">

**Author:** ![ZakPhoenixMcKracken](https://yyz2.discourse-cdn.com/flex030/user_avatar/forums.thimbleweedpark.com/zakphoenixmckracken/32/41_2.png) [@ZakPhoenixMcKracken](https://forums.thimbleweedpark.com/u/ZakPhoenixMcKracken)\
**Post date:** [October 30, 2017, 8:17am UTC](https://forums.thimbleweedpark.com/t/text-version-of-podcasts/221/62 "2017-10-30T08:17:13Z")

</div>

I love you guys!  
And also I love AI 😊

---

<div class="post-metadata">

**Author:** ![Someone](https://avatars.discourse-cdn.com/v4/letter/s/b77776/32.png) [@Someone](https://forums.thimbleweedpark.com/u/Someone)\
**Post date:** [October 30, 2017, 10:24am UTC](https://forums.thimbleweedpark.com/t/text-version-of-podcasts/221/63 "2017-10-30T10:24:33Z")

</div>

Great work! I will try to download the subtitles later (after work :)).

---

<div class="post-metadata">

**Author:** ![Sushi](https://yyz2.discourse-cdn.com/flex030/user_avatar/forums.thimbleweedpark.com/sushi/32/163_2.png) [@Sushi](https://forums.thimbleweedpark.com/u/Sushi)\
**Post date:** [October 30, 2017, 4:16pm UTC](https://forums.thimbleweedpark.com/t/text-version-of-podcasts/221/64 "2017-10-30T16:16:34Z")

</div>

So, I should stop my manual work, I guess? Or perhaps finish one and see if some things can be improved in the auto-extract?  
Note that I’m adding some annotations in the transcript, which no tool could do.

---

<div class="post-metadata">

**Author:** ![Sushi](https://yyz2.discourse-cdn.com/flex030/user_avatar/forums.thimbleweedpark.com/sushi/32/163_2.png) [@Sushi](https://forums.thimbleweedpark.com/u/Sushi)\
**Post date:** [October 30, 2017, 4:47pm UTC](https://forums.thimbleweedpark.com/t/text-version-of-podcasts/221/65 "2017-10-30T16:47:59Z")

</div>

Wow. That is pretty bad (edit: for such an expensive tool), imo. Without proper punctuation and cleaning of speech patterns, stuttering etc., it is quite unreadable unless used as a subtitle.  
It may look like actual text at first, but it is very tiring to make sense out of it.

---

<div class="post-metadata">

**Author:** ![besmaller](https://yyz2.discourse-cdn.com/flex030/user_avatar/forums.thimbleweedpark.com/besmaller/32/40_2.png) [@besmaller](https://forums.thimbleweedpark.com/u/besmaller)\
**Post date:** [October 30, 2017, 4:48pm UTC](https://forums.thimbleweedpark.com/t/text-version-of-podcasts/221/66 "2017-10-30T16:48:11Z")

</div>

> [@Sushi](#):
>
> So, I should stop my manual work, I guess?

Well, this process is just auto “captoning” for the video format of the podcast. I would see this as an “aid” to proper transcription, doing most of the word entry, but it should be changed to include speaker names and clean up the occasional (now more rare) transcription errors.

If you are near completion with one or more podcasts, I would finish that work, and then we can compare the effort with someone working to clean up one of the auto-created transcripts, and decide the best way to go forward.

---

<div class="post-metadata">

**Author:** ![Sushi](https://yyz2.discourse-cdn.com/flex030/user_avatar/forums.thimbleweedpark.com/sushi/32/163_2.png) [@Sushi](https://forums.thimbleweedpark.com/u/Sushi)\
**Post date:** [October 30, 2017, 4:50pm UTC](https://forums.thimbleweedpark.com/t/text-version-of-podcasts/221/67 "2017-10-30T16:50:43Z")

</div>

Great idea. I’m not a fast typer, so any base to start from will be useful. For the Friday Questions, I would make sure everyone’s unpronounceable nickname be written correctly by going back to the posts (and copy and paste the questions while I’m at it, of course)

---

<div class="post-metadata">

**Author:** ![besmaller](https://yyz2.discourse-cdn.com/flex030/user_avatar/forums.thimbleweedpark.com/besmaller/32/40_2.png) [@besmaller](https://forums.thimbleweedpark.com/u/besmaller)\
**Post date:** [October 30, 2017, 4:53pm UTC](https://forums.thimbleweedpark.com/t/text-version-of-podcasts/221/68 "2017-10-30T16:53:54Z")

</div>

> [@Sushi](#):
>
> That is pretty bad

Yes, it’s not bad for an automatic tool, but has lots of limitations. I needed to double check which post you were replying to, and I see it was the Sonix version. The Google automatic captions are actually much better than the Sonix output, but it doesn’t do speaker identification, and punctuation, etc…

---

<div class="post-metadata">

**Author:** ![BigRedButton](https://yyz2.discourse-cdn.com/flex030/user_avatar/forums.thimbleweedpark.com/bigredbutton/32/212_2.png) [@BigRedButton](https://forums.thimbleweedpark.com/u/BigRedButton)\
**Post date:** [October 30, 2017, 5:21pm UTC](https://forums.thimbleweedpark.com/t/text-version-of-podcasts/221/69 "2017-10-30T17:21:17Z")

</div>

If Sonix can write time-stamps in the output, we could merge the speaker informations from Sonix and the translation from YouTube.

---

<div class="post-metadata">

**Author:** ![Someone](https://avatars.discourse-cdn.com/v4/letter/s/b77776/32.png) [@Someone](https://forums.thimbleweedpark.com/u/Someone)\
**Post date:** [October 30, 2017, 5:33pm UTC](https://forums.thimbleweedpark.com/t/text-version-of-podcasts/221/70 "2017-10-30T17:33:04Z")

</div>

It’s working. 😉

youtube-dl is able to download all subtitles from all videos in the playlist. The results are WebVTT files that looks like this:

> **WebVTT**
>
> ```
> WEBVTT
> Kind: captions
> Language: en
> Style:
> ::cue(c.colorCCCCCC) { color: rgb(204,204,204);
> }
> ::cue(c.colorE5E5E5) { color: rgb(229,229,229);
> }
> ##
> 
> 00:00:06.550 --> 00:00:11.610 align:start position:19%
> [Music]
> 
> 00:00:09.650 --> 00:00:14.130 align:start position:19%
> hi<c.colorCCCCCC><00:00:10.650><c> I'm</c><00:00:10.889><c> Ron</c><00:00:11.099><c> Gilbert</c></c>
> 
> 00:00:11.610 --> 00:00:16.949 align:start position:19%
> I'm<c.colorE5E5E5><00:00:11.880><c> Gary</c><00:00:12.150><c> winning</c><00:00:12.480><c> and</c><00:00:12.780><c> this</c><00:00:13.530><c> is</c><00:00:13.650><c> our</c><00:00:13.799><c> first</c></c>
> 
> 00:00:14.130 --> 00:00:18.449 align:start position:19%
> stand-up<c.colorE5E5E5><00:00:15.000><c> meeting</c><00:00:15.360><c> podcast</c><00:00:15.929><c> and</c><00:00:16.350><c> a</c><00:00:16.529><c> stand-up</c></c>
> ...
> 
> ```

I can convert these files into the SRT format that looks like this:

> **SRT**
>
> ```
> 1
> 00:00:06,550 --> 00:00:11,610
> [Music]
> 
> 2
> 00:00:09,650 --> 00:00:14,130
> hi I'm Ron Gilbert
> 
> 3
> 00:00:11,610 --> 00:00:16,949
> I'm Gary winning and this is our first
> 
> 4
> 00:00:14,130 --> 00:00:18,449
> stand-up meeting podcast and a stand-up
> 
> 5
> 00:00:16,949 --> 00:00:21,720
> meeting is the thing a project usually
> 
> 6
> 00:00:18,449 --> 00:00:24,810
> has in the mornings where everybody on
> 
> 7
> 00:00:21,720 --> 00:00:27,029
> the team gets together and talks very
> ...
> 
> ```

A little sed command converts this to plain TXT:

> **Summary**
>
> ```
> [Music]
> hi I'm Ron Gilbert
> I'm Gary winning and this is our first
> stand-up meeting podcast and a stand-up
> meeting is the thing a project usually
> has in the mornings where everybody on
> the team gets together and talks very
> very briefly about what's going on in
> the project what happened what's gonna
> happen today and what happened yesterday
> usually the meetings are stand up just
> so they're quick and they don't drag on
> so we're gonna do the stand-up meeting
> podcasts they're probably gonna last
> less than five minutes and we'll
> ...
> 
> ```

Which file format do we need? All? 🙂

There are several tools to edit and translate SRT files. So theoretically someone could translate the podcast in, let’s say, Italian and then we could re-import the subtitles into the video. So everyone can hear the English podcast with Italian subtitles. But that would be a huge amount of work of course. 🙂

---

<div class="post-metadata">

**Author:** ![milanfahrnholz](https://avatars.discourse-cdn.com/v4/letter/m/cab0a1/32.png) [@milanfahrnholz](https://forums.thimbleweedpark.com/u/milanfahrnholz)\
**Post date:** [October 30, 2017, 5:42pm UTC](https://forums.thimbleweedpark.com/t/text-version-of-podcasts/221/71 "2017-10-30T17:42:41Z")

</div>

I always knew it. While Gary doesn´t show himself too much publicly, he secretly does _all_ of the winning. So much winning!

---

<div class="post-metadata">

**Author:** ![Someone](https://avatars.discourse-cdn.com/v4/letter/s/b77776/32.png) [@Someone](https://forums.thimbleweedpark.com/u/Someone)\
**Post date:** [October 30, 2017, 5:47pm UTC](https://forums.thimbleweedpark.com/t/text-version-of-podcasts/221/72 "2017-10-30T17:47:32Z")

</div>

> [@BigRedButton](#):
>
> If Sonix can write time-stamps in the output, we could merge the speaker informations from Sonix and the translation from YouTube.

So, you would pay the Sonix service?

> [@milanfahrnholz](#):
>
> I always knew it. While Gary doesn´t show himself too much publicly, he secretly does all of the winning.

I thought the same. Although I liked the Pockesphinx interpretation more, for example: “brenda our red phantom asks how much of the boot park was playboy wire frame before the real art was added.”

---

<div class="post-metadata">

**Author:** ![BigRedButton](https://yyz2.discourse-cdn.com/flex030/user_avatar/forums.thimbleweedpark.com/bigredbutton/32/212_2.png) [@BigRedButton](https://forums.thimbleweedpark.com/u/BigRedButton)\
**Post date:** [October 30, 2017, 5:59pm UTC](https://forums.thimbleweedpark.com/t/text-version-of-podcasts/221/73 "2017-10-30T17:59:14Z")

</div>

> [@Someone](#):
>
> So, you would pay the Sonix service?

No, I forgot about the costs. Sorry.

---

<div class="post-metadata">

**Author:** ![ZakPhoenixMcKracken](https://yyz2.discourse-cdn.com/flex030/user_avatar/forums.thimbleweedpark.com/zakphoenixmckracken/32/41_2.png) [@ZakPhoenixMcKracken](https://forums.thimbleweedpark.com/u/ZakPhoenixMcKracken)\
**Post date:** [October 30, 2017, 6:10pm UTC](https://forums.thimbleweedpark.com/t/text-version-of-podcasts/221/74 "2017-10-30T18:10:40Z")

</div>

> [@Sushi](#):
>
> I would make sure everyone’s unpronounceable nickname be written correctly

I hope so.

---

<div class="post-metadata">

**Author:** ![LowLevel](https://yyz2.discourse-cdn.com/flex030/user_avatar/forums.thimbleweedpark.com/lowlevel/32/48_2.png) [@LowLevel](https://forums.thimbleweedpark.com/u/LowLevel)\
**Post date:** [October 30, 2017, 7:57pm UTC](https://forums.thimbleweedpark.com/t/text-version-of-podcasts/221/75 "2017-10-30T19:57:13Z")

</div>

> [@Someone](#):
>
> So, you would pay the Sonix service?

Can we quantify the Sonix costs? How many hours of podcasts are there?

---

<div class="post-metadata">

**Author:** ![RonGilbert](https://yyz2.discourse-cdn.com/flex030/user_avatar/forums.thimbleweedpark.com/rongilbert/32/7_2.png) [@RonGilbert](https://forums.thimbleweedpark.com/u/RonGilbert)\
**Post date:** [October 30, 2017, 8:12pm UTC](https://forums.thimbleweedpark.com/t/text-version-of-podcasts/221/76 "2017-10-30T20:12:40Z")

</div>

I like the idea of this. Once the captioning issues are figured out, I’d be happy to put them on the official YouTube channel if that’s desired.

---

<div class="post-metadata">

**Author:** ![besmaller](https://yyz2.discourse-cdn.com/flex030/user_avatar/forums.thimbleweedpark.com/besmaller/32/40_2.png) [@besmaller](https://forums.thimbleweedpark.com/u/besmaller)\
**Post date:** [October 30, 2017, 8:22pm UTC](https://forums.thimbleweedpark.com/t/text-version-of-podcasts/221/77 "2017-10-30T20:22:03Z")

</div>

> [@LowLevel](#):
>
> Can we quantify the Sonix costs? How many hours of podcasts are there?

There are approximately 21 hours of audio in the podcasts. At $8/hr that’s $168. The service also costs $15/month on top of that. It would require 1 month at least. The service has a nice editing tool as well as the transcription. However, it’s clear the transcription quality is substantially worse than Google’s, and I don’t know how easy it would be to merge data. Also, looking at the one example I posted earlier, it missed speaker transitions quite a bit, especially with more than 2 speakers.

The service offers a free trial with one hour of audio transcription. If someone wanted to, they could do one of the \<1 hour podcasts, and try it out with the free trial and see how well it works, along with the online editing (like with Podcast #1, only about 14 minutes I think). Theoretically, we could import the transcription made from Sonix into the Youtube video captions, assuming the data is represented with a timestamp format.

---

<div class="post-metadata">

**Author:** ![Nor\_Treblig](https://yyz2.discourse-cdn.com/flex030/user_avatar/forums.thimbleweedpark.com/nor_treblig/32/216_2.png) [@Nor\_Treblig](https://forums.thimbleweedpark.com/u/Nor_Treblig)\
**Post date:** [October 30, 2017, 8:46pm UTC](https://forums.thimbleweedpark.com/t/text-version-of-podcasts/221/78 "2017-10-30T20:46:46Z")

</div>

> [@besmaller](#):
>
> There are approximately 21 hours of audio in the podcasts.

21 hours? This can’t be right, you are missing [half an hour](http://www.cinemapioxi.it/zak/Thimbleweed_Park_Fake_Podcasts.html)!

---

<div class="post-metadata">

**Author:** ![besmaller](https://yyz2.discourse-cdn.com/flex030/user_avatar/forums.thimbleweedpark.com/besmaller/32/40_2.png) [@besmaller](https://forums.thimbleweedpark.com/u/besmaller)\
**Post date:** [October 30, 2017, 8:50pm UTC](https://forums.thimbleweedpark.com/t/text-version-of-podcasts/221/79 "2017-10-30T20:50:12Z")

</div>

> [@Nor\_Treblig](#):
>
> 21 hours? This can’t be right, you are missing half an hour!

🙂 Good point! I’ll add these tonight, I’m curious how the Google transcription will work…

---

<div class="post-metadata">

**Author:** ![BigRedButton](https://yyz2.discourse-cdn.com/flex030/user_avatar/forums.thimbleweedpark.com/bigredbutton/32/212_2.png) [@BigRedButton](https://forums.thimbleweedpark.com/u/BigRedButton)\
**Post date:** [October 30, 2017, 9:06pm UTC](https://forums.thimbleweedpark.com/t/text-version-of-podcasts/221/80 "2017-10-30T21:06:25Z")

</div>

I think I have found the name of what we are discussing: [Speaker diarisation](https://en.wikipedia.org/wiki/Diarization).

There are some open-source speaker recognition tools out there. Though, I’m not sure about their functionality.

[Here is a link](http://alize.univ-avignon.fr/) to one tool which seems to provide diarization. The “Tutorial for LIA\_SpkSeg — Top-down Speaker Segmenting and Clustering System” may be interesting for us.

---

<div class="post-metadata">

**Author:** ![ZakPhoenixMcKracken](https://yyz2.discourse-cdn.com/flex030/user_avatar/forums.thimbleweedpark.com/zakphoenixmckracken/32/41_2.png) [@ZakPhoenixMcKracken](https://forums.thimbleweedpark.com/u/ZakPhoenixMcKracken)\
**Post date:** [October 30, 2017, 10:31pm UTC](https://forums.thimbleweedpark.com/t/text-version-of-podcasts/221/81 "2017-10-30T22:31:03Z")

</div>

> [@Nor\_Treblig](#):
>
> 21 hours? This can’t be right, you are missing half an hour!

Ahah, thank you for your mention, but I don’t think I deserve it. Maybe as a “bonus file” 😀

> [@besmaller](#):
>
> I’m curious how the Google transcription will work…

…with an English pronunciation with heavy italian accent. It’s a tough challenge for the AI !

[Previous page](https://forums.thimbleweedpark.com/t/text-version-of-podcasts/221.md?page=3)

[Next page](https://forums.thimbleweedpark.com/t/text-version-of-podcasts/221.md?page=5)
