Vidoumo

By Stuart Frisby (@mrfrisby.com)
Published:

YouTube already has Japanese subtitles. For a learner that's half a tool: you get the characters, and nothing to help you read them, they're frustrating more often than they're useful, and they were never meant for a learner, so that shouldn't come as a surprise. The usual fix for a student of Japanese is a second window with a dictionary in it, and a lot of pausing. In my experience, that just becomes too high a hurdle to overcome, and my ambition to actively learn gets reduced to me passively listening and hoping I learn through osmosis and stubbornness alone.

Vidoumo is a browser extension that does the second window's job inside the captions. Furigana sits over the kanji, English sits underneath, and every word is clickable: the video pauses, a card shows the reading, the JLPT level and the meanings, and one button saves the word. Furigana and translation each switch off, so the same video works for a beginner and for someone who only wants the dictionary now and then.

<figure class="note-figure note-figure--wide" style="width: min(100vw, 57.5rem);"> <img src="/images/vidoumo-note-demo.jpg" alt="A browser window floating on a green gradient, showing the extension running over a real video: a Japanese caption with furigana above the kanji and an English line below, a dotted underline on an already saved word, the dictionary card for 時間 open above it, and a toolbar icon showing one saved word." loading="lazy" /> <figcaption>The extension on a real clip. Furigana, the English line, a saved word with its dotted underline, and the card for 時間.</figcaption> </figure>

Saved words keep the sentence they came from and the moment in the video, so a word can take you back to where you met it. They export as CSV, which is as far as my commitment to other people's flashcard apps goes. There's a live demo on the landing page, if you'd like to try it before you install anything.

<figure class="note-figure note-figure--wide"> <img src="/images/vidoumo-note-compare.jpg" alt="The same moment twice. Above, YouTube's own Japanese caption: plain white kanji on a dark bar. Below, Vidoumo's: furigana above the kanji, the words highlighted as clickable, and the English translation underneath." loading="lazy" /> <figcaption>The same moment. YouTube's caption, then Vidoumo's.</figcaption> </figure>

Reading the Japanese

Vidoumo starts from the Japanese subtitles YouTube already has, uploaded or automatic. Japanese has no spaces, so there's no specific thing to click until something decides where the words are. That is where kuromoji comes in handy, it is a JavaScript tokeniser, which runs in the background and returns words, readings and dictionary forms. This gives us the structure on top of which we can do all of the fun things that make Vidoumo useful.

Getting the right meaning

Dictionary lookups go to Jisho, which is excellent but ranks short kana words oddly. Search はい and the first answer is 灰, ash. The word you meant, yes, is third.

Vidoumo knows two things Jisho doesn't: what kuromoji thinks the word is, here an interjection, and the English translation of the line it's in. Re-ranking Jisho's eight results on those two facts, plus the usual nudge for common words, puts ‘yes’ first.

English from the source

The English under each line comes from the best place available, tried in order.

First, the uploader's own English subtitles, if the video has them. Written by a person and timed to the video, they're what the Japanese actually meant, not what a machine guessed. If they skip too many lines to be of use, they're passed over.

Second, YouTube's own English translation of the Japanese track. It's a machine translation, but it's cut to the same timing as the Japanese, and Vidoumo rebuilds it into whole sentences and shows the one being spoken. Third, and only if neither of those works, a line goes to Google, one at a time.

Design

It's designed in what I guess for now is my house style, lifted from the palette, typography and design direction for this very website. A clickable word is a pale green chip with an underline. A word you've saved fades to a dotted underline, so what's new stays visible and what you've already saved stops vying for attention. The caption box is the one and only thing Vidoumo renders on top of the video — the captions are big and readable, but respectful of the fact that they are still secondary to the thing you're watching.

<figure class="note-figure"> <img src="/images/vidoumo-note-popup.png" style="width: 21.25rem; max-width: 100%; margin: 0 auto; border: 0; border-radius: 0;" alt="The toolbar popup: switches for Furigana and Translation, a saved word with its reading, meaning, sentence, a link to the moment in the video, and an Export CSV button." loading="lazy" /> <figcaption>The popup. The switches, and a saved word with the moment and the video it came from.</figcaption> </figure>

The landing page was the other design problem. A learning tool is hard to explain in screenshots, so the page runs the real extension over a scene from Midnight Diner. I wanted that page to be more than just a link to a download, but rather a way to explain the decisions made in the making of the product, and the excitement I felt when I used it for the first time and could see that I was onto something.

What it is and isn't

It isn't in the Chrome Web Store. It's a zip and a few steps, because I'd rather find out whether anyone wants it before filling in a form to ask Google's permission to find out.

Where a video has English subtitles, nothing is sent to Google at all. Where it doesn't, the extension falls back to YouTube's translation, and then to Google. Pairing the Japanese with the English a person wrote for it gives you idiomatic Japanese and English side by side, with the definition, pronunciation and meaning connecting the two. Maybe that's the really novel thing here.

I'm sure this has been done before, since it feels so obvious. Building it was more fun than paying someone to build a worse one for me. I have more in mind, a review mode for the words you've saved, but this is enough for a V1 while I establish that I am the entire market, and the product has therefore reached 100% product-market fit.