22 languages tuned end to end. Captions in 40+.
Skapo transcribes, captions word by word, trims filler words and cuts clips at speech boundaries in 22 languages, each with its own filler list and a caption font that covers its script. The spoken language is detected from the audio; you never have to pick it. Everything below is what the pipeline actually ships today, including the two scripts it does not.
Word-level timing
A forced-alignment model per language, so every caption word lands on the syllable it belongs to.
Filler list per language
"Uhm" in Dutch, "ähm" in German, "euh" in French. Never an English list applied to everyone.
A font for the script
Kana, kanji, Hangul, Thai, Devanagari, Cyrillic and Greek burn in as text, not empty boxes.
Boundary rules
Clips end on a finished thought in that language's rhythm, so a Japanese clip is not cut like an English one.
Tuned end to end (22)
| Language | Code | Transcription | Word captions | Filler removal | Script | Notes |
|---|---|---|---|---|---|---|
| 🇬🇧English | en | Latin | ||||
| 🇳🇱DutchNederlands | nl | Latin | Filler list covers "uhm", "eh" and "nou ja"; these are not treated as English. | |||
| 🇩🇪GermanDeutsch | de | Latin | "Ähm" and "äh" are trimmed; "er" is kept, because in German it is a pronoun. | |||
| 🇫🇷FrenchFrançais | fr | Latin | "Euh" and "bah" are trimmed. | |||
| 🇪🇸SpanishEspañol | es | Latin | ||||
| 🇵🇹PortuguesePortuguês | pt | Latin | European and Brazilian audio both transcribe; the model is not split by variant. | |||
| 🇮🇹ItalianItaliano | it | Latin | ||||
| 🇵🇱PolishPolski | pl | Latin | ||||
| 🇸🇪SwedishSvenska | sv | Latin | ||||
| 🇳🇴NorwegianNorsk | no | Latin | ||||
| 🇩🇰DanishDansk | da | Latin | ||||
| 🇫🇮FinnishSuomi | fi | Latin | ||||
| 🇪🇪EstonianEesti | et | Latin | ||||
| 🇱🇻LatvianLatviešu | lv | Latin | ||||
| 🇱🇹LithuanianLietuvių | lt | Latin | ||||
| 🇯🇵Japanese日本語 | ja | CJK | Kana and kanji render from a shipped CJK face; no Latin fallback boxes. | |||
| 🇨🇳Chinese中文 | zh | CJK | Simplified and Traditional both render. Transcription is Mandarin. | |||
| 🇰🇷Korean한국어 | ko | Hangul | ||||
| 🇮🇳Hindi | hi | Devanagari | Captions burn in as Hinglish (romanized) by default; switch to Devanagari per job. | |||
| 🇹🇭Thaiไทย | th | Thai | Thai has no spaces between words; caption line breaks follow the alignment model, not whitespace. | |||
| 🇮🇩IndonesianBahasa Indonesia | id | Latin | ||||
| 🇲🇾MalayBahasa Melayu | ms | Latin | Separate from Indonesian: different filler list and alignment model. |
Language codes are the ones the API and brand kits accept. Detection is automatic; set the code by hand only when an episode mixes languages.
Captioned, not tuned
The transcription model understands roughly 99 languages. Any of them whose script has a shipped font gets transcribed and captioned word by word, without a per-language filler list or tuned clip boundaries. That is where the “40+” comes from. Examples:
Afrikaans, Bulgarian, Catalan, Croatian, Czech, Filipino, Galician, Greek, Hungarian, Icelandic, Macedonian, Romanian, Russian, Serbian, Slovak, Slovenian, Swahili, Turkish, Ukrainian, Vietnamese, Welsh.
Scripts shipped, and not
- Latin (including all accented and Baltic letters)
- Cyrillic
- Greek
- Japanese (kana and kanji)
- Chinese (Simplified and Traditional)
- Korean (Hangul)
- Thai
- Devanagari
Not supported yet: Arabic and Hebrew. No font ships for them and right-to-left layout needs more than a font. Those jobs are refused with a clear message rather than rendered with broken captions.
Interface language is a separate axis
The website and dashboard come in 11 languages: English, Dutch, French, Portuguese, Spanish, German, Polish, Japanese, Korean, Traditional Chinese, Indonesian. That has nothing to do with the audio you process. A Dutch-speaking editor on the Dutch dashboard can clip a Japanese interview and get Japanese captions.
Machine-readable version of this page for assistants and crawlers: /llms.txt. Language codes for structured data: en nl de fr es pt it pl sv no da fi et lv lt ja zh ko hi th id ms.
Questions people ask before trusting a language claim
Which languages does Skapo support?
22 languages are tuned end to end: English, Dutch, German, French, Spanish, Portuguese, Italian, Polish, Swedish, Norwegian, Danish, Finnish, Estonian, Latvian, Lithuanian, Japanese, Chinese, Korean, Hindi, Thai, Indonesian, Malay. Each of those gets a forced-alignment model for word-level caption timing, its own filler-word list, clip-boundary rules and a caption font that covers the script. Word-level captions work in 40+ languages overall.
Does Skapo remove filler words in Dutch and German?
Yes, with a separate list per language. "Uhm" and "eh" come out of Dutch audio, "ähm" and "äh" out of German. The lists are per language because a filler in one language is an ordinary word in another: "er" is hesitation in English and a pronoun in German, so a single English-trained list would cut real words.
Do Japanese, Chinese, Korean and Thai captions render correctly?
Yes. Captions burn in from a font that covers the script (kana and kanji, Simplified and Traditional Chinese, Hangul, Thai), so they read as text rather than empty boxes. Line breaks in Thai follow the alignment model instead of whitespace, since Thai does not put spaces between words.
Can I clip a podcast in a language the website is not translated into?
Yes. The interface language and the processing language are separate. The website and dashboard come in 11 languages; the audio you process can be any of the 22 tuned languages, or any captioned language, regardless of which interface language you use. A Dutch-speaking user can clip a Japanese episode.
Do I have to pick the language before uploading?
No. The spoken language is detected from the audio automatically. You can pin it by hand in the advanced options if an episode mixes languages or the detection picks the wrong one, and a brand kit can fix the language for every job in a client workspace.
What about Arabic or Hebrew?
Not supported yet. The transcription model understands both, but Skapo does not ship a caption font for them, and right-to-left layout needs more than a font. Rather than burn in broken captions, those jobs are refused with a clear message. The scripts that are shipped: Latin (including all accented and Baltic letters); Cyrillic; Greek; Japanese (kana and kanji); Chinese (Simplified and Traditional); Korean (Hangul); Thai; Devanagari.
Does Hindi burn in as Devanagari or as Hinglish?
Hinglish (romanized) by default, because that is how most Hindi short-form captions are read. Switch to Devanagari per job in the advanced options; the Devanagari face ships either way.