← Back to context

Comment by dirteater_

5 hours ago

IMO the SotA for this is https://www.speechsuper.com/. Amazon suffers for similar

> One annoyance is that for Mandarin, the percentage is calculated at the character level, whereas with English, it gives you a more granular score at the phoneme level.

This is the case for most solutions you'd find for this task. Probably because of the 1 character -> 1 syllable property. It's pretty straightforward to split the detected pinyin into initial+final and build a score from that though.