SONAR is an evaluation toolkit for multilingual ASR that goes beyond WER/CER. It combines semantic similarity, the Poseidon Score, and analysis across dialect, demographic, and metadata-based failure modes. š
Our goal is to make it easier for everyone to understand why an ASR model fails, not just how often. š You can plug in your own models + audio, extend it to new languages and datasets, or contribute directly. š ļø
MIT licensed. Would love feedback from the HF community! š¤