agentlans/chinese-japanese-classification-fasttext
Viewer • Updated • 4.56M • 40
How to use agentlans/fasttext-chinese-japanese-classifier with fastText:
from huggingface_hub import hf_hub_download
import fasttext
model = fasttext.load_model(hf_hub_download("agentlans/fasttext-chinese-japanese-classifier", "model.bin"))A lightweight FastText supervised model designed to accurately distinguish between various Chinese writing systems, Cantonese, and Japanese texts.
Standard language identification models frequently struggle to differentiate between similar scripts (such as Hanzi, Kanji, and Kana). This model is specifically trained to resolve these ambiguities.
This repository provides two versions of the trained model:
zhja_classifier.bin: The standard trained model.zhja_classifier.ftz: The compressed and quantized model (recommended for faster loading and smaller storage footprint).The model outputs one or more of the following labels:
| Label | Description |
|---|---|
__label__zhs |
Simplified Chinese |
__label__zht |
Traditional Chinese |
__label__yue |
Cantonese |
__label__ja |
Japanese |
Evaluated on a test set of N = 455,617 samples:
zhja_classifier.bin)
zhja_classifier.ftz)
You can easily use this model in Python using the fasttext library.
pip install fasttext
import fasttext
# Load the model (download zhja_classifier.bin from the repo first)
model = fasttext.load_model("zhja_classifier.bin")
# Test texts across different supported categories
texts = [
"这是一个用于区分中文和日文的分类器模型。", # Simplified Chinese
"這是一個用於區分中文和日文的分類器模型。", # Traditional Chinese
"呢個係一個用嚟區分唔同中文同日文嘅模型。", # Cantonese
"これは中国語と日本語を区別するための分類器モデルです。" # Japanese
]
for text in texts:
# Clean newlines as FastText expects single-line inputs
clean_text = text.replace("\n", " ")
predictions = model.predict(clean_text)
label = predictions[0][0].replace("__label__", "")
confidence = predictions[1][0]
print(f"Text: {text}")
print(f"Predicted Label: {label} (Confidence: {confidence:.4f})\n")
Distributed under the MIT License.