Text Classification
fastText
Chinese
Japanese
language-identification

FastText Chinese-Japanese Classifier

A lightweight FastText supervised model designed to accurately distinguish between various Chinese writing systems, Cantonese, and Japanese texts.

Standard language identification models frequently struggle to differentiate between similar scripts (such as Hanzi, Kanji, and Kana). This model is specifically trained to resolve these ambiguities.

Model Files

This repository provides two versions of the trained model:

  • zhja_classifier.bin: The standard trained model.
  • zhja_classifier.ftz: The compressed and quantized model (recommended for faster loading and smaller storage footprint).

Labels & Outputs

The model outputs one or more of the following labels:

Label Description
__label__zhs Simplified Chinese
__label__zht Traditional Chinese
__label__yue Cantonese
__label__ja Japanese

Performance Metrics

Evaluated on a test set of N = 455,617 samples:

Standard Model (zhja_classifier.bin)

  • Precision@1: 0.940
  • Recall@1: 0.924

Quantized Model (zhja_classifier.ftz)

  • Precision@1: 0.929
  • Recall@1: 0.914

Quick Start & Example

You can easily use this model in Python using the fasttext library.

1. Installation

pip install fasttext

2. Python Usage

import fasttext

# Load the model (download zhja_classifier.bin from the repo first)
model = fasttext.load_model("zhja_classifier.bin")

# Test texts across different supported categories
texts = [
    "这是一个用于区分中文和日文的分类器模型。",  # Simplified Chinese
    "這是一個用於區分中文和日文的分類器模型。",  # Traditional Chinese
    "呢個係一個用嚟區分唔同中文同日文嘅模型。",  # Cantonese
    "これは中国語と日本語を区別するための分類器モデルです。"   # Japanese
]

for text in texts:
    # Clean newlines as FastText expects single-line inputs
    clean_text = text.replace("\n", " ")
    predictions = model.predict(clean_text)
    label = predictions[0][0].replace("__label__", "")
    confidence = predictions[1][0]
    print(f"Text: {text}")
    print(f"Predicted Label: {label} (Confidence: {confidence:.4f})\n")

Limitations

  • Scope: Restricted to the supported Chinese scripts, Cantonese, and Japanese. Do not use for unrelated languages.
  • Text Length: Performance drops when handling very short texts (e.g., single-word inputs or short fragments lacking contextual cues).
  • Nuance: May occasionally misclassify highly localized slang, specialized borrowed terms, or hybrid phrasing.

License

Distributed under the MIT License.

Downloads last month
8
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train agentlans/fasttext-chinese-japanese-classifier