Nano Language Detector (EDAN20, Lab 4)

A small language identifier inspired by Google's CLD3, built for the EDAN20 Language Technology course at Lund University. It recognises eight languages: Arabic (ara), Mandarin Chinese (cmn), Danish (dan), English (eng), French (fra), Japanese (jpn), Korean (kor), and Swedish (swe).

Model

Each sentence is lowercased and split into character unigrams, bigrams, and trigrams. The n-grams are hashed with MD5 into 521 + 1031 + 1031 = 2583 buckets, and the input vector contains the relative frequency of each bucket. The network has one hidden layer:

input (2583) -> Linear + ReLU (50) -> Linear (8) -> softmax

The same architecture was trained twice: with scikit-learn (MLPClassifier, 5 epochs) and with PyTorch (Adam, cross-entropy, batches of 32, 7 epochs).

Files

File Content
nld.joblib scikit-learn MLPClassifier
nld.pth PyTorch weights (state_dict)
nld_vectorizer.joblib DictVectorizer mapping hash buckets to input columns (needed by both models)
nld_lang_codes.joblib Dictionary from class index to language code
predict.py Feature extraction and inference with both models

Usage

pip install torch scikit-learn joblib huggingface_hub
python predict.py "Hejsan grabbar!" "Salut les gars !"

The feature extraction in predict.py must be used as is: the models only work with the same hashing (MD5) and the same bucket sizes.

Training data and results

Trained on Tatoeba sentences: 25,000 per language (11,674 for Korean), 80 % for training and 20 % (37,335 sentences) for validation.

Model Accuracy Macro F1
scikit-learn 0.9912 0.9916
PyTorch 0.9924 0.9927

Most errors are between Swedish and Danish, and very short sentences can be misclassified.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support