Nano Language Detector (EDAN20, Lab 4)
A small language identifier inspired by Google's CLD3, built for the EDAN20 Language Technology course at Lund University. It recognises eight languages: Arabic (ara), Mandarin Chinese (cmn), Danish (dan), English (eng), French (fra), Japanese (jpn), Korean (kor), and Swedish (swe).
Model
Each sentence is lowercased and split into character unigrams, bigrams, and trigrams. The n-grams are hashed with MD5 into 521 + 1031 + 1031 = 2583 buckets, and the input vector contains the relative frequency of each bucket. The network has one hidden layer:
input (2583) -> Linear + ReLU (50) -> Linear (8) -> softmax
The same architecture was trained twice: with scikit-learn (MLPClassifier, 5 epochs) and with PyTorch (Adam, cross-entropy, batches of 32, 7 epochs).
Files
| File | Content |
|---|---|
nld.joblib |
scikit-learn MLPClassifier |
nld.pth |
PyTorch weights (state_dict) |
nld_vectorizer.joblib |
DictVectorizer mapping hash buckets to input columns (needed by both models) |
nld_lang_codes.joblib |
Dictionary from class index to language code |
predict.py |
Feature extraction and inference with both models |
Usage
pip install torch scikit-learn joblib huggingface_hub
python predict.py "Hejsan grabbar!" "Salut les gars !"
The feature extraction in predict.py must be used as is: the models only work with the same hashing (MD5) and the same bucket sizes.
Training data and results
Trained on Tatoeba sentences: 25,000 per language (11,674 for Korean), 80 % for training and 20 % (37,335 sentences) for validation.
| Model | Accuracy | Macro F1 |
|---|---|---|
| scikit-learn | 0.9912 | 0.9916 |
| PyTorch | 0.9924 | 0.9927 |
Most errors are between Swedish and Danish, and very short sentences can be misclassified.