๐ฎ๐ณ Qwen3.5-9B Hindi Instruct โ it stops thinking in English Ask base Qwen3.5-9B a question in Hindi and it burns hundreds of tokens thinking in English inside its think block before a single Devanagari word appears โ then code-switches in the answer. I fine-tuned it to close the think block instantly and reply in pure, native Hindi. โ Model (16-bit): pankajpandey-dev/qwen3.5-9b-hindi-instruct โ GGUF (Q4/Q5/Q8): pankajpandey-dev/qwen3.5-9b-hindi-instruct-GGUF โ Try it in the browser: pankajpandey-dev/qwen3.5-9b-hindi-demo Recipe: Unsloth + LoRA (r=16, response-only loss) on 12.9k Hindi pairs โ AI4Bharat anudesh + dolly-hi + wikiHow-hi + Aya Hindi (human-written). The Q4_K_M is 5.4 GB and runs on a plain laptop CPU. New in this run vs my earlier models: mixed in long-form native sources (wikiHow) after my last eval showed the fine-tune traded detail for conciseness โ this one keeps answers detailed and native. Part of my weekly ๐ฎ๐ณ Hindi LLM Series. Feedback welcome ๐ #Hindi #IndicNLP #Qwen #GGUF #LocalLLM #Unsloth
๐ฎ๐ณ New in my Hindi LLM Series: Gemma-4 E4B, fine-tuned for Hindi โ and it runs on your laptop's CPU. I fine-tuned Google's new Gemma-4 E4B on ~10k Hindi instruction pairs (AI4Bharat: anudesh + dolly) using Unsloth + LoRA, on a single L4 GPU. Then I ran an honest side-by-side eval: base Gemma-4 vs my fine-tune, across 25 Hindi prompts. The results were interesting ๐ โ My fine-tune is more concise โ ask for "3 tips" and it gives exactly 3. Base writes a 1,200-character essay.
โ Pure native Hindi โ base keeps slipping into English ("เคธเคเคคเฅเคฒเคฟเคค เคเคนเคพเคฐ (Eat a Balanced Diet)", "เคคเคพเคฐเคพ (Star)"). My fine-tune stays in clean Hindi.
โ Tighter instruction-following โ ask for a "short message" and it gives one, not a menu of options. โ๏ธ And to be honest: base Gemma-4 is more detailed and comprehensive. I didn't build a "smarter" model โ I built a focused, Hindi-native, edge-friendly one that runs as a 5GB GGUF (Q4) on CPU. ๐ Try it:
๐ฎ๐ณ Gemma-3-1B Hindi Instruct โ a Hindi LLM that runs fully offline, anywhere. Last week I shipped Qwen3-4B Hindi. This week I went the other direction: how tiny can a useful Hindi model get? So I fine-tuned Gemma-3-1B on quality-filtered Hindi instruction data and shipped the full GGUF ladder. โ Fine-tune (16-bit): pankajpandey-dev/gemma-3-1b-hindi-instruct โ GGUF (Q4/Q5/Q8): pankajpandey-dev/gemma-3-1b-hindi-instruct-GGUF Runs in Ollama, llama.cpp, and LM Studio. The Q4_K_M is just 806 MB โ runs on CPU, a cheap laptop, even a Raspberry Pi. What I tried this round: chrF-filtered the training data to drop weak translations, and used response-only loss so the model learns how to answer, not how to repeat prompts. Honest note: at 1B, Hindi fluency is strong but coherence is bounded by size โ it's a lightweight/edge experiment, not a 4B replacement. Gemma-3-4B Hindi is next. Part of my Hindi LLM Series โ openly-licensed Indic models for local & edge use. Feedback welcome ๐ #Hindi #IndicNLP #GGUF #LocalLLM #Gemma #EdgeAI
๐ฎ๐ณ Qwen3-4B Hindi Instruct v2 โ a Hindi LLM that runs on your own machine Most strong Hindi-capable models are either huge or cloud-only. I wanted one that's small enough to run locally but actually follows instructions in Hindi โ so I fine-tuned Qwen3-4B on 10K Hindi instruction pairs and shipped it with a full GGUF quant ladder. โ Fine-tune (16-bit): huggingface.co/pankajpandey-dev/Qwen3-4B-Hindi-Instruct-v2 โ GGUF (Q4/Q5/Q8): huggingface.co/pankajpandey-dev/Qwen3-4B-Hindi-Instruct-v2-GGUF Runs in Ollama, llama.cpp, and LM Studio. The Q4_K_M is just 2.5 GB โ fits comfortably on a laptop, CPU or GPU. Part of my Hindi LLM Series โ building openly-licensed Indic models for local and edge use. More coming (Gemma next). Feedback welcome ๐ #Hindi #IndicNLP #GGUF #LocalLLM #Qwen
๐งฌ Just uploaded K-quants of Carbon-3B for llama.cpp users! @HuggingFaceBio released the original GGUF in bf16 only โ so I added the full quant ladder for CPU/edge inference: โข Q2_K โ 1.4 GB โข Q3_K_M โ 1.8 GB โข Q4_K_M โ 2.1 GB โญ โข Q5_K_M โ 2.4 GB โข Q6_K โ 2.7 GB โข Q8_0 โ 3.5 GB ๐ pankajpandey-dev/Carbon-3B-GGUF Now you can generate DNA sequences on your laptop. Needs a llama.cpp build with PR #23410 (HybridDNATokenizer support). Huge thanks to the HuggingFaceBio team for the original model ๐ #GGUF #llamacpp #genomics #DNA
I built a little demo where you give three models (Apertus, Llama, Qwen3) the same prompt and in the end you have to guess which is which just based on their answers.
Turns out : if we predict ๐ earth we can save a lot of time looking for interesting things and less time looking at things that we expect to see.
Sentinel-2 imagery ๐ฐ๏ธbasically takes a long time to download towards earth. so our "near real time" systems are quite far from that in practical terms.
meanwhile , if we "predict" what we will see , based on what we do see , we can send down much less data in a timely way , and prioritize ๐กearth-bound response .
I'm talking about illegal fishing , logging , mining or building in nature reserves , the more of that we predict early the more we're able to stop it on time.
@retrain-pipelines v0.2.0 is out ! I'm at Station F at My booth with GOSIM Paris 2026 today & tomorrow. Come meet me for a live in-person demo and a chat !