view post Post 2332 🇹🇷 One of our small Turkish models quietly reached **500+ monthly downloads** on Hugging Face.**Werea-TR-TextRestore — only 300M parameters.**Its job is simple:istanbulda hava cok guzel→ İstanbul'da hava çok güzel.A lightweight model for restoring Turkish text:• diacritics• punctuation• casing• corrupted text**96.5% word accuracy** on real Turkish news sentences.And it runs without sending your text to a cloud API.🤗 Try the model: Werea-co/Werea-TR-TextRestore🇹🇷 Built in Türkiye. Open source.If you're working on Turkish NLP, I'd love to hear what we should build next.#TurkishNLP #HuggingFace #OpenSourceAI #NLP See translation 2 replies · 🤗 3 3 🔥 2 2 👍 1 1 + Reply
OCR on the Hub Collection Curated OCR models for documents, languages, handwriting and text in images. Browse four collections with short practical notes. • 4 items • Updated 3 days ago • 11
NeoMME Collection Meet NeoMME: a family of 260M and 800M Multimodal-Native Multilingual Encoders • 12 items • Updated 7 days ago • 30
Jina-OCR-v1: Efficient Document Parsing with Speculative Decoding and Dense Verifiable Rewards Paper • 2609.03181 • Published 9 days ago • 3
Kraken PP-OCRv6 text recognition models Collection Hub mirrors of Benjamin Kiessling's multilingual PP-OCRv6 line-recognition family for Kraken: tiny, small, and medium. • 3 items • Updated 7 days ago • 7
Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers Paper • 2608.18972 • Published 23 days ago • 11
Institutional Newspapers Collection A growing corpus of newspapers, parsed and optimized for computational access. • 6 items • Updated 9 days ago • 7
Nemotron-Personas Collection A collection of multilingual, region-specific synthetic persona datasets that support sovereign AI development across many countries and regions. • 10 items • Updated about 1 month ago • 71