you know, we keep talking about how ai needs more representation for low-resource languages. but kashmiri, with all that incredible literature, centuries of history, was pretty much invisible to modern ocr systems.
it bothered us. so we decided to stop talking about it and build something.
today, we're open-sourcing koshur pixel. it’s the largest synthetic ocr dataset for the kashmiri language: 613,000 image-text pairs covering words, sentences, even whole pages.
really grateful to @Haq_Nawaz_Malik and @FaizanIqbal__52. couldn't have done this without them.
paper: https://t.co/97Aww6BqVh
dataset: https://t.co/RsMpqZsmjU
"India cannot build Frontier Language Models"
Bangalore Paper Club edition 2 happened Last night - The theme was alternate architectures for language models. Why?
Scaling is a compute game. Architecture is an ideas game.
One needs GPUs you have to buy. The other needs questions worth asking. If the frontier gets redrawn, it won't be by stacking more layers - it'll be by rethinking the ones we have.
4 papers were presented -
1. LLaDA - Large Language Diffusion Model
2. TwoTower - new architecture for Diffusion Language Models
3. CLEGR - new benchmark for Graph-Language Models
4. Dognosis - cancer detection via Canine Olfaction
Thanks for showing up and asking questions - My personal takeaway was that there are more people working in diffusion language models domain than I originally thought. And that's the whole point of running a paper club.
If you are one of them, let's connect.
spent the day at the bangalore papers club. just going through graph llms, graphrag, llada, causal structures, v-jepa, sensory augmentation, bayesian fusion, even how to do multicancer detection using doggos.
@c_engines@dognosis@sohampetkar