What is with morons wanting to consistently lach anything to “sovereign ai model” ?
Love the research going on, but all reasonable architectures proposed are more compute hungry/similar to what current frontier llms use?
Diffusion based llms are simply great, they can utilise data for 100+ epoch easily mainly by training at different noise levels. But they take that much more compute to train, a diffusion llm talking roughly 16x more compute than autoregressive models to reach the same likelihood but provide better scaling and drastically more speed.
And they aren’t only a headache in training as during inference they provide absurdly high speed but also kick the batching system out of the window. This reduces total throughput of the model when it’s being provided on an api and such.
Two towers is essentially just same as any block diffusion, but they gain a decent speed boost. Nice but not really on the efficiency frontier.
I don’t know much about graph llms, but they are usually attached to a normal large language model and u would have to train it anyways. Again no improvement in efficiency, and ur compute cost stands. Not to mention the paper here is a benchmark which is more disconnected with Indian frontier model.
None of the research listed bypasses the compute deficit or counters the hardware deficit. Rethinking the layers still requires scaling them and training on trillions of tokens. Idk why or how this connects to Indian frontier model. There is no reason for u to connect it towards that or act like these just won’t require GPUs when u “redraw the frontier”.
"India cannot build Frontier Language Models"
Bangalore Paper Club edition 2 happened Last night - The theme was alternate architectures for language models. Why?
Scaling is a compute game. Architecture is an ideas game.
One needs GPUs you have to buy. The other needs questions worth asking. If the frontier gets redrawn, it won't be by stacking more layers - it'll be by rethinking the ones we have.
4 papers were presented -
1. LLaDA - Large Language Diffusion Model
2. TwoTower - new architecture for Diffusion Language Models
3. CLEGR - new benchmark for Graph-Language Models
4. Dognosis - cancer detection via Canine Olfaction
Thanks for showing up and asking questions - My personal takeaway was that there are more people working in diffusion language models domain than I originally thought. And that's the whole point of running a paper club.
If you are one of them, let's connect.
@shrav_10 Their is not “compensated well” for MOST overtime’s done in India. You do it off the books or aren’t paid shit for it. Overtime should be optional and paid, you agree to a 9 to 5 during the salary negotiations not to buy someone’s entire life.
I called u a self loathing moron cause YOU ARE a self loathing moron. That was me stating facts and not swearing. I don’t need professional background and shit to see flipkart and jio raised more then moonshot.
Though sadly for u i DO happen to have significant experience in ml and coding, but that is of no consequence here. My points stand on their own merit and facts, not cause of my labels.
You simply refuse to counter the point or see that Indian start up ecosystem can support ai model development.
I will not be arguing with u any further, u have no arguments and can only do whataboutry
Ok u are just fucked in the head, and won’t concede a fucking point. WHAT COUNTRY DOESN’T IMPORT ANYTHING? Even America imports so many things and runs a big deficit. That’s how economies work u moron, u specialise in things u are good at. And it doesn’t even have anything to do with it either.
The companies ARE MAKING 400 billion, now what the country imports or exports doesn’t matter as the companies making that money would be willing to pay for ai services. Cause uk the IT companies don’t import the fucking oil, they have the 400 billion and will buy things to keep themselves competitive.
A) you are wrong, flipkart is already acquired by Walmart and they have closed the startup cycle and made an exit.
B) none of the bullshit u said has anything to with the point, like wtf are u on. It was about Indian companies raising more than moonshot, which they did. Xiaomi building its own cars has nothing to here either, (as if Tata doesn’t do that AND airplane manufacturing) but the point is Indian companies has more capital then chinise company making ai model. IT WAS ABOUT THE MONEY TO COMPETE IN THE AI RACE.
Why does “averaging out” matter, it’s about the nation not the people.
India has already built its own homegrown, bullet train via BEML. It was entirely home made and has nothing to with the Japanese. Thats working in parallel. You are just a self loathing idiot, who would twist facts just to hate.
Poor country with a budget of 640 billion dollars 😭✌️, man I hope I can be that poor one day.
Is building bullet trains, providing 50 percent value to chip design supply chain or making most of the world’s generics a “poor countries” lane?
Jio raised 20 billion dollars in 2020, flipkart raised 3.6 billion too, compared to 3.5 raised by Kimi. What they raised for doesn’t even matter here, the fact is capital exist in the Indian startup ecosystem to fund these. Hell companies like tcs Wipro have 5 6 bill sitting in the bank, and more revenue then profit then companies like Xiaomi who are building decent models so they can compete entirely from internal money too.
India exports around 400 billion dollars in services, one ought to think such exports can easily pay for ai models. Especially cause they would be crucial in things like accounting or code, stuff that lead Indian service exports.
That’s just people being poor, gdp per capita is more of a measure of individual productivity. Indian government got a budget of 640 billion dollars. Just having the scale means a lot
else India would have built its own bullet train, dominated pharmaceuticals manufacturing or have 20 percent of global chip design workforce be in India.
Though all that is mainly from urban India but they still are a technically apt service oriented economy. From chip design to hft and international banking, all that shi is going on in Indian tier 1 cities.
Indian remittances make 135 billion, IT services export make 220 and merchandise export 400 billion so not really. India based on percentage actually sees decently low migration. Around 1.3 percent of Indian live outside India (global average is 3.7) which puts our rank at around 140 and .1 to .3 of 1k people leave the country each year which is again low and ranks us 120th something. We just got hella people, so amount of people living is hella high too.
The fact that u are comparing with us in capital markets and China is manufacturing says a lot. USA certainly has a better startup ecosystem, china certainly manufactures more steel and smartphones then us but we are by far the second in both. Second here means a lot.
Our economy can accommodate two maybe three ai labs, but the companies and capital are just not willing to take the risk.
The biggest issue with Sarvam atm is that they don’t aim to do genuine research or compete with global giants. They aim to make “India’s own ai” and “get good at Indian languages.
The models being good at Indian languages and adept at Indian situation is bare minimum and something we should be taking for granted. Chinese labs are good at Chinese languages, yet you can read all their technical reports and they don’t mention it once, since it’s assumed to be obvious. But Sarvam thinks of such as a goal, something that allows them to not compare to global models on benchmarks and then applaud themselves for that.
Their model Saaras v3, gets SOTA on “Global English”, performing similar to what they do on US/UK English at 6. something WER. But an Indian model of course should perform same on accented English, that should be given. Compared to world leaders currently hitting 1.3 WER, it leaves much to be desired.
Just cause u are Indian doesn’t mean u get to applaud being 2 years behind as some achievement. We don’t need an “Indian model” for the sake of it, we need a frontier model from India.
We especially saw it in their 100B parameters model, where it was absurdly behind but lauded just for being Indian.
Credit where it’s due, their saaras v3 Multi speaker is in my opinion genuinely good and competitive with the frontier. They are usable and might see some adoption world wide.
And saaras v3’s code switching is also frontier quality. The coding agent, albeit being a codex fork and doing benchmarks graph murders still has several new changes and atleast attempts to try new things/touch researchy subjects. I would wait for credible benchmarks though, theirs showed Gemini 3 pro a few points behind opus. But it’s still a step in the right direction.
Hopefully their 1 trillion parameters model actually competes with opus 5 and such, and doesn’t just get left as “Indian model”. Kimi, deepseek, qwen compete directly instead of just going “Chinese model so no need to compare with American ones”. We need to do the same.
@cheatyyyy Open Source is meant to be forked, and built upon.. that’s why it exists. Though u would have issues if they just changed the name and called it a day.
Rudimentary solution is creating a tool search, and let the ai search for the tool it needs. Claude actually uses this and allows up to 10k tools, with the ability to fetch 3 to 5 tools at once. This is drastically more efficient than pasting the tools directly since even 50 could consume 10k to 20k tokens. It works okayish and is the only solution if u don’t have access to original model weights.
However there is a flaw, as the agent would not have been trained on all the tools there and won’t even know the existence of some tools. This can lead to inefficient use, and several tools remaining practically unused. But if u have access to the original weights, then u can actually train the model using deepagent. This is a new rl framework to teach models to uses 100s of new unique tools very efficiently. https://t.co/nK8eJnbvYR the research paper that introduces it is great and also makes the ai good at long horizon tasks. Though do note that llms get saturated on tool calls rl in few thousand samples, so must be very deliberate on how u train it.
Mandating domestic models means that ur entire service sector is now crippled. Savram gets a big ass pool of customers, but the customer have to use that to compete with morons using fable 5. All while Indian “ai companies” don’t even have any incentive to reach that level or compete with international players.
We will get forced to use half ass ai, and loose the mnc jobs as well as our service sector dominance.
As Andrew ng said, people have the right to keep their code closed source, but those who want to make it open source should be allowed to do so too.
Anthropic is against open weights and open source models in general, they don’t want the option to be there in the first place. NVIDIA and micro soft even though, they don’t have everything as open source they still contribute significantly.
NVIDIA has put huge open ml reasearch and pushed true open source models with code data and everything. From Trenton to Linux kernel drivers they have pushed decent bit for open source community.
Microsoft has realised vscode, and they have also realised a lot of open weights model and very good research papers.
Julian is twisting these companies support, for people being “able to” have open source models into needing to push everything they make money from into open source. Nobody is asking Claude to open source fable 5, people just want them to stop lobbying to man open models.