Actually it's the opposite that's needed. We need to translate all engineering and science books into all major Indian languages. Students should be able to learn in their own language.
First result from my teams research taking help from Claude for quick experimentation has dropped. We have been able to create a system that is 3% lesser accurate but 7 times smaller in size for an ASR system. So we have a usable ASR system that is less than 50MB in size. More importantly we have found an approach to distill bigger models. Will clean up and write a paper on it. Will share the approach also tomorrow once the remaining experiments land.
The core of KooKoo, India's first cloud telephony platform was built in Perl. This is the heart that powers whole of Ozonetel. For 18 years we didnt dare to touch it. But the heart was becoming weak. Stressing out unable to bear the load. We had to deploy more servers to manage the load.
So when coding with AI became possible I gave my team a challenge. This was around 6 months back. Replace the heart(KooKoo) with an updated system. We did some research and decided to go with Go language. None of us were experts in Go. Hell, we did not even know the basics of Go. But hey, we have AI on our side right :).
So after 6 months yesterday, we replaced the heart in one of our clusters. And the transplant was a success. 5 times more efficiency, that means 5 times less servers. Huge gains.
By next month all of this will be ported to the new code.
For me, AI has to give clear gains when we use it. Congrats to the team.
This is the future.
Once we had a good enough tokenizer(we still need to do lots of research in this. If any one is interested, please connect, IIIT-H profs are also interested), the next step is to build the pre training model. This is where we start to get into "intelligence". The model we get out of this will start spitting out grammatically correct Telugu, though most of it wont make sense :)
More details about how we did the pretraining below.
https://t.co/DOMjdiNpU3
On one hand there is hype. On the other there is value. For the Meeseva bot we built along with Meta and the Telangana government, the hype is that its an AI enabled bot using open source AI(Llama). I have no doubt that its one of the largest deployments of open source AI in terms of the services supported as well as people impacted. Thanks to @sunil_abraham for starting this thread with us.
But the value is in the integrations and delivery. More than 300 services integrated. Using WhatsApp as a delivery channel. The value of AI here is to provide natural language queries and a generic framework that can adapt to new services without lots of code changes. Will do an engineering blog sometime later on what we did to pull of the scale of queries etc on hosted Llama models and what we are doing to optimize the models.
AI in India, especially in the academic field will take a long time to make an impact, if at all it will make an impact. Most of the work we do, I will never publish in a paper. The reasons are much better covered by @cneuralnetwork in their post. Most probably we will release the work in a blog post and maybe some papers can come out of it.
But it is hard work. Even students just want recognition with no hard work. We are attempting something at Viswam, but getting students inspired has been an uphill battle. We have introduced AI to some 20,000 students. Data collection effort fell short of what we were planning. But the process to create a Telugu LLM from scratch has started. From datasets in public domain and clearly marked. Very early stages, it's moving a little slow because it's all volunteer driven.
Why should customers search for a phone number or fill forms when they’re ready to engage right now?
With CXi Switch, conversations begin instantly – right at the peak of customer interest.
Watch the demo here: https://t.co/skTdWoUVUn
Lots of VCs funding AI companies and many startup founders moving to the US to build AI systems. But I believe there is a lot of potential for building AI for India. Everything is up for grabs.
This paper just exposed the biggest AI research scam 💀
MIT just proved AI can generate novel research papers.
Stanford confirmed it. OpenAI showcased examples. the papers passed peer review at major conferences. scored higher than human-written work on novelty and feasibility.
major AI labs started citing these as evidence that autonomous research agents are here. that LLMs can actually do science now.
except... they didn't prove that at all.
researchers at Indian Institute of Science ran the exact same AI systems - same prompts, same models, same pipeline. generated 50 research documents using Claude and GPT-4o.
but they changed one thing in how they evaluated them.
previous studies asked experts: "rate this on novelty and feasibility." experts looked at shuffled papers - some human, some AI - and judged them blind. no reason to suspect plagiarism. just scoring ideas.
this study asked: "find what this plagiarized from."
they told 13 domain experts to presume plagiarism exists. go hunting for it. find the source papers.
different question. nuclear results.
24% plagiarized. scores of 4 or 5 on a 5-point scale. verified by contacting the original paper authors.
not sloppy copy-paste that any undergrad could spot. sophisticated methodological rewording that fooled everyone... expert reviewers who literally work in these subfields, conference peer reviewers, academic integrity officers.
every automated plagiarism detector failed. Turnitin? 0% detection rate. OpenScholar with its 45 million paper database? 0%. the Semantic Scholar RAG systems these AI agents use internally to "check their own work" for plagiarism before publishing? caught 51% in the easiest possible test scenario where proposals were deliberately plagiarized from single papers.
in real-world generation where the AI is trying to be novel? way worse.
the exemplar papers everyone's been citing as proof AI can do real science?
one had perfect 1-to-1 mapping with "Generating with Confidence: Uncertainty Quantification for Black-box LLMs" published in 2023.
each component of the "novel" methodology corresponded exactly to sections in the original paper. just skillfully reworded.
"resonance graph" instead of "weighted adjacency matrix."
"semantic resonance uncertainty quantification" instead of "uncertainty quantification."
"pairwise evaluations for consistency" instead of "pairwise similarity scores."
five steps. five direct correspondences. same methodology. same scientific contribution. same insight.
zero attribution. zero citations.
the original authors (Lin et al.) confirmed the plagiarism after reviewing both documents.
this paper was showcased as an exemplar of AI-generated research. it passed through expert review in the original study. nobody caught it.
another exemplar combined two papers without credit - one on diffusion model gating mechanisms, another on multi-resolution training. repackaged as "DualDiff." authors of the source papers confirmed: definitively plagiarized.
these aren't edge cases.
human-written papers from major conferences? plagiarism rate around 2-6% based on peer review comments.
AI-generated proposals? 24%.
and this assumes the experts found everything. the authors explicitly say this is likely a lower bound because finding plagiarism is incredibly labor-intensive.
the really disturbing part?
the AI-generated proposals are less diverse than human work. they cluster together in embedding space. you can train a basic classifier with 93% accuracy to detect them just from titles and abstracts.
which means these systems aren't exploring novel research directions. they're pattern-matching within a narrow band of what "sounds like research" and skillfully remixing existing papers.
we built systems that repackage existing ideas so well, we convinced ourselves - and expert reviewers - they were breakthroughs.
YouTube has rolled out likeness detection for creators, allowing them to monitor where their facial likeness appears on the platform, and request removal of unauthorised or AI altered videos.
Most importantly, YouTube now treats “likeness” as a privacy issue, covering both video and voice. A few points:
1. At one level, this is a necessary step. Deepfakes are everywhere, and creators have little control over their own image.
People are rightfully worried about their likeness being used for scams, fake endorsements, or worse.
But how YouTube is implementing this protection raises fundamental questions.
2. On the Generative AI side, a few days ago, Breaking Bad star Bryan Cranston thanked OpenAI for preventing the creation of his deepfakes using Sora. AI companies are inconsistent: an ideal world, no one should have to go to court or rely on corporate goodwill to prevent AI misuse of their likeness.
3. Youtube is addressing this problem on the distribution side, and protecting only its creators. This mirrors how YouTube’s Content ID system works: copyright owners only get automated protection if they upload their work to YouTube. Otherwise, they’re on their own. You need to upload your profile to Youtube to get the protection. This creates an odd dynamic: people are effectively being nudged to join YouTube just to defend their identity. Protection shouldn’t be conditional on participation. But that's the way it was with Copyright as well.
4. The second problem is privacy aspect of image protection: To use likeness detection, creators must surrender the very data they’re trying to protect. The process requires submitting a government ID and a short video of your face. Google then uses this data to create “face templates” to detect AI-altered or unauthorized videos. So to guard against unauthorized facial data use, you must hand over more facial data—along with your government ID. This effectively expands googles biometric detection database… with the objective of privacy protection… go figure. If copyright protection doesn’t require an ID upload, why should privacy protection?
5. Then there’s the gray area of parody and fair use. If someone makes a parody video using an AI likeness of a creator, say, an MKBHD clone reviewing a Maruti 800, how should that be treated?
Under copyright law, parody is typically protected. But under YouTube’s new rules, that same parody could be removed as a “privacy violation.” This creates an uneven standard: human mimicry is acceptable, but AI mimicry may not be.
6. Another unresolved question: do deceased people have likeness or privacy protection?
Robin Williams’ daughter Zelda recently mentioned that she was unhappy about her father's deepfakes being sent to her.
YouTube’s policy even extends likeness protection to voices, but likely only for living creators who’ve registered through the platform. If a deceased person like Robin Williams never submitted their likeness to Google, there’s no mechanism for heirs to protect it. In India, India's largely useless Digital Personal Data Protection Act doesn’t cover deceased persons.
7. There’s the matter of recourse and transparency: YouTube already acts as a private arbiter of copyright, often taking down content beyond what India's laws require, and with minimal accountability. Now it’s extending that authority to questions of identity and likeness - again, without clear oversight. False takedowns, weaponized likeness claims, impersonation disputes are possible…Based on YouTube’s record, recourse tends to be inconsistent and opaque.
Now that YouTube has made this move, it’s worth watching what others do. Instagram will likely follow.
X probably won’t, unless compelled by lawsuits. Each platform will face the same tension: how to balance identity protection with freedom of expression.
Watch this space...
Ozonetel is recognized for AI Innovation in CX at the HT Media Bharat Nirmaan Conclave & Awards 2025!
Our CX platform – a great #MakeInIndia success story – stands out by delivering measurable impact across marketing, sales & service.
Proud to build in Bharat for the world.
In an interview with @EconomicTimes, Chaitanya (@nutanc) shares why AI alone won’t fix CX.
Read more to know how our CXi Studio is helping enterprises design end-to-end journeys where AI agents & humans work in sync – delivering measurable impact: https://t.co/Ir7M31fkeZ
New Launch: CXi Studio is here! A powerful AI orchestration tool that helps you design and automate end-to-end customer journeys – faster, more human, and at scale.
Built in Bharat for businesses everywhere. Because great CX knows no borders.
Follow us to learn more!
We're building the future of CX right here, in Bharat!
At the Bharat Nirman Conclave & Awards 2025, our very own @shalilgupta will unveil CXi Studio: a no-code AI orchestration tool that helps businesses build end-to-end customer journeys, faster & more human!
Stay tuned!
Most home loan inquiries fall through the cracks – delayed responses, generic scripts, and lost leads.
Our CXi Voice Agents are changing that, helping leading banks drive meaningful engagement.
Better conversations mean more conversions. Talk to us today!
Speak at The Fifth Elephant Open Source AI meet-up - because it is prestigious for your work, and to elevate your ideas. This is the Hyderabad version and I encourage all Hyderabad engineers to submit their work. @fifthel has one of the best communities.
🚀 Call for Submissions: Open Source AI Meet-up – Hyderabad, 1 Nov
Theme: AI in Software Lifecycle Development
💡 Talk/demo ideas: AI-assisted requirements, code generation, AI-powered QA, AIOps, secure SDLC, ethics, multi-agent workflows & sustainable AI.
📅 Deadline: 15 Sept | 🔗 https://t.co/BaNZEO0wcX
We’re all set for the 8th Edition of the Smart CX Summit & Awards 2025 in #Dubai!
Join us as we explore the future of #customerexperience – where intelligence, human-centered innovation & #AgenticAI are changing the way businesses engage with people.
See you there!