Our finance team needed to search thousands of contracts.
36 hours later, we had a production app. Now we've made it into ๐ฎ ๐๐ถ๐ป๐ด๐น๐ฒ ๐ฝ๐ฟ๐ผ๐บ๐ฝ๐.
Legal research is brutally complex by design. Finding a specific clause in thousands of contracts requires extreme precision - you can't just retrieve semantically similar text from 2022 when the user asked about 2024 agreements.
This is exactly why legal RAG is so hard to get right, and why most teams spend months building custom query logic, managing conversation state, and orchestrating retrievers.
We decided to see if we could break that timeline entirely.
When our finance team asked us to help build them an app to navigate internal contracts, we used the ๐ค๐๐ฒ๐ฟ๐ ๐๐ด๐ฒ๐ป๐ to build a production-ready application in just 36 hours.
Then, with our release of ๐ช๐ฒ๐ฎ๐๐ถ๐ฎ๐๐ฒ ๐๐ด๐ฒ๐ป๐ ๐ฆ๐ธ๐ถ๐น๐น๐, we made it possible to build in just ๐ฐ๐ฏ๐ฆ ๐ฑ๐ณ๐ฐ๐ฎ๐ฑ๐ต.
This was ๐ป๐ผ๐ just another simple chatbot demo. We used advanced embedding strategies, agentic search, and built a full-featured frontend for it with streaming and cited sources, and something we could ๐ข๐ค๐ต๐ถ๐ข๐ญ๐ญ๐บ use internally.
๐ง๐ต๐ฒ ๐ฎ๐ฟ๐ฐ๐ต๐ถ๐๐ฒ๐ฐ๐๐๐ฟ๐ฒ:
Instead of running OCR and chunking text, we used ColQwen to encode each PDF page directly as visual tokens (image patches), producing multivector representations that preserve layout and tables within the PDFs. Muvera compression then reduces memory and latency while keeping retrieval quality high.
Rather than dumping all contracts into a single collection, we split them into three: Commercial Agreements, Corporate & IP Agreements, and Operational Agreements. This schema gives the Query Agent explicit structure to route questions to the most relevant documents.
Results are also streamed back with cited source passages from underlying contracts to give users transparency and reduce hallucinations.
With ๐ช๐ฒ๐ฎ๐๐ถ๐ฎ๐๐ฒ ๐๐ด๐ฒ๐ป๐ ๐ฆ๐ธ๐ถ๐น๐น๐ installed in Claude Code or Cursor, you can literally build this entire application with a single prompt. We have an in-depth prompt included in the blog post to overcome some common mistakes we saw the agent making, and from my tests, it only takes about 12 minutes and roughly 20k tokens on Sonnet 4.6 to build.
The coding agent handles:
โข Downloading and processing the dataset
โข Creating the three collections with proper schema
โข Embedding PDFs with the multimodal model
โข Setting up the Query Agent with ask mode
โข Building the frontend interface with source display
โข Configuring async client and dependency injection
Check out the blog post: https://t.co/rNAwgkTfBt
The code and full implementation guide are in the blog post. Also, you could adapt this same approach to any document-heavy domain that requires precision: compliance, medical records, technical documentation, financial analysis, etc.
Most teams build RAG systems that only see text.
But a lot of queries ๐ณ๐ฆ๐ฒ๐ถ๐ช๐ณ๐ฆ visual understanding to answer correctly.
Most RAG systems treat PDFs the same way: OCR the text, chunk it, embed it, done. But that approach misses ๐ณ๐ถ๐ด๐๐ฟ๐ฒ๐, ๐๐ฎ๐ฏ๐น๐ฒ๐, ๐๐ฝ๐ฎ๐๐ถ๐ฎ๐น ๐น๐ฎ๐๐ผ๐๐, ๐ฎ๐ป๐ฑ ๐๐ถ๐๐๐ฎ๐น ๐ฟ๐ฒ๐น๐ฎ๐๐ถ๐ผ๐ป๐๐ต๐ถ๐ฝ๐ that are a core part of the data.
My colleagues at @weaviate_io just published IRPAPERS, a benchmark that directly compares text-based vs. image-based retrieval over 3,230 pages from 166 scientific papers.
The setup is straightforward: take the same PDFs and process them two ways. For text-based retrieval, run OCR with GPT-4.1, then embed with Arctic 2.0 + BM25 hybrid search. For image-based retrieval, embed the raw page images with ColModernVBERT multi-vector embeddings. Then test both on 180 needle-in-haystack questions targeting specific methodological details.
Text-based retrieval edges out images at the top rank (46% vs 43% Recall@1), but images match or exceed text at deeper recall levels (93% vs 91% Recall@20).
But these two approaches fail on ๐ฅ๐ช๐ง๐ง๐ฆ๐ณ๐ฆ๐ฏ๐ต ๐ฒ๐ถ๐ฆ๐ณ๐ช๐ฆ๐ด.
At Recall@1:
โข 22 queries succeed with text but fail with images
โข 18 queries succeed with images but fail with text
This complementarity is what makes ๐ ๐๐น๐๐ถ๐บ๐ผ๐ฑ๐ฎ๐น ๐๐๐ฏ๐ฟ๐ถ๐ฑ ๐ฆ๐ฒ๐ฎ๐ฟ๐ฐ๐ต so effective. By fusing scores from both text and image retrieval, they achieved 49% Recall@1 and 95% Recall@20 - beating either modality alone.
๐ง๐ต๐ถ๐ ๐ถ๐๐ป'๐ ๐ฎ ๐๐ถ๐บ๐ฝ๐น๐ฒ "๐ถ๐บ๐ฎ๐ด๐ฒ๐ ๐๐ ๐๐ฒ๐ ๐" ๐๐๐ผ๐ฟ๐. Text-based retrieval provides stronger precision and works for most content. But image-based retrieval can use visual structure and handle abstract visualizations way better. This is why once again (multimodal) hybrid search beats out either option alone - getting you the best of both worlds.
The most promising direction seems like it might be agentic systems that dynamically weight text vs image signals based on query characteristics, to emphasize image retrieval when more visually grounded information is needed, and text when you need keyword precision.
Paper: ๐ฎ๐ฟ๐ ๐ถ๐.๐ผ๐ฟ๐ด/๐ฎ๐ฏ๐/๐ฎ๐ฒ๐ฌ๐ฎ.๐ญ๐ณ๐ฒ๐ด๐ณ
Attended @VectorInstโs Remarkable 2026 on Feb 20th ๐ฅ
Two days of cutting edge AI research, industry applications and workshop sessions with breakout discussions
Thread ๐งต ๐๐ฝ
#Remarkable2026#AgenticAI#AICanada#LLMs
The breakout discussion with AI engineers & managers was highlight - internal knowledge base retrieval, orchestration challenges and how teams are solving them in production.
#Remarkable2026#AgenticAI#AICanada#LLMs
My first blog and first blog of 2026 โจ
Please read and share feedback to help me improve my technical writing skills.
๐ https://t.co/rxI5qHI1Rb
#MachineLearning#AIEngineering#Writing
> be Terence Tao
> born in Australia
> starts doing math at age 2
> taking university-level classes by age 9
> training with top math students by 10
> wins IMO bronze at 13
> silver at 14
> gold with a perfect score at 15
> youngest person in history to achieve that
> completes bachelorโs and masterโs early
> begins PhD at Princeton at 16
> becomes a UCLA professor in early 20s
> receives the Fields Medal in 2006
> makes major contributions in number theory, harmonic analysis, PDEs, and combinatorics
> co-proves the GreenโTao theorem: primes contain infinitely long arithmetic progressions
> continues teaching, researching, and writing
> considered one of the greatest mathematicians of the modern era
Leaving Meta and PyTorch
I'm stepping down from PyTorch and leaving Meta on November 17th.
tl;dr: Didn't want to be doing PyTorch forever, seemed like the perfect time to transition right after I got back from a long leave and the project built itself around me.
Eleven years at Meta. Nearly all my professional life. Making many friends for life. Almost eight years leading PyTorch, taking it from nothing to 90%+ adoption in AI. Walking away from this was one of the hardest things I've ever done. But I'm leaving with a full heart.
PyTorch handles exascale training now. It powers foundation models that are redefining intelligence. It's in production at virtually every major AI company. It's taught in classrooms from MIT to rural India. The tools I dreamed about making accessible? They are. The barrier to entry I wanted to lower? It's almost gone.
To be clear, thereโs so much more to do. As long as AI evolves at a breakneck pace, PyTorch will continue to play catch up. Obsessing over the yet-to-come sometimes makes us forget how much weโve already done.
To everyone who built this with meโwho believed research should be joyful, that tools should be elegant, that open source changes everythingโthank you. This wasn't my journey. It was ours.
What's next for me? Something small. Something new. Something I don't fully understand yet. Something uncomfortable. I could have moved to something else inside Meta. But I needed to know what's out there. I needed to do something small again. I couldn't live with the counterfactual regret of never trying something outside Meta.
It's very hard to leave. I probably have one of the AI industryโs most leveraged seats, I lead the software layer that powers the entire AI industry. Every major AI company and hardware vendor are on a speed dial. This kind of power is really hard to give up. But curiosity ultimately won out in my head.
Keep making AI delicious and accessible. I'll be watching. Probably filing issues. Definitely staying involved.
Is PyTorch going to be okay?
I don't want to be doing PyTorch forever. I don't want to be like Guido or Linusโ bound to a single thing for decades. Last November, coinciding with the birth of my daughter, I started planning my exit with Aparna. My goal was to leave PyTorch in a good and stable place.
By this August, during the second half of my parental leave, I knew: Edward, Suo, Alban, Greg, John, Joe and Jana were ready. The team faced hard people, product, technical and organizational problems and didnโt feel the need to lean back on me to solve these for them (unlike in the past). The product story they crafted for the PyTorch Conference was coherentโreally coherent. The things I'd flagged red were turning healthy. The project didn't need me anymore. Unlike 2020-2022 (when I stepped down to go do robotics and came back when Lin, Dima and Dwarak left), I have strong confidence that this time PyTorch is truly resilient. The most aligned culture carriers of PyTorch โ Greg, Alban, Ed, Jason and Joe are at the decision table now, and people with strong value alignment โ Suo, John and Jana have joined them at the table. And thereโs a long list of equally value-aligned people willing to sit at the table should any of these people leave. There are many little things that make up my confidence on the people โ John worked on Julia and open-source for a very long time (in fact we hacked a Torch.jl in 2015), Suo has been the strongest systems builder and strategic partner Iโve had for the past two years, and Jana worked on resilient core systems for a very long time, Iโve had long technical and organizational discussions with her over the past few months that give me confidence. And the product lineup and execution in 2025 should be sufficient evidence for any remaining doubt.
Iโm confident that this band of PyTorchers are going to do exceptionally well. PyTorch might change in flavor because I no longer impose my own taste from the top, but Iโm confident that the values are going to stay intact and the product is going to be awesome.
My time at Meta
The early years of FAIR were absolutely magical. I was part of a small family of absolutely brilliant people building state-of-the-art AI out in the open. From working on GANs with Emily Denton, Rob Fergus, Leon Bottou, Martin Arjovsky and the (now legendary) Alec Radford to building Starcraft bots with Gabriel Synnaeve, to building the first FAIR Cluster with Howard Mansell, to working on object detection with Adam Lerer and Piotr Dollar, to building PyTorch. It was more fun than I can describe in words. 2015 and 2016 were probably the most productive and professionally enjoyable years of my life. Iโll probably romanticize this period of my life forever.
When I joined FAIR, I had massive impostor syndrome, and the first 3 months were very very difficult. I canโt credit Andrew Tulloch enough for being the most thoughtful, kind and welcoming mentor, without whom I wouldnโt have made it. Iโm so damn bullish for Meta just from the fact that heโs back.
---
My time on PyTorch was special.
I loved every part of building itโdesigning it, managing it, being the PM, TL, comms lead, doc engineer, release engineer, squashing bugs, growth hacking, turning it into a coherent product with hundreds of people, transitioning it to industry stakeholdership โ the whole nine yards.
To the core PyTorch team at Meta: the engineers, researchers, open-source maintainers, docs writers, CI infrastructure folks, hardware partners, the community builders. To the hundreds more inside and outside Metaโthank you. You turned a library into a movement.
There are too many people to credit and thank, but I can't not mention Adam Paszke, Sam Gross, Greg Chanan, Joe Spisak, Alban Desmaison, Edward Yang, Richard Zou, Tongzhou Wang, Francisco Massa, Luca Antiga, Andreas Kรถpf, Zach DeVito, Zeming Lin, Adam Lerer, Howard Mansell and Natalia Gimelshein. And Schrep. They made the launch happen. And so many more people became centrally important later: Lu Fang, Xiaodong Wang, Junjie Bai, Nikita Shulga, Horace He, Mark Saroufim, Jason Ansel, Dmytro Dzhulgakov, Yangqing Jia, Geeta Chauhan, Will Constable, Briah Hirsh, Jane Xu, Mario Lezcano, Piotr Balecki, Yinghai Lu, Less Wright, Andrew Tulloch, Bruce Lin, Woo Kim, Helen Suk, Chris Gottbrath, Peng Wu, Joe Isaacson, Eli Uriegas, Tristan Rice, Yanan Cao, Elias Ellison, Animesh Jain, Peter Noordhuis, Tianyu Liu, Yifu Wang, Lin Qiao and hundreds more. Itโs criminal of me to not take the space to list out everyone else I should be mentioning here. PyTorch is nothing without its people โค๏ธ.
The most joyful moments of building PyTorch was meeting users eager to share their happiness, love and feedback. I remember a grad student coming to me at Neurips 2017, in a slurring emotional voice he said heโd been trying to make progress on his research for 3 years but within 3 months of using PyTorch he made so much progress that he was ready to graduate. That moment made it tangible that what we do matters, a lot, to a lot of people, even if you don't constantly hear from them. I do miss the intimacy of the PyTorch community, with a 300 person conference that felt like an extended family gathering, but I feel thatโs a small price to pay considering the scale of impact PyTorch is truly having today โ yes the Conference is now 3,000 people where market-moving deals get brokered, but itโs helping orders of magnitude more people to do their best AI work. I miss the intimacy, but I'm proud of that growth.
---
To Mark Zuckerberg and Mike Schroepfer, who believed that open-sourcing is fundamentally important and is a sound business strategy. This is so hard to understand for most people within the course of business, but weโve run lock-step on this strategy without ever having to discuss it. Without you two, neither FAIR nor PyTorch wouldโve happened. And those mean so much to me.
To Yann LeCun and Rob Fergus, for building the magical early FAIR that I so revere.
To Aparna Ramani, a leader that I find so rare at Meta in her ability to hold a really high bar for the org, technically brilliant with the span to discuss deep infra systems and industry-strategy within the same conversation and for being an absolute execution-machine! Iโve learned so much from you.
To Santosh, Kaushik, Delia, Oldham and Ben for being so welcoming to Infra. For someone coming over from FAIR with a wildly different culture, you all made me feel at home and made me part of the family, and thank you for that.
To all my managers who've championed me through the PSC video game โ Serkan, Howard, Jerome, Abhijit, Yoram, Joelle, Aparna and Damien โ I owe you a lifetime of drinks.
---
Signing off for now.
โSoumith
NEW: OpenAI gpt-oss models just dropped!
Two open-weight models:
- gpt-oss-120b
- gpt-oss-20b
> Full chain-of-thought available
> custom reasoning effort
> agentic capabilities
Apache 2.0 License
Here is all you need to know: