๐ฐ Your monthly AI bill keeps increasing.
๐ข Response times are getting slower.
But you have no idea what's causing it.
Is it the embedding model?
Is it retrieval?
Is it the router?
Is it the final generation step?
Or is it all of them?
Without proper observability, you're essentially flying blind.
As I continue building my AI Product Catalog Assistant, I recently integrated Langfuse to get end-to-end visibility across the entire RAG pipeline.
Now every query is traced from start to finish.
๐ Query routing latency & token usage
๐ Embedding generation costs
๐ Qdrant retrieval timings
๐ Retrieved chunks & similarity scores
๐ Final RAG prompts with injected context
๐ Input/output token consumption
๐ Cost and latency across every stage
The biggest benefit isn't the dashboard.
It's being able to answer questions like:
๐ Why did this response take 20+ seconds?
๐ Which step consumed the most tokens?
๐ Did retrieval return the wrong chunks?
๐ Why did the model generate this answer?
๐ Which part of the pipeline is driving up costs?
For the first time, I can see exactly what's happening behind every response instead of treating the system as a black box.
Building RAG systems becomes much easier when every stage is observable.
Current stack:
โข Langfuse
โข Gemini 2.5 Flash
โข Qdrant
โข Node.js
โข Docker
โข TypeScript
Next up:
โก Query Cost Analytics
โก Retrieval Quality Evaluation
โก Semantic Caching
Building in public and learning a lot along the way ๐
Follow @anurag19_dev for more updates.
Follow @aiwithaish@Krishnaik06@piyushgarg_dev@Hiteshdotcom for AI-related content.
#Langfuse #GenerativeAI #RAG #Observability #AIEngineering #LLMOps #BuildingInPublic
๐ฐ Your monthly AI bill keeps increasing.
๐ข Response times are getting slower.
But you have no idea what's causing it.
Is it the embedding model?
Is it retrieval?
Is it the router?
Is it the final generation step?
Or is it all of them?
Without proper observability, you're essentially flying blind.
As I continue building my AI Product Catalog Assistant, I recently integrated Langfuse to get end-to-end visibility across the entire RAG pipeline.
Now every query is traced from start to finish.
๐ Query routing latency & token usage
๐ Embedding generation costs
๐ Qdrant retrieval timings
๐ Retrieved chunks & similarity scores
๐ Final RAG prompts with injected context
๐ Input/output token consumption
๐ Cost and latency across every stage
The biggest benefit isn't the dashboard.
It's being able to answer questions like:
๐ Why did this response take 20+ seconds?
๐ Which step consumed the most tokens?
๐ Did retrieval return the wrong chunks?
๐ Why did the model generate this answer?
๐ Which part of the pipeline is driving up costs?
For the first time, I can see exactly what's happening behind every response instead of treating the system as a black box.
Building RAG systems becomes much easier when every stage is observable.
Current stack:
โข Langfuse
โข Gemini 2.5 Flash
โข Qdrant
โข Node.js
โข Docker
โข TypeScript
Next up:
โก Query Cost Analytics
โก Retrieval Quality Evaluation
โก Semantic Caching
Building in public and learning a lot along the way ๐
Follow @anurag19_dev for more updates.
Follow @aiwithaish@Krishnaik06@piyushgarg_dev@Hiteshdotcom for AI-related content.
#Langfuse #GenerativeAI #RAG #Observability #AIEngineering #LLMOps #BuildingInPublic
๐ค How do you prevent a RAG system from searching the wrong knowledge base?
As my AI Product Catalog Assistant started growing, a single vector collection was no longer enough.
Here's how I solved it ๐
Initially everything lived in one Qdrant collection:
โข Product Catalogs
โข Stone Catalogs
โข Other Documents
This worked at first.
But as more domains were added, retrieval quality became harder to control.
๐ง The solution: Intent-Based Query Routing
Every query now passes through Gemini 2.5 Flash before retrieval.
The router identifies user intent and returns structured output using Zod schemas.
Based on the detected intent, the query is routed to the correct Qdrant collection:
โ stones-docs
โ generic-docs
โ future domain collections
Only the relevant knowledge base is searched.
Benefits:
โ Better retrieval precision
โ Cleaner context
โ Fewer hallucinations
โ Easier scaling
The challenge isn't just retrieving relevant chunks.
It's retrieving them from the right source.
Building in public ๐
Follow @anurag19_dev for more updates.
GitHub Link: https://t.co/8ZpFVfqlzS
@aiwithaish@Krishnaik06@piyushgarg_dev@Hiteshdotcom Follow these accounts for learning resources.
#RAG #GenerativeAI #Qdrant #Gemini #AIEngineering
๐ค How do you prevent a RAG system from searching the wrong knowledge base?
As my AI Product Catalog Assistant started growing, a single vector collection was no longer enough.
Here's how I solved it ๐
Initially everything lived in one Qdrant collection:
โข Product Catalogs
โข Stone Catalogs
โข Other Documents
This worked at first.
But as more domains were added, retrieval quality became harder to control.
๐ง The solution: Intent-Based Query Routing
Every query now passes through Gemini 2.5 Flash before retrieval.
The router identifies user intent and returns structured output using Zod schemas.
Based on the detected intent, the query is routed to the correct Qdrant collection:
โ stones-docs
โ generic-docs
โ future domain collections
Only the relevant knowledge base is searched.
Benefits:
โ Better retrieval precision
โ Cleaner context
โ Fewer hallucinations
โ Easier scaling
The challenge isn't just retrieving relevant chunks.
It's retrieving them from the right source.
Building in public ๐
Follow @anurag19_dev for more updates.
GitHub Link: https://t.co/8ZpFVfqlzS
@aiwithaish@Krishnaik06@piyushgarg_dev@Hiteshdotcom Follow these accounts for learning resources.
#RAG #GenerativeAI #Qdrant #Gemini #AIEngineering
Started learning GenAI recently. The article below covers my key learnings and link to GitHub repo.
I followed content shared by @aiwithaish@Krishnaik06@piyushgarg_dev@Hiteshdotcom to improve my understanding in this space.
Started learning GenAI recently. The article below covers my key learnings and link to the GitHub repo.
I followed content shared by @aiwithaish@Krishnaik06@piyushgarg_dev@Hiteshdotcom to improve my understanding in this space.
Started learning GenAI recently. The article below covers my key learnings and the GitHub repo of the project.
I followed content shared by @aiwithaish@Krishnaik06@piyushgarg_dev@Hiteshdotcom to improve my understanding in this space.
๐๐ผ๐ผ๐ด๐น๐ฒ ๐ฆ๐ฒ๐ฎ๐ฟ๐ฐ๐ต just got its biggest upgrade in 25 years. ๐
At ๐๐ผ๐ผ๐ด๐น๐ฒ ๐/๐ข 2026, Google announced a massive shift into the era of ๐๐-๐ฝ๐ผ๐๐ฒ๐ฟ๐ฒ๐ฑ ๐ฆ๐ฒ๐ฎ๐ฟ๐ฐ๐ต ๐ฎ๐ด๐ฒ๐ป๐๐ and ๐ด๐ฒ๐ป๐ฒ๐ฟ๐ฎ๐๐ถ๐๐ฒ ๐จ๐. Powered globally by the new ๐๐ฒ๐บ๐ถ๐ป๐ถ 3.5 ๐๐น๐ฎ๐๐ต model, Google is transforming from a traditional query engine into an active ๐๐ผ๐ฟ๐ธ๐ณ๐น๐ผ๐ ๐ฒ๐ ๐ฒ๐ฐ๐๐๐ผ๐ฟ.
Here are the key structural changes coming to enterprise, tech, and everyday search:
๐ ๐๐ป๐ณ๐ผ๐ฟ๐บ๐ฎ๐๐ถ๐ผ๐ป ๐๐ด๐ฒ๐ป๐๐ (24/7)
Autonomous background agents that continuously scan the web (finance, sports, shopping) to notify users of complex updates like real-time real estate or product drops.
๐ ๐๐ฒ๐ป๐ฒ๐ฟ๐ฎ๐๐ถ๐๐ฒ ๐จ๐ (๐๐ป๐๐ถ๐ด๐ฟ๐ฎ๐๐ถ๐๐)
Search now codes custom layouts, interactive visuals, and simulations on the fly to explain complex data visually.
๐ ๐๐ด๐ฒ๐ป๐๐ถ๐ฐ ๐๐ผ๐ผ๐ธ๐ถ๐ป๐ด & ๐๐ผ๐ป๐ฐ๐ถ๐ฒ๐ฟ๐ด๐ฒ
Search can now find, map, and link availability for complex local services, and even call local businesses (home repair, pet care) on your behalf.
Google just announced the biggest upgrade to Search in over 25 years ๐จ
- Gemini 3.5 Flash is now the default model powering AI Mode
- New search box designed for longer, natural language queries
- Upload documents, photos, videos and even Chrome tabs directly into Search
- Smart suggestions help you build more complex questions
- AI Overviews and AI Mode now work together more seamlessly
- New agents can monitor topics and bring you updates later
- Search can book services and even call businesses on your behalf
- Custom widgets, tools and mini apps are coming to Search
- More interactive AI-powered experiences are coming directly to results
This is CRAZY!
Nano Banana for Videos!!
Edit through natural conversation
Think of Gemini Omni like Nano Banana โ but for video. Build and fine-tune your creation at any step with natural language.
Introducing Gemini Omni ๐ฎ........ Omni is our new model that can create anything from any input โ starting with video (think Nano Banana but for video). Available in the Gemini App, Flow, and YouTube, with API support coming soon!
Introducing Gemini Omni ๐ฎ........ Omni is our new model that can create anything from any input โ starting with video (think Nano Banana but for video). Available in the Gemini App, Flow, and YouTube, with API support coming soon!