๐ We're so excited to announce that Milvus 3.0 is available today! This is the biggest architectural release in the project's history, with two big shifts.
๐ ๐ฆ๐๐๐๐ง ๐ญ โ ๐๐ฎ๐ธ๐ฒ-๐ป๐ฎ๐๐ถ๐๐ฒ ๐ถ๐ป๐ณ๐ฟ๐ฎ๐๐๐ฟ๐๐ฐ๐๐๐ฟ๐ฒ
The problem: your vectors already live in the lake, but your search engine needs its own copy. So you build an ETL pipeline โ and then maintain it forever. With Milvus 3.0, you get:
1. External Collections โ index and serve data that stays in Parquet, Lance, Iceberg, or Vortex. No second copy, no sync job.
2. Loon (Storage v3) โ a columnar engine built for the point reads that follow an ANN search. I/O per read went from 9.4 MB to 0.07 MB in our benchmarks.
3. Snapshots + Spark + live schema changes โ dedupe, re-embed, and evaluate your data while the live collection keeps serving.
โก ๐ฆ๐๐๐๐ง ๐ฎ โ ๐ ๐บ๐ผ๐ฟ๐ฒ ๐ฝ๐ผ๐๐ฒ๐ฟ๐ณ๐๐น ๐ฟ๐ฒ๐๐ฟ๐ถ๐ฒ๐๐ฎ๐น ๐ฒ๐ป๐ด๐ถ๐ป๐ฒ
The problem: too much of your retrieval pipeline lives in application code. You over-fetch candidates, then sort, count, and reassemble them yourself. With Milvus 3.0, you get:
1. Server-side ORDER BY, aggregation, and faceted search โ sort by price or freshness, count by category, all inside the engine.
2. StructArray โ one row per document, many vectors inside. ColBERT and ColPali finally have a native home.
3. SINDI + BM25 compression โ a sparse index ~3ร smaller, with up to ~10ร the QPS on learned sparse embeddings.
Put together: one system, one copy of your data, and a lot less glue code.
That matters most if you're building: ๐ค AI agents whose data changes constantly ๐ Multimodal retrieval with late-interaction models ๐ Governed systems where the data has to stay where it is ๐ RAG and knowledge bases over long documents ๐๏ธ Search that mixes relevance with price, rating, or inventory
๐ Huge thanks to the Milvus community โ you filed the issues, tested the release candidates, and told us what was actually painful. This release is shaped by you.
Swipe through for the full breakdown ๐
โญ Star us: https://t.co/3ckzkUVicH
๐ Full launch blog: https://t.co/hq4qqZASuX
#Milvus #VectorDatabase #AIInfrastructure #RAG #OpenSource
๐ ๐ ๐ถ๐น๐๐๐ ๐ฏ.๐ฌ ๐ถ๐ ๐ป๐ผ๐ ๐ฎ๐๐ฎ๐ถ๐น๐ฎ๐ฏ๐น๐ฒ, ๐ฎ๐ป๐ฑ ๐ถ๐โ๐ ๐ฎ ๐บ๐ฎ๐ท๐ผ๐ฟ ๐๐๐ฒ๐ฝ ๐๐ผ๐๐ฎ๐ฟ๐ฑ ๐น๐ฎ๐ธ๐ฒ-๐ป๐ฎ๐๐ถ๐๐ฒ ๐๐ฒ๐ฐ๐๐ผ๐ฟ ๐๐ฒ๐ฎ๐ฟ๐ฐ๐ต.
๐ฅ Key highlights:
โข External Collections let teams build retrieval directly where their data already lives: in object storage, open data formats like Parquet and Vortex, and open table formats like Lance and Iceberg.
โข Loon Storage v3 improves point reads for lake-native retrieval on S3-style object storage.
โข Snapshots, Spark DataSource V2, and live schema changes bring stable collection views into batch workflows, recovery, backfill, and schema evolution.
โข ORDER BY, aggregation, and faceted search move more result processing into the retrieval engine.
โข StructArray and SINDI strengthen multi-vector, sparse, and hybrid retrieval workloads.
For builders working on RAG, agents, multimodal search, and AI data pipelines, Milvus 3.0 reduces duplicate data movement and pushes more retrieval logic into the engine itself.
It is a big step toward a cleaner architecture: lake-native storage, offline data improvement, online serving, and retrieval-time processing working together in one open-source vector database.
โญ Star us: https://t.co/3ckzkUVicH
๐ Full launch blog: https://t.co/hq4qqZASuX
๐ ๐ถ๐น๐๐๐ ๐ฏ.๐ฌ ๐ถ๐ป๐๐ฟ๐ผ๐ฑ๐๐ฐ๐ฒ๐ ๐ฆ๐๐ผ๐ฟ๐ฎ๐ด๐ฒ ๐๐ฏ, ๐ฎ๐น๐๐ผ ๐ธ๐ป๐ผ๐๐ป ๐ฎ๐ ๐๐ผ๐ผ๐ป, ๐ฎ ๐๐๐ผ๐ฟ๐ฎ๐ด๐ฒ ๐ฒ๐ป๐ด๐ถ๐ป๐ฒ ๐ฏ๐๐ถ๐น๐ ๐ณ๐ผ๐ฟ ๐๐ฒ๐ฟ๐๐ถ๐ป๐ด-๐๐๐๐น๐ฒ ๐ฎ๐ฐ๐ฐ๐ฒ๐๐ ๐ผ๐ป ๐ผ๐ฏ๐ท๐ฒ๐ฐ๐ ๐๐๐ผ๐ฟ๐ฎ๐ด๐ฒ.
Loon handles a specific problem: after ANN search returns candidate IDs, Milvus still has to fetch the fields for those rows.
On object storage, that fetch can become expensive. Analytical file formats are optimized for scans, not serving-style point reads. A query may only need a few vectors or metadata fields, but the layout can force the engine to read much more.
๐๐ผ๐ผ๐ป ๐ฟ๐ฒ๐ฑ๐๐ฐ๐ฒ๐ ๐๐ต๐ถ๐ ๐ฟ๐ฒ๐ฎ๐ฑ ๐ฎ๐บ๐ฝ๐น๐ถ๐ณ๐ถ๐ฐ๐ฎ๐๐ถ๐ผ๐ป ๐ฏ๐ ๐ผ๐ฟ๐ด๐ฎ๐ป๐ถ๐๐ถ๐ป๐ด ๐ฑ๐ฎ๐๐ฎ ๐ถ๐ป๐๐ผ ๐๐ผ๐น๐๐บ๐ป๐๐ฟ๐ผ๐๐ฝ๐ ๐๐ถ๐๐ต ๐ฎ๐น๐ถ๐ด๐ป๐ฒ๐ฑ ๐ฟ๐ผ๐ ๐๐๐. Different fields can then use layouts that match how they are actually accessed:
โข Scalar fields can be laid out for filtering.
โข Vectors and point-read-heavy fields can use layouts for narrow lookups.
โข Vector and inverted indexes stay separate from the file format.
โข Each dataset version is tracked by an immutable manifest.
In one internal benchmark, point-read I/O dropped from about 9.4 MB with Parquet to 0.07 MB with Vortex and Loon. This matters because zero-copy search still needs a storage path built for retrieval workloads.
Blog: https://t.co/d4Fgu31EfN
๐ We're so excited to announce that Milvus 3.0 is available today! This is the biggest architectural release in the project's history, with two big shifts.
๐ ๐ฆ๐๐๐๐ง ๐ญ โ ๐๐ฎ๐ธ๐ฒ-๐ป๐ฎ๐๐ถ๐๐ฒ ๐ถ๐ป๐ณ๐ฟ๐ฎ๐๐๐ฟ๐๐ฐ๐๐๐ฟ๐ฒ
The problem: your vectors already live in the lake, but your search engine needs its own copy. So you build an ETL pipeline โ and then maintain it forever. With Milvus 3.0, you get:
1. External Collections โ index and serve data that stays in Parquet, Lance, Iceberg, or Vortex. No second copy, no sync job.
2. Loon (Storage v3) โ a columnar engine built for the point reads that follow an ANN search. I/O per read went from 9.4 MB to 0.07 MB in our benchmarks.
3. Snapshots + Spark + live schema changes โ dedupe, re-embed, and evaluate your data while the live collection keeps serving.
โก ๐ฆ๐๐๐๐ง ๐ฎ โ ๐ ๐บ๐ผ๐ฟ๐ฒ ๐ฝ๐ผ๐๐ฒ๐ฟ๐ณ๐๐น ๐ฟ๐ฒ๐๐ฟ๐ถ๐ฒ๐๐ฎ๐น ๐ฒ๐ป๐ด๐ถ๐ป๐ฒ
The problem: too much of your retrieval pipeline lives in application code. You over-fetch candidates, then sort, count, and reassemble them yourself. With Milvus 3.0, you get:
1. Server-side ORDER BY, aggregation, and faceted search โ sort by price or freshness, count by category, all inside the engine.
2. StructArray โ one row per document, many vectors inside. ColBERT and ColPali finally have a native home.
3. SINDI + BM25 compression โ a sparse index ~3ร smaller, with up to ~10ร the QPS on learned sparse embeddings.
Put together: one system, one copy of your data, and a lot less glue code.
That matters most if you're building: ๐ค AI agents whose data changes constantly ๐ Multimodal retrieval with late-interaction models ๐ Governed systems where the data has to stay where it is ๐ RAG and knowledge bases over long documents ๐๏ธ Search that mixes relevance with price, rating, or inventory
๐ Huge thanks to the Milvus community โ you filed the issues, tested the release candidates, and told us what was actually painful. This release is shaped by you.
Swipe through for the full breakdown ๐
โญ Star us: https://t.co/3ckzkUVicH
๐ Full launch blog: https://t.co/hq4qqZASuX
#Milvus #VectorDatabase #AIInfrastructure #RAG #OpenSource
โญ ๐ ๐ถ๐น๐๐๐ ๐ฏ.๐ฌ ๐ถ๐ป๐๐ฟ๐ผ๐ฑ๐๐ฐ๐ฒ๐ ๐๐ ๐๐ฒ๐ฟ๐ป๐ฎ๐น ๐๐ผ๐น๐น๐ฒ๐ฐ๐๐ถ๐ผ๐ป๐, ๐ฎ ๐๐ฎ๐ ๐๐ผ ๐บ๐ฎ๐ธ๐ฒ ๐น๐ฎ๐ธ๐ฒ-๐ฟ๐ฒ๐๐ถ๐ฑ๐ฒ๐ป๐ ๐๐ฒ๐ฐ๐๐ผ๐ฟ ๐ฑ๐ฎ๐๐ฎ ๐๐ฒ๐ฎ๐ฟ๐ฐ๐ต๐ฎ๐ฏ๐น๐ฒ ๐๐ถ๐๐ต๐ผ๐๐ ๐ฐ๐ผ๐ฝ๐๐ถ๐ป๐ด ๐ถ๐ ๐ถ๐ป๐๐ผ ๐ฎ ๐๐ฒ๐ฟ๐๐ถ๐ป๐ด ๐ฑ๐ฎ๐๐ฎ๐ฏ๐ฎ๐๐ฒ.
Many teams already have embeddings and metadata in object storage: Parquet files in S3, Lance datasets, Iceberg tables, or other lakehouse formats.
Before Milvus 3.0, there were usually two ways to make that data searchable.
๐ข๐ฝ๐๐ถ๐ผ๐ป ๐ผ๐ป๐ฒ: ๐ฐ๐ผ๐ฝ๐ ๐ถ๐ ๐ถ๐ป๐๐ผ ๐ฎ ๐๐ฒ๐ฐ๐๐ผ๐ฟ ๐ฑ๐ฎ๐๐ฎ๐ฏ๐ฎ๐๐ฒ. You get low-latency ANN search, but now you have a second copy and an ETL pipeline to keep in sync.
๐ข๐ฝ๐๐ถ๐ผ๐ป ๐๐๐ผ: ๐พ๐๐ฒ๐ฟ๐ ๐๐ต๐ฒ ๐น๐ฎ๐ธ๐ฒ ๐ฑ๐ถ๐ฟ๐ฒ๐ฐ๐๐น๐. You avoid duplication, but without ANN indexes, vector search turns into a brute-force scan.
๐๐ ๐๐ฒ๐ฟ๐ป๐ฎ๐น ๐๐ผ๐น๐น๐ฒ๐ฐ๐๐ถ๐ผ๐ป๐ ๐ถ๐ป๐๐ฟ๐ผ๐ฑ๐๐ฐ๐ฒ ๐ฎ ๐๐ต๐ถ๐ฟ๐ฑ ๐ฝ๐ฎ๐๐ต.
You keep the data where it is, map external fields into a Milvus schema, and use the same Milvus search and query APIs. Milvus builds vector, BM25 inverted, JSON, and scalar indexes over the lake-resident data. The source files do not move.
For teams where the lake owns permissions and freshness, every extra copy creates sync, access-control, and debugging work.
๐ ๐ณ๐ฒ๐ ๐ฝ๐ฟ๐ฎ๐ฐ๐๐ถ๐ฐ๐ฎ๐น ๐ฑ๐ฒ๐๐ฎ๐ถ๐น๐ ๐บ๐ฎ๐๐๐ฒ๐ฟ:
โข External Collections are read-only and zero-copy.
โข Milvus can index newly added fragments instead of rebuilding the whole collection.
โข Three load modes let teams choose between lower storage cost and lower latency.
Native Milvus collections are better for write-heavy serving. External Collections are for lake datasets that need production search without another copy.
Know the details: https://t.co/NR2QfOORyC
How do you combine structured filters and keyword signals with semantic similarity in ecommerce search?
Take this query:
โpink headphones for kids with ears under $20โ
โPink headphones for kids with earsโ expresses semantic intent. โUnder $20โ is a hard price constraint. Ratings, reviews, brand, category, and visual similarity may also influence the results.
๐ฆ๐ผ ๐ฒ๐ฐ๐ผ๐บ๐บ๐ฒ๐ฟ๐ฐ๐ฒ ๐ฟ๐ฒ๐๐ฟ๐ถ๐ฒ๐๐ฎ๐น ๐ป๐ฒ๐ฒ๐ฑ๐ ๐บ๐ผ๐ฟ๐ฒ ๐๐ต๐ฎ๐ป ๐๐ฒ๐ฐ๐๐ผ๐ฟ ๐๐ฒ๐ฎ๐ฟ๐ฐ๐ต ๐ฎ๐น๐ผ๐ป๐ฒ.
Lumen, a demo built by @Simon Hearne, shows how these signals can work together in a small Amazon-style product search experience.
The app combines:
โข Query understanding to separate product intent from numeric filters
โข Dense vector search for semantic relevance
โข BM25 for exact keyword matching
โข Milvus filters for price, rating, reviews, brand, and category
โข Text and image vectors for โMore like thisโ recommendations
A diagnostics panel also makes the retrieval process visible, including the effective Milvus query, embedding dimension, fusion strategy, and latency.
That transparency matters. Search quality may feel subjective to users, but builders need to identify exactly which signal failed when the results feel wrong.
Lumen keeps the scope intentionally small: one catalog, one search flow, and one clear retrieval stack. Itโs a useful example of the components modern product discovery often needsโhybrid retrieval, metadata filtering, multimodal similarity, and an API layer that keeps database credentials server-side.
Try it: https://t.co/nMg9Lf9bBo
Repo: https://t.co/XYw6BzRVV1
Embedding model selection used to be mostly a leaderboard question.
In 2026, it feels much closer to a data-shape question.
A text knowledge base, a scanned contract, a product screenshot, a financial report, and a video archive contain different retrieval signals. Sending all of them through the same chunking and embedding pipeline can discard the information that matters most.
๐๐ผ๐ฟ ๐ป๐ผ๐ฟ๐บ๐ฎ๐น ๐๐ฒ๐ ๐ ๐๐๐, dense embeddings are still the practical starting point. Qwen3 Embedding, Jina v5 text, BGE-M3, OpenAI text-embedding-3, Voyage, and Cohere are all worth testing against your real chunks.
๐๐ผ๐ฟ ๐ฃ๐๐๐, ๐ถ๐บ๐ฎ๐ด๐ฒ๐, ๐๐ฐ๐ฟ๐ฒ๐ฒ๐ป๐๐ต๐ผ๐๐, ๐ฎ๐๐ฑ๐ถ๐ผ, ๐ฎ๐ป๐ฑ ๐๐ถ๐ฑ๐ฒ๐ผ, models like Gemini Embedding 2, Jina v5 omni, Cohere Embed 4, Qwen3-VL-Embedding, and Voyage Multimodal are making it easier to keep more original context in retrieval.
๐๐ผ๐ฟ ๐น๐ผ๐ป๐ด ๐ฑ๐ผ๐ฐ๐๐บ๐ฒ๐ป๐๐, contextualized chunk embeddings like Voyage context-4 address a painful issue: chunks often lose meaning when separated from the full document.
๐๐ผ๐ฟ ๐ฐ๐ผ๐บ๐ฝ๐น๐ฒ๐ ๐น๐ฎ๐๐ผ๐๐๐, ColPali-style multi-vector retrieval is worth a look, especially when tables, figures, page regions, or scanned details decide the answer.
Benchmarks like MTEB, MMEB, and ViDoRe are useful filters. But the final test still has to be your own documents, your own queries, and your own failure cases.
๐ ๐ถ๐น๐๐๐ ๐ฎ.๐ฒ also moves part of the embedding plumbing into the database layer: with Text Embedding Function, you can insert raw text, let Milvus call the configured embedding provider, store the vectors, and run text queries without managing embedding calls in every client.
We wrote a practical guide on how to choose embedding models for the second half of 2026.
๐ ๐๐น๐๐ถ-๐ต๐ผ๐ฝ ๐ฅ๐๐ ๐๐ถ๐๐ต๐ผ๐๐ ๐ฎ ๐ด๐ฟ๐ฎ๐ฝ๐ต ๐ฑ๐ฎ๐๐ฎ๐ฏ๐ฎ๐๐ฒ
Multi-hop retrieval is often solved by adding a separate graph store. We wanted to see how far we could get with Milvus alone โ so we built an open-source library and found out.
Vector Graph RAG achieves multi-hop reasoning using only Milvus. Neo4j, Cypher queries, and the second system once needed to operate are all things of the past (in this scenario).
Vector Graph RAG is built upon the fact that knowledge graph relations are just text. An example, (metformin, is the first-line drug for, type 2 diabetes) is a directed edge in a graph database โ but it's also a sentence you can embed and store in Milvus, alongside entities and source passages.
๐ฉ๐ฒ๐ฐ๐๐ผ๐ฟ ๐๐ฟ๐ฎ๐ฝ๐ต ๐ฅ๐๐ ๐๐๐ผ๐ฟ๐ฒ๐ ๐ฒ๐๐ฒ๐ฟ๐๐๐ต๐ถ๐ป๐ด ๐ถ๐ป ๐๐ต๐ฟ๐ฒ๐ฒ ๐ ๐ถ๐น๐๐๐ ๐ฐ๐ผ๐น๐น๐ฒ๐ฐ๐๐ถ๐ผ๐ป๐ ๐๐ถ๐๐ต ๐๐ ๐ฐ๐ฟ๐ผ๐๐-๐ฟ๐ฒ๐ณ๐ฒ๐ฟ๐ฒ๐ป๐ฐ๐ฒ๐:โข ๐๐ป๐๐ถ๐๐ถ๐ฒ๐ โ deduplicated nodes, each carrying a list of the relation IDs they participate in
โข ๐ฅ๐ฒ๐น๐ฎ๐๐ถ๐ผ๐ป๐ โ embedded triples pointing to the entity IDs and source passage IDs on each side
โข ๐ฃ๐ฎ๐๐๐ฎ๐ด๐ฒ๐ โ original document chunks with back-references to the extracted entities and relations
Those ID references are the graph structure. Subgraph expansion follows them to surface bridge entities that the question never mentions โ the step that makes multi-hop reasoning work. Then a single LLM reranking pass filters the expanded candidate pool down to what actually answers the question, and one generation call produces the answer from the full source passages.
That pipeline makes ๐ฎ ๐๐๐ ๐ฐ๐ฎ๐น๐น๐ ๐ฝ๐ฒ๐ฟ ๐พ๐๐ฒ๐ฟ๐ (rerank + generate), compared to 3-5 for IRCoT and 5-10+ for Agentic RAG. Against a 5-call iterative baseline, that works out to roughly ๐ฒ๐ฌ% ๐น๐ผ๐๐ฒ๐ฟ ๐๐ฃ๐ ๐ฐ๐ผ๐๐ ๐ฎ๐ป๐ฑ ๐ฎ-๐ฏ๐ ๐ณ๐ฎ๐๐๐ฒ๐ฟ ๐ฟ๐ฒ๐๐ฝ๐ผ๐ป๐๐ฒ๐, with predictable latency instead of spikes when an agent decides to loop again.
On the benchmarks: ๐ด๐ณ.๐ด% ๐ฎ๐๐ฒ๐ฟ๐ฎ๐ด๐ฒ ๐ฅ๐ฒ๐ฐ๐ฎ๐น๐น@๐ฑ across MuSiQue, HotpotQA, and 2WikiMultiHopQA โ against ๐ด๐ณ.๐ญ% for HippoRAG 2, under the same evaluation setup, with no graph database and no ColBERTv2.
๐๐ผ๐ ๐ฑ๐ผ ๐๐ผ๐ ๐ถ๐ป๐๐๐ฎ๐น๐น ๐๐ต๐ถ๐? You only need to type"pip install vector-graph-rag". It defaults to Milvus Lite โ a local .db file, so you don't need to configure anything.
๐๐ผ๐ฟ ๐บ๐ผ๐ฟ๐ฒ ๐ถ๐ป๐ณ๐ผ๐ฟ๐บ๐ฎ๐๐ถ๐ผ๐ป, ๐๐ฒ๐ฒ ๐ถ๐๐ ๐ด๐ถ๐๐ต๐๐ฏ ๐ฝ๐ฎ๐ด๐ฒ: https://t.co/uyBstuGBcq
If your corpus is knowledge-dense โ legal, biomedical, financial โ and your questions routinely cross 2-4 document boundaries, this is the architecture worth testing first.
Semantic search, full-text search, metadata filtering, and hybrid retrieval in one open-source system. Thanks to @DanKornas for sharing this clear overview of Milvus and the use cases it supports.
Need one retrieval layer for semantic search, full-text search, and metadata filtering?
Milvus is an open-source vector database for AI engineers building RAG, semantic search, multimodal search, and recommendation systems.
It helps you organize and search unstructured data by storing vectors alongside scalar fields, then combining vector search with metadata filters, full-text search, or hybrid retrieval.
Key features:
โข Dense and sparse retrieval โ supports semantic vectors, BM25 full-text search, and learned sparse embeddings.
โข Hybrid search โ stores dense and sparse vectors together and can rerank results from multiple searches.
โข Multiple index types โ includes HNSW, IVF, FLAT, SCANN, DiskANN, and quantized variants.
โข Flexible deployment โ runs as Milvus Lite, on one machine in Standalone mode, or in a distributed, Kubernetes-native architecture.
โข Access controls โ includes user authentication, TLS encryption, and role-based access control.
Itโs open-source (Apache 2.0 license).
Link in the reply ๐
@DanKornas Thanks for featuring Milvus, Dan! Really appreciate how clearly you covered the different retrieval options and deployment paths in one overview.
Everyone is talking about agent memory, but the term is starting to hide more than it explains.
For coding agents, it is not just โsave chat history, embed it, retrieve top-kโ anymore.
These open-source projects show that โagent memoryโ can mean very different things:
1. codebase-memory-mcp
https://t.co/3YTSbzwwEp
Memory as codebase structure: functions, classes, call chains, routes, dependencies, and impact paths.
2. MemSearch
https://t.co/OoqhEFvafa
Memory as retrieval infrastructure: Markdown as the source of truth, Milvus as the index, plus hybrid search and progressive recall.
3. agentmemory
https://t.co/DjB7SbAkuQ
Memory as workspace context that follows coding agents across Claude Code, Codex, Cursor, Gemini CLI, OpenCode, and MCP clients.
4. Basic Memory
https://t.co/5uv1wErqm2
Memory as a readable knowledge base: Markdown notes that humans can edit and agents can search.
5. ai-memory
https://t.co/RI1nwVvSZR
Memory as continuity and handoff, so work can move between agent vendors without re-explaining architecture, failed attempts, and open questions.
These projects point to a simple trend: agent memory is moving from โstore more contextโ to โretrieve the right context.โ
For coding agents, memory is becoming less like a notes file and more like a retrieval layer.
When ๐๐ถ๐บ๐ถ ๐๐ฏ came out, we tested it right away. This is our experimental sharing:
Kimi K3 has 2.8T parameters, a 1M-token context window, native visual understanding, and strong coding-agent scores. After a few internal runs, we decided to move part of our workflow from Fable to K3.
The first surprise came from marketing.
One teammate in marketing used K3 to prototype a campaign landing page in one afternoon. Not production-ready, but enough to align the idea before involving engineering.
That was the good part.
The less fun part was the token bill.
Large-context coding agents can spend a lot of tokens just figuring out where things are in a repo. Kimi's pricing makes that visible: cache-hit input is $0.30/MTok, cache-miss input is $3.00/MTok, and output is $15.00/MTok.
So we added ๐ฐ๐น๐ฎ๐๐ฑ๐ฒ-๐ฐ๐ผ๐ป๐๐ฒ๐ ๐.
It gives coding agents semantic code search, so they retrieve relevant files and snippets instead of loading whole directories into the prompt.
In our controlled evaluation, Claude Context MCP reduced tokens by around 40% at equivalent retrieval quality.
K3 made the agent stronger. claude-context made it cheaper to aim that strength at the right code.
GitHub: https://t.co/yyWCfHbgZv
๐๐ผ๐ ๐ฑ๐ผ ๐๐ผ๐ ๐ฐ๐ผ๐บ๐ฏ๐ถ๐ป๐ฒ ๐ฎ๐ด๐ฒ๐ป๐ ๐บ๐ฒ๐บ๐ผ๐ฟ๐ ๐๐ถ๐๐ต ๐๐ธ๐ถ๐น๐น๐? ๐ง๐ฟ๐ ๐ผ๐๐ฟ ๐ผ๐ฝ๐ฒ๐ป-๐๐ผ๐๐ฟ๐ฐ๐ฒ ๐๐ผ๐ผ๐น ๐บ๐ฒ๐บ๐๐ฒ๐ฟ๐ฎ๐ฐ๐ต ๐๐ต๐ฎ๐'๐ ๐๐ฝ๐ฑ๐ฎ๐๐ฒ๐ฑ ๐๐ต๐ฒ ๐ณ๐ฒ๐ฎ๐๐๐ฟ๐ฒ, ๐บ๐ฒ๐บ๐ผ๐ฟ๐-๐๐ผ-๐๐ธ๐ถ๐น๐น.
๐๐ด๐ฒ๐ป๐ ๐บ๐ฒ๐บ๐ผ๐ฟ๐ ๐๐ต๐ฎ๐'๐ ๐ฎ๐ฐ๐๐๐ฎ๐น๐น๐ ๐๐ฎ๐น๐๐ฎ๐ฏ๐น๐ฒ ๐ฐ๐ฎ๐ฝ๐๐๐ฟ๐ฒ๐ ๐๐ต๐ฟ๐ฒ๐ฒ ๐๐ต๐ถ๐ป๐ด๐: ๐๐ต๐ฎ๐ ๐๐ต๐ฒ ๐ฎ๐ด๐ฒ๐ป๐ ๐ธ๐ป๐ผ๐๐, ๐๐ต๐ฎ๐ ๐ถ๐ ๐ต๐ฎ๐ ๐ฒ๐ ๐ฝ๐ฒ๐ฟ๐ถ๐ฒ๐ป๐ฐ๐ฒ๐ฑ, ๐ฎ๐ป๐ฑ ๐ต๐ผ๐ ๐ถ๐ ๐ฑ๐ผ๐ฒ๐ ๐๐ต๐ถ๐ป๐ด๐.
Developers used to keep those in three separate systems: knowledge went into RAG, experience piled up in an agent memory layer, and procedures were written down as skill files.
But those memories โ debug sessions, API integrations, releases, user corrections โ are essentially part of the agent's skills too. Each one carries the scenario, the feedback, and the outcome. Plenty of reusable workflows are already hiding in there.
๐ฆ๐ผ ๐๐ฒ ๐ฏ๐๐ถ๐น๐ ๐ ๐ฒ๐บ๐ผ๐ฟ๐-๐๐ผ-๐ฆ๐ธ๐ถ๐น๐น ๐ถ๐ป๐๐ผ ๐ ๐ฒ๐บ๐ฆ๐ฒ๐ฎ๐ฟ๐ฐ๐ต, ๐๐ต๐ฒ ๐ฎ๐ด๐ฒ๐ป๐ ๐บ๐ฒ๐บ๐ผ๐ฟ๐ ๐น๐ถ๐ฏ๐ฟ๐ฎ๐ฟ๐ ๐๐ฒ ๐ผ๐ฝ๐ฒ๐ป-๐๐ผ๐๐ฟ๐ฐ๐ฒ๐ฑ ๐ฒ๐ฎ๐ฟ๐น๐ถ๐ฒ๐ฟ ๐๐ต๐ถ๐ ๐๐ฒ๐ฎ๐ฟ.
An opt-in background pass mines your history for the workflows that keep recurring and are worth capturing, then automatically distills them into candidate skills. If a release workflow keeps showing up, the system recognizes it can be abstracted into a skill and writes it up as a candidate file. When the workflow changes later, the candidate evolves with it.
๐๐๐ ๐ป๐ผ ๐ฎ๐ด๐ฒ๐ป๐ ๐ฒ๐๐ฒ๐ฟ ๐ฎ๐๐๐ผ-๐น๐ผ๐ฎ๐ฑ๐ ๐ฎ ๐ฐ๐ฎ๐ป๐ฑ๐ถ๐ฑ๐ฎ๐๐ฒ ๐๐ธ๐ถ๐น๐น. It stays a draft. You review it, edit it, confirm it, and then install it into the skill directory of Claude Code, Codex, or other agents you use.
๐ ๐ฒ๐บ๐ฆ๐ฒ๐ฎ๐ฟ๐ฐ๐ต ๐ถ๐ ๐ผ๐ฝ๐ฒ๐ป ๐๐ผ๐๐ฟ๐ฐ๐ฒ. ๐๐ณ ๐๐ผ๐'๐ฟ๐ฒ ๐ฎ๐น๐๐ผ ๐๐ผ๐ฟ๐ธ๐ถ๐ป๐ด ๐ผ๐ป ๐ฎ๐ด๐ฒ๐ป๐ ๐บ๐ฒ๐บ๐ผ๐ฟ๐, ๐น๐ผ๐ป๐ด-๐๐ฒ๐ฟ๐บ ๐ฐ๐ผ๐ป๐๐ฒ๐ ๐, ๐ผ๐ฟ ๐ฐ๐ฟ๐ผ๐๐-๐ฝ๐น๐ฎ๐๐ณ๐ผ๐ฟ๐บ ๐๐ ๐๐ผ๐ฟ๐ธ๐ณ๐น๐ผ๐๐, ๐ด๐ถ๐๐ฒ ๐ถ๐ ๐ฎ ๐๐ฟ๐: https://t.co/eGAavvLonV
Semantic search in news feeds, e-commerce, and recommendation systems needs to balance relevance with freshness.
๐ ๐ถ๐น๐๐๐ ๐ฝ๐ฟ๐ผ๐๐ถ๐ฑ๐ฒ๐ ๐๐ต๐ฟ๐ฒ๐ฒ ๐ฑ๐ฒ๐ฐ๐ฎ๐ ๐ณ๐๐ป๐ฐ๐๐ถ๐ผ๐ป๐ ๐ณ๐ผ๐ฟ ๐ฟ๐ฒ๐ฟ๐ฎ๐ป๐ธ๐ถ๐ป๐ด. Three decay models are available, each producing a different freshness curve:
โข ๐๐ ๐ฝ๐ผ๐ป๐ฒ๐ป๐๐ถ๐ฎ๐น: fast initial decay with a long tail. News and social feeds where recency should dominate.
โข ๐๐ฎ๐๐๐๐ถ๐ฎ๐ป: bell-shaped curve with gradual roll-off. Balanced search like restaurant recommendations, where moderate distance should reduce rather than eliminate a result's score.
โข ๐๐ถ๐ป๐ฒ๐ฎ๐ฟ: constant penalty with a clear cutoff point. Event search where anything older than two weeks should not appear at all.
In this test, the query is "artificial intelligence advancements," against seven articles spanning 1 to 120 days old, under four ranking approaches.
No decay: pure semantic similarity.
โข Ranking is determined entirely by L2 distance (lower = more similar). Time is invisible.
โข Two articles with nearly identical content but published 90 days apart score the same: 0.7090. The 90-day-old article sits at #1.
โข An article about "AI Ethics Guidelines" scores lowest because "ethics" is semantically farther from "advancements" than the other articles, even though it is also about AI.
Gaussian decay (offset: 7 days, scale: 14 days, decay: 0.5)
โข The two newest articles (1-day and 5-day) jump to the top two positions. Both fall within the 7-day offset, so they receive no time penalty and rank by their original semantic similarity.
โข A 15-day-old article drops to 0.3601 but stays visible. A 30-day-old article falls to 0.0642.
โข Everything older than 60 days: 0.0000. The bell curve penalizes gradually at first, then sharply.
Exponential decay (offset: 3 days, scale: 10 days, decay: 0.3)
โข A 1-day-old article scores 0.5979. A 5-day-old article drops to 0.4774 because it has already exceeded the 3-day offset, even though it is still relatively new.
โข At 15 days: 0.1065, far below Gaussian's 0.3601 for the same article. At 30 days: nearly zero.
โข Under this configuration, the steepest curve. It forgets faster than Gaussian and penalizes harder at every age beyond the offset.
Linear decay: the most gradual decline.
โข A 90-day-old article still scores 0.3037. A 120-day-old article: 0.2339. Both score zero under Gaussian and exponential.
โข Because the penalty is constant rather than accelerating, semantic relevance carries more weight relative to time. A highly relevant 90-day-old article outranks a weakly relevant 30-day-old one: the time penalty was not steep enough to override the stronger semantic match.
โข The slope is gentle, but linear decay does reach zero at a clear cutoff point, giving you a predictable expiration boundary.
๐๐ต๐ผ๐ผ๐๐ถ๐ป๐ด ๐๐ต๐ฒ ๐ฐ๐๐ฟ๐๐ฒ ๐บ๐ฎ๐๐๐ฒ๐ฟ๐ ๐บ๐ผ๐ฟ๐ฒ ๐๐ต๐ฎ๐ป ๐ฐ๐ต๐ผ๐ผ๐๐ถ๐ป๐ด ๐๐ผ ๐ฎ๐ฑ๐ฑ ๐ฑ๐ฒ๐ฐ๐ฎ๐.
Learn more here:https://t.co/72j2ip5E7I
๐๐ป ๐ฒ-๐ฐ๐ผ๐บ๐บ๐ฒ๐ฟ๐ฐ๐ฒ ๐ฎ๐ป๐ฑ ๐ป๐ฒ๐๐ ๐๐ฒ๐ฎ๐ฟ๐ฐ๐ต, ๐๐ฒ๐บ๐ฎ๐ป๐๐ถ๐ฐ ๐๐ถ๐บ๐ถ๐น๐ฎ๐ฟ๐ถ๐๐ ๐ฎ๐น๐ผ๐ป๐ฒ ๐ถ๐ ๐ป๐ผ๐ ๐ฒ๐ป๐ผ๐๐ด๐ต ๐๐ผ ๐ฟ๐ฎ๐ป๐ธ ๐ฟ๐ฒ๐๐๐น๐๐. Milvus's Time-aware Ranking Functions (also called Decay Functions) rerank retrieval results by applying time-decay functions, so recent content ranks higher without sacrificing relevance.
Milvus's Decay Functions combining vector similarity with configurable time-decay to dynamically adjust each document's relevance score. ๐ง๐ต๐ฒ ๐๐ฐ๐ผ๐ฟ๐ถ๐ป๐ด ๐๐ผ๐ฟ๐ธ๐ ๐ถ๐ป ๐๐ต๐ฟ๐ฒ๐ฒ ๐๐๐ฒ๐ฝ๐:
โข ๐๐ผ๐ป๐๐ฒ๐ฟ๐ ๐ฑ๐ถ๐๐๐ฎ๐ป๐ฐ๐ฒ๐ ๐ถ๐ป๐๐ผ ๐ต๐ถ๐ด๐ต๐ฒ๐ฟ-๐ถ๐-๐ฏ๐ฒ๐๐๐ฒ๐ฟ ๐๐ฐ๐ผ๐ฟ๐ฒ๐. For L2 and JACCARD (lower = more similar): normalized_score = 1.0 โ (2 ร arctan(score)) / ฯ. COSINE, IP, and BM25 scores are used directly.
โข ๐๐ผ๐บ๐ฝ๐๐๐ฒ ๐ฎ ๐ฑ๐ฒ๐ฐ๐ฎ๐ ๐๐ฐ๐ผ๐ฟ๐ฒ from a numeric field like a timestamp, using one of three decay models: exponential, Gaussian, or linear. The result is a value between 0 and 1, representing how close the document is to a reference point (typically "now").
โข ๐ ๐๐น๐๐ถ๐ฝ๐น๐: final_score = normalized_similarity_score ร decay_score. In hybrid search with multiple vector fields, Milvus takes the maximum normalized score first: final_score = max(normalized_scores) ร decay_score.
Each model takes a configurable origin, scale (the distance beyond the offset at which the score drops to the decay value), and an optional offset (a no-penalty zone around the origin where the score stays at ๐ญ.๐ฌ).
๐๐ฒ๐ฎ๐ฟ๐ป ๐บ๐ผ๐ฟ๐ฒ ๐ต๐ฒ๐ฟ๐ฒ: https://t.co/PLK2ouE4Dt
๐๐ฟ๐ฒ๐ฝ-๐๐๐๐น๐ฒ ๐๐ฒ๐ฎ๐ฟ๐ฐ๐ต ๐ถ๐ป ๐๐น๐ฎ๐๐ฑ๐ฒ ๐๐ผ๐ฑ๐ฒ ๐๐ผ๐ฟ๐ธ๐ ๐ณ๐ผ๐ฟ ๐ฐ๐ผ๐ฑ๐ฒ ๐ฟ๐ฒ๐๐ฟ๐ถ๐ฒ๐๐ฎ๐น, ๐ฏ๐๐ ๐ถ๐ ๐ฏ๐๐ฟ๐ป๐ ๐๐ต๐ฟ๐ผ๐๐ด๐ต ๐๐ผ๐ธ๐ฒ๐ป๐ ๐๐ป๐ป๐ฒ๐ฐ๐ฒ๐๐๐ฎ๐ฟ๐ถ๐น๐. ๐ช๐ฒ ๐ฏ๐๐ถ๐น๐ ๐๐น๐ฎ๐๐ฑ๐ฒ ๐๐ผ๐ป๐๐ฒ๐ ๐, ๐ฎ๐ป ๐ผ๐ฝ๐ฒ๐ป-๐๐ผ๐๐ฟ๐ฐ๐ฒ ๐ ๐๐ฃ ๐๐ผ๐ผ๐น ๐๐ต๐ฎ๐ ๐ฐ๐๐๐ ๐๐ผ๐ธ๐ฒ๐ป ๐๐๐ฎ๐ด๐ฒ ๐ฏ๐ ~๐ฐ๐ฌ%.
Grep-style code retrieval has drawn plenty of community discussion. For pure recall, literal matching gets the job done. But in a large codebase, it forces the model to wade through massive amounts of irrelevant code to find the few lines that matter.
๐ง๐ต๐ฟ๐ฒ๐ฒ ๐๐ต๐ถ๐ป๐ด๐ ๐บ๐ฎ๐ธ๐ฒ ๐๐ต๐ถ๐ ๐ฒ๐ ๐ฝ๐ฒ๐ป๐๐ถ๐๐ฒ ๐ฎ๐ ๐๐ฐ๐ฎ๐น๐ฒ:
โข ๐๐ป๐ณ๐ผ๐ฟ๐บ๐ฎ๐๐ถ๐ผ๐ป ๐ผ๐๐ฒ๐ฟ๐น๐ผ๐ฎ๐ฑ. Large repos produce hundreds of literal matches for common terms. Most are noise, but they all consume tokens and inference time.
โข ๐ฆ๐ฒ๐บ๐ฎ๐ป๐๐ถ๐ฐ ๐ฏ๐น๐ถ๐ป๐ฑ๐ป๐ฒ๐๐. Grep matches characters, not meaning. compute_final_cost() and calculate_total_price() do the same thing, but a string match won't connect them.
โข ๐๐ป๐ฐ๐ผ๐บ๐ฝ๐น๐ฒ๐๐ฒ ๐ฐ๐ผ๐ป๐๐ฒ๐ ๐. Line-level matches strip away the surrounding class and method structure. The model compensates by reading more files, burning more tokens.
๐ช๐ฒ ๐ผ๐ฝ๐ฒ๐ป-๐๐ผ๐๐ฟ๐ฐ๐ฒ๐ฑ ๐๐น๐ฎ๐๐ฑ๐ฒ ๐๐ผ๐ป๐๐ฒ๐ ๐ ๐๐ผ ๐ฎ๐ฑ๐ฑ๐ฟ๐ฒ๐๐ ๐ฒ๐ ๐ฎ๐ฐ๐๐น๐ ๐๐ต๐ถ๐. ๐๐ ๐ต๐ถ๐ #๐ญ ๐ผ๐ป ๐๐ถ๐๐๐๐ฏ'๐ ๐ฑ๐ฎ๐ถ๐น๐ ๐๐ฟ๐ฒ๐ป๐ฑ๐ถ๐ป๐ด ๐ฐ๐ต๐ฎ๐ฟ๐ ๐ฎ๐ป๐ฑ ๐ฐ๐๐ฟ๐ฟ๐ฒ๐ป๐๐น๐ ๐ต๐ฎ๐ ๐ผ๐๐ฒ๐ฟ ๐ญ๐ฎ,๐ฌ๐ฌ๐ฌ ๐๐๐ฎ๐ฟ๐.
๐๐ป ๐ผ๐๐ฟ ๐ฏ๐ฒ๐ป๐ฐ๐ต๐บ๐ฎ๐ฟ๐ธ, ๐๐น๐ฎ๐๐ฑ๐ฒ ๐๐ผ๐ป๐๐ฒ๐ ๐ ๐ฟ๐ฒ๐ฑ๐๐ฐ๐ฒ๐ฑ ๐๐ผ๐ธ๐ฒ๐ป ๐ฐ๐ผ๐ป๐๐๐บ๐ฝ๐๐ถ๐ผ๐ป ๐ฏ๐ ~๐ฐ๐ฌ% ๐ฎ๐ ๐ฒ๐พ๐๐ถ๐๐ฎ๐น๐ฒ๐ป๐ ๐ฟ๐ฒ๐๐ฟ๐ถ๐ฒ๐๐ฎ๐น ๐พ๐๐ฎ๐น๐ถ๐๐.
It's a code retrieval MCP server that integrates a vector database and embedding model into the search layer. It plugs into Claude Code and is also compatible with Codex CLI, Gemini CLI, Qwen Code, Cline, Cursor, Windsurf, and other MCP-compatible agents.
๐๐ ๐ฑ๐ผ๐ฒ๐ ๐๐ต๐ฟ๐ฒ๐ฒ ๐๐ต๐ถ๐ป๐ด๐ ๐ฑ๐ถ๐ณ๐ณ๐ฒ๐ฟ๐ฒ๐ป๐๐น๐:
โข ๐ฆ๐บ๐ฎ๐ฟ๐ ๐ณ๐ถ๐น๐๐ฒ๐ฟ๐ถ๐ป๐ด. Vector similarity ranks code by relevance, surfacing the most related results first instead of every literal match.
โข ๐๐ผ๐ป๐ฐ๐ฒ๐ฝ๐ ๐บ๐ฎ๐๐ฐ๐ต๐ถ๐ป๐ด. Dense retrieval can match conceptually related code even when function names differ entirely.
โข ๐๐ผ๐ป๐๐ฒ๐ ๐-๐ฎ๐๐ฎ๐ฟ๐ฒ ๐ฟ๐ฒ๐๐ฟ๐ถ๐ฒ๐๐ฎ๐น. Results are chunked around complete syntactic units (functions, classes, methods), giving the model enough structural context to reason about behavior.
Under the hood, Claude Context uses MCP as its interface layer and Milvus for codebase indexing and hybrid search (BM25 + dense vectors). Embeddings support OpenAI, VoyageAI (which ships a code-specific model), and Ollama for local deployment when privacy matters.
Code chunking uses AST parsing by default, splitting along semantic boundaries instead of cutting at arbitrary line counts. For files AST can't parse, LangChain's text splitter handles the fallback.
For incremental updates, Claude Context uses Merkle-tree change detection. If the root hash hasn't changed, the index stays untouched. When it changes, the system identifies exactly which files were modified and re-embeds only those.
Change one function, and you don't re-index the entire project.
๐ https://t.co/yyWCfHbgZv
๐ Full benchmark and architecture: https://t.co/Sevmvp37HT
Using a demo to estimate the real production cost of a vector database is where cost planning goes wrong for many teams.
The demo phase can give you a reasonable answer on two things:
โข Storage: embeddings, raw content, metadata, scalar fields, and index files.
โข Online serving: the compute needed for the current query volume and latency target.
But a demo usually cannot answer three other cost questions:
โข Updates: how often data is added, deleted, re-embedded, re-indexed, or checked for recall regressions.
โข Workload mismatch: whether offline analysis, evaluation, or batch jobs are running on always-on infrastructure.
โข Duplication: how many copies of the same AI data end up across vector search, lakehouse, training, evaluation, and governance systems.
Different data scales and usage frequencies require different resource models.
For example, one autonomous driving workload needed vector search over a 1B-row collection for analysis.
๐ ๐ฑ๐ฒ๐ฑ๐ถ๐ฐ๐ฎ๐๐ฒ๐ฑ ๐ฐ๐น๐๐๐๐ฒ๐ฟ was estimated at about $7,000 per month. ๐ฆ๐ฒ๐ฟ๐๐ฒ๐ฟ๐น๐ฒ๐๐ was about $10,800 per month. But the workload only ran for a few hours each month. With ๐ญ๐ถ๐น๐น๐ถ๐ ๐๐น๐ผ๐๐ฑ ๐ข๐ป-๐๐ฒ๐บ๐ฎ๐ป๐ฑ ๐ฆ๐ฒ๐ฎ๐ฟ๐ฐ๐ต, the cost came in under $500 per month.
In enterprise environments, however, these workload patterns usually coexist. A single system may have steady online traffic, offline batch processing, evaluation jobs, and sudden demand spikes.
The goal is not to choose one resource model for everything. It is to accurately identify each workload and match it to the right instance type.
Use dedicated capacity for steady, latency-sensitive traffic; serverless for unpredictable or spiky online demand; and on-demand compute for low-frequency analysis, evaluation, and batch jobs.
Duplicate text in your LLM pre-training corpus drives up training cost and reweights parts of the training data. ๐ ๐ถ๐ป๐๐ฎ๐๐ต ๐๐ฆ๐ ๐บ๐ฎ๐ธ๐ฒ๐ ๐๐ฐ๐ฎ๐น๐ฎ๐ฏ๐น๐ฒ ๐ฑ๐ฒ๐ฑ๐๐ฝ ๐ฝ๐ฟ๐ฎ๐ฐ๐๐ถ๐ฐ๐ฎ๐น, ๐ฎ๐ป๐ฑ ๐ ๐ถ๐น๐๐๐ ๐ฎ.๐ฒ ๐ฎ๐ฑ๐ฑ๐ฒ๐ฑ ๐ถ๐ ๐ฎ๐ ๐ฎ ๐ป๐ฎ๐๐ถ๐๐ฒ ๐ถ๐ป๐ฑ๐ฒ๐ ๐๐๐ฝ๐ฒ.
The usual dedup options break down quickly:
โข ๐๐ ๐ฎ๐ฐ๐ ๐บ๐ฎ๐๐ฐ๐ต๐ถ๐ป๐ด is too strict for text dedup. Near-identical documents often differ by boilerplate, formatting, punctuation, or minor edits, and exact comparison misses all of them.
โข ๐ฃ๐ฎ๐ถ๐ฟ๐๐ถ๐๐ฒ ๐ฐ๐ผ๐บ๐ฝ๐ฎ๐ฟ๐ถ๐๐ผ๐ป catches more, but the computation explodes with data volume. At millions of documents, it's infeasible.
๐ ๐ถ๐ป๐๐ฎ๐๐ต ๐๐ฆ๐ works well as a coarse-pass filter. It's approximate deduplication with a tunable recall/precision tradeoff, fast enough to run inline during ingestion to check candidates before storing redundant content.
๐ ๐ถ๐น๐๐๐ ๐ฎ.๐ฒ ๐ฎ๐ฑ๐ฑ๐ฒ๐ฑ ๐ป๐ฎ๐๐ถ๐๐ฒ ๐ ๐ถ๐ป๐๐ฎ๐๐ต ๐๐ฆ๐ ๐ถ๐ป๐ฑ๐ฒ๐ ๐ถ๐ป๐ด ๐๐ถ๐๐ต ๐ฎ ๐ฑ๐ฒ๐ฑ๐ถ๐ฐ๐ฎ๐๐ฒ๐ฑ ๐บ๐ฒ๐๐ฟ๐ถ๐ฐ ๐๐๐ฝ๐ฒ, ๐น๐ฒ๐๐๐ถ๐ป๐ด ๐๐ฒ๐ฎ๐บ๐ ๐ถ๐ป๐ฑ๐ฒ๐ ๐ฎ๐ป๐ฑ ๐๐ฒ๐ฎ๐ฟ๐ฐ๐ต ๐ ๐ถ๐ป๐๐ฎ๐๐ต ๐๐ถ๐ด๐ป๐ฎ๐๐๐ฟ๐ฒ๐ ๐ณ๐ผ๐ฟ ๐ฑ๐๐ฝ๐น๐ถ๐ฐ๐ฎ๐๐ฒ ๐ฎ๐ป๐ฑ ๐ป๐ฒ๐ฎ๐ฟ-๐ฑ๐๐ฝ๐น๐ถ๐ฐ๐ฎ๐๐ฒ ๐ฑ๐ฒ๐๐ฒ๐ฐ๐๐ถ๐ผ๐ป ๐ฎ๐ ๐ฐ๐ผ๐น๐น๐ฒ๐ฐ๐๐ถ๐ผ๐ป ๐๐ฐ๐ฎ๐น๐ฒ.
๐ช๐ต๐ฒ๐ฟ๐ฒ ๐๐ต๐ถ๐ ๐ต๐ฒ๐น๐ฝ๐:
โข ๐๐๐ ๐๐ฟ๐ฎ๐ถ๐ป๐ถ๐ป๐ด ๐ฑ๐ฎ๐๐ฎ ๐ฐ๐น๐ฒ๐ฎ๐ป๐ถ๐ป๐ด: removing duplicate text from crawled corpora before training begins
โข ๐๐ฎ๐ฟ๐ด๐ฒ-๐๐ฐ๐ฎ๐น๐ฒ ๐ฝ๐น๐ฎ๐ด๐ถ๐ฎ๐ฟ๐ถ๐๐บ ๐๐ฐ๐ฟ๐ฒ๐ฒ๐ป๐ถ๐ป๐ด: detecting substantial text overlap in academic or content pipelines as a coarse-pass filter
๐๐ถ๐บ๐ถ๐๐ฎ๐๐ถ๐ผ๐ป: MinHash operates on token-set overlap, not word order or semantics. It catches verbatim and near-verbatim duplication but won't flag paraphrases. For semantic or paraphrase-level deduplication, teams layer in SimHash, TF-IDF with cosine similarity, or embedding-based methods.
In Milvus, MinHash signatures are stored as ๐ฏ๐ถ๐ป๐ฎ๐ฟ๐ ๐๐ฒ๐ฐ๐๐ผ๐ฟ๐ approximating ๐๐ฎ๐ฐ๐ฐ๐ฎ๐ฟ๐ฑ ๐๐ถ๐บ๐ถ๐น๐ฎ๐ฟ๐ถ๐๐ between document token sets. Standard metrics like Jaccard, L2, or cosine can't be applied directly to these signatures, so Milvus introduced a dedicated metric called ๐ ๐๐๐๐๐๐๐ฅ๐. For higher accuracy, Milvus also supports a refined search mode that recomputes exact Jaccard using stored token sets, reducing false positives.
The dedup pipeline:
โข ๐ฆ๐ฒ๐๐๐ฝ: Create a collection with a ๐ ๐๐ก๐๐๐ฆ๐_๐๐ฆ๐ index, using the MHJACCARD metric and configurable LSH parameters (e.g., band count)
โข ๐ฆ๐ถ๐ด๐ป๐ฎ๐๐๐ฟ๐ฒ ๐ด๐ฒ๐ป๐ฒ๐ฟ๐ฎ๐๐ถ๐ผ๐ป: For each incoming document, tokenize or shingle it and generate a MinHash signature (e.g., using datasketch)
โข ๐๐๐ฝ๐น๐ถ๐ฐ๐ฎ๐๐ฒ ๐ฐ๐ต๐ฒ๐ฐ๐ธ: Search the collection with the new signature to find existing near-duplicates
โข ๐๐ป๐๐ฒ๐ฟ๐: Store only documents that don't exceed your similarity threshold, alongside metadata
An AI model company used this pipeline to deduplicate ๐ญ๐ฌ ๐ฏ๐ถ๐น๐น๐ถ๐ผ๐ป ๐ฑ๐ผ๐ฐ๐๐บ๐ฒ๐ป๐๐ for LLM training, more than doubling throughput and cutting costs ๐ฏ-๐ฑ๐ compared to their previous MapReduce-based approach.
More info here: https://t.co/3JQyKaXSCi
Debugging a slow Milvus cluster once meant bouncing between the console, Grafana, and several monitoring views just to find out where to look.
๐ง๐ผ๐ฑ๐ฎ๐, ๐๐ถ๐๐ต ๐๐๐๐ ๐ฏ.๐ฌ ๐ฏ๐ฒ๐๐ฎ, ๐ฐ๐น๐๐๐๐ฒ๐ฟ ๐ผ๐ฝ๐ฒ๐ฟ๐ฎ๐๐ถ๐ผ๐ป๐ ๐ฎ๐ฟ๐ฒ ๐ฎ๐น๐น ๐ผ๐ป ๐ผ๐ป๐ฒ ๐๐ฐ๐ฟ๐ฒ๐ฒ๐ป. ๐ฌ๐ผ๐ ๐ฐ๐ฎ๐ป ๐ฐ๐ต๐ฒ๐ฐ๐ธ:
โข ๐ข๐๐ฒ๐ฟ๐๐ถ๐ฒ๐: version, deployment mode, node count, database count, collection count, and total entity count
โข ๐ฅ๐ฒ๐ฎ๐น-๐๐ถ๐บ๐ฒ ๐บ๐ฒ๐๐ฟ๐ถ๐ฐ๐: QPS, insert/delete rates, query latency, cache hit rate, and storage, pulled from a Prometheus endpoint without opening Grafana separately
โข ๐ฆ๐น๐ผ๐-๐พ๐๐ฒ๐ฟ๐ ๐ฎ๐ป๐ฎ๐น๐๐๐ถ๐: slow queries by type, duration, collection, timestamp, and user
โข ๐ง๐ผ๐ฝ๐ผ๐น๐ผ๐ด๐ ๐๐ถ๐ฒ๐: node topology and coordinator component status
โข ๐๐ฎ๐ฐ๐ธ๐๐ฝ ๐ฎ๐ป๐ฑ ๐ฟ๐ฒ๐๐๐ผ๐ฟ๐ฒ: full or incremental backups to S3, MinIO, GCS, or Azure
For teams running several Milvus clusters, the multi-cluster sidebar gives you a quick read on cluster state before you start digging. The point is that the first pass of cluster inspection no longer has to start with five tabs and a guessing game.
๐ง๐ต๐ฒ ๐ณ๐๐น๐น ๐๐๐๐ ๐ฏ.๐ฌ ๐ฟ๐ฒ๐ฏ๐๐ถ๐น๐ฑ, ๐ถ๐ป๐ฐ๐น๐๐ฑ๐ถ๐ป๐ด ๐๐ต๐ฒ ๐บ๐๐น๐๐ถ-๐ฐ๐น๐๐๐๐ฒ๐ฟ ๐๐ถ๐ฑ๐ฒ๐ฏ๐ฎ๐ฟ ๐ฎ๐ป๐ฑ ๐ฏ๐๐ถ๐น๐-๐ถ๐ป ๐๐ ๐ฎ๐ด๐ฒ๐ป๐, ๐ถ๐ ๐ผ๐ป ๐๐ต๐ฒ ๐ ๐ถ๐น๐๐๐ ๐ฏ๐น๐ผ๐ด: https://t.co/LLfUcMHsnF