@UnstructuredIO had a massive release last Friday. Over the past month we've doubled down on our core library and focused on further enhancing the core user experience. Highlights below:
➡️Destination Connectors Are Here!
We’ve finalized the base abstraction for our new destination connectors, and users can now write results directly to a Databricks Delta Table from the ingest CLI. This functionality allows users to couple over 20 existing source connectors with destination locations (e.g., vector databases). Over the coming weeks, we’ll be rolling out numerous downstream connectors. Let us know which ones you’d like prioritized!
➡️Better Table Extraction
We’ve worked to enhance table extraction in .html, .epub., .md, .rst, .odt, and .msg. At the core of this change is the .html partition functionality, which is leveraged by the other affected doc types. Stay tuned for a huge leap forward in table extraction from PDFs over the coming weeks.
➡️Chunking
Previously, users were responsible for their own chunking after partitioning elements, often required for downstream applications (e.g. LangChain). Now, individual document elements may be combined into right-sized chunks for better downstream results. This enables users to use partitioned results effectively in downstream applications (e.g. RAG architecture apps) without additional post-processing.
➡️Element Classification
We’ve implemented a wide range of enhancements, fixes, and features to enhance the accuracy of our classification of document elements. For example, the partition_pdf with fast strategy previously classified numbered list item lines as separate elements. This enhancement leverages the x and y coordinates and bounding box sizes to help decide whether a chunk of text continues the previous List Item element.
➡️Languages
Foundational changes have been made to enhance language support, allowing users to input document languages using standard language codes compatible with Tesseract.
➡️Accelerated OCR
We’ve reduced by 50% the number of calls to Tesseract when partitioning a PDF or image.
➡️Hierarchy Detection for More Effective RAG
Additional metadata attached to individual document elements helps with RAG and over the past several weeks our team has been working on approaches to detect hierarchies among document elements. In this release, we’re rolling out our initial functionality. Read more in our release notes and please let us know how we can further develop this functionality to help with retrieval.
➡️New Document and Source Metadata for More Effective RAG
We’ve now added document and data source properties to upstream connectors. This new functionality not only records key attributes about the source document but also the source document application, e.g. a link to a GDrive doc, Salesforce record, etc. New metadata fields include Document Hierarchy, Date Created, Date Modified, Version, Source URL, and Record Locator.
https://t.co/phV7vZRMGZ
💪Deep dive on retrieval
Excited to have the opportunity to talk about the advanced retrieval tools we're building:
- smart text splitting
- self-query
- retrieval agents
Even more excited to be joined by @UnstructuredIO@trychroma... should be fun
https://t.co/8OD0OTKPGK
@recursecenter I'm very interested in attending to learn more about the retreat! According to the Calendly, there arr no times in December, January, or February. Is there an issue with the link?
@spbail I've seen this, too. Sometimes, people don't want to acknowledge or let go of their coping mechanisms. It's understandable because doing that is hard.
@spbail How do you stay energized and alert after cutting out coffee? I've been swapping out coffee with matcha lately but haven't cut out coffee completely yet
@spbail When I'm in a meeting and can't follow the discussion, my instinct has been to feel discomfort. I'm trying to shift to this attitude and enjoy the opportunity to learn.