An LLM’s understanding of text, built through modeling word relationships, might approximate real text and real word relationships very closely. But it’s not because an LLM understands the meaning of individual words. *This* meaning is essentially impossible for it to infer.
Large language models, trained purely on text to predict the next word (token), are not taught the meaning of any individual word—nor do they learn it.
Kensho just released PubTables-v2 on Hugging Face
The first large-scale dataset for multi-page table extraction
with 135k+ cropped tables, 548k+ full-page tables, and 9k+ documents
More than 4 years after the release of PubTables-1M, PubTables-v2 is finally here! Among the highlights: this is the first dataset for document table extraction containing thousands of multi-page tables, some spanning up to 13 pages!
This is achieved via prompting in natural language, which provides a way to interpolate new tasks in between tasks the model has previously seen.
While still a shallow form of generalization, this has undoubtedly been a massive leap in the capabilities of these models.
In deep learning today generalization is powered by interpolation.
Before LLMs, generalization was constrained to handling novel input for a single narrowly-defined task.
The true breakthrough of LLMs in deep learning was to generalize beyond new inputs to entirely new *tasks*.
@tdietterich@ylecun@Michael_J_Black This is exactly how I remember it, as well. As someone who has always been a huge fan of Yann and his work, his consistent distortion and/or misunderstanding of the backlash to Galactica’s own overhyped claims has been painful to read.
Looking forward to #CVPR2024 in Seattle! If you’re attending, reach out and let’s grab ☕️.
My team at Kensho is hiring a new grad MLE for visual document understanding. We don’t have a booth this year but if you track me down I have some free Kensho swag to hand out. 🧢✍️🛍️
Today we released the weights for three new pre-trained Table Transformer (TATR) models, for recognizing and extracting data from tables in unstructured documents. Check them out here: https://t.co/mxhYFoVWPI #icdar2023
@fchollet François, your clarity of thought and willingness to share in resistance to hype and groupthink is unmatched and immensely valuable. Thank you.
We are excited to announce our work “PubTables-1M: Towards comprehensive table extraction from unstructured documents” has been accepted at #CVPR2022! @RohithPesala
Paper, code, and data are all publicly released: https://t.co/mxhYFpcZRI
PubTables-1M: a large-scale dataset consisting of ~1 million tables from PubMed Central Open Access scientific articles.
Can be used for training/evaluating machine learning models for the tasks of table detection, table structure recognition, and more.
https://t.co/4x0DJd4FxX