Open source document redaction tool to identify and remove personal information (PII) with an easy to use GUI @Gradio. Uses local models @spacy_io to identify PII, or ties into AWS services such as Textract and Comprehend for best results. Try it out here: https://t.co/ULDnNP7RbC
Works with PDF, images, and tabular data as CSV/XLSX files. You can also chat over the documentation with a free chatbot here @huggingface : https://t.co/P0YtVidswF
Open source document redaction tool to identify and remove personal information (PII) with an easy to use GUI @Gradio. Uses local models @spacy_io to identify PII, or ties into AWS services such as Textract and Comprehend for best results. Try it out here: https://t.co/ULDnNP7RbC
The app has comprehensive redaction review features to review and modify redactions. Also, identify duplicate pages in your documents or export to Adobe Acrobat to modify suggested redactions there. @github repo and user guide: https://t.co/md82gGV6m7
This augmented topics table is then presented to the LLM for the next batch, which grows iteratively as it progresses through the dataset. You can see the prompts used, and other settings on the LLM Settings tab. 7/7
Large language model topic modelling live on @huggingface spaces. Identify the key topics and subtopics in tabular data, alongside sentiment and summaries. Built on @gradio. Try it out here (zero GPU space): https://t.co/eGaR9iTTqi 1/7
How does it work? The LLM is called iteratively on a 'batch' of open text responses from the data. The model assigns each response to existing or new topics (if none are relevant). The LLM returns a markdown table the topics alongside relevant response rows. 6/7
Document redaction app on @huggingface:
✅ Open source
✅ Redacts PDFs, images, tabular data such as CSVs/Excel
✅ Export all OCR results
✅ Modify suggested redactions with a point and click interface
Powered by @gradio and @spacy_io. Try it here: https://t.co/ULDnNP7RbC
If you remember our Applied LLMs course, you'll love this. Today, we are making all these resources available for free to everyone! 📚
We did extra work to add learning tracks, resources, and notes to each lesson to maximize your learning. Link in next tweet
Llama 405B is here, and it comes with more than expected! 🚨 @AIatMeta Llama 3.1 comes in 3 sizes, 8B, 70B, and 405B, and speaks 8 languages! 🌍 Llama 3.1 405B matches or beats the Openai GPT-4o across many text benchmarks.
New and improvements of 3.1✨:
🧮 8B, 70B & 405B versions as Instruct and Base with 128k context
🌐 Multilingual, supports 8 languages, including English, German, French, and more.
🔠 Trained on >15T Tokens & fine-tuned on 25M human and synthetic samples
📃 Commercial friendly license with allowance to use model outputs to improve other LLMs
⚖️ Quantized versions in FP8, AWQ, and GPTQ for efficient inference.
🚀 Llama 3 405B matches and beast GPT-4o on many benchmarks
🧑🏻💻 8B & 70B improved Coding and instruction, following up to 12%
⚒️ Supports Tool use and Function Calling
🤖 Llama 3.1 405B available on Hugging Face Inference API and in HuggingChat
🤗 Available on @huggingface
🔜 1-click deployments on Hugging Face, Amazon SageMaker, Google Cloud
Blog: https://t.co/UzquOLNaUi
Model Collection: https://t.co/7YR9Bzw1Ea
Big Kudos to Meta for releasing Llama 3.1, including 405B. This will help everyone accelerate and adopt AI more easily and faster. ❤️
⏰New blog post : “Chunking for RAG: best practices”
Learn about the importance of chunking, common methods, and smart chunking strategies. Bring your RAG system's performance to the next level with our Serverless API that is free to get started.
https://t.co/t0hydwztPb
Please try it out! As well as running on @huggingface publicly, you can duplicate it to your own private space to avoid queues, or clone the repo to run it locally on your computer.
New updates to the Topic Modeller app made with @gradio and based on the excellent BERTopic from @MaartenGr. Try the @huggingface space free here: https://t.co/qpLNzsD7mI
The underlying embedding model in 'high-quality' mode has been improved (now uses @mixedbreadai), and the LLM topic representation is now based on a quantised Microsoft Phi3-mini model, which seems to give better results.