JATS conversion is a huge pain point that slows down science - costs dollars *per page*, and takes weeks. We can do it for cents per page in less than 5 minutes, at close to human accuracy.
Our pipeline scores 92.6% in our human-matched benchmark (more on this next week).
We're dropping a new Datalab processor: PDF β JATS XML. ππ
Drop in a scientific paper (including scans of decades-old print journals) and get DTD-validated JATS 1.2 XML, the figure images it references, and a conformance verdict for the run. π§΅
To get to 94% from 65%, we first fix the scorer to strip meta fields - gets to 91%, then we make the scorer fair on ambiguous fields - gets to 93.6%.
Read more here - https://t.co/zJd9Trclt8 .
It's frustrating to see benchmarks with basic mistakes like this that get spun into press releases/marketing, etc.
Another reason you should always run your own evals, and not trust any vendor benchmarks (including ours).
@raychelmayy We are - this is the closest we have now - https://t.co/mSXDD8hqtS .
But feel free to DM, we're open to adding someone in a more pure support/cs role also.
Marker started as a small project to parse my PDFs. Today Datalab projects have 72k GitHub stars and 6M downloads a month.
We're hiring a founding open-source engineer to build the next chapter.
Our OSS projects like surya, marker, chandra, and lift (you'll work on these and more) are here - https://t.co/zKWCeHzGK2 and here - https://t.co/awg7drSUbX .
Apply at https://t.co/JBvShi50Gz - onsite in NYC (remote okay in exceptional cases), salary range is 225k-300k USD, with equity.
Feel free to DM me if you're interested.
Thanks to everyone who has tried marker and given feedback on it over the last couple of years. We'll be making more improvements, especially to accuracy, in the next few weeks.
Marker 2 is out now - up to 5x faster and more accurate than mineru, docling, and liteparse with similar configs.
Converts pdfs, images, docx to markdown. CPU + GPU compatible, up to 27 pages/s.
If you want the highest accuracy, use Chandra (https://t.co/HPgq7j5WZJ), or the datalab API at https://t.co/wsVqmf8wG7 .
Pipeline systems like marker are great for customizability and speed. They blend text from the document with OCR, so you don't need to OCR everything.
Vendor benchmarks (including ours) can be gamed in many ways.
You should trust your own evals on your own documents - we built a tool, Forge Evals, that helps you do that - https://t.co/QPJkztFo8Q.
Recently, a competitor released a dense table extraction benchmark. It's setup to be very hard for anyone but Reducto to win.
@datalabto now has the high score.
I'm thankful to Reducto for taking the time to benchmark us. There aren't any good public extraction benchmarks, and this is at least a step in that direction.
And while recall is their chosen headline metric (which we lead), they do still lead Datalab in leaf accuracy.