calling it a good stopping point for a while. AI cleanup and the training pipeline are both paused since the server hosting them got suspended out of nowhere (still appealing, no timeline). From here it’s bug fixes, not new features.
I’m gonna focus on my other project
OCD mode used to score the whole page once and pick a single language for everything. Now it checks confidence line by line so a stray Japanese line inside an English page gets caught and read with the right model instead of coming out mangled.
v3.1.0 is live probably the biggest update LiteDoc’s had. Fixed a bug that was silently killing table/figure extraction, added a CLI so you can run it from your terminal, and I’ve got a training pipeline running 24/7 to keep tuning the parser over time.
v3 fixes the big one: scanned PDFs now come out in the right reading order (OCR line-banding rewrite).
Also new: optional AI cleanup (beta) — runs on community-funded servers, you only pay for what it actually repairs. Clean docs = free
https://t.co/3qxCKtAeoi
Literally two days before my files got corrupted, I was in a car accident and then out of nowhere windows decided to shit itself and overwrite the files
I used dmde to try to recover my data, but most of the things that I got was corrupted files
And the backups were outdated
All of my projects are gone months of work
Ts is taking a lot on me mentally
It’s literally a hit after hit
I honestly don’t know what to do
Unfortunately, LiteDoc won’t get new updates anytime soon All of my data got corrupted and deleted by Windows The training data and the whole system to train the algorithm are gone.
I’m trying my best to recover what I can
New update is out now
V2.6.0
I tried to train it but I spent most of the time debugging the code it has some improvements
I’m gonna try to train it using python instead of js
Stay tuned!
I'm also stepping back from actively hunting bugs for now since the engine is stable.
If you hit any parsing issues or need features, open a GitHub issue or DM me. Otherwise, I'm just letting it run.
LiteDoc v2.2.0 is out.
Just shipped a major rewrite of the layout engine so multi-column papers finally extract perfectly without text interleaving. Also added native math equation extraction (no more hallucinated OCR) and a fallback layer for corrupted PDF fonts.
I’m still working on the pdf engine trying to find creative ways to make it better
But no matter what I do it still depends on the quality of the PDF that u r trying to transfer
And fuck all of the clowns on Reddit who are shitting on this project mfs project is in development
The PDF parser is actively evolving, so expect edge cases.
I'll be pushing updates to patch formatting bugs as they come up. Extracting Markdown from PDFs is a notoriously complex process, and this tool relies entirely on mathematical layout analysis and geometry zero LLM.
LiteDoc v2.0 is officially LIVE.
I completely nuked the old engine and rebuilt it with a recursive XY-Cut layout analyzer. Now it handles dense multi column academic papers and mobile screens without breaking a sweat.
Full notes & GitHub link: https://t.co/TJr3ges0EJ
OVO OUT✌️
My goblinMD first release was sloppy cuz I wasn’t really Passionate about the project so I took the most very useful feature and turned into Browser-based tool try it at https://t.co/3qxCKtAeoi
built GoblinMD — an offline desktop app to pack code folders & PDFs into token-saving prompts for LLMs.
strips comments, docstrings & space
extracts PDF text & visual diagrams to drop in
offline token count & Git diff filter
free & open source:
https://t.co/oyNGBu4hcq