MSc project update: my RAG system looked slightly worse than the baseline at retrieval.
Then I ran a test, the gap wasn't real. They tie.
Which reframed everything: the real difference isn't ranking, it's whether the system knows when to say "I can't help."
#RAG#MachineLearning
INSTEAD OF WATCHING AN HOUR OF NETFLIX TONIGHT.
This 60-minute Cambridge lecture by Demis Hassabis will teach you more about the future of AI than most people will learn in the next 5 years.
Bookmark it and give it an hour, no matter what.
4/ Takeaway: building RAG is easy to start, hard to evaluate. Knowing when your system should say nothing — and how to measure if it's trustworthy — is the real work. Writing up now.
MSc project update: I've finished evaluating my classifier-assisted RAG system for university IT support.
Biggest lesson: a small local LLM can generate grounded, cited answers well — but is too weak to reliably judge answer quality.
Also: keyword search is hard to beat
3/ The surprise: I ran it on a small local model for cost/privacy. It generated grounded, cited answers well — but was too unreliable to act as an automated judge of answer quality. So I leaned on a manual rubric and reported the gap honestly.
2/ BM25 is fast and genuinely hard to beat at ranking on a small corpus — but it can't recognise an out-of-scope question, so it always returns something. Adding an intent classifier + rejection layer fixed that: the system refuses safely.
Finished the evaluation stage of my MSc project: a classifier-assisted RAG system for university IT support, benchmarked against a BM25 baseline. A few things I learned
shipped a PDF toolkit 🛠️
merge • split • compress • OCR • watermark • password-protect • convert (PDF ↔ images/Word/text)
built with Next.js 15 + pdf-lib, live preview on every tool
try it → https://t.co/95bhdhkc2x
My fraud model catches 84% of fraud on its own data.
fed it transactions from a different source: 6%. zero errors. completely silent failure.
so I built the system that catches this: PSI drift detection, auto retraining, shadow deployments.
live demo: https://t.co/cjxY4kqVVy
I built Cognify: upload any PDF → get an adaptive 3-phase exam that finds your weak spots and drills them.
AI-generated questions on Claude Haiku, ~5¢ per session, rate-limited free tier, BYOK for power users, automatic fallback if the API dies.
MSc project follow-up 🚀
I’m building a classifier-assisted RAG system for University IT support.
So far:
✅ FastAPI backend
✅ Streamlit interface
✅ SQLite logging
✅ BM25 baseline comparison
Learning: RAG is easy to demo, but harder to evaluate
#RAG#LLM#AI#DataScience