OmniRetrieval
Why force all knowledge into one format? This unified retriever meets text, SQL tables, RDF graphs, and property graphs on their own terms—routing natural language to 309 distinct knowledge bases across 13 datasets.
What if your retriever could speak every language your data speaks? 🌐
Your answer might live in a document 📄, a SQL table 🗃️, an RDF knowledge graph 🔗, or a property graph 🕸️, and OmniRetrieval reaches into all of them, meeting each source in its own native query language instead of flattening everything into one lossy space.
Paper: https://t.co/dI6IvBwfWW
📢 New preprint out on contextual integrity (CI) and a new Product-of-Experts (PoE) view of self-distillation!
Introducing SelfCI, a novel self-distillation framework that operationalizes CI by optimizing for the intersection of task utility and minimal disclosure.
🧵👇
Can LLM agents build memory before seeing any user task?
Memory is usually built from human tasks or deployment interactions. New tool environments often have neither, creating cold-start gap.
Introducing PREPING: building agent memory without tasks.
https://t.co/bTV24GP4qc
ThinkSafe: Self-Generated Safety Alignment for Reasoning Models
A framework that restores safety in reasoning models without external teachers. It unlocks latent harm-identification knowledge via lightweight refusal steering, improving safety while preserving reasoning—with lower compute than GRPO.