I am a researcher and entrepreneur currently working on continual learning architecture. Previously founded Vedabio (YC S21). PhD BioE from UIUC - 50+ patents
Catastrophic forgetting in continual learning has an architectural solution.
I’m excited to release TFGN: a new transformer overlay that lets LLMs keep learning new domains (Prose → Python → Math → Biomedical → Chinese → JavaScript, 1B tokens per domain) with no replay buffer and no task IDs, while preserving prior knowledge at LLM scale.
Tested from ~398M to ~9B parameters.
Tightest result: BWT = −0.007 on LLaMA 3.1 8B Retrofit.
Full paper (arXiv): https://t.co/NVQABvRsAu
Readable explainer: https://t.co/VIXycvOV6i
Would love your thoughts, this is my first major ML paper.
What's coming next:
→ Public API (demo initially)— train TFGN on your own data/codebase without forgetting
→ Surgical unlearning demo
→ Safety & alignment applications (§8.5 of the paper points at this)
→ Robotics: continual skill acquisition
→ 70B+ validation — the milestone that converts research-scale to production-ready
I have self-funded this work and all compute so far, and am now identifying funding opportunities to enable future research, and deployment in this direction.
If this intersects your work in continual learning, alignment, agentic systems, or enterprise AI I'd love your feedback, your pushback, and your questions.
Catastrophic forgetting in continual learning has an architectural solution.
I’m excited to release TFGN: a new transformer overlay that lets LLMs keep learning new domains (Prose → Python → Math → Biomedical → Chinese → JavaScript, 1B tokens per domain) with no replay buffer and no task IDs, while preserving prior knowledge at LLM scale.
Tested from ~398M to ~9B parameters.
Tightest result: BWT = −0.007 on LLaMA 3.1 8B Retrofit.
Full paper (arXiv): https://t.co/NVQABvRsAu
Readable explainer: https://t.co/VIXycvOV6i
Would love your thoughts, this is my first major ML paper.
Two further capabilities fall out of the same substrate, with no changes to it:
• Extension A — a closed-loop self-regulating layer that decides when to update vs. consolidate from the model's own internal signals.
• Extension B — operator-level plan vectors that reshape the effective weights at 99.96% cosine fidelity.
Details in the paper.