π Hi from Papersβ.Dataβ.Code
We track ML research from arXiv, GitHub, and Hugging Face, surfacing what's worth your time.
Here you'll see:
π Daily ML picks: papers, repos, datasets
π Weekly TL;DR every Friday
π§΅ Monthly, quarterly, yearly threads at horizon turns
π± Tracking growth and impact of selected works
Selected, not collected.
Full feed β https://t.co/nGRf9ZUha0
Digests β https://t.co/wlpnEMU8qu
π ML Weekly Recap Β· Sep 28 β Oct 04
β‘ Trends
βΈ Test-time context or future-state conditioning improves long-horizon agents and world models
βΈ Efficiency-focused inference and attention redesigns cut latency without major quality loss
βΈ Large specialized benchmarks emphasize zero-shot generalization, grounding, and reproducible evaluation
π§ TL;DR
π TabFM: A Zero-Shot Foundation Model for Tabular Data
Weihao Kong, Erez Louidor Ilan, Shuxin Nie et al.
Zero-shot tabular FM beats tuned AutoGluon across 51 datasets.
π https://t.co/5FmajKO6dr
β CLM
Practical low-latency decision scoring matches Jev on real agent tasks.
β https://t.co/ZodnYGZQkF
π‘ Models increasingly externalize structure to gain efficiency, generalization, and controllability.
β https://t.co/wlpnEMTAAW
π Context Language Models
by Rulin Shao (@RulinShao), Shannon Zejiang Shen et al.
Rather than external memory policies, the LM edits its own context file in place. This makes context management a native behavior that can be prompted or RL-trained.
Key points:
β’ BrowseComp-Plus: +11.4% accuracy with 21.5% fewer FLOPs vs best baseline
β’ 12-hour EdgeBench: +5% score with 59% fewer FLOPs vs summarization
β’ RL on Qwen3.5-9B: 28.8% -> 42.5% on BrowseComp-Plus with 12% fewer FLOPs
Making context management intrinsic to the model yields better long-horizon agent behavior than fixed external memory policies, while also lowering compute.
π False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents
by Meijia Chen, Hao Li et al.
Cross-fitted proposer feedback replaces same-source evaluator reuse with source-excluded scoring, while leaving the main solver training rule unchanged.
Key points:
β’ Avg Cover-EM 48.8/51.2, +8.8/+8.4 pts vs Dr. Zero at 4B/9B
β’ False-agreement mass drops 6.1%β3.0% and 8.8%β3.7% with CrossFit
β’ Beats Search-R1 by 8.7/7.8 pts average across 7 benchmarks
The main failure in self-evolving search is feedback provenance: excluding same-source training from the evaluator sharply reduces shared-error reinforcement.
π§΅ The 3 #ML Datasets that shaped Q3 2026
1. secemp9/arxiv-complete
Unlike paper-text corpora, it combines 3.15M arXiv papers with source files, resolved TeX, and 5.03M-version history in one auditable 16.08 TB snapshot.
2. microsoft/XL-DocBench
Unlike broad corpora, it tests evidence-grounded QA over documents up to 2,935 pages with page-level evidence, cross-document questions, abstention, and only 38.36% top accuracy.
3. hamzabagirsakci/turkish-court-decisions
It offers 11 million CC0 full-text Turkish court decisions spanning 1962β2026 across five court systems.
Full breakdown β
π₯ XL-DocBench by microsoft
Contains 1,345 expert-verified QA records over 292 long professional documents, with page-level evidence for single-doc and cross-doc QA.
Key points:
β’ 1,345 total questions: 1,191 single-doc + 154 cross-doc
β’ 292 documents from 6 professional domains
β’ 429 multimodal-evidence questions; 188 require None answers
Growth since tracking: 24 days through Sep 27
π₯ 1k (+1k) Β· β€οΈ 7 (+1) Β· π 0 (+0)
π https://t.co/cIbIgWXrUw
π https://t.co/lKbbSJMHHA
π PETARSeg-11K by UW-Madison-Dept-Radiology
Contains 11K-scale PET/CT data with lesion-level correspondences between localized abnormalities and free-text radiology findings.
Key points:
β’ 11K-scale dataset
β’ PET/CT imaging paired with text
β’ Lesion-level spatial grounding annotations
Useful for training and evaluating lesion-grounded PET/CT report generation and vision-language models.
π₯ 126 downloads
π TabFM: A Zero-Shot Foundation Model for Tabular Data
by Weihao Kong, Erez Louidor Ilan et al.
A synthetic-data-trained in-context transformer replaces per-dataset tabular training with single-pass zero-shot prediction on real tables.
Key points:
β’ Overall TabArena: 1785 Elo, above AutoGluon 1.5 extreme at 1676
β’ Classification: 1768.6 Elo vs TabPFN-3 1641.9 and AutoGluon 1669.7
β’ Regression: 2055.2 Elo vs EXAONE-Tabular 1973.1 and TabPFN-3 1866.6
Shows that large tabular models trained purely on synthetic causal tables can transfer zero-shot to real-world datasets at or above tuned AutoML.
π§΅ The 3 #ML Repos that shaped Q3 2026
1. deepseek-ai/DeepSpec
Unlike single-method decoding repos, it unifies data prep, multi-GPU training, and evaluation across DSpark, DFlash, Eagle3, and four target-model setups.
2. jaredpalmer/kev
Unlike recent typed-decision repos, it packs many question types into one calibrated, decode-free pass with 2.0Γ speedup and 4e-6 output parity.
3. Tencent/WeMM-Embedding
Unlike single-modality embedders, it unifies five input types in one space and retains 98.7% image/video performance at 256 dimensions.
Full breakdown β
π₯ jaredpalmer/kev
Reads one state once and scores many typed questions in parallel with exact question isolation, outputting probabilities instead of generated text.
Key points:
β’ Packed vs separate max delta is 3.7e-6, with packed 2.0Γ faster
β’ kev-4b gets 0.806 locked-test OOD accuracy and serves in ~1 s on M5
β’ Supports noul, choice, and score questions with 2-255 options
Growth since tracking: 3 days through Sep 27
β 7.3k (+1.3k) Β· π΄ 437 (+117) Β· π 0 (+0)
β https://t.co/So4ZnA7arL
π EmbRACE by mxlin043
Contains 3,421 egocentric human demonstrations in 55 environments and a 686-task benchmark in 7 more, for closed-loop embodied navigation/manipulation.
Key points:
β’ 3,421 dataset tasks + 686 benchmark tasks
β’ 48,264 rationale-labeled steps; 10,552 benchmark steps
β’ 62 UE environments total; frames are 640x480; CC-BY-NC-4.0
Useful for training and evaluating embodied agents that must act from egocentric vision and language in closed loop.
π₯ 76 downloads