Introducing Gacha Decoding ๐ฒ: we show how to diversify LM responses, beating prior work by >11x in sample efficiency.
Reroll your LM for varied ideas, data, RL envs, {your use here}!
We get diversity from instruction following + external RNG, not token entropy. Here's how ๐งต
โผ๏ธThe Bitter Lesson for context management: Giving LMs unrestricted control over their context beats human-designed SOTA!
Introducing ๐ฉตContext Language Models (CLMs)๐ฉต
- Natively manage their own context
- Treat context as a file
- Learn policies in CLM weights, no harness
Introducing โ ๐๐ป๐ฐ๐ต๐ผ๐ฟ๐ฒ๐ฑ ๐๐ฒ๐ฐ๐ผ๐ฑ๐ถ๐ป๐ด: a copyright mitigation strategy for any language model! With @uwnlp
LMs today reproduce copyrighted textโraising concerns for creator consent and potential legal (and ๐ธ ๐ธ) liabilities for AI developers. ๐ซ
๐๐ป๐ฐ๐ต๐ผ๐ฟ๐ฒ๐ฑ ๐๐ฒ๐ฐ๐ผ๐ฑ๐ถ๐ป๐ด relies on two off-the-shelf LMs:
๐งผA ๐๐ฎ๐ณ๐ฒ ๐๐ trained only on permissively licensed text,
โ ๏ธA higher-utility ๐ฟ๐ถ๐๐ธ๐ ๐๐ trained on any data.
The ๐ฟ๐ถ๐๐ธ๐ ๐๐ drives generation, but the ๐๐ฎ๐ณ๐ฒ ๐๐ acts as an anchor. If the ๐ฟ๐ถ๐๐ธ๐ ๐๐ drifts into memorization, the ๐๐ฎ๐ณ๐ฒ ๐๐ pulls it back โฉ๏ธ.
๐คWe provide a formal guarantee: outputs stays within a user-set budget of the ๐๐ฎ๐ณ๐ฒ ๐๐ .
Details below! ๐
[1/โ]
๐Introducing โStochasTok: Improving Fine-Grained Subword Understanding in LLMsโ!๐
LLMs are incredible but still struggle disproportionately with subword tasks, e.g., for character counts, wordplay, multi-digit numbers, fixing typosโฆ Enter StochasTok, led by @anyaasims!
[1/]
Join us to hear about the BANKSY spatial omics clustering method at the Sydney Precision Data Science Centre's bioinformatics seminar on Monday, 12 May, at 1 pm Sydney time. Link: https://t.co/PoJgi7kbZn
Just out in Cell: the Asian Immune Diversity Atlas (AIDA) is a 5-nation single-cell blood atlas. We show that humans are remarkably diverse in gene expression and cell phenotype, with disease implications #healthcaredisparities#PrecisionMedicine [https://t.co/Rz3ENKb1vm]. [1/14]
Thatโs a wrap at @icmlconf 2024! Tons of exciting work, old friends and new, and many fan girl moments!
Main takeaway: fewer but more generalizable, efficient foundation models, and watch out for xLSTMs (I have no idea what they are)!
Wonderful to see BANKSY incorporated into this @satijalab spatial data analysis workflow. Looking forward to trying out the other parts. Well deserved recognition for BANKSYโs developers @vipul1891@NigelChouS@jleechung!
BANKSY
A #SpatialOMics R/Python analysis framework using Transcriptomic Neighborhood๐ค (Not spatial coordinates Only)
1โฃDistance-weighted mean gene expression of each cell's neighborhood
2โฃAzimuthal Gabor filterโถ๏ธGradient of gene expression in each cell's neighborhood
Work with SlideSeq VISIUM MERFISH VERAFISH CosMx CODEX
vs Giotto MERINGUE FICT SpiceMIX BayesSpace STAGATE GraphST SpaGCN
Can Perform nonspatial & spatial clustering by adjusting "ฮปโโโ[0,โ1], to weight the contributions of the cell-transcriptome matrix and the neighbor expression matrices"
Can handle ~2 M cells in ~1 h!
@vipul1891@khchenlab@shyam_lab@NatureGenet 2024
https://t.co/UW9HYJtva7
1. BANKSY is out! https://t.co/uy8myuVA5p
Our spatial clustering algo applies to any spatial RNA/protein/... assay, scales to 2M cells, detects both cell typing and spatial domains and facilitates spatial batch correction.
๐งต ๐