Congrats to the Causica team @MSFTResearchCam . We got two papers accepted in #ICML 24, contributing to our efforts on integrating #Causality with modern #FoundationModels. 1/3
TL;DR: For the potential outcome framework, self-attention = causal inference via optimal balancing [1]; For the SCM framework, SCM learning = fixed-point problem given TO = a specific causally-consistent attention mechanism [2].
Both discoveries contributes towards enabling transformer-like architecture to directly solve causal reasoning tasks, even when presented with an unseen dataset in a foundational setting. @pwnic@JiaqiZhangVic@ScetbonM@agrinh
[1] Towards Causal Foundation Model: on Duality between Causal Inference and Attention (https://t.co/oXzvfSSZ8e)
[2] FiP: a Fixed-Point Approach for Causal Generative Modeling (https://t.co/FwyawE2a92)
Models learned via gradient-descent can be very robust and nonrobust, depending on problem structure. Checkout new preprint with @ScetbonM (to appear in AISTATS 2023): "Robust Linear Regression: Gradient-Descent, Early-Stopping, and Beyond" https://t.co/52nPkKe0g0. Ping @MetaAI
#AISTATS2023 accepted papers:
1) "Robust Linear Regression : Gradient-descent, Early-stopping, and Beyond", by
@ScetbonM & E.D. Arxiv preprint coming soon.
2) "Origins of Low-Dimensional Adversarial Perturbations", by E.D., C. Guo, &
@MGoibert. Preprint https://t.co/rlYobnYNRr
Triangular Flows for Generative Modeling: Statistical Consistency, Smoothness Classes, and Fast Rates
https://t.co/gPiG9JXaN7
by Nicholas J. Irons et al. including @ScetbonM#Estimator#Statistics
Congratulations to @SchreuderNico and @echzhen for being selected as runner-up for best student paper award #UAI2021
Classification with abstention but without disparities
Nicolas Schreuder, Evgenii Chzhen
https://t.co/kuvsVprhEH
Comparing distributions: Kernels estimate good representations, l1 distances give good tests
A simple summary of our #NeurIPS2019 work
https://t.co/xT7cp88iW6
Given two set of observations, how to know if they are drawn from the same distribution? Short answer in the thread..
Comparing distributions: ℓ1 geometry improves kernel two-sample testing
Our #NeurIPS2019 paper with @ScetbonM
https://t.co/BErD9jFgiA
Revisit metrics between distributions based on kernel embedding (inspired by MMD) and show benefits of ℓ1 distances of embeddings: |μ₁ - μ₂|
The K-SVD algorithm (2006) was, for a short while, the state-of-the-art in denoising. But over the years it has been surpassed by many newcomers. Here @ScetbonM, Miki Elad, and I take it deeper and make it a lot better.
https://t.co/V7pfdz6P7t