Need scalable and efficient large language models for long sequences? Check our SPADE models in https://t.co/D190LCZw7U. By leveraging a state space layer, SPADE complements the lack of long-range dependency issue in transformer models using local attentions. (1/3)
Paper Title: "Fine-Tuning Pre-trained Language Model with Weak Supervision: A Contrastive-Regularized Self-Training Approach". Joint work with @yueyu30308379 @SimiaoZuo@jiang_haoming, Wendi Ren and @chaozhangcs
Working on my last rebuttal for EMNLP. Then I saw the following comment from a reviewer:
"The increase of the classification accuracy does not necessarily mean that wrong predictions are corrected, as the accuracy may depend on many other factors".
Struggling with fine-tuning BERT models? Overfit your tasks again?
Check our recent work on ACL 2020 "SMART: Robust and Efficient Fine-Tuning for Pre-trained Natural Language Models through Principled Regularized Optimization" https://t.co/JO2pxb8SiD
Frustrated by your LSTM-based point process models? Fail to capture long-term dependencies? Cannot scale to large data well?
Check our recent paper on ICML 2020 -- “Transformer Hawkes Process” at https://t.co/X5kpOfVTpj
@SimiaoZuo@jiang_haoming