ChatGLM2-6B
model on @huggingface: https://t.co/5eMmyUrgar
ChatGLM 2-6B introduces the following new features :
Stronger Performance: Based on the development experience of the first-generation ChatGLM model, we have fully upgraded the base model of ChatGLM2-6B. ChatGLM2-6B uses the hybrid objective function of GLM, and has undergone pre-training with 1.4T bilingual tokens and human preference alignment training. The evaluation results show that, compared to the first-generation model, ChatGLM2-6B has achieved substantial improvements in performance on datasets like MMLU (+23%), CEval (+33%), GSM8K (+571%), BBH (+60%), showing strong competitiveness among models of the same size.
Longer Context: Based on FlashAttention technique, we have extended the context length of the base model from 2K in ChatGLM-6B to 32K, and trained with a context length of 8K during the dialogue alignment, allowing for more rounds of dialogue. However, the current version of ChatGLM2-6B has limited understanding of single-round ultra-long documents, which we will focus on optimizing in future iterations.
More Efficient Inference: Based on Multi-Query Attention technique, ChatGLM2-6B has more efficient inference speed and lower GPU memory usage: under the official implementation, the inference speed has increased by 42% compared to the first generation; under INT4 quantization, the dialogue length supported by 6G GPU memory has increased from 1K to 8K.
Textbooks Are All You Need
paper page: https://t.co/E6D7tLK7mv
introduce phi-1, a new large language model for code, with significantly smaller size than competing models: phi-1 is a Transformer-based model with 1.3B parameters, trained for 4 days on 8 A100s, using a selection of ``textbook quality" data from the web (6B tokens) and synthetically generated textbooks and exercises with GPT-3.5 (1B tokens). Despite this small scale, phi-1 attains pass@1 accuracy 50.6% on HumanEval and 55.5% on MBPP. It also displays surprising emergent properties compared to phi-1-base, our model before our finetuning stage on a dataset of coding exercises, and phi-1-small, a smaller model with 350M parameters trained with the same pipeline as phi-1 that still achieves 45% on HumanEval.
OpenLLaMA 13B Released
model: https://t.co/n1vUb2I9wx
present a permissively licensed open source reproduction of Meta AI's LLaMA large language model. We are releasing 3B, 7B and 13B models trained on 1T tokens. We provide PyTorch and JAX weights of pre-trained OpenLLaMA models, as well as evaluation results and comparison against the original LLaMA models.
LLM Agents in Group Settings w/ Humans & AI
Pairwise-trained LLMs lack:
-being able to decide when to talk
-being coherent grounded on multiple characters
-Build new data of group conversations
-SoTA on consistency, engagingness, identity, sensibility
https://t.co/EFvfZ48UcL
Tractable Control of LLMs
-Sampling from Pr(text|α) is intractable for even simplest constraints α
-Use distilled hidden Markov models to control generations from LLM
-SoTA performance on challenging benchmark for constrained text generation, CommonGen
https://t.co/MQO7IUfnvD
[CL] Pretrain on just structure: Understanding linguistic inductive biases using transfer learning
I Papadimitriou, D Jurafsky [Stanford University] (2023)
https://t.co/f5HrYx6FO0
LlamaAcademy: Fine-tuning LLMs to Learn How to Talk to APIs
Pipeline:
-Crawling
-GPT-4 data gen
-Fine-tuning Vicuna-13B on synthetic data
LLM can then read new API docs (Stripe Notion etc), gen code
Instead of hosting API docs, host API implementation
https://t.co/Uq2oZG8uty
Can Large Language Models (#ChatGPT) transform Computational Social Science?
Our recent work shows how they might (in partnership w/ experts).
We evaluate on 24 #CSS tasks + draw a roadmap 🚗🗺️ to guide #LLM-augmented social science 🚀
Paper: https://t.co/bGCn83PWie
🧵 thread
"The Little Book of Deep Learning"
Consider this as a beta version rough on the edges.
Comments are welcome.
https://t.co/BgNLBplHDt
@unige_en@sciences_UNIGE
Presenting TiDE, a time-series dense encoder for long-term time-series forecasting that enjoys the simplicity and speed of linear models while also being able to handle covariates and non-linear dependencies. Learn more →https://t.co/9y8kmN2lV2