Introducing StarChat2 - Your new coding buddy! 🙌 We are excited to release StarChat2 a fine-tuned @BigCodeProject Starcoder2 model with enhanced assistant and copilot skills to answer all your coding questions. 🚀🎉
StarChat2 can help you:
🙋���♂️ Answer coding questions in over 216 languages, including Python, Java, C++ and more!
🧠 Explain concepts and help debug your code.
📊 Generate sample code for data visualizations and plots in Python
💬 Iterate together to solve your coding errors
WOW. Stability release a image-to-3d model with Tripo AI. Runs under low inference budgets (even without GPU) and have MIT License with weights and source code available for download.
Super proud of what the @BigCodeProject community achieved. Building the best in class code LLMs in an open and collaborative way is no easy feat and is the result of the hard work of many community members!
Copilot Evaluation Harness
Evaluating LLM-Guided Software Programming
The integration of Large Language Models (LLMs) into Development Environments (IDEs) has become a focal point in modern software development. LLMs such as OpenAI GPT-3.5/4 and Code Llama offer the potential to significantly augment developer productivity by serving as intelligent, chat-driven programming assistants. However, utilizing LLMs out of the box is unlikely to be optimal for any given scenario. Rather, each system requires the LLM to be honed to its set of heuristics to ensure the best performance. In this paper, we introduce the Copilot evaluation harness: a set of data and tools for evaluating LLM-guided IDE interactions, covering various programming scenarios and languages. We propose our metrics as a more robust and information-dense evaluation than previous state of the art evaluation systems. We design and compute both static and execution based success metrics for scenarios encompassing a wide range of developer tasks, including code generation from natural language (generate), documentation generation from code (doc), test case generation (test), bug-fixing (fix), and workspace understanding and query resolution (workspace). These success metrics are designed to evaluate the performance of LLMs within a given IDE and its respective parameter space. Our learnings from evaluating three common LLMs using these metrics can inform the development and validation of future scenarios in LLM guided IDEs.
@GoogleDeepMind@ancadianadragan@UCBerkeley So this is great news but the question remains, how much autonomy will you have to implement safety and will DM support you when priorities conflict ☺️
Welcome MLX-Swift! 🍎
Bringing MLX closer to iOS/ Mac developers. Allows you to leverage the metal backend efficiently across Apple devices!
Comes with a ready-to-use example for inferring with llama 7B quantised checkpoints. 🔥
GG @awnihannun and team! Y'all are setting the bar for maintaining quality whilst shipping fast! ⚡
🚨How good is your LLM at game theory?
Check out GTBench, to evaluate LLMs on logic+strategy games!
LLM-vs-LLM competition in 10 board+card games e.g Kuhn Poker, Tic-Tac-Toe, Breakthrough etc ➡️ LLMs overall lag on GT tasks & code-pretraining helps.
https://t.co/5T2uBlpqmN
🧵
Open Source Video to Sound Effects 🎞️👂has landed on @huggingface 🤗
Follow @fffiloni there (https://t.co/6fnQqZnK2F) to keep track of updates and feature request
Demo link available in quoted post 😉
Great series. I really enjoyed learning about the life of a great physicist and interesting character. More of these types of episodes in the pipeline?
Freakonomics Radio is No. 1 in the Documentary category on @ApplePodcasts! Check out our series “The Curious, Brilliant, Vanishing Mr. Feynman” https://t.co/yPgv4xYveb
#BreakingNews Three survivors have escaped and found freedom from sex trafficking in South Africa!
After being recruited, these survivors endured unimaginable exploitation. Thank you for partnering with us to provide care, safety and support for survivors all across the world.
4x faster Llama inference! 🔥
> leverages static cache.
> uses torch compile for decoder models.
> very minimum code changes required.
> coming to mistral and other models soon.
> opens possibility to unlock even more speed-ups.
massive kudos to @art_zucker for working on this beast! 🤗
left: dynamic cache, transformers 4.37
right: static cache + torch compile, transformers (main)
Argus-3D
Pushing Auto-regressive Models for 3D Shape Generation at Capacity and Scalability
Auto-regressive models have achieved impressive results in 2D image generation by modeling joint distributions in grid space. In this paper, we extend auto-regressive models to 3D domains, and seek a stronger ability of 3D shape generation by improving auto-regressive models at capacity and scalability simultaneously. Firstly, we leverage an ensemble of publicly available 3D datasets to facilitate the training of large-scale models. It consists of a comprehensive collection of approximately 900,000 objects, with multiple properties of meshes, points, voxels, rendered images, and text captions. This diverse labeled dataset, termed Objaverse-Mix, empowers our model to learn from a wide range of object variations. However, directly applying 3D auto-regression encounters critical challenges of high computational demands on volumetric grids and ambiguous auto-regressive order along grid dimensions, resulting in inferior quality of 3D shapes. To this end, we then present a novel framework Argus3D in terms of capacity. Concretely, our approach introduces discrete representation learning based on a latent vector instead of volumetric grids, which not only reduces computational costs but also preserves essential geometric details by learning the joint distributions in a more tractable order. The capacity of conditional generation can thus be realized by simply concatenating various conditioning inputs to the latent vector, such as point clouds, categories, images, and texts. In addition, thanks to the simplicity of our model architecture, we naturally scale up our approach to a larger model with an impressive 3.6 billion parameters, further enhancing the quality of versatile 3D generation. Extensive experiments on four generation tasks demonstrate that Argus3D can synthesize diverse and faithful shapes across multiple categories, achieving remarkable performance.
Chain-of-Thought Reasoning Without Prompting
paper page: https://t.co/o5fcJqa20L
In enhancing the reasoning capabilities of large language models (LLMs), prior research primarily focuses on specific prompting techniques such as few-shot or zero-shot chain-of-thought (CoT) prompting. These methods, while effective, often involve manually intensive prompt engineering. Our study takes a novel approach by asking: Can LLMs reason effectively without prompting? Our findings reveal that, intriguingly, CoT reasoning paths can be elicited from pre-trained LLMs by simply altering the decoding process. Rather than conventional greedy decoding, we investigate the top-k alternative tokens, uncovering that CoT paths are frequently inherent in these sequences. This approach not only bypasses the confounders of prompting but also allows us to assess the LLMs' intrinsic reasoning abilities. Moreover, we observe that the presence of a CoT in the decoding path correlates with a higher confidence in the model's decoded answer. This confidence metric effectively differentiates between CoT and non-CoT paths. Extensive empirical studies on various reasoning benchmarks show that the proposed CoT-decoding substantially outperforms the standard greedy decoding.