π CTO | Visionary AI Leader | Pioneering Generative AI Innovation π | Mastermind in LLMs & AI Agents π€ | Empowering the Future of Tech π‘ #AILeader#NLP#AI
Iβm excited to host an exclusive webinar βLaunching Your AI Journey: From Curious to Capableβ
π Date: May 21st
π Time: 7:00 - 8:00 PM IST
π Register now: https://t.co/xcHQtzb9mX
Exclusive Webinar! π
Join us for a live demo & see how Aracor AI transforms contract review!
π Feb 27, 1 PM EST | π₯οΈ Online (Zoom)
π Exclusive trial access for attendees!
π Register now: https://t.co/M4CDRX5NRu
#AracorAI#aiduediligence#futureoflaw#ailegalassistant
Iβm truly honored to have my journey and recent milestone featured in this news article.
Check out the feature here - https://t.co/8XlYb76jyG
Letβs keep building on this momentum, innovating together, and dreaming even bigger as we strive for greater achievements!
#AI#CTO
Meet Aracor's New CTO, @leslyarun
Lesly shares his vision for transforming VC & SaaS with AI innovation:
βοΈ Advanced OCR
βοΈ Optimized Deal Rooms
βοΈ Privacy First
βοΈ Streamlined Redaction
#AracorAI#LegalTech#AIInnovation
@AracorLegal Grateful to step into the role of CTO! A huge thank you to @KatyaLawyer for believing in me and giving me this incredible opportunity. Excited for the challenges ahead and to drive innovation with an amazing team! Letβs make it happen! #Leadership#CTO#Innovation#AI
@AracorLegal Honored and Thrilled to Step into the Role of CTO
I am truly grateful for the opportunity to take on this exciting new chapter as CTO at @AracorLegal. This role represents not just a promotion but a chance to lead, innovate, and make a meaningful impact on the future of tech.
Exciting news! π
Lesly Arun is now Aracor's CTO! π With a stellar AI background at AstraZeneca & projects with Verizon & Schneider Electric, Leslyβs vision will lead Aracor into the future of legal tech. Join us in celebrating this milestone! π
#Leadership#AIInnovation#CTO
Thank you to everyone who attended the DSPy meetup and those of you who registered! We are so excited to share the recording π
TL;DR of the talks:
1. @simigd from @cohere gave an overview of Cohereβs latest developments in RAG from Command R+ to embeddings and more
2. @cshorten30 from @weaviate previewed gfl-dspy, building Generative Feedback Loops with DSPy
3. @mikeldking from @arize presented the evolution of observability and ops with Arize and DSPy
4. @lateinteraction from @stanfordnlp gave a deep-dive overview of DSPy, from Language Programs, the new category of ML models, to the DSPy programming model, DSPy optimizers, comparison to PyTorch, and so much more
Video: https://t.co/9tPEQKIbtH
LLaMA-70b inferencing using only a single GPU and achieving 1.69x-2.65x higher normalized inference throughput than the FP16 baseline. with Six-bit quantization (FP6) π₯
Deepspeed has just recently released this Paper and also integrated the FP6 quantization -
"FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design" β¨
FP6 quantization is a practical alternative to further democratize the deployment of LLMs without significantly sacrificing model quality on complex tasks and various model sizes.
π Six-bit quantization (FP6) can effectively reduce the size of large language models (LLMs) and preserve the model quality consistently across varied applications. However, existing systems do not provide Tensor Core support for FP6 quantization and struggle to achieve practical performance improvements during LLM inference.
π It is challenging to support FP6 quantization on GPUs due to (1) unfriendly memory access of model weights with irregular bit-width and (2) high runtime overhead of weight de-quantization.
π To address these problems, this paper proposes TC-FPx, the first full-stack GPU kernel design scheme with unified Tensor Core support of float-point weights for various quantization bit-width.
π TC-FPx breaks the limitations of the underlying GPU hardware, allowing the GPU to support linear layer calculations on model weights of arbitrary bit width. By increasing the number of bit-width options for efficient quantization, TC-FPx significantly mitigates the "memory wall" challenges of LLM inference. In TC-FPx, Tensor Cores are utilized for intensive computation of matrix multiplications, while SIMT cores are effectively leveraged for weight dequantization, transforming the x-bit model weights to FP16 type during runtime before feeding them to Tensor Cores.
π This paper, integrates TC-FPx kernel into an existing inference system, providing new end-to-end support (called FP6-LLM) for quantized LLM inference, where better trade-offs between inference cost and model quality are achieved.
π They also found that INT4 quantization heavily relies on Fine-Grained Quantization (FGQ) methods to maintain high model quality, whereas our FP6 quantization already works well on coarse-grained quantization.
π On average, the TC-FPx kernel demonstrates a 2.1-fold improvement in processing speed over the FP16 cuBLAS benchmark during memory-intensive General Matrix Multiply (GEMM) operations on NVIDIA A100 GPUs.
π Note, TC-FPx kernel only supports NVIDIA Ampere GPUs and is only tested and verified on A100 GPUs
Tired of resource-heavy LLMs?
MobiLlama offers a powerful alternative for resource-constrained environments, achieving competitive performance with significantly lower training costs.
#SLM#MachineLearning#AI#OpenSource#Transparency#chatgpt#nlp
Are high GPT-4 API costs and lengthy prompts hindering your work?
Discover the power of Prompt Compression with LLMLingua & LongLLMLingua.
A Thread π§΅
#AI#MachineLearning#Innovation#ChatGPT#GPT4