BREAKING: Email platforms are officially obsolete.
Mailchimp charges $300+ for what this open-source tool gives you free.
BillionMail gives you unlimited sending, full control, & no monthly fees.
Hereâs how to set it up in 5 minutes:
Today, @OpenAI launched their Agents SDK developer framework.
Until now, agents have mostly been chat-oriented (chat in, chat out), but agents will increasingly be action-oriented (data in, action out).
We, @stripe, would love to show you a few financial agents weâve built! —ïž
HOLY SHITT, Microsoft dropped an open-source Multimodal (supports Audio, Vision and Text) Phi 4 - MIT licensed! đ„
> Beats Gemini 2.0 Flash, GPT4o, Whisper, SeamlessM4T v2
> Models on Hugging Face hub, integrated with/ Transformers!
Phi-4-Multimodal:
> Modalities: Integrates text, vision, and speech/audio
> Architecture: Uses "Mixture of LoRAs" to add modality-specific adapters without fine-tuning the base model
> Vision Modality: SigLIP-400M image encoder, 2-layer MLP projector, dynamic multi-crop strategy
> Speech/Audio Modality: 3-layer convolution, 24 conformer blocks, 80ms token rate
> Performance: Ranks first on OpenASR leaderboard, supports vision+language, vision+speech, and speech/audio tasks, outperforming larger models
Phi-4-Mini:
> Parameters: 3.8 billion
> Architecture: 32 Transformer layers, 3,072 hidden state size, Group Query Attention (GQA) with 24 query heads and 8 key/value heads
> Vocabulary: 200K tokens for multilingual support.
Training Data: High-quality web and synthetic data, emphasizing math and coding
> Performance: Outperforms similar-sized models and matches larger models (e.g., DeepSeek-Rl-Distill-Qwen-7B) on math and coding tasks
Training Pipeline:
> Language Training: Pre-training on 5 trillion tokens, post-training with function calling, summarization, and instruction-following data
> Multimodal Training: Vision training (4 stages), speech/audio training (2 stages), and joint vision-speech training
> Reasoning Training: Pre-trained on 60B CoT tokens, fine-tuned on 200K high-quality CoT samples, and DPO-trained on 300K preference samples
Vision Benchmarks:
> Outperforms Phi-3.5-Vision, Qwen2.5-VL, InternVL2.5, and matches Gemini and GPT-4o on tasks like chart understanding and OCR
> Vision-Speech Benchmarks: Significantly outperforms InternOmni and Gemini-2.0-Flash
Speech Benchmarks:
> ASR: Achieves SOTA on CommonVoice, FLEURS, and Open ASR Leaderboard, surpassing WhisperV3 and SeamlessM4T
> AST: Best performance on CoVoST2, comparable to GPT-4o on FLEURS
> Speech Summarization: First open-source model with this capability, close to GPT-4o in quality
Language Benchmarks:
> Outperforms similar-sized models (Llama-3.2, Ministral) and matches larger models (Qwen2.5-7B) on math, reasoning, and coding tasks
> Coding: Strong performance on HumanEval, MBPP, and BigCodeBench
Reasoning Benchmarks:
> Reasoning-enhanced Phi-4-Mini outperforms DeepSeek-Rl-Distill-Llama-8B and matches DeepSeek-Rl-Distill-Qwen-7B on AIME, MATH-500, and GPQA Diamond
I got early access to ChatGPT Operator.
It's OpenAI's new AI agent that autonomously takes action across the web on your behalf.
The 9 most impressive use cases Iâve tried (videos sped up):
1. Ordering dinner ingredients based on a picture and a recipe
AWS released a new Multi-Agent AI framework!
A flexible and powerful framework for managing multiple AI agents and handling complex conversations.
It let's you dynamically route LLM queries to the most suitable agent based on context.
100% Open Source.
Can you launch your startup in 1 week?
Absolutely!
Actually, you should.
I've built 14+ products and scaled one to $1,000,000 ARR
Let me teach you how to build one in a week:
Every highly repetitive legal service that is not built into a piece of software will be built into a piece of software. How much of the legal industry is software waiting to happen?