A 744-billion-parameter AI model running on a regular computer with only 25 GB of RAM—and no GPU. 🤯
Colibrì is a lightweight, open-source inference engine written in pure C that can run GLM-5.2 on consumer hardware.
Instead of loading the entire model into memory, it keeps RAM usage low by streaming only the required Mixture-of-Experts components directly from an SSD while generating each token.
The impressive part:
⚡ Pure C
📦 Zero runtime dependencies
💾 Around 25 GB of RAM
🧠 744B total parameters, with roughly 40B activated
🚫 No expensive GPU required
There is an important trade-off: this is not fast inference, and the quantized model files still require hundreds of gigabytes of SSD storage. But as a technical proof of concept, it is remarkable.
It shows that running enormous AI models locally may depend as much on clever memory management as it does on expensive hardware.
Tiny engine. Massive model. Very smart engineering. 🐦💻
Repository link below 👇
The model is officially listed as a 744B-parameter MoE with about 40B active
A serious open-source alternative to Vercel and AWS has arrived. 🚀
OpenShip lets you build, deploy, manage, and scale applications on infrastructure you fully control—all from one platform.
No vendor lock-in, more ownership, and 100% open source. 🔥
Introducing OpenShip.
An open-source application platform for building, deploying, operating, and scaling applications on infrastructure you own.
Replace deployment tools, managed services, and infrastructure workflows with one open-source platform.
Available today:
- Mail Server (Built in. One click).
• Runs on your own VPS
• Unlimited domains
• Unlimited mailboxes
• Modern webmail included
• Connect with Gmail, Outlook, Apple Mail, Thunderbird, or any IMAP/SMTP client
• Send email directly from your applications using SMTP
• No mailbox subscriptions or API fees
• High email deliverability
- Deployment:
• Deploy any stack
• Git-based deployments
• Zero-downtime deployments
• One-click rollbacks
• Development, Staging, and Production environments
• Multi-branch deployments with isolated environments • Deploy to VPSs, dedicated servers, cloud VMs, or your homelab
- Services:
Provision the services your applications need in one click.
Replace multiple managed providers with services running on your own infrastructure.
• Supabase
• PostgreSQL
• MySQL
• MariaDB
• MongoDB
• Redis
• MinIO
• Meilisearch
• Qdrant
• RabbitMQ
• Kafka
• ClickHouse
• Elasticsearch
...and deploy any other service alongside your applications.
- Operate:
• Live deployment logs
• Live request logs
• Real-time traffic analytics
• Automated backups
• Monitoring
• Secrets management
• Domains
• Automatic SSL
• Environment variables
• Scheduled jobs
• Health checks
-Multi Environments:
• Separate Development, Staging, and Production environments
• Deploy every branch independently
• Test changes before production
• Isolated services, secrets, and configuration per environment
- Security & Teams:
• Team management
• Role-based access control
• IP allow/block rules
• Rate limiting
• Security rules
- Developer Experience:
• Web dashboard
• Native desktop application
• CLI
• REST API
• MCP support for AI agents, just add the mcp and your agent can do the work for you
• Manage your infrastructure without living in SSH
- Coming Soon:
• Multi-server clustering for applications and databases
• One-click load balancing
• Horizontal scaling across multiple servers
• Built-in high availability and failover
• Scale from a single VPS to a cluster with the same workflow
• Just add servers. OpenShip handles the rest.
Open source.
Europe’s open-source AI race is heating up. 🇪🇺🔥
The Soofi alliance has introduced Soofi S 30B-A3B, a new model reportedly ranking among the strongest fully open models worldwide.
High capability, strong efficiency, and a serious challenge to the current leaders. 🚀
China just dropped Ling-2.6-1T, an open-source coding model with 72.2% on SWE-bench Verified, a 256K context window, and Claude Code support—while using far fewer tokens. 🚀
China just changed the game 🤯
They just dropped an open-sourced model that burns only 1% of the tokens compared to your favorite US models.
It's called Ling-2.6-1T and it goes toe-to-toe with Claude and GPT on agentic coding.
→ 1 trillion params
→ 72.2% on SWE-bench Verified
→ 256K context window
→ Free on OpenRouter
→ Plugs directly into Claude Code
100% Open Source.
A new Gemini 3.5 Pro benchmark leak is turning heads across the AI community. 👀
According to the reports, Google delayed the release until July 17 to retrain the model on a stronger architecture rather than shipping a minor upgrade.
The early results suggest a major jump in mathematical reasoning and SVG generation.
If these numbers are accurate, Gemini 3.5 Pro could be one of Google’s most significant model upgrades yet. 🚀
👀 Gemini 3.5 Pro benchmark leak just dropped.
Reportedly:
• Beating Claude Fable 5 in internal evals
• Beating GPT-5.6 in internal evals
• Massive gains over Gemini 3.1 Pro
• Public rollout preparations underway
• July 17 launch target
If these numbers survive real-world testing, Google isn't catching up.
Google is taking the lead.
Fable 5 may have just become the model everyone is chasing.
Brian just compressed GLM-5.2’s 753B weights from 1.4TB down to 980GB—with no quantization, retraining, or loss in accuracy. 🤯
A 30% reduction while staying bit-exact makes running it on a 3× DGX Spark setup feel far more realistic.
I removed 423 GB from GLM‑5.2 without changing the model.
1,403 GB → 980 GB.
753B weights.
Bit for bit exact.
No quantization or retraining.
The weights remain compressed in VRAM instead of rebuilding the full model first.
Full writeup and repo in the next post.
This free, open-source tool turns OpenStreetMap data into clean, editable SVG map layers—ready to use in Illustrator, Inkscape, or Affinity Designer. 🗺️
The latest update also brings full Windows support, easier installation across macOS, Windows, and Linux, plus a redesigned interface for smoother exports and reusable map presets.
A very useful tool for designers, developers, and anyone who works with custom maps. ⚡
You may not need to spend thousands on marketing consultants anymore. 💡
This free AI tool lets you present your strategy to a simulated council of 12 legendary marketers, including Ogilvy and Hormozi.
It even assigns a “devil’s advocate” to challenge your plan, uncover weak points, and expose risky assumptions before you launch and waste money in the market. 🧠
Marketing Skills v2.8.0 is live 🟢
New skill: /marketing-council — pitch your marketing to a simulated board of 12 legendary marketers.
+ Seth Godin, David Ogilvy, Alex Hormozi, April Dunford & more
+ a designated dissenter in every session
47 skills. Free & open source 👇
An Anthropic engineer just shared a ready-made coding-agent setup with 49 tested skills. ⚡🧠
It follows a simple workflow:
Plan → Build → Verify
Compatible with Claude Code, Cursor, Codex, and Gemini, it could help one developer work with the structure and speed of a much larger team. 🚀
Anthropic is making it harder and harder to justify paying for Claude. 🤨
After July 13, access to Fable 5 will be removed, while weekly usage limits will reportedly be reduced by 33%.
At that point, keeping the subscription starts to feel like paying more for less—especially when competitors offer powerful models at lower prices and with fewer restrictions. ⚠️
Why continue paying for tighter limits and less value? 🤷♂️
A 744-billion-parameter AI model running on a regular computer with only 25 GB of RAM—and no GPU. 🤯
Colibrì is a lightweight, open-source inference engine written in pure C that can run GLM-5.2 on consumer hardware.
Instead of loading the entire model into memory, it keeps RAM usage low by streaming only the required Mixture-of-Experts components directly from an SSD while generating each token.
The impressive part:
⚡ Pure C
📦 Zero runtime dependencies
💾 Around 25 GB of RAM
🧠 744B total parameters, with roughly 40B activated
🚫 No expensive GPU required
There is an important trade-off: this is not fast inference, and the quantized model files still require hundreds of gigabytes of SSD storage. But as a technical proof of concept, it is remarkable.
It shows that running enormous AI models locally may depend as much on clever memory management as it does on expensive hardware.
Tiny engine. Massive model. Very smart engineering. 🐦💻
Repository link below 👇
The model is officially listed as a 744B-parameter MoE with about 40B active
Charge your laptops. GPT-5.6 is entering the room. 🔥
GPT-5.6 Sol, along with Terra and Luna, is expected to launch publicly this Thursday.
Preview access is also being expanded globally in the meantime.
The AI season is getting brighter. 🌞
#GPT56#ChatGPT#OpenAI#AI#Sol #Terra #Luna