Building a BYOK AI chat app.
You bring your OpenAI, xAI, or OpenRouter key. I handle the rest. One clean interface for any model you want to use.
Documenting the whole thing here as I build it.
https://t.co/qnJfotCHBL
@rozhkov_@cheatyyyy Exactly! Also, there's always a free alternative to paid tools. Maybe, as a gesture of good faith, add a banner on the app's website saying "This is a fork of [original app]."
@konstipaulus@giffboake even if there are , so what ? if you can market your product better or make it user friendly why won't you get a piece of the pie
all this arguing over who serves glm-5.3-flash cheapest. meanwhile Modal hands out $30/mo in free credits and hosts it too. even at ~3x the price that's still ~180m tokens a month for $0.
https://t.co/om7DWYtmr5
GLM-5.3-Flash scores 57 on the Artificial Analysis Intelligence Index. At $0.09 Cost per Task, it sits comfortably on the Intelligence vs. Cost per Task Pareto frontier
@Zai_org has released GLM-5.3-Flash, a smaller and cheaper sibling to GLM-5.3 at 320B total parameters and just 18B active parameters. GLM-5.3-Flash supports low/high/max reasoning efforts, and scores 57 evaluation on the Artificial Analysis Intelligence Index with max reasoning effort. This places the model only 3 points behind GLM-5.3 at 60 and in line with GPT-5.6 Terra and Muse Spark 1.2.
On Z AI's first-party API, GLM-5.3-Flash is priced at $0.15 / 1M input tokens and $0.50 / 1M output tokens, just over 10% of the price of GLM-5.3. Cached input tokens are priced at $0.026 / 1M tokens, an 80% discount. Its Cost per Task on the Intelligence Index is $0.09, compared to $0.68 for GLM-5.3 (max), and it sits on the Pareto frontier for Intelligence vs. Cost per Task.
Key results:
➤ GLM-5.3-Flash is 3 points behind GLM-5.3 (max) on the Artificial Analysis Intelligence Index, at ~7.5x lower Cost per Task. At $0.09 per Intelligence Index task against $0.68 for GLM-5.3, it sits on the Pareto frontier for Intelligence vs. Cost per Task. It ties GPT-5.6 Terra ($0.51) and Muse Spark 1.2 ($0.40) at 57 while costing ~5.7x and ~4.4x less per task.
➤ GLM-5.3-Flash is less token efficient, but its low per-token pricing means this does not translate into a high Cost per Task. The model used 149M output tokens to run the Intelligence Index, ~11% fewer than GLM-5.3 at 168M, but more than Kimi K3 (133M) and Qwen3.8 2.4T A95B (136M) which score the same on the Intelligence Index. Reasoning tokens account for 134M of the 149M total (~90%).
➤ GLM-5.3-Flash matches GLM-5.3 on real-world agentic work on GDPval-AA v2. With an Elo of 1770, the model is tied within the margin of error for GLM-5.3 and Grok 4.6. This places it behind only Claude Opus 5 (xhigh and max). On Terminal-Bench v2.1 it also matches GLM-5.3 (84.3% vs 83.9%), and on τ³-Banking it trails by 3.1 p.p. at 47.2%.
➤ GLM-5.3-Flash demonstrates good real-world knowledge and hallucination rate, scoring +7 on AA-Omniscience. Its AA-Omniscience Accuracy is 28%, 6 p.p. below GLM-5.3 (max) at 34% and well below GPT-5.6 Terra at 47%. However, with a Hallucination Rate of 28%, it is an improvement over GLM-5.3 at 30%. In real-world knowledge, GLM-5.3-Flash knows less than the bigger models and frontier proprietary models in its Intelligence Index tier with an accuracy of 28%.
Additional model details:
➤ Pricing: On Z AI's first-party API, $0.15 / 1M input tokens and $0.50 / 1M output tokens . Cached input tokens are priced at $0.03/ 1M tokens, an 80% discount.
➤ Accessibility: Accessible through Z AI's first-party API at launch.
➤ Size: 320B total parameters with 18B active parameters
➤ License: MIT
➤ Context Window: 400k
GLM-5.3-Flash just went public, after eating 20T+ tokens as the anonymous "Ox Alpha" on OpenRouter.
First test: the Earth-collides-with-Mars Three.js challenge that's circulating on X.
@EstebanCervi@vasielv also , I am not saying that his product will find customers for sure , I am just arguing your point of "if i can do it this, why would i use your way"
@EstebanCervi@vasielv That’s not my point. My point is that if he can save people time and deliver good results, he will find customers. I’m a developer, so I enjoy building things myself, but most people don't feel that way.
@vasielv@EstebanCervi I’m working on this myself, and it takes time to get good results. If your product delivers great results right out of the box and saves people time, there's definitely a market for it. Don't listen to people who ask why anyone would use your product when other tools exist.