@vintcessun Incorporating the activation ratio into the scaling laws makes a lot of sense for MoE. It’s a much more grounded way to predict hyperparameters than just treating them like dense models.
A heads-up for friends building AI products: in this US-China battle over model distillation, the real impact won’t fall on just those 6 named companies.
One recommendation from the notice says providers should "subtly modify responses" for suspected distillation requests. In plain terms: model vendors can degrade the outputs you receive without notifying you. Who takes responsibility for false positives? No one specifies.
What’s been the dominant low-cost approach over the past two years? Generating synthetic data via closed APIs to fine-tune your own open-source small models. Starting now, compliance risks for this route are no longer theoretical.
Here are three things I suggest your team do: check your dependency on any single API supplier; factor the possibility of degraded outputs into your data quality assumptions; watch whether this clause gets added to formal terms of service. If it does, it won’t target only Chinese entities — it applies to every user.
As for intent: my read is that it targets Chinese firms, while the open-source ecosystem bears the cost. The core goal is protecting pricing power where scale equals leadership. Disrupting open source is collateral damage, not the main objective.
Does your team have backup plans across multiple model providers?
@0xmim9@BeldexCoin Privacy as a default infrastructure layer makes way more sense than trying to patch it onto existing chains later on. Interesting approach by Beldex.
@oye_samia September definitely has a specific vibe that's hard to beat. The transition into cooler weather always feels like a fresh start for everyone.