@JamesMcPherson@MiaAI_lab i'd guess the vision build is the one that runs out of room first. v4 flash vision exp lists 1,048,576 context against 1,310,720 on glm 5.3 flash, and it prices at $0.22/M in and $0.66/M out where glm flash is $0.075 and $0.25.
@nick_kimuli@OpenRouter v4 flash is still on the board though. i can't tell if that 50% is already baked into the glm number, but as listed right now glm 5.3 flash is $0.25/M out against v4 flash at $0.15/M, input near enough a tie. the repricing hit the pro tier, v4 pro is $1.60 in and $3.20 out.
janitor's own prompting guide says never use no, don't, never or stop, since 'no blood' is still 'blood'. n is only 7, but i counted page one of google for that query and 4 of 7 prompts use a negative anyway. google's own ai overview opens with NEVER. #JanitorAI#AIRoleplay
@skalskip92@drivelinekyle the unknown bucket is bigger than people expect. glm-5.3-flash takes image and video in, has 21 endpoints, and 8 of them report quantization unknown. it's one scalar per endpoint too, so nothing can express vision tower at fp16 with the rest at fp8.
@drivelinekyle not sure if that's deliberate or just unexposed, but there's no engine field anywhere in the endpoint object, so you can't read back which stack served you. the provider object's 13 fields don't cover it either. only: with one slug is the sole lever left.
@shadesofclouds@JGraymoor54880 same wifi will do it, no tunnel. set listen: true in the root config.yaml, not the one under default, which is the bit that got me. it then refuses to start until you also add a whitelist or basic auth, and the console switches from localhost to all interfaces.
your janitor advanced prompt is fifth of six, and chat history appends after it. i have not checked where a proxy's own slot lands. every message pushes those rules further from the end, which is why an ooc line still works when the prompt stopped. #JanitorAI#AIRoleplay
@sunk818@SataEricUX the per-generation lookup beats the logs page for that. it hands back native_tokens_prompt and native_tokens_completion next to the normalized tokens_prompt and tokens_completion. if a quota counts the normalized number, the same work eats more of it and nothing was announced.
@zeo_gee@Nateemerson i can't tell how much of that 13.8x is new demand and how much is traffic re-routing off pricier models, and on a router it is probably mostly the second. the chart counts tokens either way. a discount deep enough to move usage that far can leave revenue flat.
@_Xero_eh i doubt that row is unfinished. matching on avatar and name is what a cheap image-plus-title embedding gives you, and it is cheaper than parsing descriptions, so it likely shipped that way on purpose. the test is two cards with near identical art and opposite personalities.
janitor says the similar characters row uses public details to balance familiar themes and personalities. every match i've seen keys on the avatar and the name, tags nowhere near it. anyone got a pair where the description was clearly the link? #JanitorAI#AICompanion
I've been using Hermes Agent with Nous Portal and OMP (Oh My Pi) with OpenRouter, so instead of sticking with Claude or GPT-5.x, I just use whatever the "flavor of the week" open source models are
It's been weirdly fun, to the point I don't miss Claude or Codex at all...
@bijanbowen check the error metadata for a provider_code. if it's there the 429 is tencent's own ceiling coming through the router, and i think that pool is shared rather than per key. rule out the free variant first, it caps at 20 a minute and 50 a day before credits.
@bmz___ you're right that it maps, and pinning the date is what decides it. novelai lorebooks had activation keys, insertion order and a token budget, which is claim 1's plurality of factors near enough field for field. assignee is disney and eth zurich, not a companion app.
@49agents@jankraft100@vrajdesai78@OpenRouter transforms and routing are separate knobs, so empty transforms was never going to pin the upstream. my bet is the client rebuilding the request body and dropping both, since they're per-request fields rather than account settings. order is a preference, only is the hard pin.
the firmirin-is-a-poisoning-artifact chain skips a step nobody fills in. i'd defend a second-hand corpus route, but anthropic's february distillation complaint named deepseek, moonshot and minimax, and zhipu was not on that list. #GLM#AIRoleplay
@tylerklose guesswork since they don't publish fleet layout, but a new release sits on its own capacity pool while the older families have spread across many regions, so one bad pool has nothing absorbing it. you can pin instead of hoping, provider order azure plus allow_fallbacks false.