Three years. Three products. Each one solved the problem the last one created.
Xinference started with a simple frustration: deploying an open-source model privately shouldn't take weeks of infrastructure work. So we built a platform that collapses it to one command. Your cloud, your GPUs, your data — never leaves your environment. Same OpenAI-compatible API. 9.4k+ GitHub stars. 8M+ downloads. Enterprise teams running everything from RAG pipelines to real-time trading agents.
Once teams had private inference running, the next question was obvious: "How do I build things with these models?"
So we built Xagent. Describe what you want an AI agent to do. It builds it. No code, no workflow YAML, no DAG editors. Enterprise-grade agents deployed in minutes on your own infrastructure, powered by the Xinference engine underneath.
Then came the problem I didn't expect.
Teams weren't running one model. They were running ten or twenty — different sizes, different specialties, different cost profiles. Sending everything to the biggest model is expensive. Sending everything to the smallest is unreliable. Hard-coded rules break the moment traffic shifts.
So we built XRouter. Smart model routing that learns which model handles which request most cost-effectively — in real time. Least cost. No quality degradation. Teams are saving another 30–50% on top of the savings from running open-source models privately.
The lesson: the real value in AI infrastructure is compound control. Not one clever feature — but owning the full path from request to response and optimising every layer.
Sovereignty → Automation → Cost intelligence. One stack.
Still a small team but growing fast. Visit us at https://t.co/5nWl9yP8oj.
Once an agent produces more drafts, approval becomes the bottleneck instead of production.
Batch reviews into two fixed windows, set a default for anything unreviewed after 24 hours, name the backup approver before someone takes leave.
https://t.co/OPZaMprslq
On a shared cluster the question is not whether GPU memory fills. It is what gets evicted when it does.
Without a policy, the model that loses its place is the one a quiet internal tool depends on. Decide what stays pinned in advance.
https://t.co/r9jHSEByBf
An agent has no sense of what is urgent unless you give it one.
It treats a routine newsletter and a launch-day landing page as equivalent work. Priority has to arrive as data: due date, tier, campaign. People read context, software reads fields.
https://t.co/OPZaMpqUvS
Friday: two providers retire models from their serverless catalogues. Fireworks drops DeepSeek V4 and Kimi endpoints, Baseten six Model APIs. Their deprecation calendar is your migration calendar, unless you run the weights yourself.
Production readiness means tested latency objectives, real concurrency, recovery, access controls, model provenance and rollback.
Record the target, result, environment and owner for every launch check.
https://t.co/L7D1tiPm11
This Friday, multiple managed providers are retiring models from their serverless catalogues.
When you rely on third-party endpoints, their deprecation calendar becomes your migration calendar.
Run the weights on your own infrastructure, deprecate on your own terms.
Keep audience, position, offer and creative direction human. Standardise the repeated loop around research, drafts, source checks, review and reporting.
Version the brief so reusable agents do not repeat stale strategy.
https://t.co/OPZaMprslq
A safer model rollout has three gates: task quality, representative capacity tests and a proven rollback path.
Canary a small cohort and define expansion and stop rules before traffic moves.
https://t.co/L7D1tiPm11
Score the first marketing agent on time returned, human agreement, exceptions and cost per accepted outcome.
Task count can hide rewrites and manual repair. Pilot one narrow workflow on edge cases.
https://t.co/5Fp3redBXs
For some regulated workloads, the correct public endpoint count is zero.
Test that the service works without external network access, then define the approved path for updates, monitoring and model promotion.
https://t.co/mrGVUrYiav
Before an inbound agent gets CRM access, define qualification evidence, allowed actions and escalation rules.
Measure response time, human agreement and exception rate on a lead sample before increasing autonomy.
https://t.co/saHg2K5szw
TTFT measures the wait before output starts. TPOT measures generation speed after that. Track both by percentile, model and workload.
One latency average cannot tell a platform team which control to change.
https://t.co/r9jHSEC6qN
Use fixed automation for stable handoffs. Use an agent where campaign work depends on account context, source material or reviewer feedback.
Test missing data, rejected output and revoked permissions before expanding.
https://t.co/tXYqpXi01Y
AIA moved from a self-managed vLLM and Kubernetes stack to a governed platform. New model rollout fell from 2-3 weeks to 1-2 days while serving 500K+ daily requests.
The safe path became the faster path.
https://t.co/2dPpWvhg3H
Turn weekly reporting into a governed pipeline: named sources, fixed metric rules, visible exceptions, a draft summary and human interpretation.
Measure preparation time and corrections, not dashboard size.
https://t.co/OPZaMpqUvS
An agent is useful when repeated work still contains judgment. If every step is known, use a workflow. If a person must direct each step, use an assistant.
Scope the first agent around bounded decisions and a clear handoff.
https://t.co/kZfAaJBMpl
Managed inference is easy to start and costly to unwind once data residency, model versions and pricing become product dependencies.
Compare both paths against the same quality, latency and control requirements.
https://t.co/mrGVUrYQ03
The best model for your workload?
Benchmarks measure a model on someone else's stack. Your system has different traffic, a different latency budget, and a compliance review at the end.
The process we actually use:
Split the work by task first. Extraction, classification, summarisation and code generation have different accuracy bars. One model rarely wins all four.
Build the eval set from your own traffic, including the long contexts and the adversarial instructions.
Measure serving behaviour. Time to first token, inter-token latency, throughput at your real concurrency, memory headroom.
Then price it as cost per successful task. Retries and idle capacity are where per-token maths falls apart.
The right production model is the smallest one clearing your quality, latency, governance and reliability thresholds. Everything above that line is money spent on a leaderboard position your users never see.
Full framework: https://t.co/Z7MtUJqWLW
Final day of The AI Summit Australia. Let's make it count! ⏰
The exhibition hall is open from 10 to 2 only. The conference closes at 3:35.
Find Xinference at stand H06 until the hall closes. Last chance to catch us!
@TheAISummit#TheAISummit
Day one of The AI Summit Australia is here! 👋
The exhibition hall is open from 10 to 5, with networking drinks on the floor from 5:15.
Find Xinference at stand H06. We're ready when you are!
@TheAISummit#TheAISummit