Cut your inference bill up to 67%.
Change one line — your base URL. Same OpenAI-compatible request, same model, routed to the cheapest verified seller and settled in USDG. Your code doesn't change. Your bill does.
https://t.co/ZinByUl9Xj
Cut your inference bill up to 67%.
Change one line — your base URL. Same OpenAI-compatible request, same model, routed to the cheapest verified seller and settled in USDG. Your code doesn't change. Your bill does.
https://t.co/ZinByUl9Xj
AI inference is becoming a real marketplace — and LLM Mart is live.
Instead of locking yourself into one provider, you get one OpenAI-compatible endpoint with access to 100+ models, automatically routed to the cheapest live seller for every request.
67% median discount vs direct pricing — calculated from the live order book.
But LLM Mart isn’t just for buyers.
Sellers can monetize their unused API capacity too.
Connect your own custom API endpoint + API key, and LLM Mart can probe the endpoint, detect the models it supports, and turn it into a marketplace listing.
Set your own price. Keep your margin.
Once a request comes in, you get paid per inference directly to your wallet.
> 0% platform fee on inference.
> Per-request settlement powered by x402.
> Payments settle on-chain in USDG.
> Buyers get a single API key that works with existing OpenAI-compatible tooling.
No complicated integrations.
No rebuilding your infrastructure.
Just connect your endpoint, set your price, and start selling.
Buy cheaper. Sell your surplus. Monetize your API.
Everything is live.
Everything is on-chain.
Every fill comes with a verifiable transaction.
https://t.co/Q0KAFhjban
Two new updates on LLM Mart.
1. Live API (Try it live)
A real playground for the marketplace. Connect a wallet, send a prompt — the x402 rail settles real USDG on Robinhood Chain and the answer streams from the live gateway.
We put GPT-5.6 Luna on our homepage. Free. No wallet. No signup. Just type. 3 messages on us — then mint a key and pay cents per million tokens instead of dollars.
2. Docs Updated
LLM Mart speaks x402: HTTP 402 means "pay this exact amount to continue". Request, challenge, signed USDG transfer, resend with receipt — milliseconds end to end.
Cut your inference bill up to 67%. Change one line — your base URL. Same OpenAI-compatible request, same model, routed to the cheapest verified seller and settled in USDG. Your code doesn't change. Your bill does.
Everything in one place — wallet, credits, purchases, sales, x402 payments.
Per-call payments without an account: unpaid requests get HTTP 402 with a payment challenge, your wallet signs an EIP-3009 transferWithAuthorization, and the retry carries the X-PAYMENT header. Settlement is USDG, wallet-to-wallet, no deposit step.
OpenAI-compatible inference, paid in USDG on Robinhood Chain. One signature, then it's just API calls.
1. Endpoint : Drop-in replacement for the OpenAI Chat Completions endpoint. Send your inf_ key as a bearer token. Streaming and non-streaming both supported.
(POST /api/inference/v1/chat/completions)
2. Routing : For each request the router tries, in order:
- Your priority key (if set in Router Settings)
- The cheapest healthy marketplace offer where you
- have credits with the seller
Your fallback key (if set)
3. Billing : Buyer is charged the seller's quoted price per million tokens. Credits are pre-purchased from /markets in USDG. Platform takes a 1% fee at top-up time; the rest goes straight to the seller's wallet.
When credits with a seller run out, the API returns 402 insufficient_credits.
4 Settlement : Top-ups are a single Robinhood Chain transaction with two ERC-20 transfers — seller wallet and treasury — in one buyer-signed tx. USDG never sits in the platform's custody for the seller portion.
5. Compatibility > Works out of the box with:
- curl / fetch
- OpenAI Python / TS SDK (base_url)
- Anthropic SDK (ANTHROPIC_BASE_URL)
- Claude Code (same env override)
https://t.co/N6jQaF6sKu
We're just added some models on the site, now 110+ models available on marketplace!
For each request the router tries, in order:
1. Your priority key (if set in Router Settings)
2. The cheapest healthy marketplace offer where you have credits with the seller
3. Your fallback key (if set)
https://t.co/qhdnuwWDHG
We're preparing the new models with upgraded the site to be friendly for mobile users.
thanks for believe on us! also we'll post another demo video soon, sou you can see it really works as well.
We didn't want developers to learn another completely different inference API.
PONS LLM Mart uses an OpenAI-compatible endpoint.
POST /api/inference/v1/chat/completions
Set the bearer token to your inf_ key and point your client at: https://t.co/nWQFQHZn6p
You can use:
- OpenAI Python SDK
- OpenAI TypeScript SDK
- Anthropic SDK
- Claude Code
- curl
- fetch
basically anything that lets you customize base_url
So the workflow becomes:
existing AI application
> PONS LLM Mart
> marketplace router
> healthy seller
> provider
> inference response
No application rewrite required.
AI agents can hire us directly — pay in USDG, get inference or a ready-to-use API account back.
Fully autonomous, running 24/7.
https://t.co/sazY8KPvD5