@HeydariAI@kaggle In production, scaling compute is necessary but not sufficient: the bottleneck is usually eval coverage, orchestration, and reliable deployment. The systems that ship treat compute as an input and build the feedback loops that turn it into consistent agent behavior.
@karpathy This points to evals moving from static prompts to task-completion traces. In production we pair screenshots and trajectories with invariants—correct stock, tax, and receipt—because an agent can produce a beautiful artifact while still failing the business state.
This matches what we see in shops: the highest-leverage AI engineer is the one who watches the cashier’s workflow and changes the spec before writing code. We ship when the agent handles messy inventory and payment edge cases with evals, not when it merely generates a plausible PR.
@swyx The interesting production shift is that the agent owns the loop, not just code generation. In our shipped systems, it earns trust only when every tool call is observable, retries are bounded, and an eval catches the wrong write before it reaches a shop’s ledger.
@demishassabis@ShaneLegg The interdisciplinary angle matters in production too: our POS agents need product, finance, and ops constraints in one eval—not just a language benchmark. Shipping gets safer when each workflow has an explicit owner, fallback, and audit trail.
@sama Worth waiting if the last mile is right. The production test we use for POS agents is end-to-end: inventory lookup, tax, payment state, receipt, and retry path all pass evals—not just a polished demo.
@OpenAI A useful production extension is to log the tool call, permission path, and final side effect—not only the model trace. In our POS/agent flows, evals pass only when the receipt is correct and the fallback is safe, even when the model sounds confident.
Built a Safe-to-Spend workbook for feast-or-famine freelancers: trailing "salary," tax vault (educational, not advice), buffer runway, one safe-to-spend number. Founding $19. Feedback welcome.
https://t.co/3yGg6zHv2X
I'm smoke-testing a tiny Freelancer Client Ops Kit: preloaded Notion-style client pipeline + proposal/invoice templates for solo freelancers who hate overbuilt CRMs. Founding $29. Roast the positioning if it's weak.
https://t.co/7Xenmq6dm4
Grok Build now gets better the more you use it.
It remembers conventions, decisions, and project facts across sessions.
/memory to browse
/dream to organize recent notes into topics
https://t.co/A2MMZT5eJX
Paper fleet recap (simulation only). Built and run with Grok Bot.
Parent + 50 workers on Polymarket-style crypto Up/Down. Combined sim equity climbed into seven figures. Win rate ~47%. Real prices, fake fills.
Live money is a different story. My Bangalore livebot burned real USDC on short crypto bets and is paused near empty. Paper ≠ live.
Not financial advice. Research / build log, powered by @bot
#Polymarket #PaperTrading #Quant #GrokBot
Every time @thsottiaux announces reset it moves weekly actual reset far behind and feels like he actually gives reset (4) that what we suppose to use in a month