Initial Virtual Computer in sandBox implementation: The pipeline is API/Playwright → Socat → Chromium CDP. Currently running into issues with real-time forwarding via Socat.
#ai#aiagent#llm
I'm building a general-purpose AI agent similar to Manus. Currently, I've set up a basic sandbox to execute risky commands securely. The next step is to implement browser automation via the Chrome DevTools Protocol (CDP).
@hex_agent Every Agent system has its own unique logic and strategy. You can't just blindly assume 0.8 is your system's optimal threshold; it must be derived from your own extensive data and evaluation
Vertical AI Agent for E-commerce Customer Support
This demo presents a vertical AI agent designed for e-commerce customer support workflows.
By integrating a local RAG pipeline with domain-specific knowledge bases, the agent can accurately retrieve..
#ai#aiagent
@hex_agent logs are so critical. If you're chasing maximum precision, you should have your team manually tag about 300 logs to build a gold standard dataset. Once you visualize the accuracy curve in your eval system, you'll be able to identify the threshold that works best for your setup
@hex_agent You can't nail the threshold right away. I’m sticking to a 'Tight-to-Loose' approach: prioritize human hand-offs initially to keep trust high, then use post-mortem log analysis to optimize the triggers based on real-world data
@hex_agent don't cache entire messy context, which is noisy and token-heavy, for example I enforce a structured schema (Schema A) for key entities like Order IDs and Return Policies. The AI parses the 3,000-word RAG results and populates this schema A.