I would start Jev with one bounded job: choosing the detail level of candidate test logs before diagnosis.
Known-required evidence stays in. Original logs remain retrievable. Keep a fallback that exposes the full relevant output, and count how often it is needed.
If Jev beats simple rules under the same completion and correctness checks, expand to more evidence types. If it doesn't, keep the rules.
That's the architectural opportunity: turn a hidden part of agent behavior into a replaceable, measurable decision layer.
The compelling demo would show the choice, the evidence loaded, and the task outcome—not just a smaller prompt.
Jev's most interesting job in a coding agent may be deciding what the coding model sees next.
A code fix depends on earlier choices: which files to read, which logs to expand, which evidence to carry forward.
My bet: make those choices explicit enough to test a specialized decision model against simple rules and the coding model itself.
Here's a proposed Jev-powered context-selection loop, with an input/output example and the cost trap that could make it lose.
Test the value of Jev separately from the value of the harness.
Use the same evidence store, candidate cards, options, and retrieval tools. Compare three selectors:
A. Rules based on paths, versions, and test dependencies
B. The coding model making the selection
C. Jev making the selection
Keep the coding model, tasks, starting snapshots, and execution limits fixed. Repeat runs.
Track task success, missed evidence, retrievals, end-to-end latency, and total spend across all attempts per successful task.
A shadow run can expose questionable choices. It cannot establish the cost of following them.
🚀 Your AI skill can follow every instruction and still give the wrong answer.
The definition it retrieved may be outdated. The right exception may live in a file it never opened. A recent event may be missing from its context.
Before adding another instruction, inspect what the agent actually read.
Here's a practical architecture for knowledge-heavy skills: route the question, load a coherent core, add fresh context when needed, and test every update. 🧵
🛠️ Before your next prompt patch, inspect one failed run.
Write down:
1. What the agent needed to know
2. What it actually loaded
3. Which facts were missing, irrelevant, or out of date
4. Which decision sent it down the wrong path
That gives you a specific change to test: improve a route, restore an exception, version a definition, or revise a method's trigger.
Try the smallest fix against the current baseline. A small, coherent knowledge set may work well with simple loading; extra routing can add its own failures.
What breaks most often in your skills: retrieval, freshness, or method selection?
#AIAgents #ContextEngineering
🧪 Measure the failures this architecture could introduce.
Include cases with an ambiguous domain, a missing exception, a historical date, and a question that requires more than one knowledge section.
Check the final answer AND the retrieval trace:
• Did it select the right scope and version?
• Did it retrieve the required evidence?
• Did it recognize missing information?
• What did the answer cost in tokens and latency?
Keep held-out cases so fixing the regression suite doesn't become the whole objective.
User corrections can suggest new tests. Repeated feedback still needs verification before becoming a rule; repetition alone doesn't establish truth.