Jev is a judgment primitive, not a chat model.
Feed it one thing. Ask typed questions. Read the probabilities.
Batch it like an LLM, and you’ll blame the model for your prompt.
Code + report on request.
Jev is not an LLM. Stop using it like one.
I tested the same 40 Korean sentences two ways:
• docs-style: 40/40, 1.9s, $0.0012
• LLM-style: 62%
Same model. Same questions. Here’s the cliff 🧵
Caveats:
• only 40 sentences
• bad examples were obvious
• Claude latency was inflated; costs estimated
• bad sentences mostly sat at even positions
Treat this as a sharp signal, not a final verdict.