At Osvi, a journey follows one customer across chat and calls until a job is done, like collecting a late EMI. Each journey is built from AOPs.
Sometimes it does not. The customer says "call me after 6", the agent says "Sure, we'll call you", and no call is ever booked. There is no error and no log. This post shows how we use Jev to classify AOPs and catch that miss[1], and what we found when we tested it.
What Jev is
Jev came out on 15 September 2026. It does not write text. You give it some text and a question, and it gives back an answer from the options you sent, with a confidence score.[2] It costs $0.042 per million input tokens, output is free, and TypeSafe says it answers in 70 to 500 ms.
Jev as the AOP classifier
Jev does one job: it reads the customer's latest message and decides which AOP it matches. We run this after the customer already has their reply, so the chat never waits for it.
- chat agent
- Talks to the customer and usually uses the right tool.
- AOP check
- Runs after each customer message and sends Jev the message and the AOP list.
- Jev
- Classifies the message: names the matching AOP, or none.
- journey manager
- Wakes up on the AOP and decides what to do next, like booking a call.
The question we send
// example, shortened
{
"state": { "customer_latest_message": "Can't talk now. Call me after 6." },
"questions": { "aop": {
"type": "choice",
"criteria": {
"Callback requested": "Matches when: the customer asks to be called. Does NOT match when: they only ask how callbacks work.",
"Promise to pay": "Matches when: ... Does NOT match when: ...",
"Dispute": "Matches when: ... Does NOT match when: ...",
"none": "A topic mention, a what-if, or a how-does-this-work question."
}
}}
}
// answer
{ "choice": "Callback requested", "confidence": 0.99 }
Sent to OpenRouter's decisions endpoint.[3]
Two things in this question matter most. First, each AOP says what does not count. Second, none is a real option with its own meaning, so questions like "how do returns work?" have a place to go.
Before Jev, an LLM (DeepSeek) did this job. It was told to reply with just the AOP name, but it often wrote sentences like "No match: Refund was only mentioned as a what-if". That sentence has an AOP name in it, so we needed tricky code to read it safely. Jev can only answer with an AOP from our list, so that problem is gone.
Safety rules
- No double starts. Jev sees every AOP. If it names one the agent already started, we do nothing. Each message can start an AOP only once, even on a retry.
- Jev never writes the proof. The customer's exact words come from our saved chat, not from the model.
- Humans come first. If a person takes over the chat while Jev is thinking, we stay quiet.
- If Jev fails, the LLM takes over. An error, no answer in 5 seconds, or an unknown answer sends the same question to DeepSeek. Every fallback is logged, and one setting turns Jev off.
How we tested it
We wrote 36 customer messages for three journeys (loan payments, online shopping and a clinic). Half were real requests. The rest sounded close but should match no AOP, or had nothing to do with the journey. We translated them into 47 languages, including Hinglish and Tanglish, and sent all 1,692 to both models. The right answers were set before either model ran.
| Result | Jev | DeepSeek |
|---|---|---|
| Correct answers | 99.2% | 98.2% |
| False alarms | 5 | 24 |
| Typical time | 0.35 s | 1.14 s |
| Slowest 1% | 0.89 s | 5.03 s |
| Cost per 1,000 checks | $0.025 | $0.025 |
The number we care about most is false alarms: starting an AOP the customer never asked for, like an unwanted call. In India, too many collection calls is a legal risk, so going from 24 to 5 is the real win.
The accuracy gap is smaller than it looks. On real requests, both models did about the same. Jev's lead came almost all from one message:
"If I pay tomorrow, will the late fee be waived?"
DeepSeek took this as a promise to pay in 16 of 47 languages. Jev did in 4.
The "does not match" text does the work
Jev charges only for input, so we tried shorter questions to save money.
| Question | Correct | False alarms | Cost |
|---|---|---|---|
| Full question | 99.2% | 5 | 100% |
| No "does not match" text | 96.0% | 60 | 72% |
| AOP names only | 95.5% | 59 | 60% |
Both short versions did worse than DeepSeek. Without the "does not match" text, Jev thought "How do returns work?" was a return request in almost every language. So we keep the full question.
Weak spots
Jev did worse than DeepSeek in Swahili (89% against 94%), and both found Finnish hard. When Jev was unsure (confidence 0.5 to 0.7), it was right only 8 times out of 13, but there are too few cases to act on yet. These were test messages, not real chats, and each ran only once.
Next: the phone
Jev classifies AOPs on chat today. On calls, every extra step adds silence, so Jev must run beside the reply, never in front of it. Others have found Jev good at spotting voicemail and at after-call checks, but not ready to guide a live call.[4][5] Our plan:
- Lock the model version, so these results stay true.
- Classify AOPs from call transcripts, not only chat.
- Test Jev on voicemail and after-call questions for a week before we trust it.
Takeaway
Asking Jev to pick from a list, instead of asking an LLM to write an answer, made our check faster and cut false alarms by almost 80% at the same cost. But the biggest gain came from the question itself. Telling the model what does not count mattered more than which model we used.
References
- TypeSafe. Introducing System One Models & Jev.
- TypeSafe. Models, pricing and limits; Known weaknesses of jev-1.13.
- OpenRouter. Decisions endpoint,
/api/alpha/decisions(alpha). - Pipecat. Voicemail detection.
- Hacker News. Jev launch discussion.