What it does
Jev reads the context you provide and answers a question within boundaries you set. That might mean choosing a support queue, rating how urgent a request is, or estimating whether a message contains a purchase enquiry. Your application gets a value it can use for its next step.
There are three question types. Choice picks from named options. Score evaluates against ordered levels. Both include probabilities and confidence. Noul estimates the probability of a statement being true; it has no separate confidence field. Jev does not write the customer’s reply, and it does not decide which actions your application is allowed to take.
Who it suits
Start here if your bottleneck is repeatedly reading short pieces of text and making the same kind of judgment. A small sales team sorting enquiries or a support desk assigning tickets can define a useful first task. If the job is simply checking whether an order is paid, use the order record and a normal rule. A model is useful when the wording needs interpretation.
A situation you might recognize
Consider an inbox with genuine enquiries, backlink sales pitches and existing customers asking for help. Ask for the message’s intent, then combine that answer with your customer records. A friendly email is not evidence that someone has paid you.
Keep “other” as an available category. A customer offering a partnership should not be forced into a sales-pitch category just because your initial list was too narrow. This is a suggested workflow, not a report of a ToolAI production test.
A sensible first setup
- Choose one decision with a clear owner, such as which queue receives a new email.
- Write descriptions for the categories, including ambiguous cases and an “other” option.
- Collect representative messages and have a person label them before comparing Jev’s answers.
- Start with suggested tags; review mistakes before allowing automatic routing.
- Record the model version and question wording so later changes can be evaluated fairly.
What to weigh before choosing
Typed output makes integration easier; it does not prove that a judgment is correct. A high confidence value can accompany a wrong answer. Choose review thresholds using examples from your business, and judge missed customers separately from harmless extra reviews.
The official documentation identifies English as the primary training language. Test Chinese and other languages separately. Keep arithmetic and date comparisons in code, and send only the context needed for the question. Use a generative model when the next step needs a written explanation or reply.
Accounts and running costs
The native service requires a TypeSafe account and API key. Its documented charging basis is input tokens; free output does not make the entire request free. Estimate usage from realistic message lengths, questions and retry volume. The SDKs connect to the hosted service; installing one does not provide local model weights. Check the official account terms and current prices before committing.