Laya vs Jev: Local Decision Model or Hosted API?
Laya and Jev both return a choice or a probability instead of a paragraph. Laya is a model you run yourself. Jev is TypeSafe AI's paid API. There is no shared speed test.
Laya and Jev answer the same kind of question: you hand over a message and a fixed set of choices, and you get a pick or a probability back, not a paragraph. One is a model you download and run. The other is an API you call. This is a comparison of how you use them, not a shared speed test. Laya's site publishes figures against Jev from third-party writeups, and says Jev was not run in that project. Those figures are not repeated here.
Quick answer
| Laya | Jev | |
|---|---|---|
| What it is | An open-source decision model | TypeSafe AI's hosted decision API |
| Question shapes | Choice, score, and a yes-or-no probability | Noul (yes/no probability), Choice, and Score |
| How you run it | pip install laya. Weights download from Hugging Face when you ask for a checkpoint. | Early access at typesafe.ai. After you are in, you create an API key. Official SDKs are Python 3.10+ and Node 20+. |
| License and price | Apache 2.0. No charge to download and run the weights. | Input tokens are $0.042 per million. Output tokens are free. |
| Language | An English checkpoint, plus a multilingual one documented for 100+ languages. A router picks which checkpoint to use. | The product page does not publish a language list. |
| What they say about speed | Laya's README says about 33 ms for one question on a T4, and about 7.2 ms per question when batched. That is their measurement. | The project listing gives a range of 70–500 ms. TypeSafe's own benchmark claims a much larger gap versus a chat model. That claim is theirs, on their workflow. |
Laya: the weights stay on your machine
Laya reads a message and answers the questions you set. It does not draft the reply. The project's quickstart uses one billing email about a duplicate March charge. The documented department answer is billing. Urgency is a score on a scale you define. "Does the user threaten to cancel" comes back as a probability, not a sentence.
Three checkpoints ship under convaiinnovations/laya. The English one is 421M parameters with a 512-token context. The multilingual one is 322M parameters. A third checkpoint is fine-tuned on invoices, security incidents, customer service, and agent traces. On Laya's own typed-decisions benchmark, the base English checkpoint scores 0.362 and the fine-tuned one scores 0.766. Those numbers compare Laya to itself, not to Jev.
Jev: a hosted call, with a price per million tokens
Jev is the API from TypeSafe AI. You send a situation and the questions. A support note such as "Help! My payouts have been failing for 3 days" can be asked, at the same time, whether it is urgent and which team should take it. The documented shape of the answer is a probability and a choice, not a written explanation.
It is in early access: join the waitlist at typesafe.ai, then create a key in the console. Input is $0.042 per million tokens and output tokens are free. LangChain has a classifier integration. The team says a Jev call is meant to sit next to a chat model: Jev for the repeated yes-or-no and pick-one steps, the chat model for open-ended writing.
Which one to use
- You want the model on your own machine, under Apache 2.0, including non-English text. Use Laya. Plan to fine-tune if the base English checkpoint is not enough. The 0.766 figure is after fine-tuning on their benchmark, not the out-of-the-box score.
- You want a hosted API and you are willing to wait for early access. Use Jev. Budget the $0.042 per million input tokens, and treat the 70–500 ms range as their stated range, not a number from a test against Laya.
- You want a winner on speed or accuracy. Neither side has published a run of both models on the same questions. Skip the head-to-head charts until that exists.