The setup and API details below come from TypeSafe’s documentation, launch post, and community write-ups as of September 28, 2026. Jev launched in early access two weeks ago, so check the official SDK reference before shipping.
Jev is the first public model in a class TypeSafe calls System One models. TypeSafe AI’s Jev is a model designed to return typed decisions with probabilities instead of text. This post covers what it is, what it isn’t, when it earns a place in your architecture, and how to use it, with code.
TypeSafe’s founder, Diogo Almeida, who previously worked on instruction-following research at OpenAI, describes Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out.
You send two things:
- State: the context, as a string, a JSON object, or an array (example – a support ticket, an order record).
- Questions: typed questions about that state.
You get back one typed answer per question, with probabilities attached. There is no generated text, so there is nothing to parse.
The three primitives
| Primitive | Asks | Returns |
|---|---|---|
| Choice | Which of these options fits? (up to 255 options) | The selected option, a probability per option, and a confidence |
| Score | Where does this fall on this ordered rubric? (2 to 10 levels) | A probability-weighted score, the distribution, and a confidence |
| Noul | Is this statement true? | A probability from 0 to 1 (no separate confidence field) |
Every question in a request is evaluated in parallel and in isolation against the same state. Adding questions barely changes latency, and questions can’t contaminate each other the way they can in a long chat context.
How it differs under the hood (per TypeSafe)
- Training: Chat models are tuned with RLHF, which optimizes for human preference. Jev uses what TypeSafe calls Reinforcement Learning for Calibrated Decisions (RLCD), which optimizes for honest probabilities.
- Sampling: LLMs generate one token at a time. Jev produces all its outputs in a single parallel pass.
- Output space: The possible outputs are declared in advance, so the answer always matches the schema.
What Jev is not
It is not a chatbot or a text generator. No summaries, no explanations, no code, no reasoning traces. If you need prose, use an LLM.
It is not “an LLM with JSON mode.” Structured-output features constrain what an autoregressive model emits, but the model underneath is still generating tokens and can still be confidently wrong. Jev’s answer format is the model’s native output .
It is not infallible. This is the most important one. “Type-safe” means your code never receives a malformed answer. It does not mean the answer is correct. Jev can be wrong, and TypeSafe publishes a “jaggedness” page covering weak spots such as arithmetic, dates, distractors, and adversarial input. The value is that it reports how sure it is, so you can decide what to do about it.
It is not a replacement for deterministic code. The docs’ own advice, echoed across community guides: don’t ask the model something your code can compute exactly, and don’t hide several judgments inside one question.
It doesn’t see images, audio, or video. It reads text and JSON only at the moment [9].
It is not an agent. It makes one decision per question. Your code owns the loop, the tools, and the side effects.
Why use it?
- It fits the shape of the problem. Most “AI in the workflow” tasks are classify, route, score, pick-from-a-list, or branch. Those are decisions, not essays or long prompts.
- Calibrated confidence enables real automation. If a model can do a task 95% of the time but can’t tell you which 5% it will miss, you can’t automate the task. Confidence you can threshold on lets you act autonomously above a line and escalate to a human below it.
- Speed and cost (vendor claims). TypeSafe lists $42 per billion input tokens ($0.042 per million) with output tokens free, and 70 to 500ms end-to-end latency. Their headline is 193.6x faster and 444.6x cheaper on their own workflow evals, which they say are toward the high end of real-world gains. If the numbers are even 10x better than your current LLM calls, though, that changes what you can put in a hot path: per-row processing, per-request routing, or real-time UX.
- Composability. Because each answer is a typed value with a probability, you combine decisions with ordinary code: weights, thresholds,
ifstatements. When priorities change, you change a coefficient instead of rewriting a prompt.
Getting started
Get an API key from the TypeSafe console (early access is waitlisted at typesafe.ai). Jev is also reachable without the waitlist through gateways such as OpenRouter, Vercel AI Gateway, and Cloudflare Workers AI. Then:
|
1 2 |
pip install typesafe-sdk export TYPESAFE_API_KEY=... |
The client reads TYPESAFE_API_KEY from the environment and defaults to jev-latest.
Your first call
|
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 |
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient ticket = ( "I was charged twice for my subscription and need the duplicate " "refunded today. This is the second time this has happened." ) with TypeSafeClient() as client: response = client.system_one( state={"ticket": ticket}, questions={ "department": Choice( instructions="Which team should handle this ticket?", criteria={ "billing": "Payment, refund, or subscription issues", "technical": "Bugs or integration problems", "sales": "Pricing or plan questions", }, ), "frustration": Score( instructions="How frustrated does the customer appear?", criteria=[ "Calm, just stating facts", "Frustrated but civil", "Very angry, strong language", ], ), "is_urgent": Noul( instructions="The message conveys time pressure.", ), }, ) print(response.model) # versioned ID that answered, e.g. "jev-1.13.0" print(response.answers["department"].choice) # e.g. "billing" print(response.answers["frustration"].score) # e.g. 1.0 print(response.answers["is_urgent"].noul) # e.g. 0.97 |
Three judgments, one request, no parsing. The SDK also exposes typed accessors like response.choices[...] and response.nouls[...], each carrying probabilities and, for Choice and Score, confidence. Field names come from the current SDK, so confirm against the reference for your version.
If you are not on Python or JavaScript, the underlying REST call is a single POST:
|
1 2 3 4 5 6 7 8 9 10 11 12 |
curl -X POST https://api.typesafe.ai/v1/systemone \ -H "Authorization: Bearer $TYPESAFE_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "state": "I was charged twice and need the duplicate refunded today.", "questions": { "is_urgent": { "type": "noul", "instructions": "The message conveys time pressure." } } }' |
Community clients already exist for Go, Rust, Ruby, .NET, PHP, Elixir, Haskell, and Scala, plus framework adapters for LangChain, Pydantic AI, LiteLLM, Effect, and others.
Writing good questions
The docs and community experience converge on a few rules [3][10]:
- One judgment per question. Don’t hide “is this urgent and about billing?” in one Noul. Ask two.
- Phrase Noul so that true means “yes.” A Noul where the affirmative maps to “no” underperforms.
- Treat criteria as part of the instruction. Contradictory instructions and criteria confuse the model. Be specific about what each option means, and include an “other” option so the model isn’t forced into a bad fit.
- Don’t ask for things code can compute. Dates, arithmetic, and exact matching belong in code.
- Send structured state. JSON with named fields is the intended input style, not just a wall of text.
Architecture patterns
Pattern 1: Confidence-gated automation
This is the core idea. Decide what “sure enough” means, act if it is met, and escalate if not.
|
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 |
from dataclasses import dataclass from typing import Optional @dataclass class Routing: team: Optional[str] needs_human: bool reason: str CONFIDENCE_FLOOR = 0.80 # tune per workflow, against your own labeled data def route_ticket(client, ticket_text: str) -> Routing: r = client.system_one( state={"ticket": ticket_text}, questions={ "team": Choice( instructions="Which team should own this ticket?", criteria={ "billing": "Payments, refunds, invoices", "technical": "Bugs, outages, integrations", "account": "Login, permissions, profile changes", "other": "None of the above clearly fits", }, ), }, ) team = r.choices["team"] if team.choice == "other": return Routing(None, True, "no clear category") if team.confidence < CONFIDENCE_FLOOR: return Routing(None, True, f"low confidence ({team.confidence:.2f})") return Routing(team.choice, False, "auto-routed") |
Two things to notice. Confidence is a separate signal from the winning option’s probability: the probabilities describe the distribution over options, while confidence describes how much the model trusts that distribution, so a top pick at 0.6 can come with either high or low confidence. And the threshold lives in your code, where you can version it, test it, and change it per workflow.
Tuning advice: log every response, including the versioned model ID it reports, and set thresholds by measuring precision on a labeled sample. Pin a specific model (for example jev-1.13.0) once you’ve tuned against it, since jev-latest moves when TypeSafe ships a new release.
|
1 |
client = TypeSafeClient(model="jev-1.13.0") |
Pattern 2: Decompose, then combine in code
Suppose you want to prioritize inbound sales leads. The naive approach asks, “How good is this lead?” The System One approach asks small questions and combines them yourself.
|
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 |
LEAD_WEIGHTS = {"fit": 0.4, "intent": 0.35, "budget": 0.25} def score_lead(client, lead: dict) -> float: r = client.system_one( state=lead, questions={ "fit": Score( instructions="How well does this company match our ideal customer " "profile (B2B SaaS, 50-500 employees)?", criteria=["Poor fit", "Partial fit", "Strong fit"], ), "intent": Score( instructions="How strong is the buying intent in the message?", criteria=["Browsing", "Evaluating", "Ready to buy"], ), "budget": Noul( instructions="The lead has signaled they have budget allocated.", ), }, ) fit = r.answers["fit"].score / 2 # Score is zero-based; normalize yourself intent = r.answers["intent"].score / 2 budget = r.answers["budget"].noul return ( LEAD_WEIGHTS["fit"] * fit + LEAD_WEIGHTS["intent"] * intent + LEAD_WEIGHTS["budget"] * budget ) |
Note that Score is on the rubric’s own zero-based scale and is not automatically normalized to 0 to 1. Also, when the sales team says budget matters more this quarter, you edit a dictionary instead of re-prompting.
Pattern 3: A cheap router in front of an expensive model
Use a fast decision to decide whether you need a slow, expensive one. LiteLLM and LangChain already ship Jev-backed routing middleware built on this idea.
|
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 |
MODELS = { "simple": "small-fast-llm", "standard": "mid-tier-llm", "reasoning": "frontier-reasoning-llm", } def pick_model(client, user_request: str) -> str: r = client.system_one( state={"request": user_request}, questions={ "tier": Choice( instructions="What level of model does this request require?", criteria={ "simple": "Lookup, rewording, short factual answer", "standard": "Multi-paragraph writing or straightforward code", "reasoning": "Multi-step reasoning, math, or complex debugging", }, ), }, ) tier = r.choices["tier"] # When unsure, spend more rather than risk a bad answer if tier.confidence < 0.7: return MODELS["reasoning"] return MODELS[tier.choice] |
At sub-second latency and fractions of a cent per call, the router pays for itself if it downshifts even a modest share of traffic. Note the asymmetric fallback: uncertainty routes to the safer option.
Pattern 4: Guardrails on agent tool calls
Before an agent executes something, ask a narrow question. Jev is fast enough to sit in the loop, and Composio’s integration exposes exactly this kind of destructive-action gate.
|
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 |
def guard_tool_call(client, tool_name: str, args: dict, user_intent: str) -> str: r = client.system_one( state={"tool": tool_name, "arguments": args, "user_intent": user_intent}, questions={ "irreversible": Noul( instructions="Executing this tool call causes an irreversible " "change, such as deleting data or sending money." ), "matches_intent": Noul( instructions="This tool call matches what the user asked for." ), }, ) if r.answers["irreversible"].noul > 0.3: return "require_confirmation" if r.answers["matches_intent"].noul < 0.8: return "block_and_review" return "allow" |
This complements your deterministic checks and does not replace them. Authentication, authorization, resource ownership, and input validation stay in code [8]. Jev handles the fuzzy layer on top.
Pattern 5: Map over lots of data
Because calls are cheap and quick, you can afford to apply a judgment to every row instead of a sample.
|
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 |
from concurrent.futures import ThreadPoolExecutor from functools import partial def classify_review(client, review: str) -> dict: r = client.system_one( state={"review": review}, questions={ "topic": Choice( instructions="What is this review mainly about?", criteria={"shipping": None, "quality": None, "price": None, "support": None}, ), "negative": Noul(instructions="The reviewer is dissatisfied."), }, ) return { "topic": r.choices["topic"].choice, "topic_conf": r.choices["topic"].confidence, "negative": r.answers["negative"].noul, } with TypeSafeClient() as client, ThreadPoolExecutor(max_workers=16) as pool: results = list(pool.map(partial(classify_review, client), reviews)) |
You get typed features you can drop straight into a dataframe or warehouse table. Respect rate limits and use the SDK’s retry handling. One community integration even exposes the primitives as SQL table functions in DuckDB so you can LATERAL JOIN a judgment onto a table.
Pattern 6: Use an LLM and Jev together
Jev doesn’t generate, so pair it with something that does when you need both.
- Extraction: if you need a value pulled from free text, get candidates with a regex or an LLM and let Jev pick among them.
- LLM-output verification: score or guardrail a model’s draft answer (groundedness, policy compliance, tone) before it reaches a user. TypeSafe lists “verify everything” as a headline use case.
Limits
- Early days. Launched September 2026. Expect API and SDK changes, and expect the community ecosystem to be uneven.
- Vendor-reported benchmarks. The speed and cost claims are plausible given the architecture, but the evaluation workflows were built by TypeSafe, and TypeSafe acknowledges the pricing can’t yet be proven sustainable. Run your own bake-off. TypeSafe does publish the full workflows, queries, and disagreements on its evals site, which makes an independent replication feasible.
- Bounded input. State and all questions share roughly 64,000 tokens, and the state plus the longest single question must fit in about 32,000. Choice tops out at 255 options and Score at 10 levels. For larger choice sets, score candidates independently and then choose, which TypeSafe does in its own Wikirace demo.
- Text only. No vision yet.
- Jagged edges. Arithmetic, dates, distractor-heavy input, and adversarial state are weak spots per the vendor’s own documentation.
When to reach for it
Use Jev when the task is a bounded judgment made many times inside software, where latency or cost matters and where knowing your uncertainty lets you automate safely: triage, routing, scoring, moderation, guardrails, feature extraction from text, and real-time decisions.
Reach for an LLM when you need generated language, open-ended reasoning, code, or a conversation with a human.
References
-
- Diogo Almeida, “Introducing System One Models & Jev,” TypeSafe AI blog, September 15, 2026: https://typesafe.ai/blog/introducing-system-one-models-and-jev
- OpenRouter, “Jev Documentation”: https://openrouter.ai/docs/guides/community/jev
- “Jev AI TypeSafe AI: A Practical Guide to Typed Decisions in Production,” Hugging Face community blog: https://huggingface.co/blog/sora-2/jev-ai-typesafe-ai-a-practical-guide-to-typed-deci
- Flavio Copes, “A deep dive into Jev, TypeSafe’s System One model”: https://flaviocopes.com/jev/
- “How to Use Jev: A practical guide to TypeSafe’s System One model,” DEV Community: https://dev.to/valyuai/how-to-use-jev-a-practical-guide-to-typesafes-system-one-model-g5e
- awesome-jev, community-maintained list of SDKs, integrations, and limitations (unofficial): https://github.com/cobanov/awesome-jev/wiki and https://github.com/kraayenjon/awesome-jev/wiki
- Haskell
jevpackage README (notes on Score scaling and Noul semantics): https://hackage-content.haskell.org/package/jev - Langfuse, “TypeSafe integration”: https://langfuse.com/integrations/model-providers/typesafe.md
- HoneyHive, “How to trace TypeSafe with HoneyHive”: https://docs.honeyhive.ai/v2/integrations/typesafe.md
- LangChain, “What Is Jev? A Guide to TypeSafe AI’s System One Model”: https://www.langchain.com/blog/building-a-harness-with-jev