TypeSafe Jev (System One Decision Model)
TypeSafe Jev is a System One decision model delivering sub-100ms latency, zero-token generation overhead, typed scoring, and guaranteed $0 output token billing.
Last updated
typesafe/jev is an enterprise System One probabilistic decision model engineered specifically for high-throughput classification, intent routing, policy verification, and structured decision scoring. Unlike traditional autoregressive generative models that produce conversational text token-by-token, Jev operates as a non-generative decision engine that delivers sub-100ms latency, zero-token generation overhead, and a predictable $0.00 output token billing contract.
Architectural Paradigm: System 1 vs. System 2
Modern AI architectures are adopting dual-process cognitive designs inspired by Daniel Kahneman's Thinking, Fast and Slow:
| Attribute | System One: TypeSafe Jev (typesafe/jev) | System Two: Generative LLMs (Claude Opus, GPT-5, Llama 405B) |
|---|---|---|
| Primary Function | Fast, reflexive classification, intent routing, policy enforcement | Deep multi-step reasoning, creative synthesis, open-ended dialog |
| Execution Model | Single forward pass probabilistic decision scoring | Autoregressive token-by-token sequence generation |
| Typical Latency | 35ms – 80ms (sub-100ms guaranteed) | 400ms – 5,000ms+ |
| Output Token Risk | $0.00 (Zero generation overhead) | High variance based on output token volume |
| Determinism | Strict schema validation, typed logits, calibrated probabilities | Probabilistic generation subject to drift or hallucination |
By deploying TypeSafe Jev at the gateway boundary as a System One gatekeeper, applications inspect incoming user intent, route to domain-specific agent loops, and reject invalid requests before expensive System Two reasoning models execute.
Incoming Request
│
▼
┌────────────────────────┐
│ TypeSafe Jev (Sys 1) │ ──► Latency: ~40ms | Cost: $0.20/1M in, $0.00 out
│ Typed Decision Score │
└────────────────────────┘
│
├──► Intent: "Billing Inquiry" ──► Route to Lightweight Agent (GPT-4o-mini)
├──► Intent: "Complex Code RCA" ──► Route to Frontier LLM (Claude Sonnet 3.7 / Opus)
└──► Intent: "Policy Violation" ──► Immediate Terminal Block ($0 spent on downstream)Pricing Contract
TypeSafe Jev operates on an asymmetric, budget-safe pricing contract designed for ultra-high-volume evaluation pipelines:
| Metric | Rate | Notes |
|---|---|---|
| Input Tokens | $0.20 / 1M tokens | Billed at flat provider list price with zero markup (Rule #28) |
| Output Tokens | $0.00 / 1M tokens | Guaranteed $0.00 output charge; no token generation waste |
| Cache Read Tokens | $0.05 / 1M tokens | Billed when identical input contexts hit the nRouter gateway cache |
| Context Window | 128,000 tokens | Large context capacity for processing long documentation or multi-turn history |
| Max Output Tokens | 1,024 tokens | Compact, strictly typed decision payloads |
Because output tokens are billed at $0.00, your organization never faces budget overruns caused by runaway generation loops or verbose explanations.
Key Capabilities
1. Sub-100ms Empirical Latency
Jev eliminates the autoregressive decoding loop that accounts for 85–95% of LLM latency. While standard generative models require hundreds of sequential GPU matrix multiplications to emit tokens one by one, Jev evaluates input embeddings and decision heads in parallel, achieving p50 latencies under 40ms.
2. Zero-Token Generation Overhead
Generative models frequently emit conversational filler ("Sure! Here is the JSON you requested:") or invalid markdown backticks that waste tokens and require brittle client-side regex parsers. Jev outputs pure, validated decision schemas with calibrated confidence logits.
3. Typed Decision Scoring & Schema Validation
Jev returns structured decision objects conforming to your defined schema, including classification categories, calibrated confidence scores (0.0 to 1.0), and intent labels.
Model Comparison: Decision Tasks
Comparing TypeSafe Jev against general-purpose generative models for intent classification and decision workflows:
| Feature / Metric | TypeSafe Jev (typesafe/jev) | OpenAI GPT-4o-mini | Anthropic Claude 3.5 Haiku | Google Gemini 2.0 Flash |
|---|---|---|---|---|
| Model Type | System 1 Decision | Generative LLM | Generative LLM | Generative LLM |
| p50 Latency | 35ms | 140ms | 160ms | 120ms |
| Input Cost (per 1M) | $0.20 | $0.15 | $0.80 | $0.10 |
| Output Cost (per 1M) | $0.00 | $0.60 | $4.00 | $0.40 |
| Output Generation Waste | None (0 tokens) | 10–50 tokens/call | 10–50 tokens/call | 10–50 tokens/call |
| Cost per 1M Decisions* | $200 | $450 – $750 | $2,800 – $4,800 | $300 – $500 |
| Type Safety | Native Typed Scoring | Prompt-guided JSON | Prompt-guided JSON | Prompt-guided JSON |
*Estimated based on 1,000 input tokens and 50–100 generated output tokens per decision on general LLMs.
Request & Response Format
TypeSafe Jev is called through the standard /v1/chat/completions endpoint using either the full model ID typesafe/jev or the alias jev.
POST https://api.nrouter.ai/v1/chat/completionsIntent Classification Example
curl https://api.nrouter.ai/v1/chat/completions \
-H "Authorization: Bearer $NROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "typesafe/jev",
"messages": [
{
"role": "system",
"content": "You are a customer routing classifier. Classify the user query into: [BILLING, TECHNICAL_SUPPORT, SALES, CHITCHAT]."
},
{
"role": "user",
"content": "How do I upgrade our organization subscription to the Enterprise tier?"
}
],
"response_format": { "type": "json_object" }
}'Standard Decision Payload Response
{
"id": "chatcmpl-jev-84a1e9",
"object": "chat.completion",
"created": 1709251200,
"model": "typesafe/jev",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "{\"decision\":\"SALES\",\"confidence\":0.984,\"category\":\"subscription_upgrade\",\"route_target\":\"agent_sales_deals\"}"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 58,
"completion_tokens": 28,
"total_tokens": 86
}
}Response headers reflect the sub-100ms execution and $0.00 output charge:
x-nr-request-id: req_jev_0192e4
x-nr-model: typesafe/jev
x-nr-latency-ms: 38
x-nr-request-cost: $0.000011SDK Integration Examples
import os
import json
from openai import OpenAI
client = OpenAI(
api_key=os.environ.get("NROUTER_API_KEY"),
base_url="https://api.nrouter.ai/v1",
)
def classify_and_route(user_prompt: str) -> dict:
response = client.chat.completions.create(
model="typesafe/jev",
messages=[
{
"role": "system",
"content": "Evaluate query intent. Output JSON with fields: intent, confidence, requires_code_agent.",
},
{"role": "user", "content": user_prompt},
],
response_format={"type": "json_object"},
temperature=0.0,
)
return json.loads(response.choices[0].message.content)
decision = classify_and_route("Fix a memory leak in our C++ thread pool implementation.")
print("Decision:", decision)
# Output: {'intent': 'code_debugging', 'confidence': 0.992, 'requires_code_agent': True}Best Practices for Decision Pipelines
- Gate Heavy Models: Run
typesafe/jevbefore dispatching to frontier models like Claude Sonnet or GPT-5. If a request is simple chit-chat, routing, or an invalid query, resolve or reject it in ~35ms for pennies. - Deterministic Temperature: Always specify
temperature: 0.0for decision scoring to ensure maximum consistency across classification distributions. - Use with Smart Routers: Combine Jev with nRouter's Smart Routers (
nrouter/auto) or custom fallbacks. Use Jev's outputroute_targetto dynamically select the model identifier for your next execution turn. - Zero-Token Guardrails: Because Jev charges $0 for output tokens, high-frequency security filtering and compliance auditing can be run on 100% of user traffic without unpredictable cost inflation.
Next Steps
- Model Catalog — Compare Jev against generative models
- Meta Llama Guide — Use Jev to route between Llama 70B and 405B
- Router Settings — Configure intelligent routing and caching
- Observability Guide — Monitor latency distributions and costs
Meta Llama Models
Access Meta Llama 3.3 70B, 3.1 405B, and Llama 4 Scout via nRouter's unified gateway with real-time streaming, guardrails, and zero-markup list pricing.
A/B Testing
Split live production traffic between model variants to compare output quality, latency, and cost across providers in real time with zero code changes.