Introduction
What if you could run an AI model that makes thousands of split-second decisions for your application every day, costs next to nothing, and completely eliminates the slow, expensive bottleneck of text generation?
In this lesson, you will learn how Jev—a specialized model designed by Typesafe AI—shifts the paradigm of how application code interacts with artificial intelligence.
The Bottleneck of Traditional LLMs
Most mainstream large language models (LLMs) today behave like System 2 thinkers. Based on the cognitive science concepts popularized by Daniel Kahneman, System 2 thinking is slow, deliberate, and effortful—like calculating 17 × 24 in your head. Traditional LLMs are designed to generate conversational text for human consumption, which makes them slow, expensive, and difficult for application code to parse.
If you try to use a standard LLM to categorize 2,000 support tickets, it might return a wordy paragraph instead of a simple label. Even if you force it into a strict JSON schema to get a single word like “billing,” you face another hurdle: unreliable confidence scores.
Traditional models are trained using Reinforcement Learning from Human Feedback (RLHF). This tuning process rewards answers that human testers prefer, which does not necessarily align with mathematical accuracy. Consequently, if you ask a standard LLM how sure it is, its confidence score is essentially a guess.
To solve this, Typesafe AI introduced Jev, a model built from the ground up for System 1 thinking. System 1 is fast, automatic, and intuitive—like instantly recognizing a friend’s face. Jev is optimized to be consumed by code, not humans.
How Jev Thinks: Inputs, Outputs, and Question Types
Instead of sending conversational prompts, your application interacts with Jev by passing it two specific structures:
- The State: The raw text or data that needs to be judged, such as a customer support ticket.
- The Questions: The specific dimensions you want evaluated, along with defined criteria.
Jev evaluates these inputs and outputs a clean probability spread and a confidence score instead of text. Jev achieves highly calibrated probabilities through Reinforcement Learning for Calibrated Decisions (RLCD). Instead of human ratings, RLCD uses self-generated training data where the correct answer is already known. This allows the model’s confidence scores to map directly to real-world accuracy—an 80% probability means the answer will be correct exactly 80% of the time.
When querying Jev, you can use three distinct question types:
- Choice: Selects one option from a predefined list, such as routing to Billing, Technical, or Sales.
- Score: Places the input on an ordered scale, such as rating customer frustration from Calm to Annoyed to Furious.
- Null: Evaluates a plain yes-or-no statement and returns a single probability number between 0 and 1, such as “Is the customer asking for a refund?”
Architectural Patterns: Speculative Fan-Out and Gated Routing
Traditional LLMs charge a premium for output tokens because text generation is the most resource-intensive part of running AI infrastructure. Generating text requires the GPU to repeatedly load its entire weight library into memory for every single token produced, creating a massive memory-bandwidth bottleneck.
Because Jev only outputs numbers and probabilities, its output tokens are completely free. Its input cost is also incredibly low, sitting at about 4 cents per million tokens ($42 per billion).
This unique pricing and architecture enables developers to leverage highly efficient patterns:
- Speculative Fan-Out: Since Jev answers all questions in a single, parallel pass, you can ask multiple questions at once for virtually the same cost. You can ask whether the customer is angry, whether they mentioned a competitor, and whether they are planning to cancel—all in one call.
- Confidence-Gated Routing: You can build automated gates in your code based on Jev’s calibrated confidence scores. For example, if confidence is above 0.9, automatically route the ticket. If it is between 0.5 and 0.9, ask the user to confirm. If it is below 0.5, escalate the ticket to a human agent or a heavier System 2 LLM.
Summary
Key Takeaways
- System 1 vs. System 2: Traditional LLMs are slow, conversational System 2 processors. Jev is a fast, numeric System 1 processor designed specifically to be called by application code.
- RLCD vs. RLHF: Jev is trained via RLCD (Reinforcement Learning for Calibrated Decisions), ensuring its probability outputs are mathematically aligned with actual accuracy, unlike human-preferred RLHF models.
- Zero-Cost Outputs: Because Jev returns numbers rather than generating text, there is no text-generation bottleneck, making output tokens entirely free.
- Advanced Routing Patterns: Developers can use speculative fan-out to ask multiple questions in a single pass and build confidence-gated routing systems to automate high-volume workflows safely.
- Explicit Limitations: Jev excels at high-volume micro-judgments of text but cannot count, do math, or process dates chronologically.