
On September 15, a new model dropped and it immediately racked up over a thousand points on Hacker News. It’s called Jeff. Not GPT-5, not Claude 4, Jeff. A couple days after the launch, Vercel had already added it to their AI gateway. The company behind it is Typesafe AI, founded by someone who co-invented ChatGPT. So the pedigree is there, but the approach is completely different. Jeff isn’t a text generation model. It’s a decision-making model.
When you send a prompt to a normal large language model, it generates a response token by token. That’s slow. In one example, a standard LLM took 8 seconds to answer a question. Give the same question to Jeff and it doesn’t generate a long text response. You provide it with a set of options, Jeff scores each one, and whichever option gets the highest score is the decision it returns. That’s it. No paragraphs, no explanations, just a choice. Because it skips the text generation entirely, the response time is drastically faster. In a quick demo, Jeff played Doom and was making 10 decisions per second. That’s fast enough to play a real-time game.
What Makes Jeff Fundamentally Different
The core difference is in what the model outputs. A standard LLM call is a conversation: you ask, it answers, and the answer takes time because every word is predicted one after another. Jeff doesn’t answer. It decides.
You frame the problem as a set of possible actions or categories. Jeff evaluates each one and assigns a score. The highest-scoring option wins. That’s the entire output. This isn’t a model that’s been fine-tuned to be terse; it’s architected from the ground up to output decisions, not prose.
Speed and Cost Structure
Because Jeff doesn’t generate text, the cost structure is unusual. Output tokens are completely free. You pay only for input tokens. At the time of launch, the pricing was $0.042 per 1 million input tokens. Compare that to models like OpenAI’s O1 or Anthropic’s Opus, where you’re looking at $5, $10, or more for the same volume. The difference is enormous.
The speed advantage comes from the same design choice. A model that only needs to score a handful of options can return a result in a fraction of a second. The Doom demo showed 10 decisions per second. For any application where latency matters, and there are a lot of them, that’s a game changer.
Intelligence and Hallucination
Jeff’s intelligence score sits in the same range as models like Gemini 1.5 Flash or Claude 3.5 Sonnet. It’s not competing with the deep frontier reasoning models like O1 or Opus, but it’s not trying to. The hallucination rate is effectively zero. When a model’s only job is to pick from options you’ve already defined, there’s nothing to hallucinate. It can’t invent facts because it can’t invent text.
What Jeff Can and Cannot Do
This is where you need to be clear-eyed about the tradeoffs. Jeff can pick an action, classify an input, score a set of candidates, and return a result faster than any text-generating model. That’s the strength.
What it cannot do is equally important. Jeff cannot write a sentence. It cannot explain its reasoning. It cannot write code. It cannot reason step by step. If you need a model that walks through a complex problem and shows its work, Jeff is the wrong tool. Those deep reasoning tasks belong to the frontier models that take their time.
Think of it this way: Jeff has a similar intelligence score to Sonnet or Flash, but it’s not a deep thinker. It’s a fast reactor. You don’t throw away your other models when you add Jeff to your stack. You use Jeff for the fast decision points and let the heavier models handle the deep work.
Real Use Cases People Are Already Building
There’s a public repository curating projects built with Jeff, and at the time of recording, it already had over 40 entries. A large portion of them fall into two categories: classification and routing.
Classification and Routing
Say you’re running a code review pipeline. A piece of code comes in and you need to decide which model or which tool should handle it. You could write a bunch of rules, or you could let Jeff make that call instantly. It scores the options and routes the task to the right place. The same pattern works for any multi-model setup where you want to save tokens by not sending every request to your most expensive model. Jeff acts as a fast, cheap dispatcher.
Games and Simulations
The Doom demo isn’t just a party trick. The same capability applies to drones, Pokémon bots, or any simulation where an agent needs to make rapid decisions from a constrained set of actions. You define the possible moves, Jeff picks the best one, and the loop runs at 10 decisions per second.
Financial and Short-Term Trading
One of the most interesting applications is day trading. Short-term trading requires fast decisions based on incoming data. You define the possible trades as options, feed in the market context, and Jeff returns a scored decision. The speed means you can react to price movements without waiting seconds for a text model to finish writing a paragraph about its reasoning. The decision lands immediately.
Browser Automation
Browser automation with AI has been possible for a while, but it’s slow. When you use ChatGPT’s browser tool or Claude’s computer use, the model has to process the page, generate a plan, and output text describing each action. That consumes tokens and time.
Jeff changes the equation. The browser-use library has already integrated Jeff. You feed it all the possible operations on a page, click here, type there, scroll, and Jeff scores them. It decides what to do next without generating a sentence. In one demo, someone gave Jeff a source and destination, and it navigated a flight booking site and searched for flights in under 10 seconds. That’s end-to-end browser interaction at a speed that feels closer to a macro than an AI agent.
How to Try Jeff Right Now
There are two paths to access Jeff, and both involve a waitlist or a specific platform.
The Official Waitlist
Typesafe AI has a waitlist for direct access. You sign up and wait. That’s the route for getting your hands on the raw model and building whatever you want around it.
Vercel AI Gateway
If you want to try Jeff immediately, Vercel has it available in their AI gateway. It was added on September 15, the same day as the launch. You can find it by searching the models list in the Vercel dashboard. At the time of recording, it was completely free to use through Vercel.
You can call it through the Vercel AI SDK with a straightforward integration, or you can use a simple curl request. Either way, you send your options and get back a decision.
Claude Skills and AI Agent Skills
Typesafe AI also released a set of skills for Claude and other AI agents. These are essentially pre-built instructions that teach an AI agent how to use Jeff as a decision-making subcomponent. You pass the skill to your agent, and it learns to delegate fast, small decisions to Jeff while handling the broader conversation or workflow itself. This lets you weave Jeff into existing automation without building everything from scratch.
Where Jeff Fits in a Multi-Model Workflow
The mental model that makes sense here is a tiered architecture. You keep your deep reasoning models, O1, Opus, Sonnet, for tasks that require step-by-step thinking, code generation, or long-form explanation. In front of them, or alongside them, you place Jeff for the decision points.
A practical example: a customer support pipeline. A message comes in. Jeff classifies it, refund request, technical issue, billing question, in milliseconds. Based on that classification, the message gets routed to the appropriate handler. Maybe a simple FAQ bot handles refunds, while a more expensive model handles technical troubleshooting. Jeff made the routing decision for nearly zero cost and zero latency.
The same pattern works for content moderation, lead scoring, intent detection, or any system where you have a finite set of categories and need a fast, accurate choice.
Points clés à retenir
- Jeff n’est pas un modèle de génération de texte, c’est un modèle de prise de décision qui score des options et retourne le meilleur choix.
- Les tokens de sortie sont gratuits parce que le modèle ne génère pas de texte, ce qui rend son utilisation extrêmement peu coûteuse comparée aux LLMs traditionnels.
- La vitesse est le vrai avantage : Jeff peut prendre 10 décisions par seconde, assez rapide pour jouer à Doom en temps réel.
- Le taux d’hallucination est quasiment nul puisqu’il choisit uniquement parmi les options que vous lui fournissez, sans rien inventer.
- Jeff ne peut pas écrire de phrases, expliquer son raisonnement, coder ou raisonner étape par étape, ce n’est pas un modèle frontalier de réflexion profonde.
- Les cas d’usage réels incluent la classification, le routage entre modèles, les jeux et simulations, le trading court terme et l’automatisation de navigateur.
- On peut tester Jeff tout de suite via la passerelle IA de Vercel ou s’inscrire sur la liste d’attente officielle pour un accès direct.
The thing that sticks with me about Jeff is how it forces you to rethink what you actually need from a model. A lot of tasks we throw at LLMs don’t require a paragraph of reasoning. They require a fast, accurate pick from a known set of options. Jeff strips away everything that isn’t that. The result is a model that’s cheaper, faster, and doesn’t hallucinate, because it literally can’t. It’s not a replacement for the heavy lifters, but for the right slice of the workflow, it’s exactly the right shape.
