
Un $40 million AI startup's model just got cloned and its free version is already sitting on Hugging Face. The original Jeff is locked behind an API. This free one beats it on accuracy, runs six to eight times faster, and you can run it on your own machine. Now, if you're not technical, this is basically an AI that never writes a single word. It only makes calls. You hand it a support ticket and ask, "Is this billing, a bug, or a refund?" And it answers in a blink, plus how sure it is.
So, I plugged into my own support inbox and ran 45 real tickets through it. It sorted every single one in two seconds, and it came out around a 100 times cheaper than running the same job through Claude. And here's the part most people miss. When most AI models say they're 90% sure, that number often doesn't mean much. This one was trained, so its confidence is actually honest. That is a big game changer because now we can let it handle the sure ones on its own and only send the shaky ones to a human.
What exactly is this model and why is it different
Most people think of AI as something that writes. ChatGPT writes emails, writes code, writes poems. This model doesn't write anything. It classifies. You give it a piece of text and it tells you what bucket it belongs in. That's it.
The original model is called Jeff. It was built by a startup that raised $40 million. They charge for API access. Someone cloned it, and now the clone is free on Hugging Face. And the clone is actually better. It's more accurate. It runs six to eight times faster. And you can run it locally, on your own machine, without sending data to anyone else's server.
For a lot of business use cases, this is exactly the kind of AI that matters. Not the one that generates paragraphs of fluffy text. The one that looks at a support ticket and says "billing" or "bug" or "refund" in under a second. The one that reads a customer email and routes it to the right department without a human ever touching it.
How I tested it with real tickets
I didn't just read the benchmarks and call it a day. I took 45 real support tickets from my own inbox and fed them through this model. Real tickets. Real customer problems. Real messy language.
Every single ticket got sorted in two seconds. Not some of them. Not most of them. Every single one. Two seconds.
Then I checked what it would have cost me to run the same 45 tickets through Claude. The difference was staggering. This free model did the job for about 100 times less. A hundred times. Not 20% cheaper. Not half the price. Two orders of magnitude.
And speed matters too. Six to eight times faster than the original Jeff. When you're processing thousands of tickets a day, that difference compounds. What took an hour now takes minutes. What took minutes now takes seconds.
The confidence number you can actually trust
Here's where it gets interesting. Most AI models will give you a confidence score. "I'm 92% sure this is a billing issue." But that number is often garbage. The model wasn't really trained to be honest about its uncertainty. It just spits out a probability that looks nice.
This model is different. Its confidence is actually honest. That's a direct result of how it was trained. When it says it's 90% sure, it really is 90% sure. When it's 50% sure, it's genuinely uncertain.
This changes everything about how you deploy it in a real workflow. With most models, you can't trust the confidence score enough to automate decisions. You still need a human reviewing everything because the model might be confidently wrong. With this one, you can set a threshold. Everything above 90% confidence gets handled automatically. Everything below gets routed to a human.
That's not a small optimization. That's a fundamental shift in how you run a support operation. The model handles the obvious stuff instantly. Humans only touch the edge cases. Nobody wastes time on "yes this is clearly a billing question" anymore.
What this means for support teams
Think about the typical support workflow. A ticket comes in. Someone reads it. They figure out what it's about. They route it to the right person or team. That first step, the triage step, is pure overhead. It doesn't solve the customer's problem. It just figures out who should solve it.
This model eliminates that step entirely. The ticket arrives, the model classifies it in two seconds, and it goes straight to the right place. The customer gets a faster response. The support team spends less time on administrative busywork. And because the confidence is honest, you're not introducing a bunch of misroutes that make things worse.
The cost angle is almost an afterthought, but it shouldn't be. A hundred times cheaper than Claude for the same task. If you're processing thousands of tickets a month, that's real money. Money you can put toward actually solving customer problems instead of sorting them.
Running it on your own machine
One detail that matters a lot: you can run this locally. It's not locked behind an API. You download it from Hugging Face, set it up on your own hardware, and it runs there. Your data never leaves your machine.
For any company that handles sensitive customer information, this is huge. You don't need to send support tickets to a third-party API. You don't need to negotiate data processing agreements for your AI tool. You just run it in-house.
The speed advantage compounds here too. No network latency. No API rate limits. No waiting for someone else's server to respond. The model is running on your hardware, processing your tickets, at six to eight times the speed of the original.
Why honest confidence is the real breakthrough
I keep coming back to the confidence thing because it's the part most people will overlook. Everyone focuses on accuracy and speed. Those are easy to measure and easy to brag about. But honest confidence is what makes the whole system actually usable in production.
Without it, you have two bad options. Option one: you trust the model's confidence scores and it confidently misroutes tickets. Option two: you ignore the confidence scores and have a human review everything, which defeats the purpose of automation.
With honest confidence, you get a third option. You trust the model when it's sure and escalate when it's not. The model knows what it doesn't know. That's rare in AI. It's also exactly what you need to build a system that works without constant human supervision.
The fact that this came from a cloned model, available for free, running faster than the original, is almost beside the point. The real story is that we finally have a classifier with confidence scores that mean something. That's the game changer.
Points clés à retenir
- Un modèle d'IA classificateur d'une startup à 40 millions de dollars a été cloné et est disponible gratuitement sur Hugging Face
- La version gratuite est plus précise, six à huit fois plus rapide, et fonctionne en local sur votre propre machine
- Le modèle ne génère pas de texte, il classe uniquement des tickets en catégories comme "facturation", "bug" ou "remboursement"
- Testé sur 45 vrais tickets de support, il les a tous triés en deux secondes pour un coût environ 100 fois inférieur à Claude
- Sa confiance est réellement honnête, ce qui permet de l'automatiser pour les cas certains et de n'envoyer que les cas incertains à un humain
- L'exécution locale signifie que vos données ne quittent jamais votre machine, ce qui est crucial pour les informations sensibles
This model isn't flashy. It doesn't write poetry or generate images. It just does one thing, does it fast, does it cheap, and tells you honestly when it's not sure. For anyone running a support team, that's worth more than a hundred general-purpose AI tools that can do everything but none of it reliably.
