핵심 요약
Jev는 텍스트를 생성하거나 대화하는 대신 오직 정형화된 '의사결정(분류, 점수, 참/거짓 판별)'만을 전문적으로 수행하는 새로운 AI 모델이다. 기존 대형 언어 모델 대비 속도는 20~200배 빠르고 비용은 40~400배 저렴해 초고속 자동화 시스템 구축에 최적화되어 있다. 글 작성이나 심층 분석은 불가능하지만 실시간 데이터 라벨링이나 초단위 자동 매매처럼 빠른 판단이 요구되는 작업에서 강력한 성능을 발휘한다.
주요 포인트
- **RLCD 기반 의사결정 특화 모델**: ChatGPT 공동 개발자가 만든 모델로, 텍스트 토큰을 생성하지 않고 의사결정(RLCD 기법 적용) 결과만 JSON 형태로 즉시 출력함.
- **3가지 결정 형태 제공**: 참/거짓(Yes/No와 신뢰도), 다중 선택(카테고리 분류 및 신뢰도), 점수 매기기(0~10 척도)의 세 가지 출력 방식으로 작동함.
- **압도적인 속도와 저렴한 비용**: 일반 챗 모델 대비 20~200배 빠르며 40~400배 저렴하고, 출력 토큰 비용이 발생하지 않음.
- **실시간 자동화 활용성**: 실시간 트윗 성격 분류 크롬 확장 프로그램, 초단위 비트코인 매매 판단 봇 등 빠른 실시간 처리가 필요한 분야에 적합함.
- **명확한 한계와 제약**: 텍스트 작성, 긴 글 요약, 심층 분석 기능은 지원하지 않으며 컨텍스트 창이 64k 토큰 수준으로 상대적으로 작음.
So, Jev is literally everywhere and I think it's going to change how AI automations are built. So, I came in here and tested it on 12 use cases and compared it with other AI models on things like speed and cost. I was even able to build this Chrome extension that will label X tweets as breaking or, you know, golden nuggets or AI slop in real time for me. On the back end, you can see that it uses Jev to actually instantly categorize all this stuff as soon as it enters my screen. It's not doing very well in the first hour, but I also made this Jev trader which is literally every single second analyzing if Bitcoin's going to go up or go down or stay. And then it basically places trades in real time for me because this model is so good at quick decisions. But anyways, by the end of this video, you'll understand how Jev works and where you should actually use it in your life. So, let's not waste any time and just get straight into this one. All right, we're going to start off with just like what is Jev. I'm not going to do a super super deep dive, just enough for you to understand how it works and what we're looking at in today's video. So the interesting thing about Jev is that it's an AI that makes decisions, but it doesn't write anything. It doesn't output tokens. It's not anything that you could actually have a conversation with. It just makes decisions. So here was kind of the announcement tweet from Dogo. He co-invented CHABT and then he has been building in the past 2 years this new way to train models, RLCD. And you can see what that stands for is reinforcement learning for calibrated decisions. And this is on Typesafe's blog. And by the way, if you want to actually get in here so that you can start playing around with Jev, then go to Typesafe AAI and join the wait list and then hopefully in a few hours you're able to sign in. But also, this is available through like Versel's gateway as well as open router. So if you're not in the wait list yet, then you can still go out and play with Jeff. But anyways, essentially what happens is instead of a normal chat model where you would send in a message like this, this is some sort of support ticket and the AI model would read it would reason, would think, and then output like a message or output some sort of classification. It basically just outputs these types of things which are a yes or no confidence level, a category, and sort of a score. So for the first one, is it urgent? 99% confidence is yes, it is urgent. Which team? There were probably multiple routes like technical or billing or support and it labeled it as technical. and then how frustrated. It gave it a one out of two on the frustration score or scale. But you're fully in control. You basically will set up Jev with, hey, this is essentially how you're supposed to make decisions and here is sort of like the classification criteria. So, it's three types of decisions. Like I said, the yes or no is called a new. The pick one is a choice. And then we have an actual score. And I'm going to show you guys real examples of all of these being run on these 12 use cases. So, don't worry. But I just wanted to sort of lay the foundation here. So, like I said, in this example, it's yes or no. And there's a confidence score in the team or categorization example. It's different categories as well as a confidence score and then a score from 0 to 10 on some sort of scale. And in here, one meant that they were frustrated. And the reason why this is getting so much traction is because Dio said that this is 20 to 200 times faster and 40 to 400 times cheaper with output tokens being free. And so if we look at the speed here, compared to models like Terra and Luna and Soul, this thing is going to be a lot faster. This was just one very quick test I ran. This doesn't mean that Terra is always faster than Luna, but this was just one quick example of how significantly faster Jev is. Once again, it doesn't have to output all these tokens or reason. It just boom makes a decision and outputs it in like this JSON format. And same thing from a cost perspective. If you were running thousands and thousands of decisions per day, this is what it could actually end up looking like. Now, obviously, the important thing is you're paying a lot less. So, you want to make sure that the quality is the exact same as what you'd be getting here or here to justify it. If we're just looking at the cost right now, it is significantly cheaper and it is proven that it's significantly cheaper. So, what it cannot do is write or summarize or find themes or do deep analysis. It basically just outputs decisions. And one other limitation right now is that it has a very small input context window. It's 64,000 tokens, whereas a lot of the models that we're used to using today, whether that be Claude or GBT, are more on the side of a million tokens. So, if we look at this on a use case like YouTube comments, if I fed in 5,000 YouTube comments, Jev could sort