핵심 요약
Claude Opus 5가 출시되어 터미널 코딩과 문제 해결 등 주요 벤치마크에서 최고 성능(SOTA)을 달성했다. Opus 5는 기존 Fable 5보다 우수한 성능을 보이면서도 비용은 절반 수준에 불과하여 뛰어난 가성비를 자랑한다. 실제 활용 가치가 높은 만큼 VS Code 등에서 Claude Code를 업데이트하여 직접 테스트해보는 것이 권장된다.
주요 포인트
- 터미널 코딩 벤치마크 43% 기록 (Fable 5: 33%, Opus 4.8: 21% 대비 큰 폭 상승)
- 문제 해결 벤치마크 성능이 기존 Opus 4.8(1.5%)에서 Opus 5(30%)로 급격히 향상
- 컴퓨터 사용, 비즈니스 워크플로우, 다학제 추론 등 주요 분야에서 Fable 5보다 우수한 성능 달성
- Fable 대비 절반 가격으로 제공되어 크레딧 소모를 줄이고 대폭 비용 절감 가능
- 벤치마크 수치에만 의존하지 말고 VS Code Claude Code 업데이트 후 실제 사용 환경에서 직접 테스트 권장
Okay, so Claude Opus 5 literally just dropped. Now, obviously I'm going to spend the rest of my day playing around with it, comparing it to Fable 5, and I'm going to bring you guys an experiment. So, stay tuned for that later today. But, I wanted to just real quick come in here and just show you guys this cuz I think it's really, really interesting. On several coding and knowledge work evaluations, Opus 5 is the new state of the art. So, look at this first benchmark right here. Agentic terminal coding. Opus 5 has 43%. Fable 5 had 33%, and Opus 4.8 had 21%. So, this is a major jump in agentic terminal coding above Opus 4.8, but also above Fable 5. Same thing with knowledge work. This thing crushes Opus 4.8. I don't exactly know what the novel problem-solving benchmark is solving, but look at this. It went from Opus 4.8 at 1.5 to 30% at Opus 5. So, anyways, these benchmarks are insane. You can also see here computer use, which I always would default to Codex. I always thought that GPT 5.6 or 5.5 with Codex was just the absolute best with computer use, but look at this with Opus. Opus 5, I'm going to have to try it out with computer use. Anyways, what's really interesting to me is this efficiency because as you guys know, Opus is half the price of Fable, and Fable will burn through your credits like nobody's business. But, look at this. Agentic computer use. Opus 5 performs better than Fable and for cheaper. Look at this. Agentic business workflows. Opus 5 is better than Fable and for cheaper, which is just really, really interesting to me. Multidisciplinary reasoning by effort level. Opus 5 does better than Fable and is cheaper than Fable. So, like I said, guys, I'm about to run a ton of usage credits. If you look at this right here in my Claude, I am already at my weekly Fable limit, and I'm about to go into, you know, thousands of dollars in extra usage credits, but I'm going to run a ton of experiments here, and I'm going to bring you guys my consensus. But, I just wanted to come in here and just like show you guys that real quick. It's really interesting. So, if you real quick go into your VS Code, go into Claude Code, update it, get on Opus 5, and just start running your skills and start running things and seeing how it feels. Because as we know, these benchmarks are fun to look at, but always take them with a grain of salt. It always matters on your use case, the way you talk to agents, the type of work you're doing, things like that. So, Claude Opus 5 provides greatly improved performance for the same cost as its predecessor Opus 4.8. Opus 5 excels at valuable software engineering tasks like the Frontier bench, on Cursor bench. I really want to see how it performs on the Deep suite because that's the one that's been like a pretty true like leading indicator. But, the whole cost thing is just really interesting to me compared to Fable 5. Cuz I know a lot of you guys were kind of like, "Oh man, it it sucks that we're going to be losing Fable or that, you know, we only get a certain different limit for it, but it's overkill for most things and it's really expensive. So, if Opus 5 is better for like general knowledge work, then that's a huge win. And this is really exciting to me. Opus 5 is much stronger at verifying its work and iterating carefully until it succeeds. I heard this really funny analogy on X, which is honestly pretty true. They basically said like, "Hey, Fable 5 is like a wise old owl. It's really good at planning and managing and ideating and coming up with ideas, whereas GPT-5.6-Soul is like a Rottweiler because that thing will grab onto a task and it will shake it and it won't let go until the task is done, which I definitely felt. It was spinning up more tests and it was better at verification than Fable was, right? But now, maybe Opus 5 saw that and was like, "Oh wow, we need to work in some of this GPT-5.6-Soul verification." And now, if we have that, that's huge because verification is kind of what powers all of these agentic loops that everyone's talking about. And if you're using a model that's really, really good at verification, your loops are going to be better and your outputs are going to be better. So, I'm really excited to test out Opus 5 with some verification. And here are some examples where they gave Opus 5 some real scenarios and had it find bugs, verify, take different approaches, and ultimately give a good output. But, anyways, this is now completely available to everybody on all platforms, same exact price as Opus 4.8. So, like I said, I'm about to jump in right now and start testing the heck out of it against Fable 5 and I'll bring you guys that full experiment later today. So, keep your eyes peeled for that video. But, anyways, I'll see you guys there. Thanks, everyone.