← 대시보드로
2026-07-25
3개 영상
2026-07-25 06:01 생성
Claude Opus 5 is Going to Save You Money
Nate Herk 2026-07-25
Claude Opus 5 is Going to Save You Money
↗ 유튜브에서 보기

핵심 요약

Claude Opus 5가 출시되어 터미널 코딩과 문제 해결 등 주요 벤치마크에서 최고 성능(SOTA)을 달성했다. Opus 5는 기존 Fable 5보다 우수한 성능을 보이면서도 비용은 절반 수준에 불과하여 뛰어난 가성비를 자랑한다. 실제 활용 가치가 높은 만큼 VS Code 등에서 Claude Code를 업데이트하여 직접 테스트해보는 것이 권장된다.

주요 포인트

  • 터미널 코딩 벤치마크 43% 기록 (Fable 5: 33%, Opus 4.8: 21% 대비 큰 폭 상승)
  • 문제 해결 벤치마크 성능이 기존 Opus 4.8(1.5%)에서 Opus 5(30%)로 급격히 향상
  • 컴퓨터 사용, 비즈니스 워크플로우, 다학제 추론 등 주요 분야에서 Fable 5보다 우수한 성능 달성
  • Fable 대비 절반 가격으로 제공되어 크레딧 소모를 줄이고 대폭 비용 절감 가능
  • 벤치마크 수치에만 의존하지 말고 VS Code Claude Code 업데이트 후 실제 사용 환경에서 직접 테스트 권장
Okay, so Claude Opus 5 literally just dropped. Now, obviously I'm going to spend the rest of my day playing around with it, comparing it to Fable 5, and I'm going to bring you guys an experiment. So, stay tuned for that later today. But, I wanted to just real quick come in here and just show you guys this cuz I think it's really, really interesting. On several coding and knowledge work evaluations, Opus 5 is the new state of the art. So, look at this first benchmark right here. Agentic terminal coding. Opus 5 has 43%. Fable 5 had 33%, and Opus 4.8 had 21%. So, this is a major jump in agentic terminal coding above Opus 4.8, but also above Fable 5. Same thing with knowledge work. This thing crushes Opus 4.8. I don't exactly know what the novel problem-solving benchmark is solving, but look at this. It went from Opus 4.8 at 1.5 to 30% at Opus 5. So, anyways, these benchmarks are insane. You can also see here computer use, which I always would default to Codex. I always thought that GPT 5.6 or 5.5 with Codex was just the absolute best with computer use, but look at this with Opus. Opus 5, I'm going to have to try it out with computer use. Anyways, what's really interesting to me is this efficiency because as you guys know, Opus is half the price of Fable, and Fable will burn through your credits like nobody's business. But, look at this. Agentic computer use. Opus 5 performs better than Fable and for cheaper. Look at this. Agentic business workflows. Opus 5 is better than Fable and for cheaper, which is just really, really interesting to me. Multidisciplinary reasoning by effort level. Opus 5 does better than Fable and is cheaper than Fable. So, like I said, guys, I'm about to run a ton of usage credits. If you look at this right here in my Claude, I am already at my weekly Fable limit, and I'm about to go into, you know, thousands of dollars in extra usage credits, but I'm going to run a ton of experiments here, and I'm going to bring you guys my consensus. But, I just wanted to come in here and just like show you guys that real quick. It's really interesting. So, if you real quick go into your VS Code, go into Claude Code, update it, get on Opus 5, and just start running your skills and start running things and seeing how it feels. Because as we know, these benchmarks are fun to look at, but always take them with a grain of salt. It always matters on your use case, the way you talk to agents, the type of work you're doing, things like that. So, Claude Opus 5 provides greatly improved performance for the same cost as its predecessor Opus 4.8. Opus 5 excels at valuable software engineering tasks like the Frontier bench, on Cursor bench. I really want to see how it performs on the Deep suite because that's the one that's been like a pretty true like leading indicator. But, the whole cost thing is just really interesting to me compared to Fable 5. Cuz I know a lot of you guys were kind of like, "Oh man, it it sucks that we're going to be losing Fable or that, you know, we only get a certain different limit for it, but it's overkill for most things and it's really expensive. So, if Opus 5 is better for like general knowledge work, then that's a huge win. And this is really exciting to me. Opus 5 is much stronger at verifying its work and iterating carefully until it succeeds. I heard this really funny analogy on X, which is honestly pretty true. They basically said like, "Hey, Fable 5 is like a wise old owl. It's really good at planning and managing and ideating and coming up with ideas, whereas GPT-5.6-Soul is like a Rottweiler because that thing will grab onto a task and it will shake it and it won't let go until the task is done, which I definitely felt. It was spinning up more tests and it was better at verification than Fable was, right? But now, maybe Opus 5 saw that and was like, "Oh wow, we need to work in some of this GPT-5.6-Soul verification." And now, if we have that, that's huge because verification is kind of what powers all of these agentic loops that everyone's talking about. And if you're using a model that's really, really good at verification, your loops are going to be better and your outputs are going to be better. So, I'm really excited to test out Opus 5 with some verification. And here are some examples where they gave Opus 5 some real scenarios and had it find bugs, verify, take different approaches, and ultimately give a good output. But, anyways, this is now completely available to everybody on all platforms, same exact price as Opus 4.8. So, like I said, I'm about to jump in right now and start testing the heck out of it against Fable 5 and I'll bring you guys that full experiment later today. So, keep your eyes peeled for that video. But, anyways, I'll see you guys there. Thanks, everyone.
How to Automate an Accounting Firm
Nicholas Puru 2026-07-25
How to Automate an Accounting Firm
↗ 유튜브에서 보기

핵심 요약

회계법인에서 몸값이 가장 높은 파트너들이 매달 초 반복적인 리포트 작성 작업에 시간을 낭비하지 않도록 업무를 자동화하는 솔루션을 제시한다. AI 에이전트가 퀵북스(QuickBooks)에서 수치를 가져와 리포트 템플릿과 코멘터리 초안을 자동 작성하고, 파트너는 검토 및 승인만 담당하도록 구성한다.

주요 포인트

  • 파트너급 고급 인력이 매월 초 5일 동안 반복적인 고객 리포트 재작성 업무에 묶여 있는 문제를 해결함
  • 장부 마감 즉시 AI 에이전트가 퀵북스에서 데이터를 직접 추출해 리포트 템플릿에 자동 입력함
  • 월별 변동 내역, 전월 대비 차이점, 고객 전달 필요 사항 등 리포트 분석 코멘터리 초안을 자동 생성함
  • 수치를 임의 생성하지 않고 감사 추적(Audit trail)을 통해 모든 수치의 출처 링크를 명확히 남김
  • AI는 초안만 생성하며, 파트너의 최종 검토와 승인을 거쳐야만 고객에게 리포트가 발송되는 안전장치를 둠
If I ran an accounting firm, here's exactly what I would do to automate my entire business. I'd first get my partners off of work that doesn't need a partner. Because right now, the most expensive people in your firm are spending the first 5 days of every single month just rebuilding the same client report. Watch this. So, the second the books close, an agent pulls the numbers straight from QuickBooks, it drops them into your reporting template, and it writes the first draft of the commentary. So, what moved this month, what's off versus last month, and what the client needs to be told. So, every number is pulled from the source, never just invented and every line it links back to where it came from on an audit trail. It then lands into a review queue where your partner edits the story and sends it off. Now, this always drafts and never sends anything itself. So, nothing reaches a client until a human approves it. If you want this set up for your firm, just comment need and I'll send you more information.
I Tested Opus 5 vs. Fable 5. What You Need to Know.
Nate Herk 2026-07-25
I Tested Opus 5 vs. Fable 5. What You Need to Know.
↗ 유튜브에서 보기

핵심 요약

Claude Opus 5는 Fable 5의 절반 가격이면서도 코딩 및 지식 작업 벤치마크에서 오히려 Fable 5를 앞서는 성능을 보여준다. 작성자는 Claude Code 하네스와 Chat 환경에서 두 모델의 비용, 시간, 토큰 소비 및 실제 결과물을 직접 비교 실험했다. 실무 워크플로우 대부분에서 Opus 5가 가성비와 성능 모두를 만족시키는 강력한 대안임을 확인했다.

주요 포인트

  • Opus 5는 Fable 5 대비 50% 저렴하면서 Frontier Bench, Cursor Bench 등 주요 지표에서 더 뛰어난 성과를 냄
  • 단순 벤치마크 참고에 그치지 않고 Claude Code 하네스 내부 및 Claude Chat에서 직접 비교 테스트 진행
  • AI 에이전트용 벡터 검색 원리 다이어그램(Excalidraw JSON) 생성 등 실제 프롬프트로 성능 검증
  • 사용자가 각 모델을 효율적으로 선택할 수 있도록 비용, 소요 시간, 토큰 사용량 기준 분석 제공
All right, so Claude Opus 5 is here and if you start to look at the benchmarks, it's really interesting because it shows us that for a lot of things that I care about, it's actually better than Fable and it is half the cost of Fable. Ultimately, Fable 5 is still Anthropic's most impressive and, you know, strongest model. But for a lot of these things, you know, I've realized when I'm doing knowledge work and when I'm building, you know, my videos or my research or whatever it is, Opus is more than enough power than what I need. And when you look at some of these charts, it's really interesting because it shows on things like the Frontier Bench and the Cursor Bench and this coding agent index that Opus is actually outperforming Fable and it's cheaper. And this really shocked me. So, obviously, I like to take all this stuff with a grain of salt. It's fun to look at and it's good to look at, but you want to actually get your hands dirty and run these models through your own actual workflows. So in today's video, I'm just going to break down a bunch of different experiments that I ran with Opus 5 versus Fable 5 and break down things like the cost, the time, and the tokens so that you can start to understand where you should work in these different models within your workflows. All right, so pretty much all of the experiments that I've been running today that I'm going to show you guys, I did within Claude Code, which means we're comparing the models, but also inside of the Cloud Code harness. and the variable is the same so it doesn't really change too much but I did do a few tests where I was actually in clawed chat and I was just you know seeing how they felt without a harness wrapped around and let me just show you one quick example so here I asked Fable 5 and Opus 5 to generate me an Excal diagram that accurately and visually explains how semantic search with vectorization works on a large data set for AI agents and it's interesting here because there's no skills that it can use and it doesn't have any context of me or you know any really way to verify all it did was it spit out a um JSON file of Excal for me and then I pasted it into Excal. So, here's what we got. Fable came back with this version over here where we see we've got like our indexing pipeline. We have a large database and it looks like it actually misspelled this right here, which is interesting. Large. Oh, data set. Okay, it was just like not expanded enough. Same thing over here. Vectorize. And this is part of the whole um it had no way to verify. And as you guys know, if you've been kind of building agent loops and stuff, verification is so so important. So anyways, large data set, we chunk it up. We vectorize it with an embedding model. We then get our embeddings, which is just like the numerical representation of the data. We put it into a vector database here. And then we can actually start to search. So we've got similarity and that's on, you know, points being close together. So your question lands here and it would grab the k nearest neighbors and different meaning is farther apart. We've got different clusters here. And then we come down here to the actual query. So if the user asks how I get my money back, the agent searches the knowledge base and then it does semantic search. It looks it up with the actual query vectors and we get the matches back. So pretty accurate. I will say though, Opus' layout seems a bit more organized, right? Like it's it's got boxes and it's got I mean this might not be as visual. You could argue you could argue that Fables was more visual, which you know I think that that's true, but this definitely feels more organized. it feels a little bit more detailed as well. So, that's just a very subjective exam. A lot of the stuff that I'm going to be talking about today is just really opinionated and subjective, but I'm going to still give you my own thoughts. So, in this example, I think that if I wanted to teach someone, I probably would take Opus 5's version here. Okay. So, let's start off with the first test I ran, which was basically giving them a huge codebase and having them look through any bugs and looking through like the expected behavior and some instructions like that. So, it had to do some exploration here and help us out, right? So, I set the goal and I gave it this prompt. And then what I did is I had Codeex review the output that Fable 5 gave us and that Opus gave us. So, real quick before we look at the results, Fable took about 11 minutes and it costed 5 bucks. 30, whereas Opus here took 13 minutes, so a little bit longer, but it was cheaper at $4.22. You can also see the breakdown here of input and output tokens and like what models they use and stuff like that. But, let me switch over to Codeex here. This is the actual result. So head-to-head, they both pretty much passed everything, which is great. But Codex thinks that Fable wins here because Fable's production patch is exactly the onelin