コンテンツへスキップ
BlogAI Research

Claude vs. GPT: 私たちが見つけたこと

We're model-agnostic. Here's what an internal benchmark of 500 identical financial-research tasks showed when we compared GPT-4o and Claude Sonnet on accuracy, reasoning, hallucinations, and confidence calibration.

V

Valarn

AI Research

28. April 2026
6 min read
AILLMGPT-4o
Claude vs. GPT: 私たちが見つけたこと

At Valarn, we're model-agnostic.

That means we're not tied to one AI model. We use whichever model gives our users the best research.

To make sure we're using the right tools, we regularly compare the latest AI models.

Earlier this year, we compared GPT-4o and Claude Sonnet using 500 identical financial research tasks.

One important note: this is an internal benchmark, not an academic study. These results reflect our testing and should be viewed as helpful guidance, not absolute truth.

How We Tested#

We used 500 historical research scenarios, including:

  • 200 large U.S. companies
  • 150 mid-sized U.S. companies
  • 100 sector ETFs
  • 50 international companies

Each analysis was run using the same prompts and the same historical information. Because we already knew what happened afterward, we could compare each model's research with the market's later performance.

What We Measured#

We looked at several areas:

  • How often the overall market outlook matched later market direction.
  • How well the AI explained its reasoning.
  • How often it presented incorrect or made-up information.
  • Whether its confidence scores matched its actual performance.

Overall Results#

Model30-Day Accuracy90-Day Accuracy
GPT-4o62.1%58.4%
Claude Sonnet64.8%61.7%

Both models performed well, but Claude Sonnet consistently scored higher in our testing.

Confidence Matters#

One thing stood out.

As Claude Sonnet became more confident, its research generally became more reliable.

GPT-4o didn't follow the same pattern. In fact, its highest-confidence analyses were sometimes less reliable than those with slightly lower confidence.

That's important because confidence should mean something.

Quality of Reasoning#

Three senior analysts reviewed every report.

They looked for:

  • Clear explanations
  • Strong use of supporting data
  • Logical thinking
  • Honest discussion of uncertainty
  • Consideration of different viewpoints

Overall, Claude Sonnet produced more balanced and better-supported research.

GPT-4o often provided very detailed answers, but some details sounded more certain than the available evidence supported.

Sometimes being more detailed doesn't mean being more accurate.

Hallucinations#

This was the biggest surprise.

A hallucination happens when an AI presents incorrect information as if it were true.

In our testing:

  • GPT-4o: 8.3% of reports contained at least one incorrect data point.
  • Claude Sonnet: 3.1% of reports contained at least one incorrect data point.

That's nearly a 3× difference.

For financial research, this matters.

A single incorrect earnings number or financial metric can change the entire conclusion of an analysis.

We'd rather have an AI say "I don't know" than confidently provide the wrong number.

Confidence Calibration#

Another area where Claude performed better was confidence calibration.

Simply put, when Claude reported high confidence, it was more likely to actually deserve that confidence.

That makes its confidence scores more useful when you're evaluating research.

What Changed at Valarn?#

Based on these results, we updated our platform.

1. Claude Sonnet became our default model.#

The lower hallucination rate and more reliable confidence scores made it the better choice for our research pipeline.

2. GPT is still part of the process.#

We didn't throw GPT out.

Instead, we use it where it performs well.

In some Ensemble Runs, GPT provides an additional perspective that is checked against verified information before becoming part of the final research.

This lets us benefit from GPT's detailed analysis while reducing the risk of incorrect data.

A Few Important Notes#

AI models improve quickly.

Both Claude and GPT have been updated since we ran this benchmark, so future results may look different.

Our prompts are also designed specifically for financial research. Different prompts or different tasks could produce different outcomes.

Finally, while 500 analyses provide a solid internal benchmark, they're not enough to make broad claims about every possible situation.

The Bottom Line#

There isn't a single "best" AI model.

Every model has strengths and weaknesses.

Today, our testing shows that Claude Sonnet provides more reliable financial research, fewer hallucinations, and confidence scores that were better calibrated to the reliability of their own findings.

That's why it's our primary research model today.

Tomorrow?

We'll test again.

As AI improves, Valarn will continue using the models that deliver the most accurate, transparent, and trustworthy research for our users—not simply the newest or most popular ones.

TagsAILLMGPT-4oClaudeBenchmarking
V

Valarn

AI Research

Valarn Research Team

Valarn

Try Valarn for free

Run AI-powered analysis on any stock in under 5 minutes.

Get started free