Is the newest AI model always the best one? We put four to the test

Is the newest AI model always the best one? We put four to the test

Every few months, a new AI model arrives claiming to be the smartest yet. Better benchmarks. Better reasoning. Better coding. Better everything.

It’s easy to assume that whatever launched most recently is automatically the sharpest tool for the job. But instead of relying on benchmark reports and marketing claims, we decided to put them to the test. We wanted to find out which models are actually worth using.

We selected four of today’s most talked-about AI models: 

  • Anthropic’s Claude Fable 5
  • Anthropic’s Claude Opus 4.8 
  • OpenAI’s GPT-5.6 Sol 
  • Anthropic’s Claude Sonnet 4.6 

Each received the exact same three prompts with specific goals:

  • Building a landing page
  • Writing a website’s home page copy
  • Analyzing sales data

We described the desired outcome in the prompts but did not provide layout instructions, step-by-step checklists, or any follow-up context. Instead, we let each model decide how to get there.

This approach let us see how well each model handled ambiguity on the first try. Would it ask clarifying questions or confidently make its own assumptions? And just as importantly, how much time and money would it take to reach a good result?

Here’s what we discovered.

Which model offered the most value for creating a landing page?

We first challenged each model to design and build a landing page from a single creative brief. This is the prompt we used:

Design and build a premium landing page for a pottery studio called "Primavera". The page should evoke the feeling of a serene artist studio: calm, sophisticated, and carefully paced. The landing page should have 3 pages, including a contact/booking form. I want the primary experience to be one continuous scroll, with smooth transitions between pages. Include a bold, animated hero element that acts as the visual centerpiece of the entire experience. It should be interactive and invite exploration through user input. It should evoke the tactile feeling of being in a pottery studio. It is NOT a static image or decorative effect.

We intentionally left plenty of room for interpretation to test the models. Rather than prescribing layouts or components, the prompt described just the feeling, interactions, and overall experience we wanted.

We then deployed each AI model’s output to Hostinger from Claude Code, Cursor, and Codex desktop apps via Hostinger Connector. The same deployment process was used for all models to maintain consistency in the comparison.

Sonnet 4.6’s output

The Sonnet-built website met all the requirements in the prompt, but its hero animation was the simplest of the four. It was also the only model to move away from the clay-inspired color palette on the hero section.

Even with these shortcomings, Sonnet produced a polished, well-structured website that fulfilled the brief. Despite being one of the older models in this benchmark, it proved it can still hold its own against much newer models.

Sonnet 4.6 generated website: https://teal-camel-563249.hostingersite.com/ 

GPT-5.6 Sol’s output

Sol took the most literal interpretation of “tactile.” Its hero section lets visitors drag and shape a piece of clay with a reset button, making it the clearest execution of the prompt’s request for an interactive centerpiece.

The rest of the site leaned into short, punchy microcopy and a tidy three-tier workshop list. It was polished and on-brand, but the website content was pretty comparable with the Sonnet 4.6’s produced outcome.

GPT-5.6 Sol generated website: https://primavera-ceramics-studio.hostingersite.com/ 

Opus 4.8’s output

Opus produced the most detailed landing page. It expanded the fictional backstory of the pottery studio, introduced workshop statistics, and created five tiers of workshops instead of three.

Its interactive clay-shaping hero was just as polished as Sol’s, and in terms of sheer completeness, Opus arguably responded to the brief most thoroughly.

The trade-off was that the page also felt busier, which didn’t necessarily make the experience stronger – just denser.

Opus 4.8 generated website: https://sandybrown-goshawk-747835.hostingersite.com/ 

Fable 5’s output

Fable took the most technically ambitious approach to a richer, graphics-driven experience. 

The interactive pottery animation was the highlight. Clicking and dragging shapes the clay in real time, responding smoothly to direction and pressure and closely mimicking the feel of hand-building pottery.

Sol attempted something similar, but the result didn’t look like clay being molded the way Fable’s did – it felt less like shaping a three-dimensional material and more like dragging a flat image around.

While this interaction was memorable, the overall copy and messaging weren’t as strong as Sol’s or Opus’s.

Fable 5 generated website: https://gray-albatross-412637.hostingersite.com/ 

An alternative model for graphically rich pages

In a similar model comparison I did for fun, I compared Fable 5, GPT-5.6 Sol, and Qwen 3.8 Max for creating landing pages. I simply wanted to see which one produced the best website.

I noticed that Qwen’s output consistently added more animation and interactivity than Fable 5, not just matching it. This Alibaba’s AI model even generated a favicon automatically, and overall, the design looked the most polished and elegant to my non-designer eyes.

So, if dynamic, high-level visual and animation pages are your priority, Qwen 3.8 Max is worth trying as an alternative to the most popular frontier models.

Editor

Tomas Rasymas

AI Research Lead at Hostinger

The winner

Visual and copywriting preferences are inherently subjective, so we understand that the best result will vary from person to person. 

To make the comparison more balanced, we also evaluated each model based on task completion time and cost, calculated as a dollar amount based on each model’s monthly usage limits.

Here’s how the four models compared:

AI modelTimeCostToken usage in total (input, output, cached)
Sonnet 4.6 ~14 min$2.776M
GPT-5.6 Sol~30 min$8.7411M
Opus 4.8~12 min$4.686.6M
Fable 5~25 min$9.893.7M

*Cost calculated as a dollar amount from monthly usage limits.

On raw output, this one’s a genuine toss-up – any of the newest model’s results could reasonably be called the most impressive build, depending on what design value you like the most. 

But “most impressive” and “best value” aren’t the same question, and the token and cost data make the latter easier to answer.

A quick note on the table: cost depends on both the number of tokens used and the price per token. Since models are priced differently, one can use fewer tokens yet still end up costing more if it belongs to a pricier tier.

That’s what shows up here. Sonnet and Opus used a comparable number of tokens (6M vs. 6.6M), yet Opus cost 69% more. This shows that the gap comes from the pricing tier, not from Opus being any less efficient with its tokens. 

Fable makes the same point from the other direction. It used the fewest tokens of any model in the test (3.7M) but was the most expensive of the four, because its per-token price sits well above the rest of the field. 

Sol, on the other hand, got to a higher bill the more straightforward way by using nearly twice as many tokens as any other model (11M).

Put together, Sonnet wasn’t being the most frugal with its tokens – Opus was just as frugal. But Sonnet delivered a complete, on-brief result while sitting in the lowest-cost tier of the four, which is why it came out cheapest by a wide margin, regardless of how you slice the data.

That makes Sonnet 4.6 our pick for this benchmark. It didn’t build the flashiest site, but it delivered a complete, on-brief result at a fraction of the cost of the alternatives. 

If budget and turnaround time are part of your decision, that’s the more useful signal than any single design choice.

Which AI model executed the creative writing brief best?

Our second benchmark focused on text-based work that many people fold into their day-to-day, more lightweight tasks. 

Do note that we don’t expect any AI model to match the output of experienced copywriters, but we do expect it to produce a draft worth further refining.

To see which model came closest, we asked each one to write the hero section for a premium home organizing service with this prompt:

You are writing a homepage hero copy for a home organizing service that charges more than its competitors. Write a headline, a subheadline, and a two-sentence CTA that earns the price difference without mentioning price, competitors, or using the words "transform." The copy should make someone feel understood before it makes them feel sold to.

Sonnet 4.6’s output

Headline: You know where everything goes. You just need space to get there.

Subheadline: We don’t just clear clutter — we figure out why it came back last time. Then we build a home that actually works for how you live.

CTA: Every session starts with a conversation, not a bin bag. Book a free walkthrough and we’ll map out exactly what your home needs — before anything moves.

Sonnet’s headline reframed the reader as capable, not chaotic, which is a smart move for a premium-priced service. The copy was polished and ready to use. 

The trade-off: Sonnet appended a full paragraph explaining its own choices, unprompted. This can be useful if you want the rationale, but it’s extra scrolling for a task framed as quick and everyday.

GPT-5.6 Sol’s output

Headline: Your home isn’t messy. It’s carrying too much.

Subheadline: We create thoughtful, lasting systems around the way you actually live—so your space feels calmer, functions better, and stays that way.

CTA: Let’s make room for the life you’re trying to live. Book a consultation to start with what’s weighing on you most.

Sol’s headline was the tightest and most quotable of the four. It did the same reframe as Sonnet’s in fewer words. It’s also the only model that returned just the copy, with zero commentary. For a lightweight, paste-and-use task, Sol has the most efficient output of the bunch.

Opus 4.8’s output

Headline: The clutter isn’t a character flaw. It’s just been waiting for the right system.

Subheadline: You’ve tried the bins, the weekend purges, the promises to yourself. What you haven’t had is someone who studies how you actually live—then builds something around it that still works in six months, not just the day we leave.

CTA: Tell us about the room you keep closing the door on. We’ll come see it with fresh eyes and no judgment, and show you exactly what’s possible before you commit to anything.

Opus’s copy was empathetic and well-crafted, but it ran longer than the others. The subheadline in particular did more explaining than a hero section typically needs. It read well, but a copywriter would need more refinement effort before shipping it. Like Sonnet, it came with reasoning notes attached.

Fable 5’s output

Headline: A place for everything you love. Nothing else.

Subheadline: We build systems around the way you actually live — down to the drawer you open half-asleep — so everything stays exactly where your hands expect it, long after we’ve gone.

CTA: The rooms you walk back into tonight are already shaping tomorrow. Schedule your consultation and decide what they become instead.

Fable’s headline was the most compressed and stylistically ambitious, although “Nothing else” carried a slight risk of ambiguity that a human editor would probably flag as out of context. The subheadline’s “drawer you open half-asleep” was a genuinely strong, specific detail. Like the others, except for Sol, it added explanatory notes beyond the copy itself.

The winner

All models produced headlines strong enough to survive a light edit. But for this benchmark, we’d call it for GPT-5.6 Sol, because it was the only model that understood the assignment as a lightweight, everyday task and answered accordingly. 

It handed back exactly three things: a headline, a subheadline, and a CTA. Nothing to strip out, nothing to scroll past.

That said, this is close: Sonnet’s headline reframing was arguably the sharpest of the four, and if you’re the kind of user who wants the “why” alongside the draft, Sonnet or Opus would be the better fit.

Which AI model reasoned best for data analysis?

Another common business use case for text-based AI is data analysis, so we tested the four models’ reasoning. 

We prepared a CSV file containing dummy sales data for a small business selling puzzle books and merchandise (Puzzled? Hat, Custom Puzzle Notebook, and Puzzle Grid Tote Bag). There’s also a surprise in the third month’s sales numbers, generating more revenue than usual.

We asked each model to analyze it and write actionable business recommendations by sending the file together with this prompt:

I'm attaching six months of sales data for a small ecommerce business. Give me: (1) a one-paragraph plain-English summary of what the data shows, (2) one underperforming area with a specific reason why — not a general observation, a diagnosis, and (3) one concrete action to take based on the numbers. Not a general recommendation — a specific one I can act on next week.

Sonnet 4.6’s output

Sonnet caught the March wholesale-order skew and was the only model to support its analysis with charts, plotting gross margin by product so the hat’s underwater margin is impossible to miss. Its diagnosis was clear, and its recommended action was concrete – pausing promotions until the base price is fixed. Clean and correct, but it stops at the repair.

GPT-5.6 Sol’s output

GPT-5.6 Sol flagged the March anomaly and worked from the fullest revenue picture, factoring shipping income into its summary. Its diagnosis added a useful layer: shipping charges didn’t cover shipping costs, deepening the hat’s per-unit loss beyond the pricing problem alone. 

The action was the most cautious of the four – a modest reprice calculated to sit just above break-even, plus pulling the hat from both promo codes. Sound, but purely defensive, with no upside quantified.

Opus 4.8’s output

Opus delivered the sharpest single insight of the test: every hat order was a hat-only purchase, so the product can’t even be defended as a loss leader that pulls in book sales. 

Its reprice was the boldest, deliberately set to bring the hat in line with the margins of the rest of the merchandise, with the monthly impact modeled. 

It also added a caveat that no other model raised: verify the production cost is current before repricing, because if it’s stale, the real fix is sourcing rather than price.

The continuation of Opus 4.8's data analysis reasoning output

Fable 5’s output

Fable 5 reached the same diagnosis but traced the damage furthest: several hat sales had gone out at a discount via the store’s email and social promo codes, meaning marketing spend was actively amplifying the loss. Its action paired the fix with a lever – reprice the hat (or pull it entirely), and in the same edit, swap it out of both promos for the tote bag, the store’s highest-margin merchandise item, turning every promoted sale from a loss into a gain.

The winner

We think Fable 5 delivered the strongest answer, though this round was closer than the diagnosis alone suggests. All four models caught the March wholesale distortion and correctly identified the mispriced hat. The separation came entirely from what they did with it.

Fable 5 was the only model that treated the marketing angle as part of the problem. It redirected the promo slots toward the highest-margin product and quantified the swing per sale. Opus 4.8 was a very close second, with the deepest repricing math and the smartest caveat of the test.

Sonnet 4.6, meanwhile, held its own with pattern-catching. It spotted the wholesale skew and was the only model to visualize its findings. However, its recommended action was simply the plainest of the four: a correct price fix, with nothing layered on top.

Our take on the experiment results

So, does the latest model always produce better results? Based on these benchmarks, not necessarily. 

Output quality wasn’t determined by which model was newest, but by how well each model’s strengths matched the task.

A great idea can succeed with almost any modern AI model. But even the most capable model can’t rescue a weak idea or an unclear prompt.

No matter which model you choose, these three habits consistently lead to better results:

  • Write goals, not checklists. Describe the outcome you want, not every step to get there. This gives the AI models room to reason and find higher-leverage solutions. Reserve rigid instructions for workflows that must produce the exact same result every time.
  • Define your boundaries. If you don’t specify what’s in or out of scope, the model will make those decisions for you. Be explicit about constraints, required formats, and what success looks like.
  • Manage your context. As conversations grow longer, models become more likely to lose track of earlier information. Start fresh when a task changes, and avoid cramming unrelated work into a single chat.

After all, AI is becoming less about finding the best model and more about learning how to work well with whichever one you have.

Author
The author

Larassatti D.

Larassatti Dharma is a content writer with 4+ years of experience in the web hosting industry. She has populated the internet with over 100 YouTube scripts and articles around web hosting, digital marketing, and email marketing. When she's not writing, Laras enjoys solo traveling around the globe or trying new recipes in her kitchen. Follow her on LinkedIn

Author
The Co-author

Marina Moreira

Marina is a scriptwriter for Hostinger Academy, where she writes about web hosting, AI tools, and digital marketing. Her background in theatre and storytelling helps her turn technical topics into accessible, engaging videos. When she's not writing, Marina enjoys knitting, hiking, and watching movies (not at the same time). Follow her on Linkedin.