MODEL COMPARISON · 2026
Haiku 5.5 vs Luna 6: Scores, Pricing & Real Demos
Haiku 5.5 vs Luna 6 is a close comparison at the base API price, but a less straightforward choice once you include prompt length, task reliability and response time. Here is what current documentation, an independent benchmark and public same-task demos actually show.
- Scores: Haiku 5.5 leads the AA Intelligence Index, 43 vs 38.
- Cost: Both start at $0.10 input / $0.50 output per million tokens. Luna costs less for long prompts in the examples below.
- Speed: Haiku leads in output throughput; Luna finished sooner in the cited 40-task test.
- Demos: Compare the interchange and fairground outputs side by side below.
Haiku 5.5 vs Luna 6: scores and specifications
| Metric | Claude Haiku 5.5 | GPT-6 Luna |
|---|---|---|
| AA Intelligence Index | 43 · max | 38 · max |
| AA output throughput | ≈241 tokens/s | ≈127 tokens/s |
| Context window | 1 million tokens | 1.05 million tokens |
| Maximum standard output | 128K tokens | 128K tokens |
| Base input / output, per 1M tokens | $0.10 / $0.50 | $0.10 / $0.50 |
| Long-prompt threshold | Over 100K input tokens | Over 272K input tokens |
| Long-prompt input / output, per 1M | $0.50 / $2.50 | $0.20 / $0.75 |
Artificial Analysis lists Haiku at 43 and Luna at 38 on its Intelligence Index, with both pages identifying the max configuration. Its output-throughput snapshot also favors Haiku. These figures describe that benchmark's configuration and workload, rather than a percentage of correct answers or a guarantee for your application. Sources: Haiku benchmark and Luna benchmark.
Five index points do not mean “five percent better.” A blended index compresses many tasks into one number. If your product extracts invoices, generates SQL or writes Chinese copy, the relevant question is whether the model meets those particular requirements. A model can lead the aggregate and still lose on your most frequent task. Keep the score as a starting signal, then check your own acceptance criteria.
The context and output limits come from Anthropic and OpenAI. A large context window is a capacity limit, not evidence that every fact buried in a million-token prompt will be retrieved correctly. When you evaluate document work, place important facts at different positions and ask questions whose answers you can verify against the original text.
Same-task demo screenshots: what the outputs look like
BitsMinds published a three-brief build-off using the same briefs in Claude Code and Codex. These four images are screenshots of its public rendered outputs, captured on October 9; they are not ModelCompare reruns. The interchange pair uses Haiku xhigh and Luna max; the fairground pair uses max for both. Different tools and effort settings limit causal conclusions.
Interchange: road layers, ramps and cars
Fairground: wheel, carousel and reflections
Our visual reading: Luna's interchange frame has more road detail and a clearly visible multi-lane overpass; Haiku's is simpler. Both fairground frames contain a wheel, carousel and reflected scenery, with different compositions. These are observations of the captured frames, not an automated correctness score. Click an image to inspect it at full size.
A still image cannot prove that cars follow correct paths, that motion remains smooth or that the output satisfies every part of the brief. Open the original build-off to inspect its animated demos. Separate “looks convincing” from “implements the requested behavior”: the latter needs runtime checks. The build-off also reports a rocket task without a usable Haiku output, so there is no honest two-image comparison to show for that task.
For your own frontend trial, ask both models for exactly the same component, preserve their first outputs and check them at desktop and mobile widths. Give them the same asset access and revision budget. A polished screenshot after several repair rounds answers a different question from a first-pass result. Record both if you care about how much work it takes to ship.
Does the higher score translate into better answers?
A separate Kingy AI 40-task test repeated each task twice per model. Haiku passed 60 of 80 outputs under its strict rubric, versus Luna's 51 of 80. Reported median completion times were 6.83 seconds and 2.92 seconds respectively. That supports a possible reliability-versus-latency tradeoff in that test, not a universal winner.
The test used particular clients and defaults; the two runs for each prompt are repeated observations, not 80 unrelated tasks. Its strict checks can reject an answer that is useful but violates an exact wording or ordering rule. That may be appropriate for a production parser and less representative of open-ended writing. Read the rubric before treating its pass rate as the quality of all responses.
The apparent speed conflict is useful: Haiku's higher benchmark output throughput does not contradict Luna's shorter completion time in another setup. Total latency includes queueing, prompt processing, reasoning, time to first token and output length. Throughput measures only part of that experience. For a chat product, measure time to first useful text as well as time to finish; for a background worker, throughput and total successful jobs may matter more.
API pricing: equal base rates, different long-context costs
The standard uncached text rates start at $0.10 per million input tokens and $0.50 per million output tokens for both models. The distinction appears when a single prompt exceeds the long-context threshold: 100K for Haiku, 272K for Luna. The table below applies the corresponding tier to the request and keeps output fixed so the arithmetic is comparable. Sources: Haiku documentation and OpenAI pricing.
| Input + output tokens | Haiku 5.5 | Luna 6 |
|---|---|---|
| 10,000 + 2,000 | $0.002 | $0.002 |
| 150,000 + 2,000 | $0.080 | $0.016 |
| 300,000 + 2,000 | $0.155 | $0.0615 |
Formula: input tokens × input rate / 1,000,000 + output tokens × output rate / 1,000,000. These are calculated examples, not measured bills or claims that the same document tokenizes identically with both providers. They exclude caching, batch discounts, tools, taxes and any additional billable reasoning tokens. Check the usage returned by your actual requests before forecasting a monthly budget.
The 150K example is especially relevant for document-heavy workflows: it crosses Haiku's threshold while remaining below Luna's. At 300K, both requests enter their higher tier. A million short requests do not become one long prompt simply because their monthly token total is large. Use per-request prompt length to choose the rate, then sum the requests.
For purchasing decisions, compare cost per accepted result. If a cheap first answer requires a second call, a repair step and manual checking, its nominal token price can be misleading. Log retries and rejected outputs along with billed usage. When results are acceptable from either model, prompt length may be the more decisive factor than the base rate.
Which model should you choose?
- Short, structured work: test both at the base rate. Score schema validity, factual accuracy and instruction compliance separately.
- Long documents: estimate the distribution of prompt lengths. Luna's higher tier threshold can materially change the bill.
- Frontend coding: use the screenshots to choose a trial task, then inspect the generated code, motion and responsive behavior.
- Interactive assistants: measure first useful response and full completion time at realistic concurrency.
- Editorial work: blind-review drafts for usefulness, sourcing, tone and the amount of human correction required.
A small trial can be more useful than collecting another dozen rankings. Take twenty representative tasks from your actual queue, remove private data, and write the success criteria before running either model. Include easy routine tasks, a few difficult cases and examples that have caused production errors. Keep the output budget, tool access and retry policy consistent.
Save model identifiers, dates, settings, token usage and outputs. Review answers without revealing which model produced them where possible. Compare accepted-result cost and latency alongside the pass rate. If the difference is small, choose on operational fit: integration effort, quota availability and the failure modes your team can handle. Revisit the decision when the model, pricing or workload changes.
Haiku 5.5 vs Luna 6 FAQ
Is Luna 6 the same model as GPT-6 Luna?
In this article, “Luna 6” is the search shorthand for OpenAI's GPT-6 Luna. The comparison uses Claude Haiku 5.5 and GPT-6 Luna, rather than older Haiku versions or other GPT-6 models. Use the exact API model identifier from provider documentation when testing.
Which has the higher benchmark score?
Haiku 5.5 leads the current Artificial Analysis Intelligence Index snapshot, 43 versus 38, with max configurations. That is an aggregate benchmark result, not a guarantee that Haiku wins every task.
Which is cheaper?
They share the listed base text rates. Luna is cheaper in the fixed-token long-prompt examples above because its threshold and higher-tier rates differ. Your actual bill also depends on tokenization, output length, reasoning, retries and discounts.
Which is faster?
It depends on the measurement. Haiku has higher listed output throughput in the AA snapshot; Luna completed faster in the cited 40-task test. Measure the complete request in your own client and chosen effort setting.
Which is better for coding and visual demos?
The public build-off shows meaningful differences, but two visual tasks do not establish general coding superiority. Evaluate compilation, required behavior, accessibility and maintenance as well as appearance. Use your repository's tests when you have them.
Can these screenshots prove one model always wins?
No. They document particular public outputs with stated settings. They are helpful examples, but they do not replace repeated trials, full prompts or runtime validation.



