ai-machine-learning

GPT-6 Astra vs Claude Opus 5.5 vs Grok 4.7: Bill, GDPval, and a Six-Call Meter

Written by Mert Batur
Sep 22, 2026
8 read
GPT-6 Astra vs Claude Opus 5.5 vs Grok 4.7: Bill, GDPval, and a Six-Call Meter

GPT-6 Astra vs Claude Opus 5.5 vs Grok 4.7: Bill, GDPval, and a Six-Call Meter

GPT-6 Astra, Claude Opus 5.5, and Grok 4.7 were all callable on OpenRouter at 22:34 UTC on 22 Sep 2026. Six graded calls, two tasks, one trial each, all passed. The gateway bill was $0.01626.

The launch charts do not line up.

Anthropic's board has Astra and no Grok. xAI's board has Fable 5.1 and Sol, plus Astra on one Elo chart. This page keeps shared anchors and stops where the rulers split.

The Opus-versus-Opus-5 cut stays in the Opus 5.5 post. The 4.7-versus-4.6 upgrade stays in the Grok 4.7 post. Astra versus Fable stays in the GPT-6 Astra post.

Three cards, and the prompt size that reprices them

Astra's token is 2.5× Opus 5.5 on input and output, and 5× on cache reads. Grok's xAI card is half the input and 0.3× the output, until the prompt crosses 200,000 tokens.

Opus 5.5 is the 1.0 line: $4, $20, and $0.20 per million.

Astra's model card is $10, $50, and $1.00. Grok's 21 Sep card is $2, $6, and $0.50 under 200,000 prompt tokens.

Line, per 1M tokens, under the cliffOpus 5.5GPT-6 AstraGrok 4.7, xAI card
Input$4$10$2
Output$20$50$6
Cache read$0.20$1.00$0.50
Context1,000,0001,050,000500,000
Cliffnone on this card272,000 prompt tokens: input and cache 2×, output 1.5×200,000 prompt tokens: the whole request goes to $4 / $12 / $1

Price of Astra and Grok as a multiple of Opus 5.5
Input, output, and cache read, each divided by Claude Opus 5.5. Astra is 2.5, 2.5, and 5. Grok's xAI card is 0.5, 0.3, and 2.5.

OpenRouter at 22:34 UTC matched Opus and Astra. Grok did not.

The listing was $1.60, $4.80, and $0.40, then $3.20, $9.60, and $0.80 at 200,000 prompt tokens. That is 0.8× the xAI card. Our usage.cost figures use that catalog. The Grok 4.7 write-up still prices xAI's own card at $2 and $6.

A 500,000-token cache reread, arithmetic only: Opus $0.10, Astra $0.25, Grok's xAI cache line $0.25. Cache is where Astra is the expensive one, not output. Output is where Grok's $6 against Astra's $50 is an 8.3× gap, and against Opus's $20 a 3.3× gap.

Six graded calls, and the bill each model actually sent

All six passed. Grok's two calls cost $0.0038784, Opus 5.5 cost $0.005292, Astra cost $0.00709.

Invoice answer had to be 1.96. merge_intervals had to pass hidden tests, including touching ranges.

Cached tokens were 0 on Opus and Astra. Grok reported 1,152 cached tokens on both calls, inside prompt counts of 1,310 and 1,336. 295 of its 298 invoice tokens were reasoning.

python
# 22 Sep 2026, 22:34 UTC, OpenRouter, one trial, max_tokens 1800
MODELS = [
    "anthropic/claude-opus-5.5",
    "openai/gpt-6-astra",
    "x-ai/grok-4.7",
]
TaskOpus 5.5GPT-6 AstraGrok 4.7
Invoice, must be 1.96pass, $0.001336, 46 out, 4.277spass, $0.00263, 38 out, 2.116spass, $0.002144, 298 out, 5.292s
merge_intervalspass, $0.003956, 168 out, 3.307spass, $0.00446, 69 out, 2.588spass, $0.0017344, 204 out, 4.217s
Two-task total$0.005292$0.00709$0.0038784

Six-call OpenRouter bill for Opus 5.5, Astra, and Grok 4.7
Invoice task and merge_intervals. Astra is the tall bar on both. Grok is cheapest on the function and second on the invoice.

Astra wrote the fewest tokens and sent the largest bill.

38 and 69 completion tokens, against Opus at 46 and 168. The $10/$50 card ate the savings. Those two usage.cost figures match $10 and $50 exactly. Opus matches $4 and $20.

Grok lost the invoice and won the function. 204 tokens at $4.80 landed at $0.0017344.

Astra was fastest on both draws, 2.116s and 2.588s. One trial is not a speed ranking.

The Sep 4 Astra piece said openai/gpt-6-astra returned HTTP 400. That was true at 14:40 UTC that day. It is not true of this run. The id resolved, the call returned 1.96, and the tests passed.

GDPval is the row both labs actually share

Opus 5.5 leads that shared row at 1,846. Grok 4.7 xhigh is 1,695. Astra max is 1,542.

Fable 5.1 at 1,735 is the anchor, not a fourth pick.

Anthropic prints 1,735, Astra at 1,542, and Opus 5.5 at 1,846. xAI prints the same 1,735 and 1,542, and Grok 4.7 xhigh at 1,695. Two copied numbers are why the bars share an axis.

GDPval bars for Opus 5.5, Grok 4.7, and Astra
1846, 1695, and 1542. Drawn from Anthropic's Opus cell and xAI's Grok cell, joined by the shared Fable 1735 and Astra 1542 anchors.

CellScoreWho printed it
Claude Opus 5.51,846Anthropic, 22 Sep 2026, GDPval-AA v2.1
Grok 4.7 xhigh1,695xAI launch chart, 21 Sep 2026, GDPval
GPT-6 Astra max1,542Both charts
Fable 5.1 max, anchor only1,735Both charts

Opus sits 151 Elo over Grok and 304 over Astra. Grok sits 153 over Astra and 40 behind the Fable anchor. None of that is our measurement. Our six calls were a two-decimal invoice and a merge function. They do not score office work.

Artificial Analysis put Opus 5.5 at 58 on its Intelligence Index at max effort, and said cost per task stayed level with Opus 5 because output tokens rose about 1.6×. Their Grok 4.7 note put that model at 46 on the same index. Different ruler, same direction: Opus ahead of Grok, and not a reason to ignore the token bill.

Terminal-Bench and CursorBench are not one ruler

Do not average Anthropic's 66.4% with xAI's 38.0%. They are not the same Terminal-Bench cell.

  1. Anthropic, Opus 5.5 at xhigh, Astra at high as OpenAI reported it: Terminal-Bench 4.0 is 66.4% for Opus 5.5 and 57.9% for Astra. Fable 5.1 is 55.8% on that page. CursorBench 4.0 is 57.8% for Opus 5.5 and 51.8% for Fable. Astra's CursorBench cell is blank.
  2. xAI's scoreboard: Terminal-Bench 4.0 is 38.0% for Grok 4.7 xhigh and 57.9% for Fable 5.1 Max. CursorBench is 46.3% for Grok and 51.8% for Fable. Astra is not on that board. The Fable CursorBench cell, 51.8%, matches Anthropic. The Fable Terminal-Bench cell does not: 57.9% there, 55.8% on Anthropic's page.
  3. Artificial Analysis, a third setup: Opus 5.5 at 59.6% on Terminal-Bench 4.0, level with Astra at xhigh. Grok's move on that lab's agent run was 18% to 33%, not 38.0%.

CursorBench can be set beside itself, because 51.8% shows up twice. Opus 5.5 at 57.8% leads Fable by 6.0 points on Anthropic's higher-effort figure. Grok at 46.3% trails that Fable cell by 5.5 points. Astra has no cell. Quote the lab next to the percent or the ranking is a collage.

EEBench, Harvey, and HealthBench are xAI rows. Opus 5.5 and Astra are absent. They stay in the Grok post. AutomationBench and Terminal-Bench-Science are Anthropic and OpenAI rows. Grok is absent. They stay in the Opus and Astra posts.

What each existing post still owns

This page does not replace the three write-ups it links. It only ranks the three models against each other.

  1. The Opus 5.5 post keeps the $4/$20 cut, the $0.02719 meter, the five-hour limit, and the Opus 4.8 cyber fallback. Its cover stays Anthropic's chart. The Astra columns are read here.
  2. The Grok 4.7 post keeps 4.7 against 4.6: twelve calls, 12/12, and the $2/$6 cliff. The 1,695 bar is theirs. Opus 5.5 is not.
  3. The Astra post keeps Astra against Fable 5.1, including the 4 Sep HTTP 400 and the 99.9% versus 62.7% split. Callability changed on 22 Sep. The Fable decision did not.

Which of the three to pin

Pin Grok 4.7 when the output bill is the product and the job fits in 500,000 tokens. Pin Opus 5.5 when the shared GDPval row is the product. Pin Astra when you already measured a computer-use or math job the other two posts document, and you can pay $10/$50.

  1. Short uncached calls like these six: Grok was cheapest in total, $0.0038784 against $0.005292 and $0.00709, and it was the expensive one on the invoice because reasoning ran to 295 tokens. If your short calls look like merge_intervals, the function row is the one to copy. If they look like a one-line number, Opus spent 46 tokens and won that row.
  2. Cache-heavy agents under 200,000 prompt tokens: Opus at $0.20 per million cached tokens beats both. Astra at $1.00 is 5× that line. Grok's xAI cache line is $0.50, 2.5× Opus.
  3. Prompts that cross 200,000 tokens: Grok reprices the entire request. A 250,000-token prompt on the xAI card is $4 and $12, not $2 and $6. Astra's cliff starts later, at 272,000, and it does not reprice tokens under the line the way Grok's docs describe.
  4. Office work you are willing to trust to GDPval: Opus 1,846, then Grok 1,695, then Astra 1,542, with Fable's 1,735 sitting between Opus and Grok. That order is the labs' order. It is not a score from the six calls.
  5. Terminal work: leave it off this ranking. Pick the lab whose runner you actually ship, then use that lab's number. Mixing 66.4% and 38.0% will not tell you which binary to call.

Frequently Asked Questions

Can you call GPT-6 Astra on OpenRouter now?

Yes, as of 22:34 UTC on 22 Sep 2026. The id openai/gpt-6-astra returned 1.96 on the invoice and a merge_intervals that passed the hidden tests. On 4 Sep the same id returned HTTP 400. The earlier post records that 400. This one records the call that worked. Context on the listing was 1,050,000 tokens.

Did xAI cut Grok 4.7 from $2 and $6?

Not on the launch card this page cites. xAI's table is still $2 and $6 under 200,000 prompt tokens, then $4 and $12. OpenRouter's catalog during this run was $1.60 and $4.80, then $3.20 and $9.60, which is 0.8× those lines. The six-call bill uses the OpenRouter figure. A bill you pay xAI directly still follows xAI's card until they change it.

Why is Grok missing from Anthropic's benchmark chart?

Anthropic's 22 Sep chart compares Opus 5.5 with Fable 5.1, Opus 5, GPT-6 Astra, and GPT-5.6 Sol. Grok is not a column. xAI's chart compares Grok 4.7 with Grok 4.6, Sol, and Fable. Opus 5.5 is not a column. GDPval is the overlap, because both pages print Fable at 1,735 and Astra at 1,542.

Does the six-call meter say Opus 5.5 is smarter?

No. All three passed both tasks. The meter is a bill and a token count. GDPval is the lab score, and it does put Opus 5.5 first among these three, at 1,846 against 1,695 and 1,542. A passed merge function does not confirm 1,846, and 1,846 does not predict the $0.005292 invoice.

Which model has the shorter context window?

Grok 4.7, at 500,000 tokens on the OpenRouter listing. Opus 5.5 is 1,000,000. Astra is 1,050,000. Grok also reprices at 200,000 prompt tokens, so the useful cheap window is smaller than 500,000. Astra's higher rate starts at 272,000 and applies to the tokens over that line, plus a 1.5× output multiplier.

Tags

gpt-6-astraclaude-opus-5-5grok-4-7api-pricinggdpval

Share this article

Start Your Project

Ready to build something extraordinary?

Let's turn your vision into reality. Our team is ready to help you create software that makes a difference.