Attributed compile (not original research)
Primary source: Grok 4.7 is now available on Amazon Bedrock by Suheel Farooq, Anirban Gupta, Fabio Branco, Ikenna Izugbokwe, and William Yap (AWS Machine Learning Blog), published 2026-09-28
Capability background: Introducing Grok 4.7 (SpaceXAI / xAI), published 2026-09-21
Independent numbers cited below: Artificial Analysis, as quoted in the AWS launch post
This post restates the Bedrock packaging and first-request path in my own words, with a short decision checklist for agent builders. I did not run Grok 4.7 on Bedrock for this article. Every price, model ID, and benchmark figure is taken from a named source and labeled as vendor-reported or third-party.
Yesterday Amazon Bedrock added Grok 4.7. If you already ship agents on Bedrock, that is more useful than another “new frontier model” headline: you get xAI’s coding-and-knowledge-work flagship behind the same Responses, Chat Completions, and Converse surfaces you already use, with Geo and Global cross-Region profiles.
Below is a practical walkthrough of what actually changed for builders, how to send the first request without fighting IAM, and why leaving reasoning effort on the default can quietly double your token bill.
What you are buying (and what you are not)
xAI positions Grok 4.7 as its strongest model for coding and knowledge work, with the theme of endurance rather than raw speed: longer work on hard tasks, more self-checking before the next step, and a 500K token context window. That framing comes from xAI’s September 21 launch note and is repeated in the AWS post. Treat it as product positioning, not a guarantee on your workload.
Artificial Analysis numbers, as quoted by AWS for Grok 4.7 at xhigh effort versus Grok 4.6, are the useful second reference:
| Measure (Artificial Analysis, via AWS post) | Grok 4.7 | Grok 4.6 |
|---|---|---|
| Intelligence Index | 46 | 44 |
| Coding Agent Index | 56 | 47 |
| AA-Briefcase (Elo) | 1,657 | 1,546 |
| GDPval-AA (Elo) | 1,695 | 1,605 |
| Output tokens per Intelligence Index task | ~81k | ~38k |
The last row is the one that should change how you wire defaults. Gains arrived with roughly double the output tokens per composite task in that evaluation setup. A cheaper list price does not automatically mean a cheaper finished job.
Also note what Bedrock is not claiming for you: it is not promising that Grok 4.7 beats GPT-6 Astra or Claude Fable on every agentic coding harness. Independent Terminal-Bench and Intelligence Index tables elsewhere put Astra ahead on several agentic coding scores. Bedrock’s job here is packaging and access, not settling the model war.
How Bedrock packages the model
Requests name an inference profile, not a bare foundation-model string:
| Option | Model ID | Intent |
|---|---|---|
| US Geo | us.xai.grok-4.7 |
Keep processing in the US geography |
| Global | global.xai.grok-4.7 |
Cheaper / more capacity; less control over where a call lands |
Base URL for OpenAI-compatible calls (region substituted):
https://bedrock-runtime.{region}.amazonaws.com/openai/v1
Supported surfaces, per AWS: Responses, Chat Completions, InvokeModel, and Converse. Input is text and image; output is text. Reasoning effort levels are low, medium, high, and xhigh. AWS says the default is high. Set it on purpose.
Bedrock features that matter for agents:
- Implicit prompt caching on repeated prefixes (system prompts, long reference docs)
- Guardrails by ID/version on the request
- Structured outputs constrained to a JSON Schema
- Invocation logging in CloudWatch, including reasoning token counts
Service tier is another cost dial: default (standard), priority, and flex. Check the current Bedrock pricing page for per-tier rates rather than hard-coding dollars from memory.
First request three ways
Confirm the model is enabled in the Bedrock console for the Region you will call, then install clients:
pip install openai boto3
1. OpenAI SDK + Chat Completions
Exploration path with a long-term Bedrock API key (delete it when you are done exploring):
export OPENAI_API_KEY="<Bedrock API key>"
export OPENAI_BASE_URL="https://bedrock-runtime.us-east-1.amazonaws.com/openai/v1"
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="us.xai.grok-4.7",
messages=[
{"role": "user", "content": "Summarize what Amazon Bedrock inference profiles do in three bullets."}
],
)
print(response.choices[0].message.content)
2. Responses API (better fit for tools and reasoning knobs)
response = client.responses.create(
model="us.xai.grok-4.7",
reasoning={"effort": "medium"},
input="List three failure modes for a long-running coding agent.",
)
print(response.output_text)
To pass encrypted reasoning content across turns (Responses only), AWS shows:
response = client.responses.create(
model="us.xai.grok-4.7",
reasoning={"effort": "high"},
include=["reasoning.encrypted_content"],
input="Explain quantum entanglement simply.",
)
Chat Completions does not return reasoning tokens the same way.
3. Converse with boto3 (SigV4, no bearer token)
Reasoning is always active. The first content block may be reasoning, so search for the text block instead of assuming content[0]:
import boto3
client = boto3.client("bedrock-runtime", region_name="us-east-1")
response = client.converse(
modelId="us.xai.grok-4.7",
messages=[
{"role": "user", "content": [{"text": "What is 17*23? Number only."}]}
],
inferenceConfig={"maxTokens": 3000},
additionalModelRequestFields={"reasoning_effort": "xhigh"},
)
blocks = response["output"]["message"]["content"]
text = next(b["text"] for b in blocks if "text" in b)
print(text)
IAM gotchas that burn the first afternoon
AWS is explicit about three resources on bedrock:InvokeModel:
- Your account’s default project ARN
- The inference profile you name (
us.xai.grok-4.7orglobal.xai.grok-4.7) - The underlying foundation model ARN, wildcarded across Regions because cross-Region profiles route outside the calling Region
A policy that lists only the US profile does not cover the Global profile. Bearer-token OpenAI-compatible calls also need bedrock:CallWithBearerToken; Converse with SigV4 does not.
For production, prefer short-lived bearer tokens from IAM (aws-bedrock-token-generator) over a standing long-term Bedrock API key.
A short decision checklist for agent builders
Use this instead of “just turn on xhigh everywhere”:
-
Pick the profile first. Residency or US latency needs →
us.xai.grok-4.7. Cost/throughput first →global.xai.grok-4.7. -
Set effort per route. Short extraction/classification →
lowormedium. Multi-step plans and long trajectories where an early error compounds →high/xhigh. Do not inherithighon high-volume chat. - Measure tokens on your tasks. Artificial Analysis’s ~81k vs ~38k output-token gap is a warning light, not your bill. Log input, cache hits, reasoning, and answer tokens from Bedrock invocation logs.
- Reuse prefixes. Agents that resend the same system prompt every turn should benefit from implicit prompt caching; confirm it in logs before celebrating list-price math.
- Attach Guardrails if the agent runs unattended. Especially for multi-step writes and user-facing replies.
- Do not mix harness rankings. A Coding Agent Index score from one vendor harness is not interchangeable with Terminal-Bench from another. Compare apples to apples, or run a 20-task golden set of your own.
Why this post exists
Model launches move weekly. The durable skill is not memorizing Elo tables. It is knowing which knobs (profile, effort, cache, service tier, auth path) actually change reliability and cost when you drop a new model into an existing Bedrock agent. Grok 4.7 on Bedrock is a packaging change you can act on today; the scoreboard will keep shifting.
If you try it, start with one production-like agent trajectory at medium and high, compare token counts and failure modes, then promote the winner. That experiment beats another screenshot of a leaderboard.
Byline: YongBo Yu — Toronto AI engineer (agents, LLM workflows). GitHub: YongBoYu1.
Top comments (0)