GLM 5.3 Is on Amazon Bedrock: Zhipu's 753B Coding Model Without the Self-Hosted Stack
AWS made GLM 5.3 generally available on Amazon Bedrock on October 5, 2026 — managed APIs, cross-Region inference, prompt caching, and $1.68 per 1M input tokens for eligible enterprise customers. What shipped, what it costs, and how to call it.
Zhipu’s GLM 5.3, the post-training-tuned coding model we covered at launch, is now generally available on Amazon Bedrock. AWS posted the general-availability notice and a hands-on AWS Machine Learning blog walkthrough on October 5, 2026. The model itself is not the news; the deployment option is. Meeting agentic coding demands with open-weight models has historically meant provisioning and operating your own inference infrastructure, as the launch post puts it. Bedrock adds a managed alternative that runs under your own AWS account’s controls — gated to eligible enterprise customers, which is the first thing to check.
What actually shipped
GLM 5.3 is a mixture-of-experts model with 753B total parameters and roughly 40B active per token, built on the same base model as GLM 5.2 with the gains attributed to scaled post-training. On Bedrock it exposes a one-million-token context window, up to 128K output tokens, and reasoning that is always enabled with selectable effort levels — a per-request dial that trades latency and token consumption against task performance.
Access runs through cross-Region inference profiles — us.zai.glm-5.3 and global.zai.glm-5.3 — so you send requests to a source Region and Bedrock routes them for processing. Availability is limited to what AWS calls eligible enterprise customers, so confirm model access in the Bedrock console before planning a cutover.

What it costs
The Bedrock pricing page lists GLM 5.3 on the Standard tier, per one million tokens:
| Inference profile | Input | Output | Cache read | Cache write (30 min) |
|---|---|---|---|---|
| Global CRIS | $1.68 | $5.28 | $0.312 | $2.10 |
| US CRIS | $1.848 | $5.808 | $0.3432 | $2.31 |
For comparison, GLM 5 on Bedrock runs $1.00 in and $3.20 out per 1M tokens in US Regions. Two tier multipliers matter for agentic work: Flex sits 50% below Standard for latency-tolerant overnight runs, and Priority adds 75% for latency-critical lanes. The GLM 5.3 rates sit in the page’s Z AI section; the capture below shows the page’s model-provider selector, where Z AI is one of the tabs.

Calling it in five minutes
The prerequisites are an AWS account with Bedrock access, IAM permissions, and Python 3.10 or later for the code examples. Install the OpenAI SDK and the AWS token helper, then point the client at the Bedrock runtime endpoint:
pip install -U openai aws-bedrock-token-generator
from aws_bedrock_token_generator import provide_token
from openai import OpenAI
region = "us-west-2" # your source AWS Region
client = OpenAI(
api_key=provide_token(region=region),
base_url=f"https://bedrock-runtime.{region}.amazonaws.com/openai/v1",
)
resp = client.responses.create(
model="global.zai.glm-5.3",
input="Refactor this Python function to be iterative instead of recursive: ...",
)
print(resp.output_text)
AWS recommends the OpenAI-compatible Responses and Chat Completions endpoints for new applications — they carry the most complete feature set — with the native Invoke and Converse APIs alongside. Your IAM principal needs bedrock:InvokeModel, bedrock:InvokeModelWithResponseStream, and bedrock:CallWithBearerToken, and AWS strongly prefers short-lived credentials over long-lived API keys. No-code evaluation works too: the Bedrock console playground lists GLM 5.3.
Cache the prefix or pay for it twice
Agentic loops resend stable context every turn — system prompt, tool definitions, repository files. GLM 5.3 caches implicitly by default, but explicit mode lets you mark the reusable prefixes and raise hit rates. Extending the client above, with SYSTEM_PROMPT and USER_INPUT standing in for your own strings:
resp = client.responses.create(
model="global.zai.glm-5.3",
# Enable explicit caching mode:
extra_body={"prompt_cache_options": {"mode": "explicit"}},
input=[
{"type": "message", "role": "system", "content": [
{"type": "input_text", "text": SYSTEM_PROMPT,
# A static system prompt is a good caching target:
"prompt_cache_breakpoint": {"mode": "explicit"}}]},
{"type": "message", "role": "user", "content": [
{"type": "input_text", "text": USER_INPUT,
# Layered breakpoints are allowed too:
"prompt_cache_breakpoint": {"mode": "explicit"}}]},
],
)
if resp.usage.input_tokens_details.cached_tokens:
print("cache hit")
Two rules decide whether this works: each breakpoint must cover at least 1,024 tokens to be eligible, and breakpoints can be layered for staged cache. The economics are blunt — on Global CRIS a cached input token reads at $0.312 per 1M versus $1.68 fresh, under a fifth of the price, before counting the latency win. On multi-hour runs that resend a large repository context each turn, caching is the line item that decides whether the lane is cheap or expensive. It is the same reuse discipline our context engineering playbook argues for.
The demo workload: authorized security testing
AWS’s flagship walkthrough pairs the model with Strix, an open-source AI penetration-testing agent that maps a target’s threat surface, fans sub-agents out across vulnerability categories, and validates each finding with a working proof of concept — the step that keeps triage time down. Z.ai reported a CyberGym score of 84.5 at release, and Strix documentation currently defaults to GLM 5.3.
The post’s guardrails deserve repeating: only test applications you own or have explicit written permission to test. The demo target is a deliberately vulnerable local app — OWASP Juice Shop in a Docker container — which is the right blast radius for a first run.
One rough edge: Strix reaches Bedrock through LiteLLM, which does not yet resolve bedrock/global.zai.glm-5.3. The documented workaround pins the Converse route and the inference-profile ARN. In the shell form below, AWS_REGION and AWS_ACCOUNT_ID must be set to your real values — leave them empty and expansion silently produces an invalid ARN:
export STRIX_LLM="bedrock/converse/arn:aws:bedrock:${AWS_REGION}:${AWS_ACCOUNT_ID}:inference-profile/global.zai.glm-5.3"
Why this matters if you run a fleet
A new managed lane for a frontier-class coding model is, above all, a routing and continuity event. It adds a row to the failover matrix in Managed Agents Comparison: Bedrock, Claude, OpenAI and gives the provider-cutoff drill one more answer. If you self-host GLM weights for agents — the path mapped in open-weight models that can drive a harness — price your GPU hours per completed task, as in free weights, paid tasks, and compare that against the metered Bedrock bill before deciding which lane keeps the job.
Caveats worth keeping
- Access is enterprise-gated; “eligible enterprise customers” is doing quiet work in the announcement.
- Benchmark claims are vendor-reported. The 50% gain over GLM 5.2 is Z.ai’s internal coding benchmark, and AWS notes no direct comparison to GLM 5 was published because the benchmark suite itself was updated.
- Always-on reasoning plus a million-token window is a token-burn combination. Set effort deliberately and cache the prefix.
- A strong cyber model on a managed endpoint renews the allowlist question for which agents may use it — see cyber capability gates becoming SKUs.
The short version: a frontier open-weight coding model gains one more deployment path — a region-routed API under your AWS account, alongside self-hosting and third-party inference — and the caching math is what makes long agentic runs plausible on it.