Short answer
Stay on Amazon Bedrock if your data already lives in AWS and you want several model families behind one API: AWS states that "With Amazon Bedrock, your content is not used to improve the base models and is not shared with any model providers".[2] Look at alternatives when Bedrock's model clock starts to cost you — once the Legacy period begins, "new customers can't adopt the model, and existing customers may lose access after 15 days of inactivity", and "Model lifecycle dates are specific to Amazon Bedrock and may differ from dates published by model providers (such as Anthropic or Cohere)".[3] The four alternatives are Azure OpenAI in Microsoft Foundry, Google Vertex AI, a model provider's direct API, and an open-weight model you run yourself.
Amazon Bedrock is AWS's managed service for foundation models from several developers behind one API, with a layer of product features around them. What follows is what AWS publishes on its own pages: the prices, the data terms, and the lifecycle rules that decide how long a model stays callable.
AWS lists the features around the models: "Amazon Bedrock Playground, Agents, Knowledge Bases, Prompt Management, Prompt Flows, Guardrails, and Model Evaluation", plus the Converse API, "a unified API that abstracts FM differences and enables model switching with a single parameter change". Custom Model Import brings your own weights in, but "this feature only supports Llama 2/3, Mistral, and Flan architectures".[2]
The pricing page lists on-demand prices per million tokens behind a Region selector, so the figures below are the US ones — US East (N. Virginia), US East (Ohio) and US West (Oregon): Amazon Nova Lite at $0.06 input and $0.24 output, Amazon Nova Pro at $0.80 and $3.20. Other Regions are priced differently; Nova Pro output is $4.72 per million in Europe (Paris). AWS offers select models "for batch inference at a 50% lower price compared to on-demand inference pricing".[1]
The same page carries Claude 3.5 Sonnet twice, at different prices, and a reader can easily budget from the wrong row. The live-feed row reads $3.00 input and $15.00 output per million tokens, $1.50 and $7.50 in batch. A separate static row labelled "Claude 3.5 Sonnet (Public Extended Access, Effective 1 Dec 2025)" reads $6.00 and $30.00, $3.00 and $15.00 in batch; a "Claude 3.5 Sonnet v2" row carries the same label and the same four numbers.[1]
Provisioned Throughput is sold on different terms by different model families, and the difference decides how long you are locked in. The Anthropic tables are headed "Price per hour per 1K Input TPM with 1-Month Commitment" and "… with 3-Month Commitment", over rows for Claude Opus 4.6, Opus 4.5, Haiku 4.5, Sonnet 4.6 and Sonnet 4.5. The Amazon Nova, Amazon Titan and Meta tables are headed "Price per hour per model unit with no commitment", "… for 1-month commitment" and "… for 6-month commitment" — a no-commitment option exists, but only for those families.[1]
The Anthropic section does two things at once: it publishes Claude Provisioned Throughput rates outright in two Reserved Tier tables, and then adds "For Provisioned Throughput pricing, please reach out to your account team." Access to the newest Claude models is gated separately — "Access to Claude Mythos 5.1, Claude Mythos 5, and Claude Mythos Preview is gated and requires approval. Contact your Anthropic account team to request access on Bedrock."[1]
On data, AWS states: "With Amazon Bedrock, your content is not used to improve the base models and is not shared with any model providers." It also states that "Any customer content processed by Amazon Bedrock is encrypted and stored at rest in the AWS Region where you are using Amazon Bedrock", and that "You can use AWS PrivateLink with Amazon Bedrock to establish private connectivity between your FMs and your Amazon Virtual Private Cloud (Amazon VPC) without exposing your traffic to the Internet."[2]
Bedrock has two lifecycle policies, split by launch date. The current one opens: "This page describes the model lifecycle policy for models launched on Amazon Bedrock on or after September 7, 2026", and sets three states — Active, Legacy and end-of-life. "Once the Legacy period begins, new customers can't adopt the model, and existing customers may lose access after 15 days of inactivity", you "can't create a new Provisioned Throughput for models in the Legacy state", and AWS warns that "Model lifecycle dates are specific to Amazon Bedrock and may differ from dates published by model providers (such as Anthropic or Cohere)".[3] Models launched before that date are governed by the legacy policy instead.[4]
Only the legacy policy explains why an old model can cost more, and it applies the rule "for models with EOL dates after February 1, 2026": "After a minimum of 3 months in the Legacy state, a model will enter the public extended access portion of the Legacy period. During this public extended access period, active users of a Legacy model can continue to use it until the EOL date (for a minimum of 3 months), but you should expect higher pricing, which will be set by the model provider." The mechanism is published, but the status of any one model has to be checked rather than assumed: the same page's table of models "currently in the Legacy state or … pending end-of-life (EOL)", which "does not include Active models or models that have already passed their EOL date", lists Claude Opus 4.1, Claude Sonnet 4 and Claude 3 Haiku — and no Claude 3.5 Sonnet, although the pricing page carries an extended-access row for it.[4][1]
Cross-Region inference profiles come in two kinds. Geographic profiles keep data residency "Within geographic boundaries (such as US, EU, and APAC)"; global profiles route worldwide, and AWS puts their cost at "Approximately 10% savings" against the geographic "Standard pricing", noting that "There's no additional routing cost for using cross-Region inference." One limit matters if you also buy capacity: "Inference profiles currently don't support Provisioned Throughput."[5]
Anthropic draws the same line from its side: "Partner-operated platforms (Amazon Bedrock and Google Cloud) set their own retirement schedules, so a model's lifecycle status and dates can differ." On its own platforms it commits to "providing at least 60 days' notice before model retirement for publicly released models".[6]
The combination to notice: a model you depend on moves to Legacy on Bedrock's clock rather than its maker's, an idle workload can lose access within fifteen days of that, and a model kept alive through public extended access is billed at the higher of the two rows on the pricing page.
| Option | Where the model runs | Retirement clock | Data use, as published |
|---|---|---|---|
| Amazon Bedrock | AWS Regions; geographic or global cross-Region inference profiles[5] | Active, Legacy, end of life; in Legacy "new customers can't adopt the model, and existing customers may lose access after 15 days of inactivity"[3]. Models launched before 7 September 2026 follow the legacy policy, where public extended access keeps a model callable at "higher pricing, which will be set by the model provider"[4] | "your content is not used to improve the base models and is not shared with any model providers"; content "encrypted and stored at rest in the AWS Region where you are using Amazon Bedrock"[2] |
| Azure OpenAI in Microsoft Foundry | Microsoft Azure Regions | "Retirement date (18 months out) is set programmatically"; "At 12 months from launch… New customers can't access the model."; "GA model retirement notice | At least 60 days before retirement". Exception: "Generally available models from Anthropic, DeepSeek, Fireworks, and Mistral AI follow a 12-month lifecycle instead of the standard 18-month lifecycle."[7] | prompts, completions, embeddings and training data "are NOT available to OpenAI or other providers of Models sold by Azure" and "are NOT used to train any generative AI foundation models without your permission or instruction"[8] |
| Google Vertex AI (Gemini Enterprise Agent Platform) | Google Cloud | "Models available for at least 12 months after release", with a retirement date listed per model[9] | "Google won't use your data to train or fine-tune any AI/ML models without your prior permission or instruction. This applies to all managed models on Gemini Enterprise Agent Platform, including GA and pre-GA models."[10] |
| Direct model API (OpenAI, Anthropic) | The model developer's infrastructure | OpenAI: "Generally available models: At least 6 months."[11] Anthropic: "at least 60 days' notice before model retirement for publicly released models"[6] | "Anthropic may not train models on Customer Content from Services."[13] "As of March 1, 2023, data sent to the OpenAI API is not used to train or improve OpenAI models (unless you explicitly opt in to share data with us)."[12] |
| Open-weight model with vLLM on your infrastructure | Your servers or your own cloud account | None imposed by a vendor: you upgrade when you choose | No provider receives the prompts. vLLM's repository carries the Apache-2.0 licence[14] and serves an "OpenAI-compatible API server, plus Anthropic Messages API and gRPC support"[15] |
Self-hosting moves the vendor's duties to you: accelerator capacity, patching, security of the serving layer and on-call. The trade-offs of all three deployment families, with the compliance evidence each one produces, are in choosing an LLM deployment for a closed perimeter; the closest managed substitute has a page of its own, Azure OpenAI alternatives.
| Item | What to do | Why |
|---|---|---|
| Model IDs and inference profiles in code | List every model ID and every inference profile each service calls, and where they are configured | Profiles and Regions differ between options, and a hard-coded ID breaks silently[5] |
| Provisioned Throughput commitments | Note which model family each commitment belongs to, and its end date, before planning a cutover | The terms differ by family: Anthropic's tables offer 1-month and 3-month commitments, while Amazon Nova, Amazon Titan and Meta are sold with no commitment, 1-month or 6-month[1]. New capacity also can't be created for a model already in Legacy[3] |
| Knowledge Bases | Export source documents, chunking settings and metadata; plan the new vector store and the re-ingestion | Retrieval built inside Bedrock Knowledge Bases[2] has to be rebuilt elsewhere |
| Guardrails | Write every policy, blocked topic and filter down as a specification | Guardrails are a Bedrock feature[2]; the rules must be reimplemented in code or in the next platform |
| Custom and fine-tuned models | Keep training data, hyperparameters and evaluation results under your own control | The FAQ describes importing models into Bedrock with Custom Model Import but does not address exporting a customised model out of it[2] |
| Agents, prompts and flows | Export prompt versions and flow definitions into your own repository | AWS offers "an option to easily export existing Bedrock Agents configurations as code" compatible with Strands and AgentCore; the FAQ says nothing similar about Prompt Management or Prompt Flows[2] |
| Evaluation set | Score the target model on your own documents or tickets before the cutover | Output quality differs by model; the same prompt is not the same result |
| Parallel run | Route a share of traffic to the new option and compare outputs and cost per request | A cutover without a comparison replaces a known error rate with an unknown one |
A managed cloud keeps capacity, patching and model hosting on its side and leaves you contracts, Regions and quotas. A direct API leaves you the provider's retention terms and retirement schedule; the two largest are set side by side in OpenAI API vs Claude API. Self-hosting leaves you everything, including the choice of when to upgrade: the trade-off against a managed API is in OpenAI API vs a self-hosted LLM, the serving engines in vLLM vs Ollama, and what running one on your own servers involves in adding an LLM to internal systems. A document workload built inside an existing perimeter is described in an LLM document pipeline inside your own infrastructure.
About amBrain
AWS states that "With Amazon Bedrock, your content is not used to improve the base models and is not shared with any model providers", and that "Any customer content processed by Amazon Bedrock is encrypted and stored at rest in the AWS Region where you are using Amazon Bedrock."[2]
"Once the Legacy period begins, new customers can't adopt the model, and existing customers may lose access after 15 days of inactivity", and no new Provisioned Throughput can be created for it.[3] Bedrock's dates "may differ from dates published by model providers (such as Anthropic or Cohere)", so track both. Models launched before 7 September 2026 follow a separate legacy policy, under which a model with an EOL date after 1 February 2026 spends "a minimum of 3 months in the Legacy state" before public extended access begins.[4]
Because of public extended access: "active users of a Legacy model can continue to use it until the EOL date (for a minimum of 3 months), but you should expect higher pricing, which will be set by the model provider."[4] The pricing page shows both prices for Claude 3.5 Sonnet — a live-feed row at $3.00 input and $15.00 output per million tokens, and a static "Claude 3.5 Sonnet (Public Extended Access, Effective 1 Dec 2025)" row at $6.00 and $30.00.[1] Which one applies to you depends on the model's state, which is published on its model card rather than inferred from the price table.[3]
Yes, if the perimeter and the volume justify running servers. vLLM exposes an "OpenAI-compatible API server, plus Anthropic Messages API and gRPC support"[15], which keeps client-side changes small; the work is in capacity, security, evaluation and on-call. Knowledge Bases, Guardrails and Prompt Flows have no equivalent on the other side and have to be rebuilt.[2]
Both are managed clouds with their own retirement clocks. Microsoft sets a retirement date "18 months out… programmatically" and closes a model to new customers at 12 months, except that models "from Anthropic, DeepSeek, Fireworks, and Mistral AI follow a 12-month lifecycle".[7] Google lists models "available for at least 12 months after release".[9] Choose by where your data, your contracts and your private connectivity already are.
Disclosure: this page is published by amBrain, and the last row of the comparison table is the kind of work amBrain does. Every vendor fact on this page carries a source with the date it was read; the prices were read on 2026-09-23, are the US-Region figures unless another Region is named, and change often — check the provider's own page on the day you decide. Amazon Bedrock, AWS, Azure OpenAI, Microsoft Foundry, Vertex AI, Gemini, OpenAI, Anthropic, Claude, Llama, Mistral and vLLM are trademarks of their owners and are used only to identify the products discussed.
Related