amBrain

AWS Bedrock alternatives: other managed clouds, direct model APIs or a self-hosted LLM

Published Sep 23, 2026Facts checked: Sep 23, 2026

Short answer

Stay on Amazon Bedrock if your data already lives in AWS and you want several model families behind one API: AWS states that "With Amazon Bedrock, your content is not used to improve the base models and is not shared with any model providers".[2] Look at alternatives when Bedrock's model clock starts to cost you — once the Legacy period begins, "new customers can't adopt the model, and existing customers may lose access after 15 days of inactivity", and "Model lifecycle dates are specific to Amazon Bedrock and may differ from dates published by model providers (such as Anthropic or Cohere)".[3] The four alternatives are Azure OpenAI in Microsoft Foundry, Google Vertex AI, a model provider's direct API, and an open-weight model you run yourself.

On this page

What Bedrock is and what AWS publishes

Amazon Bedrock is AWS's managed service for foundation models from several developers behind one API, with a layer of product features around them. What follows is what AWS publishes on its own pages: the prices, the data terms, and the lifecycle rules that decide how long a model stays callable.

AWS lists the features around the models: "Amazon Bedrock Playground, Agents, Knowledge Bases, Prompt Management, Prompt Flows, Guardrails, and Model Evaluation", plus the Converse API, "a unified API that abstracts FM differences and enables model switching with a single parameter change". Custom Model Import brings your own weights in, but "this feature only supports Llama 2/3, Mistral, and Flan architectures".[2]

The pricing page lists on-demand prices per million tokens behind a Region selector, so the figures below are the US ones — US East (N. Virginia), US East (Ohio) and US West (Oregon): Amazon Nova Lite at $0.06 input and $0.24 output, Amazon Nova Pro at $0.80 and $3.20. Other Regions are priced differently; Nova Pro output is $4.72 per million in Europe (Paris). AWS offers select models "for batch inference at a 50% lower price compared to on-demand inference pricing".[1]

The same page carries Claude 3.5 Sonnet twice, at different prices, and a reader can easily budget from the wrong row. The live-feed row reads $3.00 input and $15.00 output per million tokens, $1.50 and $7.50 in batch. A separate static row labelled "Claude 3.5 Sonnet (Public Extended Access, Effective 1 Dec 2025)" reads $6.00 and $30.00, $3.00 and $15.00 in batch; a "Claude 3.5 Sonnet v2" row carries the same label and the same four numbers.[1]

Provisioned Throughput is sold on different terms by different model families, and the difference decides how long you are locked in. The Anthropic tables are headed "Price per hour per 1K Input TPM with 1-Month Commitment" and "… with 3-Month Commitment", over rows for Claude Opus 4.6, Opus 4.5, Haiku 4.5, Sonnet 4.6 and Sonnet 4.5. The Amazon Nova, Amazon Titan and Meta tables are headed "Price per hour per model unit with no commitment", "… for 1-month commitment" and "… for 6-month commitment" — a no-commitment option exists, but only for those families.[1]

The Anthropic section does two things at once: it publishes Claude Provisioned Throughput rates outright in two Reserved Tier tables, and then adds "For Provisioned Throughput pricing, please reach out to your account team." Access to the newest Claude models is gated separately — "Access to Claude Mythos 5.1, Claude Mythos 5, and Claude Mythos Preview is gated and requires approval. Contact your Anthropic account team to request access on Bedrock."[1]

On data, AWS states: "With Amazon Bedrock, your content is not used to improve the base models and is not shared with any model providers." It also states that "Any customer content processed by Amazon Bedrock is encrypted and stored at rest in the AWS Region where you are using Amazon Bedrock", and that "You can use AWS PrivateLink with Amazon Bedrock to establish private connectivity between your FMs and your Amazon Virtual Private Cloud (Amazon VPC) without exposing your traffic to the Internet."[2]

Bedrock has two lifecycle policies, split by launch date. The current one opens: "This page describes the model lifecycle policy for models launched on Amazon Bedrock on or after September 7, 2026", and sets three states — Active, Legacy and end-of-life. "Once the Legacy period begins, new customers can't adopt the model, and existing customers may lose access after 15 days of inactivity", you "can't create a new Provisioned Throughput for models in the Legacy state", and AWS warns that "Model lifecycle dates are specific to Amazon Bedrock and may differ from dates published by model providers (such as Anthropic or Cohere)".[3] Models launched before that date are governed by the legacy policy instead.[4]

Only the legacy policy explains why an old model can cost more, and it applies the rule "for models with EOL dates after February 1, 2026": "After a minimum of 3 months in the Legacy state, a model will enter the public extended access portion of the Legacy period. During this public extended access period, active users of a Legacy model can continue to use it until the EOL date (for a minimum of 3 months), but you should expect higher pricing, which will be set by the model provider." The mechanism is published, but the status of any one model has to be checked rather than assumed: the same page's table of models "currently in the Legacy state or … pending end-of-life (EOL)", which "does not include Active models or models that have already passed their EOL date", lists Claude Opus 4.1, Claude Sonnet 4 and Claude 3 Haiku — and no Claude 3.5 Sonnet, although the pricing page carries an extended-access row for it.[4][1]

Cross-Region inference profiles come in two kinds. Geographic profiles keep data residency "Within geographic boundaries (such as US, EU, and APAC)"; global profiles route worldwide, and AWS puts their cost at "Approximately 10% savings" against the geographic "Standard pricing", noting that "There's no additional routing cost for using cross-Region inference." One limit matters if you also buy capacity: "Inference profiles currently don't support Provisioned Throughput."[5]

Anthropic draws the same line from its side: "Partner-operated platforms (Amazon Bedrock and Google Cloud) set their own retirement schedules, so a model's lifecycle status and dates can differ." On its own platforms it commits to "providing at least 60 days' notice before model retirement for publicly released models".[6]

The combination to notice: a model you depend on moves to Legacy on Bedrock's clock rather than its maker's, an idle workload can lose access within fifteen days of that, and a model kept alive through public extended access is billed at the higher of the two rows on the pricing page.

Bedrock and four alternatives

Where the model runs, whose clock retires it, and what the provider publishes about data
OptionWhere the model runsRetirement clockData use, as published
Amazon BedrockAWS Regions; geographic or global cross-Region inference profiles[5]Active, Legacy, end of life; in Legacy "new customers can't adopt the model, and existing customers may lose access after 15 days of inactivity"[3]. Models launched before 7 September 2026 follow the legacy policy, where public extended access keeps a model callable at "higher pricing, which will be set by the model provider"[4]"your content is not used to improve the base models and is not shared with any model providers"; content "encrypted and stored at rest in the AWS Region where you are using Amazon Bedrock"[2]
Azure OpenAI in Microsoft FoundryMicrosoft Azure Regions"Retirement date (18 months out) is set programmatically"; "At 12 months from launch… New customers can't access the model."; "GA model retirement notice | At least 60 days before retirement". Exception: "Generally available models from Anthropic, DeepSeek, Fireworks, and Mistral AI follow a 12-month lifecycle instead of the standard 18-month lifecycle."[7]prompts, completions, embeddings and training data "are NOT available to OpenAI or other providers of Models sold by Azure" and "are NOT used to train any generative AI foundation models without your permission or instruction"[8]
Google Vertex AI (Gemini Enterprise Agent Platform)Google Cloud"Models available for at least 12 months after release", with a retirement date listed per model[9]"Google won't use your data to train or fine-tune any AI/ML models without your prior permission or instruction. This applies to all managed models on Gemini Enterprise Agent Platform, including GA and pre-GA models."[10]
Direct model API (OpenAI, Anthropic)The model developer's infrastructureOpenAI: "Generally available models: At least 6 months."[11] Anthropic: "at least 60 days' notice before model retirement for publicly released models"[6]"Anthropic may not train models on Customer Content from Services."[13] "As of March 1, 2023, data sent to the OpenAI API is not used to train or improve OpenAI models (unless you explicitly opt in to share data with us)."[12]
Open-weight model with vLLM on your infrastructureYour servers or your own cloud accountNone imposed by a vendor: you upgrade when you chooseNo provider receives the prompts. vLLM's repository carries the Apache-2.0 licence[14] and serves an "OpenAI-compatible API server, plus Anthropic Messages API and gRPC support"[15]

Self-hosting moves the vendor's duties to you: accelerator capacity, patching, security of the serving layer and on-call. The trade-offs of all three deployment families, with the compliance evidence each one produces, are in choosing an LLM deployment for a closed perimeter; the closest managed substitute has a page of its own, Azure OpenAI alternatives.

Stay on Bedrock if…

  1. 1.Your workloads and data already live in AWS. Private connectivity from your VPC "without exposing your traffic to the Internet"[2] is part of a security case you would otherwise have to rebuild somewhere else.
  2. 2.You use several model families behind one API. One bill, one set of IAM policies and "model switching with a single parameter change"[2] are worth more to you than having the newest model on its first day.
  3. 3.Your traffic is steady enough to keep models in use. The Legacy rule about losing access "after 15 days of inactivity"[3] never bites a workload that runs daily — provided you also have a routine for re-testing replacements.

Consider leaving Bedrock if…

  1. 1.Symptom: a model you rely on reaches Legacy and its successor has not passed your evaluation set. Bedrock sets that clock itself, and its dates "may differ from dates published by model providers"[3]; Anthropic says the same from the other side about "Partner-operated platforms"[6].
  2. 2.Symptom: your perimeter excludes any external operator. Encryption at rest in your Region and PrivateLink[2] are not the same thing as inference on hardware you run; when only the latter is acceptable, no managed cloud qualifies.
  3. 3.Symptom: per-token pricing has become your largest line. Compare it against Provisioned Throughput on the terms your model family is sold on[1] and against accelerators you own or rent — as a measurement on your own traffic, not an estimate.

What to settle before moving a workload off Bedrock

What to inventory and secure before a cutover
ItemWhat to doWhy
Model IDs and inference profiles in codeList every model ID and every inference profile each service calls, and where they are configuredProfiles and Regions differ between options, and a hard-coded ID breaks silently[5]
Provisioned Throughput commitmentsNote which model family each commitment belongs to, and its end date, before planning a cutoverThe terms differ by family: Anthropic's tables offer 1-month and 3-month commitments, while Amazon Nova, Amazon Titan and Meta are sold with no commitment, 1-month or 6-month[1]. New capacity also can't be created for a model already in Legacy[3]
Knowledge BasesExport source documents, chunking settings and metadata; plan the new vector store and the re-ingestionRetrieval built inside Bedrock Knowledge Bases[2] has to be rebuilt elsewhere
GuardrailsWrite every policy, blocked topic and filter down as a specificationGuardrails are a Bedrock feature[2]; the rules must be reimplemented in code or in the next platform
Custom and fine-tuned modelsKeep training data, hyperparameters and evaluation results under your own controlThe FAQ describes importing models into Bedrock with Custom Model Import but does not address exporting a customised model out of it[2]
Agents, prompts and flowsExport prompt versions and flow definitions into your own repositoryAWS offers "an option to easily export existing Bedrock Agents configurations as code" compatible with Strands and AgentCore; the FAQ says nothing similar about Prompt Management or Prompt Flows[2]
Evaluation setScore the target model on your own documents or tickets before the cutoverOutput quality differs by model; the same prompt is not the same result
Parallel runRoute a share of traffic to the new option and compare outputs and cost per requestA cutover without a comparison replaces a known error rate with an unknown one

Where each option leaves the work

A managed cloud keeps capacity, patching and model hosting on its side and leaves you contracts, Regions and quotas. A direct API leaves you the provider's retention terms and retirement schedule; the two largest are set side by side in OpenAI API vs Claude API. Self-hosting leaves you everything, including the choice of when to upgrade: the trade-off against a managed API is in OpenAI API vs a self-hosted LLM, the serving engines in vLLM vs Ollama, and what running one on your own servers involves in adding an LLM to internal systems. A document workload built inside an existing perimeter is described in an LLM document pipeline inside your own infrastructure.

About amBrain

  • We have taken an LLM integration to production inside a client's FinTech perimeter: extracting and normalising unstructured broker and venue notices — corporate actions, instrument and margin changes — into structured records the trading system consumes.
  • The client keeps full ownership of the product and the code, except our reusable components.

Frequently asked questions

AWS states that "With Amazon Bedrock, your content is not used to improve the base models and is not shared with any model providers", and that "Any customer content processed by Amazon Bedrock is encrypted and stored at rest in the AWS Region where you are using Amazon Bedrock."[2]

"Once the Legacy period begins, new customers can't adopt the model, and existing customers may lose access after 15 days of inactivity", and no new Provisioned Throughput can be created for it.[3] Bedrock's dates "may differ from dates published by model providers (such as Anthropic or Cohere)", so track both. Models launched before 7 September 2026 follow a separate legacy policy, under which a model with an EOL date after 1 February 2026 spends "a minimum of 3 months in the Legacy state" before public extended access begins.[4]

Because of public extended access: "active users of a Legacy model can continue to use it until the EOL date (for a minimum of 3 months), but you should expect higher pricing, which will be set by the model provider."[4] The pricing page shows both prices for Claude 3.5 Sonnet — a live-feed row at $3.00 input and $15.00 output per million tokens, and a static "Claude 3.5 Sonnet (Public Extended Access, Effective 1 Dec 2025)" row at $6.00 and $30.00.[1] Which one applies to you depends on the model's state, which is published on its model card rather than inferred from the price table.[3]

Yes, if the perimeter and the volume justify running servers. vLLM exposes an "OpenAI-compatible API server, plus Anthropic Messages API and gRPC support"[15], which keeps client-side changes small; the work is in capacity, security, evaluation and on-call. Knowledge Bases, Guardrails and Prompt Flows have no equivalent on the other side and have to be rebuilt.[2]

Both are managed clouds with their own retirement clocks. Microsoft sets a retirement date "18 months out… programmatically" and closes a model to new customers at 12 months, except that models "from Anthropic, DeepSeek, Fireworks, and Mistral AI follow a 12-month lifecycle".[7] Google lists models "available for at least 12 months after release".[9] Choose by where your data, your contracts and your private connectivity already are.

Disclosure: this page is published by amBrain, and the last row of the comparison table is the kind of work amBrain does. Every vendor fact on this page carries a source with the date it was read; the prices were read on 2026-09-23, are the US-Region figures unless another Region is named, and change often — check the provider's own page on the day you decide. Amazon Bedrock, AWS, Azure OpenAI, Microsoft Foundry, Vertex AI, Gemini, OpenAI, Anthropic, Claude, Llama, Mistral and vLLM are trademarks of their owners and are used only to identify the products discussed.

Sources

  1. [1]Amazon Web Services, Amazon Bedrock Pricing (on-demand, batch and Provisioned Throughput tables behind a Region selector; the numbers are loaded from AWS's pricing feed when the page renders). Vendor's own page · checked Sep 23, 2026
  2. [2]Amazon Web Services, Build Generative AI Applications with Foundation Models — Amazon Bedrock FAQs (feature list, Custom Model Import, data use, encryption at rest and PrivateLink answers). Vendor's own page · checked Sep 23, 2026
  3. [3]Amazon Web Services, Model lifecycle (policy for models launched on Amazon Bedrock on or after September 7, 2026: Active, Legacy and EOL states). Official documentation · checked Sep 23, 2026
  4. [4]Amazon Web Services, Model lifecycle (Legacy) (policy for models launched before September 7, 2026; the public extended access period and the table of Legacy and pending-EOL models). Official documentation · checked Sep 23, 2026
  5. [5]Amazon Web Services, Route model inference requests across AWS Regions with cross-Region inference (geographic against global inference profiles, their cost and their limits). Official documentation · checked Sep 23, 2026
  6. [6]Anthropic, Model deprecations (which platforms the dates apply to, and the notice period). Official documentation · checked Sep 23, 2026
  7. [7]Microsoft, Foundry Models lifecycle and support policy (the 18-month pattern, the 12-month exception for some providers, and notification timings). Official documentation · checked Sep 23, 2026
  8. [8]Microsoft, Data, privacy, and security for Foundry Models sold by Azure in Microsoft Foundry. Official documentation · checked Sep 23, 2026
  9. [9]Google Cloud, Model versions and lifecycle (the table of models available for at least 12 months after release, with retirement dates). Official documentation · checked Sep 23, 2026
  10. [10]Google Cloud, Gemini Enterprise Agent Platform and zero data retention (the training restriction and what it covers). Official documentation · checked Sep 23, 2026
  11. [11]OpenAI, Deprecations (minimum notice periods by model class). Official documentation · checked Sep 23, 2026
  12. [12]OpenAI, Data controls in the OpenAI platform. Official documentation · checked Sep 23, 2026
  13. [13]Anthropic, Commercial Terms of Service. Vendor's published terms · checked Sep 23, 2026
  14. [14]vLLM project, vllm-project/vllm source repository and README (GitHub's own licence field reads Apache-2.0; the docs page does not carry the licence). Vendor's own page · checked Sep 23, 2026
  15. [15]vLLM project, vLLM documentation (feature list, including the API servers it exposes). Official documentation · checked Sep 23, 2026

Free project intro call, 30 minutes.

We reply within 24 hours.

Book a Call