>
Foundation Model

Amazon Nova — The Complete Guide

Nova is Amazon's own family of foundation models, available through Bedrock. Its proposition is cost-efficiency and AWS integration rather than frontier capability.

Foundation ModelAWS BedrockCost-efficientUpdated May 2026
Visit Amazon Nova ↗aws.amazon.com
SimpleStart here

What is Amazon Nova?

Nova is Amazon's set of AI models, offered through AWS Bedrock alongside models from Anthropic, Meta, Mistral and others.

Amazon's position is not that Nova is the most capable model available — it is that Nova is competitively priced and sits natively inside AWS.

Who it is for

Organisations already running on AWS, where using Bedrock means no new vendor relationship, no separate billing, and existing IAM and VPC controls apply automatically.

For teams not on AWS, there is little reason to choose Nova over alternatives.

WorkingBuild it

The tiers

Nova ships in several tiers spanning very fast and cheap models for simple tasks up to more capable multimodal models handling text, images and video inputs.

This mirrors the tiering strategy every major provider now uses: route simple work to cheap models, reserve expensive ones for hard problems.

Bedrock as the real product

Bedrock matters more than Nova specifically. It provides one API across many model providers, with AWS-native security, logging, guardrails and private networking.

That means you can build against Bedrock and switch the underlying model — including to Claude or Llama — without re-architecting.

Honest positioning

Nova does not lead benchmarks and Amazon does not market it as doing so. It competes on price-performance and integration.

If you are on AWS, benchmark Nova against Claude on Bedrock for your actual workload. The cost difference is real and sometimes the capability difference is not material for the task.

DeepGo deeper

Guardrails and governance

Bedrock Guardrails apply content filtering, topic restriction and PII redaction at the platform level, independent of which model is called.

For regulated environments this is a genuine advantage — policy is enforced consistently rather than depending on each model's own behaviour.

Fine-tuning and customisation

Bedrock supports fine-tuning and continued pre-training on several models including Nova, with the customised model kept private within your account.

Knowledge bases and agents are also available as managed services, which removes some RAG engineering at the cost of flexibility.

Placing It Against the Alternatives

Model choice is a routing decision rather than a ranking one, and the useful question is which part of your traffic this is right for.

Route by task, not by preference. Most requests in most applications are not hard. Classification, extraction and formatting rarely need the most capable available option, and sending everything to the top tier is the largest and most common overspend.

Test on your own evaluation set, not on published benchmarks. A benchmark measures a task that is not yours, and the ordering between models frequently reverses on specific work.

Weigh the things that are not capability. Where the data goes and under whose terms. Latency at your percentile, not the average. Whether the model can change underneath you, and whether that matters for reproducibility. Rate limits at your peak rather than your mean.

Assume you will move. Keep the provider behind an interface, keep prompts in version control, and keep an evaluation set that runs against any of them. The cost of switching is paid once at design time or repeatedly afterwards.

And re-check on a schedule. This ordering changes faster than any procurement cycle, so a decision made a year ago and never revisited is a decision that has quietly expired.

Evaluating It Against Your Own Work

Vendor demonstrations are built on material the tool handles well, so the only evaluation that predicts anything is one run on your own inputs.

Assemble twenty real examples before the trial starts, including the awkward ones — the messy input, the edge case, the one that went wrong last month. A set of clean examples measures a situation you do not have.

Define what good looks like in writing, before you see any output. Deciding afterwards is choosing the answer rather than measuring it, and it is what makes most tool trials inconclusive.

Time the whole task, not the tool. A tool that halves the generation step and adds a verification step has not saved anything. Measure the end-to-end time including checking and correction, because that is the number your team experiences.

Have two people run the same examples. Tolerance for a given failure varies more between people than between tools, and a decision made by one enthusiast rarely survives contact with the team.

And price the failure, not just the licence. What does a wrong output cost here — a correction, an apology, a customer? That number decides how much checking you need, which is usually the real cost of adoption.

Ask an AI about this page

Opens your assistant with this page as the source, and a question rather than a summary. It will ask what you are building before it answers.

ChatGPTClaudeGeminiPerplexityGrok

Nothing is sent from here. The link carries only this page’s title and address.