Generative AI platform
Amazon Bedrock Provisioned Throughput
Purchases predictable model inference capacity for supported Bedrock models.
Key points
- Provides reserved Bedrock model inference capacity for workloads with predictable demand.
- Can improve consistency for applications that need steady throughput.
- Applies to supported models and is configured separately from on-demand inference.
- Is useful when production traffic should not rely only on shared on-demand capacity.
- Capacity planning still depends on model choice, request size, and expected traffic.
When to use it
- Choose provisioned throughput for production generative AI workloads with steady or high usage.
- Use it when predictable performance matters more than occasional ad hoc model calls.
- Use on-demand inference instead for experiments, prototypes, or irregular traffic.
Exam tips
- Provisioned throughput is a Bedrock capacity option, not a SageMaker training feature.
- The exam may contrast cost-flexible on-demand inference with predictable provisioned capacity.
- Provisioned throughput does not replace Guardrails, Knowledge Bases, or Agents.
- Use it only for supported models and regions; keep answers conservative when exact availability is not stated.