watsonx.ai Runtime service plans
You use watsonx.ai Runtime resources, which are measured in capacity unit hours (CUH), when you train AutoAI models, run machine learning models, or score deployed models. You use watsonx.ai Runtime resources, measured by tokens consumed or at an hourly rate, when you run inferencing services with foundation models. This topic describes the various plans you can choose, what services are included, and how computing resources are calculated.
Choosing a watsonx.ai Runtime plan
watsonx.ai Runtime plans govern how you are billed for models you train and deploy with watsonx.ai Runtime and for prompts you use with foundation models. You must select a plan for each watsonx.ai Runtime instance you create.
Select one of the following plans based on your needs:
- Essentials plan
- A pay-as-you-go plan that gives you the flexibility to build, deploy, and manage models to match your needs.
- Standard plan
- A high-capacity enterprise plan that is designed to support all AI needs of an organization. The Standard plan incurs a monthly instance fee of $1,110 USD per month that includes a block of 2500 capacity unit hours (CUH). Any CUH usage above
this amount and all other usage is metered on a pay-as-you-go basis at the Standard plan rate.
Note:
The instance fee for the watsonx.ai Runtime Standard plan is billed regardless of CUH usage. For example, if you only consume resource units, you are still charged the instance fee. The fee is pro-rated if the plan is canceled.
For plan details and pricing, see IBM Cloud catalog: watsonx.ai Runtime.
Resource usage is measured in capacity unit hours (CUH) or resource units (RUs). For details on computing resource allocation and consumption, see Monitoring account resource usage. For resource usage pricing, see Billing details for generative AI assets.
| Plan features | Lite | Essentials | Standard |
|---|---|---|---|
| watsonx.ai Runtime usage in CUH | 20 CUH per month | CUH billing based on CUH rate multiplied by hours of consumption | 2500 CUH per month |
| Max parallel Decision Optimization batch jobs per deployment | 2 | 5 | 100 |
| Deployment jobs retained per space | 100 | 1000 | 3000 |
| Deployment time to idle | 1 day | 3 days | 3 days |
Learn more
- For more information on tracking computing resource allocation and consumption, see Runtime usage.
- IBM Cloud catalog: watsonx.ai Runtime