We recommend new projects start with resources from the AWS provider.
published on Monday, Sep 28, 2026 by Pulumi
We recommend new projects start with resources from the AWS provider.
published on Monday, Sep 28, 2026 by Pulumi
Resource Type definition for AWS::SageMaker::InferenceComponent
Create InferenceComponent Resource
Resources are created with functions called constructors. To learn more about declaring and configuring resources, see Resources.
Constructor syntax
new InferenceComponent(name: string, args: InferenceComponentArgs, opts?: CustomResourceOptions);@overload
def InferenceComponent(resource_name: str,
args: InferenceComponentArgs,
opts: Optional[ResourceOptions] = None)
@overload
def InferenceComponent(resource_name: str,
opts: Optional[ResourceOptions] = None,
endpoint_name: Optional[str] = None,
deployment_config: Optional[InferenceComponentDeploymentConfigArgs] = None,
endpoint_arn: Optional[str] = None,
inference_component_name: Optional[str] = None,
runtime_config: Optional[InferenceComponentRuntimeConfigArgs] = None,
specification: Optional[InferenceComponentSpecificationArgs] = None,
specifications: Optional[Sequence[InferenceComponentSpecificationForInstanceTypeArgs]] = None,
tags: Optional[Sequence[_root_inputs.TagArgs]] = None,
variant_name: Optional[str] = None)func NewInferenceComponent(ctx *Context, name string, args InferenceComponentArgs, opts ...ResourceOption) (*InferenceComponent, error)public InferenceComponent(string name, InferenceComponentArgs args, CustomResourceOptions? opts = null)
public InferenceComponent(String name, InferenceComponentArgs args)
public InferenceComponent(String name, InferenceComponentArgs args, CustomResourceOptions options)
type: aws-native:sagemaker:InferenceComponent
properties: # The arguments to resource properties.
options: # Bag of options to control resource's behavior.
resource "aws-native_sagemaker_inference_component" "name" {
# resource properties
}Parameters
- name string
- The unique name of the resource.
- args InferenceComponentArgs
- The arguments to resource properties.
- opts CustomResourceOptions
- Bag of options to control resource's behavior.
- resource_name str
- The unique name of the resource.
- args InferenceComponentArgs
- The arguments to resource properties.
- opts ResourceOptions
- Bag of options to control resource's behavior.
- ctx Context
- Context object for the current deployment.
- name string
- The unique name of the resource.
- args InferenceComponentArgs
- The arguments to resource properties.
- opts ResourceOption
- Bag of options to control resource's behavior.
- name string
- The unique name of the resource.
- args InferenceComponentArgs
- The arguments to resource properties.
- opts CustomResourceOptions
- Bag of options to control resource's behavior.
- name String
- The unique name of the resource.
- args InferenceComponentArgs
- The arguments to resource properties.
- options CustomResourceOptions
- Bag of options to control resource's behavior.
InferenceComponent Resource Properties
To learn more about resource properties and how to use them, see Inputs and Outputs in the Architecture and Concepts docs.
Inputs
In Python, inputs that are objects can be passed either as argument classes or as dictionary literals.
The InferenceComponent resource accepts the following input properties:
- Endpoint
Name string - The name of the endpoint that hosts the inference component.
- Deployment
Config Pulumi.Aws Native. Sage Maker. Inputs. Inference Component Deployment Config - The deployment configuration for an endpoint, which contains the desired deployment strategy and rollback configurations.
- Endpoint
Arn string - The Amazon Resource Name (ARN) of the endpoint that hosts the inference component.
- Inference
Component stringName - The name of the inference component.
- Runtime
Config Pulumi.Aws Native. Sage Maker. Inputs. Inference Component Runtime Config - Specification
Pulumi.
Aws Native. Sage Maker. Inputs. Inference Component Specification - Specifications
List<Pulumi.
Aws Native. Sage Maker. Inputs. Inference Component Specification For Instance Type> -
List<Pulumi.
Aws Native. Inputs. Tag> - Variant
Name string - The name of the production variant that hosts the inference component.
- Endpoint
Name string - The name of the endpoint that hosts the inference component.
- Deployment
Config InferenceComponent Deployment Config Args - The deployment configuration for an endpoint, which contains the desired deployment strategy and rollback configurations.
- Endpoint
Arn string - The Amazon Resource Name (ARN) of the endpoint that hosts the inference component.
- Inference
Component stringName - The name of the inference component.
- Runtime
Config InferenceComponent Runtime Config Args - Specification
Inference
Component Specification Args - Specifications
[]Inference
Component Specification For Instance Type Args -
Tag
Args - Variant
Name string - The name of the production variant that hosts the inference component.
- endpoint_
name string - The name of the endpoint that hosts the inference component.
- deployment_
config object - The deployment configuration for an endpoint, which contains the desired deployment strategy and rollback configurations.
- endpoint_
arn string - The Amazon Resource Name (ARN) of the endpoint that hosts the inference component.
- inference_
component_ stringname - The name of the inference component.
- runtime_
config object - specification object
- specifications list(object)
- list(object)
- variant_
name string - The name of the production variant that hosts the inference component.
- endpoint
Name String - The name of the endpoint that hosts the inference component.
- deployment
Config InferenceComponent Deployment Config - The deployment configuration for an endpoint, which contains the desired deployment strategy and rollback configurations.
- endpoint
Arn String - The Amazon Resource Name (ARN) of the endpoint that hosts the inference component.
- inference
Component StringName - The name of the inference component.
- runtime
Config InferenceComponent Runtime Config - specification
Inference
Component Specification - specifications
List<Inference
Component Specification For Instance Type> - List<Tag>
- variant
Name String - The name of the production variant that hosts the inference component.
- endpoint
Name string - The name of the endpoint that hosts the inference component.
- deployment
Config InferenceComponent Deployment Config - The deployment configuration for an endpoint, which contains the desired deployment strategy and rollback configurations.
- endpoint
Arn string - The Amazon Resource Name (ARN) of the endpoint that hosts the inference component.
- inference
Component stringName - The name of the inference component.
- runtime
Config InferenceComponent Runtime Config - specification
Inference
Component Specification - specifications
Inference
Component Specification For Instance Type[] - Tag[]
- variant
Name string - The name of the production variant that hosts the inference component.
- endpoint_
name str - The name of the endpoint that hosts the inference component.
- deployment_
config InferenceComponent Deployment Config Args - The deployment configuration for an endpoint, which contains the desired deployment strategy and rollback configurations.
- endpoint_
arn str - The Amazon Resource Name (ARN) of the endpoint that hosts the inference component.
- inference_
component_ strname - The name of the inference component.
- runtime_
config InferenceComponent Runtime Config Args - specification
Inference
Component Specification Args - specifications
Sequence[Inference
Component Specification For Instance Type Args] -
Sequence[Tag
Args] - variant_
name str - The name of the production variant that hosts the inference component.
- endpoint
Name String - The name of the endpoint that hosts the inference component.
- deployment
Config Property Map - The deployment configuration for an endpoint, which contains the desired deployment strategy and rollback configurations.
- endpoint
Arn String - The Amazon Resource Name (ARN) of the endpoint that hosts the inference component.
- inference
Component StringName - The name of the inference component.
- runtime
Config Property Map - specification Property Map
- specifications List<Property Map>
- List<Property Map>
- variant
Name String - The name of the production variant that hosts the inference component.
Outputs
All input properties are implicitly available as output properties. Additionally, the InferenceComponent resource produces the following output properties:
- Creation
Time string - The time when the inference component was created.
- Failure
Reason string - Id string
- The provider-assigned unique ID for this managed resource.
- Inference
Component stringArn - The Amazon Resource Name (ARN) of the inference component.
- Inference
Component Pulumi.Status Aws Native. Sage Maker. Inference Component Status - The status of the inference component.
- Last
Modified stringTime - The time when the inference component was last updated.
- Creation
Time string - The time when the inference component was created.
- Failure
Reason string - Id string
- The provider-assigned unique ID for this managed resource.
- Inference
Component stringArn - The Amazon Resource Name (ARN) of the inference component.
- Inference
Component InferenceStatus Component Status - The status of the inference component.
- Last
Modified stringTime - The time when the inference component was last updated.
- creation_
time string - The time when the inference component was created.
- failure_
reason string - id string
- The provider-assigned unique ID for this managed resource.
- inference_
component_ stringarn - The Amazon Resource Name (ARN) of the inference component.
- inference_
component_ "Instatus Service" | "Creating" | "Updating" | "Failed" | "Deleting" - The status of the inference component.
- last_
modified_ stringtime - The time when the inference component was last updated.
- creation
Time String - The time when the inference component was created.
- failure
Reason String - id String
- The provider-assigned unique ID for this managed resource.
- inference
Component StringArn - The Amazon Resource Name (ARN) of the inference component.
- inference
Component InferenceStatus Component Status - The status of the inference component.
- last
Modified StringTime - The time when the inference component was last updated.
- creation
Time string - The time when the inference component was created.
- failure
Reason string - id string
- The provider-assigned unique ID for this managed resource.
- inference
Component stringArn - The Amazon Resource Name (ARN) of the inference component.
- inference
Component InferenceStatus Component Status - The status of the inference component.
- last
Modified stringTime - The time when the inference component was last updated.
- creation_
time str - The time when the inference component was created.
- failure_
reason str - id str
- The provider-assigned unique ID for this managed resource.
- inference_
component_ strarn - The Amazon Resource Name (ARN) of the inference component.
- inference_
component_ Inferencestatus Component Status - The status of the inference component.
- last_
modified_ strtime - The time when the inference component was last updated.
- creation
Time String - The time when the inference component was created.
- failure
Reason String - id String
- The provider-assigned unique ID for this managed resource.
- inference
Component StringArn - The Amazon Resource Name (ARN) of the inference component.
- inference
Component "InStatus Service" | "Creating" | "Updating" | "Failed" | "Deleting" - The status of the inference component.
- last
Modified StringTime - The time when the inference component was last updated.
Supporting Types
InferenceComponentAlarm, InferenceComponentAlarmArgs
- Alarm
Name string - The name of a CloudWatch alarm in your account.
- Alarm
Name string - The name of a CloudWatch alarm in your account.
- alarm_
name string - The name of a CloudWatch alarm in your account.
- alarm
Name String - The name of a CloudWatch alarm in your account.
- alarm
Name string - The name of a CloudWatch alarm in your account.
- alarm_
name str - The name of a CloudWatch alarm in your account.
- alarm
Name String - The name of a CloudWatch alarm in your account.
InferenceComponentAutoRollbackConfiguration, InferenceComponentAutoRollbackConfigurationArgs
InferenceComponentAvailabilityZoneBalance, InferenceComponentAvailabilityZoneBalanceArgs
Configuration for balancing inference component copies across Availability Zones- Enforcement
Mode Pulumi.Aws Native. Sage Maker. Inference Component Availability Zone Balance Enforcement Mode - Max
Imbalance int - The maximum allowed difference in the number of inference component copies between any two Availability Zones
- Enforcement
Mode InferenceComponent Availability Zone Balance Enforcement Mode - Max
Imbalance int - The maximum allowed difference in the number of inference component copies between any two Availability Zones
- enforcement_
mode "PERMISSIVE" - max_
imbalance number - The maximum allowed difference in the number of inference component copies between any two Availability Zones
- enforcement
Mode InferenceComponent Availability Zone Balance Enforcement Mode - max
Imbalance Integer - The maximum allowed difference in the number of inference component copies between any two Availability Zones
- enforcement
Mode InferenceComponent Availability Zone Balance Enforcement Mode - max
Imbalance number - The maximum allowed difference in the number of inference component copies between any two Availability Zones
- enforcement_
mode InferenceComponent Availability Zone Balance Enforcement Mode - max_
imbalance int - The maximum allowed difference in the number of inference component copies between any two Availability Zones
- enforcement
Mode "PERMISSIVE" - max
Imbalance Number - The maximum allowed difference in the number of inference component copies between any two Availability Zones
InferenceComponentAvailabilityZoneBalanceEnforcementMode, InferenceComponentAvailabilityZoneBalanceEnforcementModeArgs
- Permissive
PERMISSIVE
- Inference
Component Availability Zone Balance Enforcement Mode Permissive PERMISSIVE
- "PERMISSIVE"
PERMISSIVE
- Permissive
PERMISSIVE
- Permissive
PERMISSIVE
- PERMISSIVE
PERMISSIVE
- "PERMISSIVE"
PERMISSIVE
InferenceComponentCapacitySize, InferenceComponentCapacitySizeArgs
Capacity size configuration for the inference component- Type
Pulumi.
Aws Native. Sage Maker. Inference Component Capacity Size Type - Specifies the endpoint capacity type.
- COPY_COUNT - The endpoint activates based on the number of inference component copies.
- CAPACITY_PERCENT - The endpoint activates based on the specified percentage of capacity.
- Value int
- Defines the capacity size, either as a number of inference component copies or a capacity percentage.
- Type
Inference
Component Capacity Size Type - Specifies the endpoint capacity type.
- COPY_COUNT - The endpoint activates based on the number of inference component copies.
- CAPACITY_PERCENT - The endpoint activates based on the specified percentage of capacity.
- Value int
- Defines the capacity size, either as a number of inference component copies or a capacity percentage.
- type "COPY_COUNT" | "CAPACITY_PERCENT"
- Specifies the endpoint capacity type.
- COPY_COUNT - The endpoint activates based on the number of inference component copies.
- CAPACITY_PERCENT - The endpoint activates based on the specified percentage of capacity.
- value number
- Defines the capacity size, either as a number of inference component copies or a capacity percentage.
- type
Inference
Component Capacity Size Type - Specifies the endpoint capacity type.
- COPY_COUNT - The endpoint activates based on the number of inference component copies.
- CAPACITY_PERCENT - The endpoint activates based on the specified percentage of capacity.
- value Integer
- Defines the capacity size, either as a number of inference component copies or a capacity percentage.
- type
Inference
Component Capacity Size Type - Specifies the endpoint capacity type.
- COPY_COUNT - The endpoint activates based on the number of inference component copies.
- CAPACITY_PERCENT - The endpoint activates based on the specified percentage of capacity.
- value number
- Defines the capacity size, either as a number of inference component copies or a capacity percentage.
- type
Inference
Component Capacity Size Type - Specifies the endpoint capacity type.
- COPY_COUNT - The endpoint activates based on the number of inference component copies.
- CAPACITY_PERCENT - The endpoint activates based on the specified percentage of capacity.
- value int
- Defines the capacity size, either as a number of inference component copies or a capacity percentage.
- type "COPY_COUNT" | "CAPACITY_PERCENT"
- Specifies the endpoint capacity type.
- COPY_COUNT - The endpoint activates based on the number of inference component copies.
- CAPACITY_PERCENT - The endpoint activates based on the specified percentage of capacity.
- value Number
- Defines the capacity size, either as a number of inference component copies or a capacity percentage.
InferenceComponentCapacitySizeType, InferenceComponentCapacitySizeTypeArgs
- Copy
Count COPY_COUNT- Capacity
Percent CAPACITY_PERCENT
- Inference
Component Capacity Size Type Copy Count COPY_COUNT- Inference
Component Capacity Size Type Capacity Percent CAPACITY_PERCENT
- "COPY_COUNT"
COPY_COUNT- "CAPACITY_PERCENT"
CAPACITY_PERCENT
- Copy
Count COPY_COUNT- Capacity
Percent CAPACITY_PERCENT
- Copy
Count COPY_COUNT- Capacity
Percent CAPACITY_PERCENT
- COPY_COUNT
COPY_COUNT- CAPACITY_PERCENT
CAPACITY_PERCENT
- "COPY_COUNT"
COPY_COUNT- "CAPACITY_PERCENT"
CAPACITY_PERCENT
InferenceComponentComputeResourceRequirements, InferenceComponentComputeResourceRequirementsArgs
- Max
Memory intRequired In Mb - The maximum MB of memory to allocate to run a model that you assign to an inference component.
- Min
Memory intRequired In Mb - The minimum MB of memory to allocate to run a model that you assign to an inference component.
- Number
Of doubleAccelerator Devices Required - The number of accelerators to allocate to run a model that you assign to an inference component. Accelerators include GPUs and AWS Inferentia.
- Number
Of doubleCpu Cores Required - The number of CPU cores to allocate to run a model that you assign to an inference component.
- Max
Memory intRequired In Mb - The maximum MB of memory to allocate to run a model that you assign to an inference component.
- Min
Memory intRequired In Mb - The minimum MB of memory to allocate to run a model that you assign to an inference component.
- Number
Of float64Accelerator Devices Required - The number of accelerators to allocate to run a model that you assign to an inference component. Accelerators include GPUs and AWS Inferentia.
- Number
Of float64Cpu Cores Required - The number of CPU cores to allocate to run a model that you assign to an inference component.
- max_
memory_ numberrequired_ in_ mb - The maximum MB of memory to allocate to run a model that you assign to an inference component.
- min_
memory_ numberrequired_ in_ mb - The minimum MB of memory to allocate to run a model that you assign to an inference component.
- number_
of_ numberaccelerator_ devices_ required - The number of accelerators to allocate to run a model that you assign to an inference component. Accelerators include GPUs and AWS Inferentia.
- number_
of_ numbercpu_ cores_ required - The number of CPU cores to allocate to run a model that you assign to an inference component.
- max
Memory IntegerRequired In Mb - The maximum MB of memory to allocate to run a model that you assign to an inference component.
- min
Memory IntegerRequired In Mb - The minimum MB of memory to allocate to run a model that you assign to an inference component.
- number
Of DoubleAccelerator Devices Required - The number of accelerators to allocate to run a model that you assign to an inference component. Accelerators include GPUs and AWS Inferentia.
- number
Of DoubleCpu Cores Required - The number of CPU cores to allocate to run a model that you assign to an inference component.
- max
Memory numberRequired In Mb - The maximum MB of memory to allocate to run a model that you assign to an inference component.
- min
Memory numberRequired In Mb - The minimum MB of memory to allocate to run a model that you assign to an inference component.
- number
Of numberAccelerator Devices Required - The number of accelerators to allocate to run a model that you assign to an inference component. Accelerators include GPUs and AWS Inferentia.
- number
Of numberCpu Cores Required - The number of CPU cores to allocate to run a model that you assign to an inference component.
- max_
memory_ intrequired_ in_ mb - The maximum MB of memory to allocate to run a model that you assign to an inference component.
- min_
memory_ intrequired_ in_ mb - The minimum MB of memory to allocate to run a model that you assign to an inference component.
- number_
of_ floataccelerator_ devices_ required - The number of accelerators to allocate to run a model that you assign to an inference component. Accelerators include GPUs and AWS Inferentia.
- number_
of_ floatcpu_ cores_ required - The number of CPU cores to allocate to run a model that you assign to an inference component.
- max
Memory NumberRequired In Mb - The maximum MB of memory to allocate to run a model that you assign to an inference component.
- min
Memory NumberRequired In Mb - The minimum MB of memory to allocate to run a model that you assign to an inference component.
- number
Of NumberAccelerator Devices Required - The number of accelerators to allocate to run a model that you assign to an inference component. Accelerators include GPUs and AWS Inferentia.
- number
Of NumberCpu Cores Required - The number of CPU cores to allocate to run a model that you assign to an inference component.
InferenceComponentContainerMetricsConfig, InferenceComponentContainerMetricsConfigArgs
The configuration for container metrics scrapingInferenceComponentContainerSpecification, InferenceComponentContainerSpecificationArgs
- Artifact
Url string - The Amazon S3 path where the model artifacts, which result from model training, are stored. This path must point to a single gzip compressed tar archive (.tar.gz suffix).
- Container
Metrics Pulumi.Config Aws Native. Sage Maker. Inputs. Inference Component Container Metrics Config - Deployed
Image Pulumi.Aws Native. Sage Maker. Inputs. Inference Component Deployed Image - Environment Dictionary<string, string>
- The environment variables to set in the Docker container. Each key and value in the Environment string-to-string map can have length of up to 1024. We support up to 16 entries in the map.
- Image string
- The Amazon Elastic Container Registry (Amazon ECR) path where the Docker image for the model is stored.
- Artifact
Url string - The Amazon S3 path where the model artifacts, which result from model training, are stored. This path must point to a single gzip compressed tar archive (.tar.gz suffix).
- Container
Metrics InferenceConfig Component Container Metrics Config - Deployed
Image InferenceComponent Deployed Image - Environment map[string]string
- The environment variables to set in the Docker container. Each key and value in the Environment string-to-string map can have length of up to 1024. We support up to 16 entries in the map.
- Image string
- The Amazon Elastic Container Registry (Amazon ECR) path where the Docker image for the model is stored.
- artifact_
url string - The Amazon S3 path where the model artifacts, which result from model training, are stored. This path must point to a single gzip compressed tar archive (.tar.gz suffix).
- container_
metrics_ objectconfig - deployed_
image object - environment map(string)
- The environment variables to set in the Docker container. Each key and value in the Environment string-to-string map can have length of up to 1024. We support up to 16 entries in the map.
- image string
- The Amazon Elastic Container Registry (Amazon ECR) path where the Docker image for the model is stored.
- artifact
Url String - The Amazon S3 path where the model artifacts, which result from model training, are stored. This path must point to a single gzip compressed tar archive (.tar.gz suffix).
- container
Metrics InferenceConfig Component Container Metrics Config - deployed
Image InferenceComponent Deployed Image - environment Map<String,String>
- The environment variables to set in the Docker container. Each key and value in the Environment string-to-string map can have length of up to 1024. We support up to 16 entries in the map.
- image String
- The Amazon Elastic Container Registry (Amazon ECR) path where the Docker image for the model is stored.
- artifact
Url string - The Amazon S3 path where the model artifacts, which result from model training, are stored. This path must point to a single gzip compressed tar archive (.tar.gz suffix).
- container
Metrics InferenceConfig Component Container Metrics Config - deployed
Image InferenceComponent Deployed Image - environment {[key: string]: string}
- The environment variables to set in the Docker container. Each key and value in the Environment string-to-string map can have length of up to 1024. We support up to 16 entries in the map.
- image string
- The Amazon Elastic Container Registry (Amazon ECR) path where the Docker image for the model is stored.
- artifact_
url str - The Amazon S3 path where the model artifacts, which result from model training, are stored. This path must point to a single gzip compressed tar archive (.tar.gz suffix).
- container_
metrics_ Inferenceconfig Component Container Metrics Config - deployed_
image InferenceComponent Deployed Image - environment Mapping[str, str]
- The environment variables to set in the Docker container. Each key and value in the Environment string-to-string map can have length of up to 1024. We support up to 16 entries in the map.
- image str
- The Amazon Elastic Container Registry (Amazon ECR) path where the Docker image for the model is stored.
- artifact
Url String - The Amazon S3 path where the model artifacts, which result from model training, are stored. This path must point to a single gzip compressed tar archive (.tar.gz suffix).
- container
Metrics Property MapConfig - deployed
Image Property Map - environment Map<String>
- The environment variables to set in the Docker container. Each key and value in the Environment string-to-string map can have length of up to 1024. We support up to 16 entries in the map.
- image String
- The Amazon Elastic Container Registry (Amazon ECR) path where the Docker image for the model is stored.
InferenceComponentContainerSpecificationForInstanceType, InferenceComponentContainerSpecificationForInstanceTypeArgs
Container specification for one Specifications entry. Distinct from InferenceComponentContainerSpecification: DescribeInferenceComponent returns no per-entry DeployedImage (VERIFIED in us-west-2), so DeployedImage is intentionally omitted here and this definition can never be aggregated into a plural READ response. The singular InferenceComponentContainerSpecification keeps DeployedImage - the service DOES return it there.- Artifact
Url string - Container
Metrics Pulumi.Config Aws Native. Sage Maker. Inputs. Inference Component Container Metrics Config - Environment Dictionary<string, string>
- Image string
- Artifact
Url string - Container
Metrics InferenceConfig Component Container Metrics Config - Environment map[string]string
- Image string
- artifact_
url string - container_
metrics_ objectconfig - environment map(string)
- image string
- artifact
Url String - container
Metrics InferenceConfig Component Container Metrics Config - environment Map<String,String>
- image String
- artifact
Url string - container
Metrics InferenceConfig Component Container Metrics Config - environment {[key: string]: string}
- image string
- artifact_
url str - container_
metrics_ Inferenceconfig Component Container Metrics Config - environment Mapping[str, str]
- image str
- artifact
Url String - container
Metrics Property MapConfig - environment Map<String>
- image String
InferenceComponentDataCacheConfig, InferenceComponentDataCacheConfigArgs
Settings that affect how the inference component caches data- Enable
Caching bool - Whether the endpoint caches the model artifacts and container image on each instance it provisions for the inference component
- Enable
Caching bool - Whether the endpoint caches the model artifacts and container image on each instance it provisions for the inference component
- enable_
caching bool - Whether the endpoint caches the model artifacts and container image on each instance it provisions for the inference component
- enable
Caching Boolean - Whether the endpoint caches the model artifacts and container image on each instance it provisions for the inference component
- enable
Caching boolean - Whether the endpoint caches the model artifacts and container image on each instance it provisions for the inference component
- enable_
caching bool - Whether the endpoint caches the model artifacts and container image on each instance it provisions for the inference component
- enable
Caching Boolean - Whether the endpoint caches the model artifacts and container image on each instance it provisions for the inference component
InferenceComponentDeployedImage, InferenceComponentDeployedImageArgs
- Resolution
Time string - The date and time when the image path for the model resolved to the
ResolvedImage - Resolved
Image string - The specific digest path of the image hosted in this
ProductionVariant. - Specified
Image string - The image path you specified when you created the model.
- Resolution
Time string - The date and time when the image path for the model resolved to the
ResolvedImage - Resolved
Image string - The specific digest path of the image hosted in this
ProductionVariant. - Specified
Image string - The image path you specified when you created the model.
- resolution_
time string - The date and time when the image path for the model resolved to the
ResolvedImage - resolved_
image string - The specific digest path of the image hosted in this
ProductionVariant. - specified_
image string - The image path you specified when you created the model.
- resolution
Time String - The date and time when the image path for the model resolved to the
ResolvedImage - resolved
Image String - The specific digest path of the image hosted in this
ProductionVariant. - specified
Image String - The image path you specified when you created the model.
- resolution
Time string - The date and time when the image path for the model resolved to the
ResolvedImage - resolved
Image string - The specific digest path of the image hosted in this
ProductionVariant. - specified
Image string - The image path you specified when you created the model.
- resolution_
time str - The date and time when the image path for the model resolved to the
ResolvedImage - resolved_
image str - The specific digest path of the image hosted in this
ProductionVariant. - specified_
image str - The image path you specified when you created the model.
- resolution
Time String - The date and time when the image path for the model resolved to the
ResolvedImage - resolved
Image String - The specific digest path of the image hosted in this
ProductionVariant. - specified
Image String - The image path you specified when you created the model.
InferenceComponentDeploymentConfig, InferenceComponentDeploymentConfigArgs
The deployment config for the inference component- Auto
Rollback Pulumi.Configuration Aws Native. Sage Maker. Inputs. Inference Component Auto Rollback Configuration - Rolling
Update Pulumi.Policy Aws Native. Sage Maker. Inputs. Inference Component Rolling Update Policy - Specifies a rolling deployment strategy for updating a SageMaker AI endpoint.
- Auto
Rollback InferenceConfiguration Component Auto Rollback Configuration - Rolling
Update InferencePolicy Component Rolling Update Policy - Specifies a rolling deployment strategy for updating a SageMaker AI endpoint.
- auto_
rollback_ objectconfiguration - rolling_
update_ objectpolicy - Specifies a rolling deployment strategy for updating a SageMaker AI endpoint.
- auto
Rollback InferenceConfiguration Component Auto Rollback Configuration - rolling
Update InferencePolicy Component Rolling Update Policy - Specifies a rolling deployment strategy for updating a SageMaker AI endpoint.
- auto
Rollback InferenceConfiguration Component Auto Rollback Configuration - rolling
Update InferencePolicy Component Rolling Update Policy - Specifies a rolling deployment strategy for updating a SageMaker AI endpoint.
- auto_
rollback_ Inferenceconfiguration Component Auto Rollback Configuration - rolling_
update_ Inferencepolicy Component Rolling Update Policy - Specifies a rolling deployment strategy for updating a SageMaker AI endpoint.
- auto
Rollback Property MapConfiguration - rolling
Update Property MapPolicy - Specifies a rolling deployment strategy for updating a SageMaker AI endpoint.
InferenceComponentMetricsEndpoint, InferenceComponentMetricsEndpointArgs
A metrics endpoint exposed by the container- Metrics
Endpoint stringPath - The path to the Prometheus formatted metrics endpoint exposed by the container
- Metric
Publish intFrequency In Seconds - The interval, in seconds, at which container metrics scraped from the endpoint are published to Amazon CloudWatch. Valid values per the SageMaker API Reference are 10, 30, 60, 120, 180, 240 and 300; the service validates the value.
- Metrics
Endpoint stringPath - The path to the Prometheus formatted metrics endpoint exposed by the container
- Metric
Publish intFrequency In Seconds - The interval, in seconds, at which container metrics scraped from the endpoint are published to Amazon CloudWatch. Valid values per the SageMaker API Reference are 10, 30, 60, 120, 180, 240 and 300; the service validates the value.
- metrics_
endpoint_ stringpath - The path to the Prometheus formatted metrics endpoint exposed by the container
- metric_
publish_ numberfrequency_ in_ seconds - The interval, in seconds, at which container metrics scraped from the endpoint are published to Amazon CloudWatch. Valid values per the SageMaker API Reference are 10, 30, 60, 120, 180, 240 and 300; the service validates the value.
- metrics
Endpoint StringPath - The path to the Prometheus formatted metrics endpoint exposed by the container
- metric
Publish IntegerFrequency In Seconds - The interval, in seconds, at which container metrics scraped from the endpoint are published to Amazon CloudWatch. Valid values per the SageMaker API Reference are 10, 30, 60, 120, 180, 240 and 300; the service validates the value.
- metrics
Endpoint stringPath - The path to the Prometheus formatted metrics endpoint exposed by the container
- metric
Publish numberFrequency In Seconds - The interval, in seconds, at which container metrics scraped from the endpoint are published to Amazon CloudWatch. Valid values per the SageMaker API Reference are 10, 30, 60, 120, 180, 240 and 300; the service validates the value.
- metrics_
endpoint_ strpath - The path to the Prometheus formatted metrics endpoint exposed by the container
- metric_
publish_ intfrequency_ in_ seconds - The interval, in seconds, at which container metrics scraped from the endpoint are published to Amazon CloudWatch. Valid values per the SageMaker API Reference are 10, 30, 60, 120, 180, 240 and 300; the service validates the value.
- metrics
Endpoint StringPath - The path to the Prometheus formatted metrics endpoint exposed by the container
- metric
Publish NumberFrequency In Seconds - The interval, in seconds, at which container metrics scraped from the endpoint are published to Amazon CloudWatch. Valid values per the SageMaker API Reference are 10, 30, 60, 120, 180, 240 and 300; the service validates the value.
InferenceComponentPlacementStatus, InferenceComponentPlacementStatusArgs
The number of inference component copies currently placed on instances of a given type- Current
Copy intCount - Instance
Type string
- Current
Copy intCount - Instance
Type string
- current_
copy_ numbercount - instance_
type string
- current
Copy IntegerCount - instance
Type String
- current
Copy numberCount - instance
Type string
- current_
copy_ intcount - instance_
type str
- current
Copy NumberCount - instance
Type String
InferenceComponentPlacementStrategy, InferenceComponentPlacementStrategyArgs
- Spread
SPREAD- Binpack
BINPACK
- Inference
Component Placement Strategy Spread SPREAD- Inference
Component Placement Strategy Binpack BINPACK
- "SPREAD"
SPREAD- "BINPACK"
BINPACK
- Spread
SPREAD- Binpack
BINPACK
- Spread
SPREAD- Binpack
BINPACK
- SPREAD
SPREAD- BINPACK
BINPACK
- "SPREAD"
SPREAD- "BINPACK"
BINPACK
InferenceComponentRollingUpdatePolicy, InferenceComponentRollingUpdatePolicyArgs
The rolling update policy for the inference component- Maximum
Batch Pulumi.Size Aws Native. Sage Maker. Inputs. Inference Component Capacity Size - The batch size for each rolling step in the deployment process. For each step, SageMaker AI provisions capacity on the new endpoint fleet, routes traffic to that fleet, and terminates capacity on the old endpoint fleet. The value must be between 5% to 50% of the copy count of the inference component.
- Maximum
Execution intTimeout In Seconds - The time limit for the total deployment. Exceeding this limit causes a timeout.
- Rollback
Maximum Pulumi.Batch Size Aws Native. Sage Maker. Inputs. Inference Component Capacity Size - The batch size for a rollback to the old endpoint fleet. If this field is absent, the value is set to the default, which is 100% of the total capacity. When the default is used, SageMaker AI provisions the entire capacity of the old fleet at once during rollback.
- Wait
Interval intIn Seconds - The length of the baking period, during which SageMaker AI monitors alarms for each batch on the new fleet.
- Maximum
Batch InferenceSize Component Capacity Size - The batch size for each rolling step in the deployment process. For each step, SageMaker AI provisions capacity on the new endpoint fleet, routes traffic to that fleet, and terminates capacity on the old endpoint fleet. The value must be between 5% to 50% of the copy count of the inference component.
- Maximum
Execution intTimeout In Seconds - The time limit for the total deployment. Exceeding this limit causes a timeout.
- Rollback
Maximum InferenceBatch Size Component Capacity Size - The batch size for a rollback to the old endpoint fleet. If this field is absent, the value is set to the default, which is 100% of the total capacity. When the default is used, SageMaker AI provisions the entire capacity of the old fleet at once during rollback.
- Wait
Interval intIn Seconds - The length of the baking period, during which SageMaker AI monitors alarms for each batch on the new fleet.
- maximum_
batch_ objectsize - The batch size for each rolling step in the deployment process. For each step, SageMaker AI provisions capacity on the new endpoint fleet, routes traffic to that fleet, and terminates capacity on the old endpoint fleet. The value must be between 5% to 50% of the copy count of the inference component.
- maximum_
execution_ numbertimeout_ in_ seconds - The time limit for the total deployment. Exceeding this limit causes a timeout.
- rollback_
maximum_ objectbatch_ size - The batch size for a rollback to the old endpoint fleet. If this field is absent, the value is set to the default, which is 100% of the total capacity. When the default is used, SageMaker AI provisions the entire capacity of the old fleet at once during rollback.
- wait_
interval_ numberin_ seconds - The length of the baking period, during which SageMaker AI monitors alarms for each batch on the new fleet.
- maximum
Batch InferenceSize Component Capacity Size - The batch size for each rolling step in the deployment process. For each step, SageMaker AI provisions capacity on the new endpoint fleet, routes traffic to that fleet, and terminates capacity on the old endpoint fleet. The value must be between 5% to 50% of the copy count of the inference component.
- maximum
Execution IntegerTimeout In Seconds - The time limit for the total deployment. Exceeding this limit causes a timeout.
- rollback
Maximum InferenceBatch Size Component Capacity Size - The batch size for a rollback to the old endpoint fleet. If this field is absent, the value is set to the default, which is 100% of the total capacity. When the default is used, SageMaker AI provisions the entire capacity of the old fleet at once during rollback.
- wait
Interval IntegerIn Seconds - The length of the baking period, during which SageMaker AI monitors alarms for each batch on the new fleet.
- maximum
Batch InferenceSize Component Capacity Size - The batch size for each rolling step in the deployment process. For each step, SageMaker AI provisions capacity on the new endpoint fleet, routes traffic to that fleet, and terminates capacity on the old endpoint fleet. The value must be between 5% to 50% of the copy count of the inference component.
- maximum
Execution numberTimeout In Seconds - The time limit for the total deployment. Exceeding this limit causes a timeout.
- rollback
Maximum InferenceBatch Size Component Capacity Size - The batch size for a rollback to the old endpoint fleet. If this field is absent, the value is set to the default, which is 100% of the total capacity. When the default is used, SageMaker AI provisions the entire capacity of the old fleet at once during rollback.
- wait
Interval numberIn Seconds - The length of the baking period, during which SageMaker AI monitors alarms for each batch on the new fleet.
- maximum_
batch_ Inferencesize Component Capacity Size - The batch size for each rolling step in the deployment process. For each step, SageMaker AI provisions capacity on the new endpoint fleet, routes traffic to that fleet, and terminates capacity on the old endpoint fleet. The value must be between 5% to 50% of the copy count of the inference component.
- maximum_
execution_ inttimeout_ in_ seconds - The time limit for the total deployment. Exceeding this limit causes a timeout.
- rollback_
maximum_ Inferencebatch_ size Component Capacity Size - The batch size for a rollback to the old endpoint fleet. If this field is absent, the value is set to the default, which is 100% of the total capacity. When the default is used, SageMaker AI provisions the entire capacity of the old fleet at once during rollback.
- wait_
interval_ intin_ seconds - The length of the baking period, during which SageMaker AI monitors alarms for each batch on the new fleet.
- maximum
Batch Property MapSize - The batch size for each rolling step in the deployment process. For each step, SageMaker AI provisions capacity on the new endpoint fleet, routes traffic to that fleet, and terminates capacity on the old endpoint fleet. The value must be between 5% to 50% of the copy count of the inference component.
- maximum
Execution NumberTimeout In Seconds - The time limit for the total deployment. Exceeding this limit causes a timeout.
- rollback
Maximum Property MapBatch Size - The batch size for a rollback to the old endpoint fleet. If this field is absent, the value is set to the default, which is 100% of the total capacity. When the default is used, SageMaker AI provisions the entire capacity of the old fleet at once during rollback.
- wait
Interval NumberIn Seconds - The length of the baking period, during which SageMaker AI monitors alarms for each batch on the new fleet.
InferenceComponentRuntimeConfig, InferenceComponentRuntimeConfigArgs
The runtime config for the inference component- Copy
Count int - The number of runtime copies of the model container to deploy with the inference component. Each copy can serve inference requests.
- Current
Copy intCount - Desired
Copy intCount - Placement
Status List<Pulumi.Aws Native. Sage Maker. Inputs. Inference Component Placement Status> - The placement status of the inference component across instance types
- Copy
Count int - The number of runtime copies of the model container to deploy with the inference component. Each copy can serve inference requests.
- Current
Copy intCount - Desired
Copy intCount - Placement
Status []InferenceComponent Placement Status - The placement status of the inference component across instance types
- copy_
count number - The number of runtime copies of the model container to deploy with the inference component. Each copy can serve inference requests.
- current_
copy_ numbercount - desired_
copy_ numbercount - placement_
status list(object) - The placement status of the inference component across instance types
- copy
Count Integer - The number of runtime copies of the model container to deploy with the inference component. Each copy can serve inference requests.
- current
Copy IntegerCount - desired
Copy IntegerCount - placement
Status List<InferenceComponent Placement Status> - The placement status of the inference component across instance types
- copy
Count number - The number of runtime copies of the model container to deploy with the inference component. Each copy can serve inference requests.
- current
Copy numberCount - desired
Copy numberCount - placement
Status InferenceComponent Placement Status[] - The placement status of the inference component across instance types
- copy_
count int - The number of runtime copies of the model container to deploy with the inference component. Each copy can serve inference requests.
- current_
copy_ intcount - desired_
copy_ intcount - placement_
status Sequence[InferenceComponent Placement Status] - The placement status of the inference component across instance types
- copy
Count Number - The number of runtime copies of the model container to deploy with the inference component. Each copy can serve inference requests.
- current
Copy NumberCount - desired
Copy NumberCount - placement
Status List<Property Map> - The placement status of the inference component across instance types
InferenceComponentSchedulingConfig, InferenceComponentSchedulingConfigArgs
The scheduling configuration that determines how inference component copies are placed across available instancesInferenceComponentSpecification, InferenceComponentSpecificationArgs
The specification for the inference component, for an endpoint with a single instance type. Specify exactly one of Specification or Specifications. InstanceType is not accepted here; use Specifications for per instance type configuration.- Base
Inference stringComponent Name The name of an existing inference component that is to contain the inference component that you're creating with your request.
Specify this parameter only if your request is meant to create an adapter inference component. An adapter inference component contains the path to an adapter model. The purpose of the adapter model is to tailor the inference output of a base foundation model, which is hosted by the base inference component. The adapter inference component uses the compute resources that you assigned to the base inference component.
When you create an adapter inference component, use the
Containerparameter to specify the location of the adapter artifacts. In the parameter value, use theArtifactUrlparameter of theInferenceComponentContainerSpecificationdata type.Before you can create an adapter inference component, you must have an existing inference component that contains the foundation model that you want to adapt.
- Compute
Resource Pulumi.Requirements Aws Native. Sage Maker. Inputs. Inference Component Compute Resource Requirements The compute resources allocated to run the model, plus any adapter models, that you assign to the inference component.
Omit this parameter if your request is meant to create an adapter inference component. An adapter inference component is loaded by a base inference component, and it uses the compute resources of the base inference component.
- Container
Pulumi.
Aws Native. Sage Maker. Inputs. Inference Component Container Specification - Defines a container that provides the runtime environment for a model that you deploy with an inference component.
- Current
Data Pulumi.Cache Config Aws Native. Sage Maker. Inputs. Inference Component Data Cache Config - The data caching configuration actually in effect, including a value the service chose rather than the template: SageMaker enables caching automatically on instance types with more than 232 GiB of local NVMe storage, whether or not DataCacheConfig was set. Returned by Describe and not settable; set DataCacheConfig instead.
- Data
Cache Pulumi.Config Aws Native. Sage Maker. Inputs. Inference Component Data Cache Config - Model
Name string - The name of an existing SageMaker AI model object in your account that you want to deploy with the inference component.
- Scheduling
Config Pulumi.Aws Native. Sage Maker. Inputs. Inference Component Scheduling Config - Startup
Parameters Pulumi.Aws Native. Sage Maker. Inputs. Inference Component Startup Parameters - Settings that take effect while the model container starts up.
- Base
Inference stringComponent Name The name of an existing inference component that is to contain the inference component that you're creating with your request.
Specify this parameter only if your request is meant to create an adapter inference component. An adapter inference component contains the path to an adapter model. The purpose of the adapter model is to tailor the inference output of a base foundation model, which is hosted by the base inference component. The adapter inference component uses the compute resources that you assigned to the base inference component.
When you create an adapter inference component, use the
Containerparameter to specify the location of the adapter artifacts. In the parameter value, use theArtifactUrlparameter of theInferenceComponentContainerSpecificationdata type.Before you can create an adapter inference component, you must have an existing inference component that contains the foundation model that you want to adapt.
- Compute
Resource InferenceRequirements Component Compute Resource Requirements The compute resources allocated to run the model, plus any adapter models, that you assign to the inference component.
Omit this parameter if your request is meant to create an adapter inference component. An adapter inference component is loaded by a base inference component, and it uses the compute resources of the base inference component.
- Container
Inference
Component Container Specification - Defines a container that provides the runtime environment for a model that you deploy with an inference component.
- Current
Data InferenceCache Config Component Data Cache Config - The data caching configuration actually in effect, including a value the service chose rather than the template: SageMaker enables caching automatically on instance types with more than 232 GiB of local NVMe storage, whether or not DataCacheConfig was set. Returned by Describe and not settable; set DataCacheConfig instead.
- Data
Cache InferenceConfig Component Data Cache Config - Model
Name string - The name of an existing SageMaker AI model object in your account that you want to deploy with the inference component.
- Scheduling
Config InferenceComponent Scheduling Config - Startup
Parameters InferenceComponent Startup Parameters - Settings that take effect while the model container starts up.
- base_
inference_ stringcomponent_ name The name of an existing inference component that is to contain the inference component that you're creating with your request.
Specify this parameter only if your request is meant to create an adapter inference component. An adapter inference component contains the path to an adapter model. The purpose of the adapter model is to tailor the inference output of a base foundation model, which is hosted by the base inference component. The adapter inference component uses the compute resources that you assigned to the base inference component.
When you create an adapter inference component, use the
Containerparameter to specify the location of the adapter artifacts. In the parameter value, use theArtifactUrlparameter of theInferenceComponentContainerSpecificationdata type.Before you can create an adapter inference component, you must have an existing inference component that contains the foundation model that you want to adapt.
- compute_
resource_ objectrequirements The compute resources allocated to run the model, plus any adapter models, that you assign to the inference component.
Omit this parameter if your request is meant to create an adapter inference component. An adapter inference component is loaded by a base inference component, and it uses the compute resources of the base inference component.
- container object
- Defines a container that provides the runtime environment for a model that you deploy with an inference component.
- current_
data_ objectcache_ config - The data caching configuration actually in effect, including a value the service chose rather than the template: SageMaker enables caching automatically on instance types with more than 232 GiB of local NVMe storage, whether or not DataCacheConfig was set. Returned by Describe and not settable; set DataCacheConfig instead.
- data_
cache_ objectconfig - model_
name string - The name of an existing SageMaker AI model object in your account that you want to deploy with the inference component.
- scheduling_
config object - startup_
parameters object - Settings that take effect while the model container starts up.
- base
Inference StringComponent Name The name of an existing inference component that is to contain the inference component that you're creating with your request.
Specify this parameter only if your request is meant to create an adapter inference component. An adapter inference component contains the path to an adapter model. The purpose of the adapter model is to tailor the inference output of a base foundation model, which is hosted by the base inference component. The adapter inference component uses the compute resources that you assigned to the base inference component.
When you create an adapter inference component, use the
Containerparameter to specify the location of the adapter artifacts. In the parameter value, use theArtifactUrlparameter of theInferenceComponentContainerSpecificationdata type.Before you can create an adapter inference component, you must have an existing inference component that contains the foundation model that you want to adapt.
- compute
Resource InferenceRequirements Component Compute Resource Requirements The compute resources allocated to run the model, plus any adapter models, that you assign to the inference component.
Omit this parameter if your request is meant to create an adapter inference component. An adapter inference component is loaded by a base inference component, and it uses the compute resources of the base inference component.
- container
Inference
Component Container Specification - Defines a container that provides the runtime environment for a model that you deploy with an inference component.
- current
Data InferenceCache Config Component Data Cache Config - The data caching configuration actually in effect, including a value the service chose rather than the template: SageMaker enables caching automatically on instance types with more than 232 GiB of local NVMe storage, whether or not DataCacheConfig was set. Returned by Describe and not settable; set DataCacheConfig instead.
- data
Cache InferenceConfig Component Data Cache Config - model
Name String - The name of an existing SageMaker AI model object in your account that you want to deploy with the inference component.
- scheduling
Config InferenceComponent Scheduling Config - startup
Parameters InferenceComponent Startup Parameters - Settings that take effect while the model container starts up.
- base
Inference stringComponent Name The name of an existing inference component that is to contain the inference component that you're creating with your request.
Specify this parameter only if your request is meant to create an adapter inference component. An adapter inference component contains the path to an adapter model. The purpose of the adapter model is to tailor the inference output of a base foundation model, which is hosted by the base inference component. The adapter inference component uses the compute resources that you assigned to the base inference component.
When you create an adapter inference component, use the
Containerparameter to specify the location of the adapter artifacts. In the parameter value, use theArtifactUrlparameter of theInferenceComponentContainerSpecificationdata type.Before you can create an adapter inference component, you must have an existing inference component that contains the foundation model that you want to adapt.
- compute
Resource InferenceRequirements Component Compute Resource Requirements The compute resources allocated to run the model, plus any adapter models, that you assign to the inference component.
Omit this parameter if your request is meant to create an adapter inference component. An adapter inference component is loaded by a base inference component, and it uses the compute resources of the base inference component.
- container
Inference
Component Container Specification - Defines a container that provides the runtime environment for a model that you deploy with an inference component.
- current
Data InferenceCache Config Component Data Cache Config - The data caching configuration actually in effect, including a value the service chose rather than the template: SageMaker enables caching automatically on instance types with more than 232 GiB of local NVMe storage, whether or not DataCacheConfig was set. Returned by Describe and not settable; set DataCacheConfig instead.
- data
Cache InferenceConfig Component Data Cache Config - model
Name string - The name of an existing SageMaker AI model object in your account that you want to deploy with the inference component.
- scheduling
Config InferenceComponent Scheduling Config - startup
Parameters InferenceComponent Startup Parameters - Settings that take effect while the model container starts up.
- base_
inference_ strcomponent_ name The name of an existing inference component that is to contain the inference component that you're creating with your request.
Specify this parameter only if your request is meant to create an adapter inference component. An adapter inference component contains the path to an adapter model. The purpose of the adapter model is to tailor the inference output of a base foundation model, which is hosted by the base inference component. The adapter inference component uses the compute resources that you assigned to the base inference component.
When you create an adapter inference component, use the
Containerparameter to specify the location of the adapter artifacts. In the parameter value, use theArtifactUrlparameter of theInferenceComponentContainerSpecificationdata type.Before you can create an adapter inference component, you must have an existing inference component that contains the foundation model that you want to adapt.
- compute_
resource_ Inferencerequirements Component Compute Resource Requirements The compute resources allocated to run the model, plus any adapter models, that you assign to the inference component.
Omit this parameter if your request is meant to create an adapter inference component. An adapter inference component is loaded by a base inference component, and it uses the compute resources of the base inference component.
- container
Inference
Component Container Specification - Defines a container that provides the runtime environment for a model that you deploy with an inference component.
- current_
data_ Inferencecache_ config Component Data Cache Config - The data caching configuration actually in effect, including a value the service chose rather than the template: SageMaker enables caching automatically on instance types with more than 232 GiB of local NVMe storage, whether or not DataCacheConfig was set. Returned by Describe and not settable; set DataCacheConfig instead.
- data_
cache_ Inferenceconfig Component Data Cache Config - model_
name str - The name of an existing SageMaker AI model object in your account that you want to deploy with the inference component.
- scheduling_
config InferenceComponent Scheduling Config - startup_
parameters InferenceComponent Startup Parameters - Settings that take effect while the model container starts up.
- base
Inference StringComponent Name The name of an existing inference component that is to contain the inference component that you're creating with your request.
Specify this parameter only if your request is meant to create an adapter inference component. An adapter inference component contains the path to an adapter model. The purpose of the adapter model is to tailor the inference output of a base foundation model, which is hosted by the base inference component. The adapter inference component uses the compute resources that you assigned to the base inference component.
When you create an adapter inference component, use the
Containerparameter to specify the location of the adapter artifacts. In the parameter value, use theArtifactUrlparameter of theInferenceComponentContainerSpecificationdata type.Before you can create an adapter inference component, you must have an existing inference component that contains the foundation model that you want to adapt.
- compute
Resource Property MapRequirements The compute resources allocated to run the model, plus any adapter models, that you assign to the inference component.
Omit this parameter if your request is meant to create an adapter inference component. An adapter inference component is loaded by a base inference component, and it uses the compute resources of the base inference component.
- container Property Map
- Defines a container that provides the runtime environment for a model that you deploy with an inference component.
- current
Data Property MapCache Config - The data caching configuration actually in effect, including a value the service chose rather than the template: SageMaker enables caching automatically on instance types with more than 232 GiB of local NVMe storage, whether or not DataCacheConfig was set. Returned by Describe and not settable; set DataCacheConfig instead.
- data
Cache Property MapConfig - model
Name String - The name of an existing SageMaker AI model object in your account that you want to deploy with the inference component.
- scheduling
Config Property Map - startup
Parameters Property Map - Settings that take effect while the model container starts up.
InferenceComponentSpecificationForInstanceType, InferenceComponentSpecificationForInstanceTypeArgs
A specification for one instance type, for use in Specifications. InstanceType is required here, and is not accepted on the singular Specification. BaseInferenceComponentName is not accepted here either: adapter inference components are supported only on the singular Specification.- Instance
Type string - Compute
Resource Pulumi.Requirements Aws Native. Sage Maker. Inputs. Inference Component Compute Resource Requirements - Container
Pulumi.
Aws Native. Sage Maker. Inputs. Inference Component Container Specification For Instance Type - Current
Data Pulumi.Cache Config Aws Native. Sage Maker. Inputs. Inference Component Data Cache Config - The data caching configuration actually in effect for this instance type, including a value the service chose rather than the template: SageMaker enables caching automatically on instance types with more than 232 GiB of local NVMe storage, whether or not DataCacheConfig was set. Returned by Describe and not settable; set DataCacheConfig instead.
- Data
Cache Pulumi.Config Aws Native. Sage Maker. Inputs. Inference Component Data Cache Config - Model
Name string - Scheduling
Config Pulumi.Aws Native. Sage Maker. Inputs. Inference Component Scheduling Config - Startup
Parameters Pulumi.Aws Native. Sage Maker. Inputs. Inference Component Startup Parameters
- Instance
Type string - Compute
Resource InferenceRequirements Component Compute Resource Requirements - Container
Inference
Component Container Specification For Instance Type - Current
Data InferenceCache Config Component Data Cache Config - The data caching configuration actually in effect for this instance type, including a value the service chose rather than the template: SageMaker enables caching automatically on instance types with more than 232 GiB of local NVMe storage, whether or not DataCacheConfig was set. Returned by Describe and not settable; set DataCacheConfig instead.
- Data
Cache InferenceConfig Component Data Cache Config - Model
Name string - Scheduling
Config InferenceComponent Scheduling Config - Startup
Parameters InferenceComponent Startup Parameters
- instance_
type string - compute_
resource_ objectrequirements - container object
- current_
data_ objectcache_ config - The data caching configuration actually in effect for this instance type, including a value the service chose rather than the template: SageMaker enables caching automatically on instance types with more than 232 GiB of local NVMe storage, whether or not DataCacheConfig was set. Returned by Describe and not settable; set DataCacheConfig instead.
- data_
cache_ objectconfig - model_
name string - scheduling_
config object - startup_
parameters object
- instance
Type String - compute
Resource InferenceRequirements Component Compute Resource Requirements - container
Inference
Component Container Specification For Instance Type - current
Data InferenceCache Config Component Data Cache Config - The data caching configuration actually in effect for this instance type, including a value the service chose rather than the template: SageMaker enables caching automatically on instance types with more than 232 GiB of local NVMe storage, whether or not DataCacheConfig was set. Returned by Describe and not settable; set DataCacheConfig instead.
- data
Cache InferenceConfig Component Data Cache Config - model
Name String - scheduling
Config InferenceComponent Scheduling Config - startup
Parameters InferenceComponent Startup Parameters
- instance
Type string - compute
Resource InferenceRequirements Component Compute Resource Requirements - container
Inference
Component Container Specification For Instance Type - current
Data InferenceCache Config Component Data Cache Config - The data caching configuration actually in effect for this instance type, including a value the service chose rather than the template: SageMaker enables caching automatically on instance types with more than 232 GiB of local NVMe storage, whether or not DataCacheConfig was set. Returned by Describe and not settable; set DataCacheConfig instead.
- data
Cache InferenceConfig Component Data Cache Config - model
Name string - scheduling
Config InferenceComponent Scheduling Config - startup
Parameters InferenceComponent Startup Parameters
- instance_
type str - compute_
resource_ Inferencerequirements Component Compute Resource Requirements - container
Inference
Component Container Specification For Instance Type - current_
data_ Inferencecache_ config Component Data Cache Config - The data caching configuration actually in effect for this instance type, including a value the service chose rather than the template: SageMaker enables caching automatically on instance types with more than 232 GiB of local NVMe storage, whether or not DataCacheConfig was set. Returned by Describe and not settable; set DataCacheConfig instead.
- data_
cache_ Inferenceconfig Component Data Cache Config - model_
name str - scheduling_
config InferenceComponent Scheduling Config - startup_
parameters InferenceComponent Startup Parameters
- instance
Type String - compute
Resource Property MapRequirements - container Property Map
- current
Data Property MapCache Config - The data caching configuration actually in effect for this instance type, including a value the service chose rather than the template: SageMaker enables caching automatically on instance types with more than 232 GiB of local NVMe storage, whether or not DataCacheConfig was set. Returned by Describe and not settable; set DataCacheConfig instead.
- data
Cache Property MapConfig - model
Name String - scheduling
Config Property Map - startup
Parameters Property Map
InferenceComponentStartupParameters, InferenceComponentStartupParametersArgs
- Container
Startup intHealth Check Timeout In Seconds - The timeout value, in seconds, for your inference container to pass health check by Amazon S3 Hosting. For more information about health check, see How Your Container Should Respond to Health Check (Ping) Requests .
- Model
Data intDownload Timeout In Seconds - The timeout value, in seconds, to download and extract the model that you want to host from Amazon S3 to the individual inference instance associated with this inference component.
- Container
Startup intHealth Check Timeout In Seconds - The timeout value, in seconds, for your inference container to pass health check by Amazon S3 Hosting. For more information about health check, see How Your Container Should Respond to Health Check (Ping) Requests .
- Model
Data intDownload Timeout In Seconds - The timeout value, in seconds, to download and extract the model that you want to host from Amazon S3 to the individual inference instance associated with this inference component.
- container_
startup_ numberhealth_ check_ timeout_ in_ seconds - The timeout value, in seconds, for your inference container to pass health check by Amazon S3 Hosting. For more information about health check, see How Your Container Should Respond to Health Check (Ping) Requests .
- model_
data_ numberdownload_ timeout_ in_ seconds - The timeout value, in seconds, to download and extract the model that you want to host from Amazon S3 to the individual inference instance associated with this inference component.
- container
Startup IntegerHealth Check Timeout In Seconds - The timeout value, in seconds, for your inference container to pass health check by Amazon S3 Hosting. For more information about health check, see How Your Container Should Respond to Health Check (Ping) Requests .
- model
Data IntegerDownload Timeout In Seconds - The timeout value, in seconds, to download and extract the model that you want to host from Amazon S3 to the individual inference instance associated with this inference component.
- container
Startup numberHealth Check Timeout In Seconds - The timeout value, in seconds, for your inference container to pass health check by Amazon S3 Hosting. For more information about health check, see How Your Container Should Respond to Health Check (Ping) Requests .
- model
Data numberDownload Timeout In Seconds - The timeout value, in seconds, to download and extract the model that you want to host from Amazon S3 to the individual inference instance associated with this inference component.
- container_
startup_ inthealth_ check_ timeout_ in_ seconds - The timeout value, in seconds, for your inference container to pass health check by Amazon S3 Hosting. For more information about health check, see How Your Container Should Respond to Health Check (Ping) Requests .
- model_
data_ intdownload_ timeout_ in_ seconds - The timeout value, in seconds, to download and extract the model that you want to host from Amazon S3 to the individual inference instance associated with this inference component.
- container
Startup NumberHealth Check Timeout In Seconds - The timeout value, in seconds, for your inference container to pass health check by Amazon S3 Hosting. For more information about health check, see How Your Container Should Respond to Health Check (Ping) Requests .
- model
Data NumberDownload Timeout In Seconds - The timeout value, in seconds, to download and extract the model that you want to host from Amazon S3 to the individual inference instance associated with this inference component.
InferenceComponentStatus, InferenceComponentStatusArgs
- In
Service InService- Creating
Creating- Updating
Updating- Failed
Failed- Deleting
Deleting
- Inference
Component Status In Service InService- Inference
Component Status Creating Creating- Inference
Component Status Updating Updating- Inference
Component Status Failed Failed- Inference
Component Status Deleting Deleting
- "In
Service" InService- "Creating"
Creating- "Updating"
Updating- "Failed"
Failed- "Deleting"
Deleting
- In
Service InService- Creating
Creating- Updating
Updating- Failed
Failed- Deleting
Deleting
- In
Service InService- Creating
Creating- Updating
Updating- Failed
Failed- Deleting
Deleting
- IN_SERVICE
InService- CREATING
Creating- UPDATING
Updating- FAILED
Failed- DELETING
Deleting
- "In
Service" InService- "Creating"
Creating- "Updating"
Updating- "Failed"
Failed- "Deleting"
Deleting
Tag, TagArgs
A set of tags to apply to the resource.Package Details
- Repository
- AWS Native pulumi/pulumi-aws-native
- License
- Apache-2.0
We recommend new projects start with resources from the AWS provider.
published on Monday, Sep 28, 2026 by Pulumi