1. Registry
  2. Packages
  3. AWS Cloud Control
  4. API Docs
  5. sagemaker
  6. getInferenceComponent

We recommend new projects start with resources from the AWS provider.

Viewing docs for AWS Cloud Control v1.81.0
published on Monday, Sep 28, 2026 by Pulumi
aws-native logo aws-native logo

We recommend new projects start with resources from the AWS provider.

Viewing docs for AWS Cloud Control v1.81.0
published on Monday, Sep 28, 2026 by Pulumi

    Resource Type definition for AWS::SageMaker::InferenceComponent

    Using getInferenceComponent

    Two invocation forms are available. The direct form accepts plain arguments and either blocks until the result value is available, or returns a Promise-wrapped result. The output form accepts Input-wrapped arguments and returns an Output-wrapped result.

    function getInferenceComponent(args: GetInferenceComponentArgs, opts?: InvokeOptions): Promise<GetInferenceComponentResult>
    function getInferenceComponentOutput(args: GetInferenceComponentOutputArgs, opts?: InvokeOutputOptions): Output<GetInferenceComponentResult>
    def get_inference_component(inference_component_arn: Optional[str] = None,
                                opts: Optional[InvokeOptions] = None) -> GetInferenceComponentResult
    def get_inference_component_output(inference_component_arn: pulumi.Input[Optional[str]] = None,
                                opts: Optional[InvokeOutputOptions] = None) -> Output[GetInferenceComponentResult]
    func LookupInferenceComponent(ctx *Context, args *LookupInferenceComponentArgs, opts ...InvokeOption) (*LookupInferenceComponentResult, error)
    func LookupInferenceComponentOutput(ctx *Context, args *LookupInferenceComponentOutputArgs, opts ...InvokeOption) LookupInferenceComponentResultOutput

    > Note: This function is named LookupInferenceComponent in the Go SDK.

    public static class GetInferenceComponent 
    {
        public static Task<GetInferenceComponentResult> InvokeAsync(GetInferenceComponentArgs args, InvokeOptions? opts = null)
        public static Output<GetInferenceComponentResult> Invoke(GetInferenceComponentInvokeArgs args, InvokeOptions? opts = null)
        public static Output<GetInferenceComponentResult> Invoke(GetInferenceComponentInvokeArgs args, InvokeOutputOptions opts)
    }
    public static CompletableFuture<GetInferenceComponentResult> getInferenceComponent(GetInferenceComponentArgs args, InvokeOptions options)
    public static Output<GetInferenceComponentResult> getInferenceComponent(GetInferenceComponentArgs args, InvokeOptions options)
    public static Output<GetInferenceComponentResult> getInferenceComponent(GetInferenceComponentArgs args, InvokeOutputOptions options)
    
    fn::invoke:
      function: aws-native:sagemaker:getInferenceComponent
      arguments:
        # arguments dictionary
    data "aws-native_sagemaker_get_inference_component" "name" {
        # arguments
    }

    The following arguments are supported:

    InferenceComponentArn string
    The Amazon Resource Name (ARN) of the inference component.
    InferenceComponentArn string
    The Amazon Resource Name (ARN) of the inference component.
    inference_component_arn string
    The Amazon Resource Name (ARN) of the inference component.
    inferenceComponentArn String
    The Amazon Resource Name (ARN) of the inference component.
    inferenceComponentArn string
    The Amazon Resource Name (ARN) of the inference component.
    inference_component_arn str
    The Amazon Resource Name (ARN) of the inference component.
    inferenceComponentArn String
    The Amazon Resource Name (ARN) of the inference component.

    getInferenceComponent Result

    The following output properties are available:

    CreationTime string
    The time when the inference component was created.
    EndpointArn string
    The Amazon Resource Name (ARN) of the endpoint that hosts the inference component.
    EndpointName string
    The name of the endpoint that hosts the inference component.
    FailureReason string
    InferenceComponentArn string
    The Amazon Resource Name (ARN) of the inference component.
    InferenceComponentName string
    The name of the inference component.
    InferenceComponentStatus Pulumi.AwsNative.SageMaker.InferenceComponentStatus
    The status of the inference component.
    LastModifiedTime string
    The time when the inference component was last updated.
    RuntimeConfig Pulumi.AwsNative.SageMaker.Outputs.InferenceComponentRuntimeConfig
    Specification Pulumi.AwsNative.SageMaker.Outputs.InferenceComponentSpecification
    Specifications List<Pulumi.AwsNative.SageMaker.Outputs.InferenceComponentSpecificationForInstanceType>
    Tags List<Pulumi.AwsNative.Outputs.Tag>
    VariantName string
    The name of the production variant that hosts the inference component.
    CreationTime string
    The time when the inference component was created.
    EndpointArn string
    The Amazon Resource Name (ARN) of the endpoint that hosts the inference component.
    EndpointName string
    The name of the endpoint that hosts the inference component.
    FailureReason string
    InferenceComponentArn string
    The Amazon Resource Name (ARN) of the inference component.
    InferenceComponentName string
    The name of the inference component.
    InferenceComponentStatus InferenceComponentStatus
    The status of the inference component.
    LastModifiedTime string
    The time when the inference component was last updated.
    RuntimeConfig InferenceComponentRuntimeConfig
    Specification InferenceComponentSpecification
    Specifications []InferenceComponentSpecificationForInstanceType
    Tags Tag
    VariantName string
    The name of the production variant that hosts the inference component.
    creation_time string
    The time when the inference component was created.
    endpoint_arn string
    The Amazon Resource Name (ARN) of the endpoint that hosts the inference component.
    endpoint_name string
    The name of the endpoint that hosts the inference component.
    failure_reason string
    inference_component_arn string
    The Amazon Resource Name (ARN) of the inference component.
    inference_component_name string
    The name of the inference component.
    inference_component_status "InService" | "Creating" | "Updating" | "Failed" | "Deleting"
    The status of the inference component.
    last_modified_time string
    The time when the inference component was last updated.
    runtime_config object
    specification object
    specifications list(object)
    tags list(object)
    variant_name string
    The name of the production variant that hosts the inference component.
    creationTime String
    The time when the inference component was created.
    endpointArn String
    The Amazon Resource Name (ARN) of the endpoint that hosts the inference component.
    endpointName String
    The name of the endpoint that hosts the inference component.
    failureReason String
    inferenceComponentArn String
    The Amazon Resource Name (ARN) of the inference component.
    inferenceComponentName String
    The name of the inference component.
    inferenceComponentStatus InferenceComponentStatus
    The status of the inference component.
    lastModifiedTime String
    The time when the inference component was last updated.
    runtimeConfig InferenceComponentRuntimeConfig
    specification InferenceComponentSpecification
    specifications List<InferenceComponentSpecificationForInstanceType>
    tags List<Tag>
    variantName String
    The name of the production variant that hosts the inference component.
    creationTime string
    The time when the inference component was created.
    endpointArn string
    The Amazon Resource Name (ARN) of the endpoint that hosts the inference component.
    endpointName string
    The name of the endpoint that hosts the inference component.
    failureReason string
    inferenceComponentArn string
    The Amazon Resource Name (ARN) of the inference component.
    inferenceComponentName string
    The name of the inference component.
    inferenceComponentStatus InferenceComponentStatus
    The status of the inference component.
    lastModifiedTime string
    The time when the inference component was last updated.
    runtimeConfig InferenceComponentRuntimeConfig
    specification InferenceComponentSpecification
    specifications InferenceComponentSpecificationForInstanceType[]
    tags Tag[]
    variantName string
    The name of the production variant that hosts the inference component.
    creation_time str
    The time when the inference component was created.
    endpoint_arn str
    The Amazon Resource Name (ARN) of the endpoint that hosts the inference component.
    endpoint_name str
    The name of the endpoint that hosts the inference component.
    failure_reason str
    inference_component_arn str
    The Amazon Resource Name (ARN) of the inference component.
    inference_component_name str
    The name of the inference component.
    inference_component_status InferenceComponentStatus
    The status of the inference component.
    last_modified_time str
    The time when the inference component was last updated.
    runtime_config InferenceComponentRuntimeConfig
    specification InferenceComponentSpecification
    specifications Sequence[InferenceComponentSpecificationForInstanceType]
    tags Sequence[root_Tag]
    variant_name str
    The name of the production variant that hosts the inference component.
    creationTime String
    The time when the inference component was created.
    endpointArn String
    The Amazon Resource Name (ARN) of the endpoint that hosts the inference component.
    endpointName String
    The name of the endpoint that hosts the inference component.
    failureReason String
    inferenceComponentArn String
    The Amazon Resource Name (ARN) of the inference component.
    inferenceComponentName String
    The name of the inference component.
    inferenceComponentStatus "InService" | "Creating" | "Updating" | "Failed" | "Deleting"
    The status of the inference component.
    lastModifiedTime String
    The time when the inference component was last updated.
    runtimeConfig Property Map
    specification Property Map
    specifications List<Property Map>
    tags List<Property Map>
    variantName String
    The name of the production variant that hosts the inference component.

    Supporting Types

    InferenceComponentAvailabilityZoneBalance

    EnforcementMode Pulumi.AwsNative.SageMaker.InferenceComponentAvailabilityZoneBalanceEnforcementMode
    MaxImbalance int
    The maximum allowed difference in the number of inference component copies between any two Availability Zones
    EnforcementMode InferenceComponentAvailabilityZoneBalanceEnforcementMode
    MaxImbalance int
    The maximum allowed difference in the number of inference component copies between any two Availability Zones
    enforcement_mode "PERMISSIVE"
    max_imbalance number
    The maximum allowed difference in the number of inference component copies between any two Availability Zones
    enforcementMode InferenceComponentAvailabilityZoneBalanceEnforcementMode
    maxImbalance Integer
    The maximum allowed difference in the number of inference component copies between any two Availability Zones
    enforcementMode InferenceComponentAvailabilityZoneBalanceEnforcementMode
    maxImbalance number
    The maximum allowed difference in the number of inference component copies between any two Availability Zones
    enforcement_mode InferenceComponentAvailabilityZoneBalanceEnforcementMode
    max_imbalance int
    The maximum allowed difference in the number of inference component copies between any two Availability Zones
    enforcementMode "PERMISSIVE"
    maxImbalance Number
    The maximum allowed difference in the number of inference component copies between any two Availability Zones

    InferenceComponentAvailabilityZoneBalanceEnforcementMode

    InferenceComponentComputeResourceRequirements

    MaxMemoryRequiredInMb int
    The maximum MB of memory to allocate to run a model that you assign to an inference component.
    MinMemoryRequiredInMb int
    The minimum MB of memory to allocate to run a model that you assign to an inference component.
    NumberOfAcceleratorDevicesRequired double
    The number of accelerators to allocate to run a model that you assign to an inference component. Accelerators include GPUs and AWS Inferentia.
    NumberOfCpuCoresRequired double
    The number of CPU cores to allocate to run a model that you assign to an inference component.
    MaxMemoryRequiredInMb int
    The maximum MB of memory to allocate to run a model that you assign to an inference component.
    MinMemoryRequiredInMb int
    The minimum MB of memory to allocate to run a model that you assign to an inference component.
    NumberOfAcceleratorDevicesRequired float64
    The number of accelerators to allocate to run a model that you assign to an inference component. Accelerators include GPUs and AWS Inferentia.
    NumberOfCpuCoresRequired float64
    The number of CPU cores to allocate to run a model that you assign to an inference component.
    max_memory_required_in_mb number
    The maximum MB of memory to allocate to run a model that you assign to an inference component.
    min_memory_required_in_mb number
    The minimum MB of memory to allocate to run a model that you assign to an inference component.
    number_of_accelerator_devices_required number
    The number of accelerators to allocate to run a model that you assign to an inference component. Accelerators include GPUs and AWS Inferentia.
    number_of_cpu_cores_required number
    The number of CPU cores to allocate to run a model that you assign to an inference component.
    maxMemoryRequiredInMb Integer
    The maximum MB of memory to allocate to run a model that you assign to an inference component.
    minMemoryRequiredInMb Integer
    The minimum MB of memory to allocate to run a model that you assign to an inference component.
    numberOfAcceleratorDevicesRequired Double
    The number of accelerators to allocate to run a model that you assign to an inference component. Accelerators include GPUs and AWS Inferentia.
    numberOfCpuCoresRequired Double
    The number of CPU cores to allocate to run a model that you assign to an inference component.
    maxMemoryRequiredInMb number
    The maximum MB of memory to allocate to run a model that you assign to an inference component.
    minMemoryRequiredInMb number
    The minimum MB of memory to allocate to run a model that you assign to an inference component.
    numberOfAcceleratorDevicesRequired number
    The number of accelerators to allocate to run a model that you assign to an inference component. Accelerators include GPUs and AWS Inferentia.
    numberOfCpuCoresRequired number
    The number of CPU cores to allocate to run a model that you assign to an inference component.
    max_memory_required_in_mb int
    The maximum MB of memory to allocate to run a model that you assign to an inference component.
    min_memory_required_in_mb int
    The minimum MB of memory to allocate to run a model that you assign to an inference component.
    number_of_accelerator_devices_required float
    The number of accelerators to allocate to run a model that you assign to an inference component. Accelerators include GPUs and AWS Inferentia.
    number_of_cpu_cores_required float
    The number of CPU cores to allocate to run a model that you assign to an inference component.
    maxMemoryRequiredInMb Number
    The maximum MB of memory to allocate to run a model that you assign to an inference component.
    minMemoryRequiredInMb Number
    The minimum MB of memory to allocate to run a model that you assign to an inference component.
    numberOfAcceleratorDevicesRequired Number
    The number of accelerators to allocate to run a model that you assign to an inference component. Accelerators include GPUs and AWS Inferentia.
    numberOfCpuCoresRequired Number
    The number of CPU cores to allocate to run a model that you assign to an inference component.

    InferenceComponentContainerMetricsConfig

    InferenceComponentContainerSpecification

    ArtifactUrl string
    The Amazon S3 path where the model artifacts, which result from model training, are stored. This path must point to a single gzip compressed tar archive (.tar.gz suffix).
    ContainerMetricsConfig Pulumi.AwsNative.SageMaker.Inputs.InferenceComponentContainerMetricsConfig
    DeployedImage Pulumi.AwsNative.SageMaker.Inputs.InferenceComponentDeployedImage
    Environment Dictionary<string, string>
    The environment variables to set in the Docker container. Each key and value in the Environment string-to-string map can have length of up to 1024. We support up to 16 entries in the map.
    Image string
    The Amazon Elastic Container Registry (Amazon ECR) path where the Docker image for the model is stored.
    ArtifactUrl string
    The Amazon S3 path where the model artifacts, which result from model training, are stored. This path must point to a single gzip compressed tar archive (.tar.gz suffix).
    ContainerMetricsConfig InferenceComponentContainerMetricsConfig
    DeployedImage InferenceComponentDeployedImage
    Environment map[string]string
    The environment variables to set in the Docker container. Each key and value in the Environment string-to-string map can have length of up to 1024. We support up to 16 entries in the map.
    Image string
    The Amazon Elastic Container Registry (Amazon ECR) path where the Docker image for the model is stored.
    artifact_url string
    The Amazon S3 path where the model artifacts, which result from model training, are stored. This path must point to a single gzip compressed tar archive (.tar.gz suffix).
    container_metrics_config object
    deployed_image object
    environment map(string)
    The environment variables to set in the Docker container. Each key and value in the Environment string-to-string map can have length of up to 1024. We support up to 16 entries in the map.
    image string
    The Amazon Elastic Container Registry (Amazon ECR) path where the Docker image for the model is stored.
    artifactUrl String
    The Amazon S3 path where the model artifacts, which result from model training, are stored. This path must point to a single gzip compressed tar archive (.tar.gz suffix).
    containerMetricsConfig InferenceComponentContainerMetricsConfig
    deployedImage InferenceComponentDeployedImage
    environment Map<String,String>
    The environment variables to set in the Docker container. Each key and value in the Environment string-to-string map can have length of up to 1024. We support up to 16 entries in the map.
    image String
    The Amazon Elastic Container Registry (Amazon ECR) path where the Docker image for the model is stored.
    artifactUrl string
    The Amazon S3 path where the model artifacts, which result from model training, are stored. This path must point to a single gzip compressed tar archive (.tar.gz suffix).
    containerMetricsConfig InferenceComponentContainerMetricsConfig
    deployedImage InferenceComponentDeployedImage
    environment {[key: string]: string}
    The environment variables to set in the Docker container. Each key and value in the Environment string-to-string map can have length of up to 1024. We support up to 16 entries in the map.
    image string
    The Amazon Elastic Container Registry (Amazon ECR) path where the Docker image for the model is stored.
    artifact_url str
    The Amazon S3 path where the model artifacts, which result from model training, are stored. This path must point to a single gzip compressed tar archive (.tar.gz suffix).
    container_metrics_config InferenceComponentContainerMetricsConfig
    deployed_image InferenceComponentDeployedImage
    environment Mapping[str, str]
    The environment variables to set in the Docker container. Each key and value in the Environment string-to-string map can have length of up to 1024. We support up to 16 entries in the map.
    image str
    The Amazon Elastic Container Registry (Amazon ECR) path where the Docker image for the model is stored.
    artifactUrl String
    The Amazon S3 path where the model artifacts, which result from model training, are stored. This path must point to a single gzip compressed tar archive (.tar.gz suffix).
    containerMetricsConfig Property Map
    deployedImage Property Map
    environment Map<String>
    The environment variables to set in the Docker container. Each key and value in the Environment string-to-string map can have length of up to 1024. We support up to 16 entries in the map.
    image String
    The Amazon Elastic Container Registry (Amazon ECR) path where the Docker image for the model is stored.

    InferenceComponentContainerSpecificationForInstanceType

    InferenceComponentDataCacheConfig

    EnableCaching bool
    Whether the endpoint caches the model artifacts and container image on each instance it provisions for the inference component
    EnableCaching bool
    Whether the endpoint caches the model artifacts and container image on each instance it provisions for the inference component
    enable_caching bool
    Whether the endpoint caches the model artifacts and container image on each instance it provisions for the inference component
    enableCaching Boolean
    Whether the endpoint caches the model artifacts and container image on each instance it provisions for the inference component
    enableCaching boolean
    Whether the endpoint caches the model artifacts and container image on each instance it provisions for the inference component
    enable_caching bool
    Whether the endpoint caches the model artifacts and container image on each instance it provisions for the inference component
    enableCaching Boolean
    Whether the endpoint caches the model artifacts and container image on each instance it provisions for the inference component

    InferenceComponentDeployedImage

    ResolutionTime string
    The date and time when the image path for the model resolved to the ResolvedImage
    ResolvedImage string
    The specific digest path of the image hosted in this ProductionVariant .
    SpecifiedImage string
    The image path you specified when you created the model.
    ResolutionTime string
    The date and time when the image path for the model resolved to the ResolvedImage
    ResolvedImage string
    The specific digest path of the image hosted in this ProductionVariant .
    SpecifiedImage string
    The image path you specified when you created the model.
    resolution_time string
    The date and time when the image path for the model resolved to the ResolvedImage
    resolved_image string
    The specific digest path of the image hosted in this ProductionVariant .
    specified_image string
    The image path you specified when you created the model.
    resolutionTime String
    The date and time when the image path for the model resolved to the ResolvedImage
    resolvedImage String
    The specific digest path of the image hosted in this ProductionVariant .
    specifiedImage String
    The image path you specified when you created the model.
    resolutionTime string
    The date and time when the image path for the model resolved to the ResolvedImage
    resolvedImage string
    The specific digest path of the image hosted in this ProductionVariant .
    specifiedImage string
    The image path you specified when you created the model.
    resolution_time str
    The date and time when the image path for the model resolved to the ResolvedImage
    resolved_image str
    The specific digest path of the image hosted in this ProductionVariant .
    specified_image str
    The image path you specified when you created the model.
    resolutionTime String
    The date and time when the image path for the model resolved to the ResolvedImage
    resolvedImage String
    The specific digest path of the image hosted in this ProductionVariant .
    specifiedImage String
    The image path you specified when you created the model.

    InferenceComponentMetricsEndpoint

    MetricsEndpointPath string
    The path to the Prometheus formatted metrics endpoint exposed by the container
    MetricPublishFrequencyInSeconds int
    The interval, in seconds, at which container metrics scraped from the endpoint are published to Amazon CloudWatch. Valid values per the SageMaker API Reference are 10, 30, 60, 120, 180, 240 and 300; the service validates the value.
    MetricsEndpointPath string
    The path to the Prometheus formatted metrics endpoint exposed by the container
    MetricPublishFrequencyInSeconds int
    The interval, in seconds, at which container metrics scraped from the endpoint are published to Amazon CloudWatch. Valid values per the SageMaker API Reference are 10, 30, 60, 120, 180, 240 and 300; the service validates the value.
    metrics_endpoint_path string
    The path to the Prometheus formatted metrics endpoint exposed by the container
    metric_publish_frequency_in_seconds number
    The interval, in seconds, at which container metrics scraped from the endpoint are published to Amazon CloudWatch. Valid values per the SageMaker API Reference are 10, 30, 60, 120, 180, 240 and 300; the service validates the value.
    metricsEndpointPath String
    The path to the Prometheus formatted metrics endpoint exposed by the container
    metricPublishFrequencyInSeconds Integer
    The interval, in seconds, at which container metrics scraped from the endpoint are published to Amazon CloudWatch. Valid values per the SageMaker API Reference are 10, 30, 60, 120, 180, 240 and 300; the service validates the value.
    metricsEndpointPath string
    The path to the Prometheus formatted metrics endpoint exposed by the container
    metricPublishFrequencyInSeconds number
    The interval, in seconds, at which container metrics scraped from the endpoint are published to Amazon CloudWatch. Valid values per the SageMaker API Reference are 10, 30, 60, 120, 180, 240 and 300; the service validates the value.
    metrics_endpoint_path str
    The path to the Prometheus formatted metrics endpoint exposed by the container
    metric_publish_frequency_in_seconds int
    The interval, in seconds, at which container metrics scraped from the endpoint are published to Amazon CloudWatch. Valid values per the SageMaker API Reference are 10, 30, 60, 120, 180, 240 and 300; the service validates the value.
    metricsEndpointPath String
    The path to the Prometheus formatted metrics endpoint exposed by the container
    metricPublishFrequencyInSeconds Number
    The interval, in seconds, at which container metrics scraped from the endpoint are published to Amazon CloudWatch. Valid values per the SageMaker API Reference are 10, 30, 60, 120, 180, 240 and 300; the service validates the value.

    InferenceComponentPlacementStatus

    InferenceComponentPlacementStrategy

    InferenceComponentRuntimeConfig

    CopyCount int
    The number of runtime copies of the model container to deploy with the inference component. Each copy can serve inference requests.
    CurrentCopyCount int
    DesiredCopyCount int
    PlacementStatus List<Pulumi.AwsNative.SageMaker.Inputs.InferenceComponentPlacementStatus>
    The placement status of the inference component across instance types
    CopyCount int
    The number of runtime copies of the model container to deploy with the inference component. Each copy can serve inference requests.
    CurrentCopyCount int
    DesiredCopyCount int
    PlacementStatus []InferenceComponentPlacementStatus
    The placement status of the inference component across instance types
    copy_count number
    The number of runtime copies of the model container to deploy with the inference component. Each copy can serve inference requests.
    current_copy_count number
    desired_copy_count number
    placement_status list(object)
    The placement status of the inference component across instance types
    copyCount Integer
    The number of runtime copies of the model container to deploy with the inference component. Each copy can serve inference requests.
    currentCopyCount Integer
    desiredCopyCount Integer
    placementStatus List<InferenceComponentPlacementStatus>
    The placement status of the inference component across instance types
    copyCount number
    The number of runtime copies of the model container to deploy with the inference component. Each copy can serve inference requests.
    currentCopyCount number
    desiredCopyCount number
    placementStatus InferenceComponentPlacementStatus[]
    The placement status of the inference component across instance types
    copy_count int
    The number of runtime copies of the model container to deploy with the inference component. Each copy can serve inference requests.
    current_copy_count int
    desired_copy_count int
    placement_status Sequence[InferenceComponentPlacementStatus]
    The placement status of the inference component across instance types
    copyCount Number
    The number of runtime copies of the model container to deploy with the inference component. Each copy can serve inference requests.
    currentCopyCount Number
    desiredCopyCount Number
    placementStatus List<Property Map>
    The placement status of the inference component across instance types

    InferenceComponentSchedulingConfig

    InferenceComponentSpecification

    BaseInferenceComponentName string

    The name of an existing inference component that is to contain the inference component that you're creating with your request.

    Specify this parameter only if your request is meant to create an adapter inference component. An adapter inference component contains the path to an adapter model. The purpose of the adapter model is to tailor the inference output of a base foundation model, which is hosted by the base inference component. The adapter inference component uses the compute resources that you assigned to the base inference component.

    When you create an adapter inference component, use the Container parameter to specify the location of the adapter artifacts. In the parameter value, use the ArtifactUrl parameter of the InferenceComponentContainerSpecification data type.

    Before you can create an adapter inference component, you must have an existing inference component that contains the foundation model that you want to adapt.

    ComputeResourceRequirements Pulumi.AwsNative.SageMaker.Inputs.InferenceComponentComputeResourceRequirements

    The compute resources allocated to run the model, plus any adapter models, that you assign to the inference component.

    Omit this parameter if your request is meant to create an adapter inference component. An adapter inference component is loaded by a base inference component, and it uses the compute resources of the base inference component.

    Container Pulumi.AwsNative.SageMaker.Inputs.InferenceComponentContainerSpecification
    Defines a container that provides the runtime environment for a model that you deploy with an inference component.
    CurrentDataCacheConfig Pulumi.AwsNative.SageMaker.Inputs.InferenceComponentDataCacheConfig
    The data caching configuration actually in effect, including a value the service chose rather than the template: SageMaker enables caching automatically on instance types with more than 232 GiB of local NVMe storage, whether or not DataCacheConfig was set. Returned by Describe and not settable; set DataCacheConfig instead.
    DataCacheConfig Pulumi.AwsNative.SageMaker.Inputs.InferenceComponentDataCacheConfig
    ModelName string
    The name of an existing SageMaker AI model object in your account that you want to deploy with the inference component.
    SchedulingConfig Pulumi.AwsNative.SageMaker.Inputs.InferenceComponentSchedulingConfig
    StartupParameters Pulumi.AwsNative.SageMaker.Inputs.InferenceComponentStartupParameters
    Settings that take effect while the model container starts up.
    BaseInferenceComponentName string

    The name of an existing inference component that is to contain the inference component that you're creating with your request.

    Specify this parameter only if your request is meant to create an adapter inference component. An adapter inference component contains the path to an adapter model. The purpose of the adapter model is to tailor the inference output of a base foundation model, which is hosted by the base inference component. The adapter inference component uses the compute resources that you assigned to the base inference component.

    When you create an adapter inference component, use the Container parameter to specify the location of the adapter artifacts. In the parameter value, use the ArtifactUrl parameter of the InferenceComponentContainerSpecification data type.

    Before you can create an adapter inference component, you must have an existing inference component that contains the foundation model that you want to adapt.

    ComputeResourceRequirements InferenceComponentComputeResourceRequirements

    The compute resources allocated to run the model, plus any adapter models, that you assign to the inference component.

    Omit this parameter if your request is meant to create an adapter inference component. An adapter inference component is loaded by a base inference component, and it uses the compute resources of the base inference component.

    Container InferenceComponentContainerSpecification
    Defines a container that provides the runtime environment for a model that you deploy with an inference component.
    CurrentDataCacheConfig InferenceComponentDataCacheConfig
    The data caching configuration actually in effect, including a value the service chose rather than the template: SageMaker enables caching automatically on instance types with more than 232 GiB of local NVMe storage, whether or not DataCacheConfig was set. Returned by Describe and not settable; set DataCacheConfig instead.
    DataCacheConfig InferenceComponentDataCacheConfig
    ModelName string
    The name of an existing SageMaker AI model object in your account that you want to deploy with the inference component.
    SchedulingConfig InferenceComponentSchedulingConfig
    StartupParameters InferenceComponentStartupParameters
    Settings that take effect while the model container starts up.
    base_inference_component_name string

    The name of an existing inference component that is to contain the inference component that you're creating with your request.

    Specify this parameter only if your request is meant to create an adapter inference component. An adapter inference component contains the path to an adapter model. The purpose of the adapter model is to tailor the inference output of a base foundation model, which is hosted by the base inference component. The adapter inference component uses the compute resources that you assigned to the base inference component.

    When you create an adapter inference component, use the Container parameter to specify the location of the adapter artifacts. In the parameter value, use the ArtifactUrl parameter of the InferenceComponentContainerSpecification data type.

    Before you can create an adapter inference component, you must have an existing inference component that contains the foundation model that you want to adapt.

    compute_resource_requirements object

    The compute resources allocated to run the model, plus any adapter models, that you assign to the inference component.

    Omit this parameter if your request is meant to create an adapter inference component. An adapter inference component is loaded by a base inference component, and it uses the compute resources of the base inference component.

    container object
    Defines a container that provides the runtime environment for a model that you deploy with an inference component.
    current_data_cache_config object
    The data caching configuration actually in effect, including a value the service chose rather than the template: SageMaker enables caching automatically on instance types with more than 232 GiB of local NVMe storage, whether or not DataCacheConfig was set. Returned by Describe and not settable; set DataCacheConfig instead.
    data_cache_config object
    model_name string
    The name of an existing SageMaker AI model object in your account that you want to deploy with the inference component.
    scheduling_config object
    startup_parameters object
    Settings that take effect while the model container starts up.
    baseInferenceComponentName String

    The name of an existing inference component that is to contain the inference component that you're creating with your request.

    Specify this parameter only if your request is meant to create an adapter inference component. An adapter inference component contains the path to an adapter model. The purpose of the adapter model is to tailor the inference output of a base foundation model, which is hosted by the base inference component. The adapter inference component uses the compute resources that you assigned to the base inference component.

    When you create an adapter inference component, use the Container parameter to specify the location of the adapter artifacts. In the parameter value, use the ArtifactUrl parameter of the InferenceComponentContainerSpecification data type.

    Before you can create an adapter inference component, you must have an existing inference component that contains the foundation model that you want to adapt.

    computeResourceRequirements InferenceComponentComputeResourceRequirements

    The compute resources allocated to run the model, plus any adapter models, that you assign to the inference component.

    Omit this parameter if your request is meant to create an adapter inference component. An adapter inference component is loaded by a base inference component, and it uses the compute resources of the base inference component.

    container InferenceComponentContainerSpecification
    Defines a container that provides the runtime environment for a model that you deploy with an inference component.
    currentDataCacheConfig InferenceComponentDataCacheConfig
    The data caching configuration actually in effect, including a value the service chose rather than the template: SageMaker enables caching automatically on instance types with more than 232 GiB of local NVMe storage, whether or not DataCacheConfig was set. Returned by Describe and not settable; set DataCacheConfig instead.
    dataCacheConfig InferenceComponentDataCacheConfig
    modelName String
    The name of an existing SageMaker AI model object in your account that you want to deploy with the inference component.
    schedulingConfig InferenceComponentSchedulingConfig
    startupParameters InferenceComponentStartupParameters
    Settings that take effect while the model container starts up.
    baseInferenceComponentName string

    The name of an existing inference component that is to contain the inference component that you're creating with your request.

    Specify this parameter only if your request is meant to create an adapter inference component. An adapter inference component contains the path to an adapter model. The purpose of the adapter model is to tailor the inference output of a base foundation model, which is hosted by the base inference component. The adapter inference component uses the compute resources that you assigned to the base inference component.

    When you create an adapter inference component, use the Container parameter to specify the location of the adapter artifacts. In the parameter value, use the ArtifactUrl parameter of the InferenceComponentContainerSpecification data type.

    Before you can create an adapter inference component, you must have an existing inference component that contains the foundation model that you want to adapt.

    computeResourceRequirements InferenceComponentComputeResourceRequirements

    The compute resources allocated to run the model, plus any adapter models, that you assign to the inference component.

    Omit this parameter if your request is meant to create an adapter inference component. An adapter inference component is loaded by a base inference component, and it uses the compute resources of the base inference component.

    container InferenceComponentContainerSpecification
    Defines a container that provides the runtime environment for a model that you deploy with an inference component.
    currentDataCacheConfig InferenceComponentDataCacheConfig
    The data caching configuration actually in effect, including a value the service chose rather than the template: SageMaker enables caching automatically on instance types with more than 232 GiB of local NVMe storage, whether or not DataCacheConfig was set. Returned by Describe and not settable; set DataCacheConfig instead.
    dataCacheConfig InferenceComponentDataCacheConfig
    modelName string
    The name of an existing SageMaker AI model object in your account that you want to deploy with the inference component.
    schedulingConfig InferenceComponentSchedulingConfig
    startupParameters InferenceComponentStartupParameters
    Settings that take effect while the model container starts up.
    base_inference_component_name str

    The name of an existing inference component that is to contain the inference component that you're creating with your request.

    Specify this parameter only if your request is meant to create an adapter inference component. An adapter inference component contains the path to an adapter model. The purpose of the adapter model is to tailor the inference output of a base foundation model, which is hosted by the base inference component. The adapter inference component uses the compute resources that you assigned to the base inference component.

    When you create an adapter inference component, use the Container parameter to specify the location of the adapter artifacts. In the parameter value, use the ArtifactUrl parameter of the InferenceComponentContainerSpecification data type.

    Before you can create an adapter inference component, you must have an existing inference component that contains the foundation model that you want to adapt.

    compute_resource_requirements InferenceComponentComputeResourceRequirements

    The compute resources allocated to run the model, plus any adapter models, that you assign to the inference component.

    Omit this parameter if your request is meant to create an adapter inference component. An adapter inference component is loaded by a base inference component, and it uses the compute resources of the base inference component.

    container InferenceComponentContainerSpecification
    Defines a container that provides the runtime environment for a model that you deploy with an inference component.
    current_data_cache_config InferenceComponentDataCacheConfig
    The data caching configuration actually in effect, including a value the service chose rather than the template: SageMaker enables caching automatically on instance types with more than 232 GiB of local NVMe storage, whether or not DataCacheConfig was set. Returned by Describe and not settable; set DataCacheConfig instead.
    data_cache_config InferenceComponentDataCacheConfig
    model_name str
    The name of an existing SageMaker AI model object in your account that you want to deploy with the inference component.
    scheduling_config InferenceComponentSchedulingConfig
    startup_parameters InferenceComponentStartupParameters
    Settings that take effect while the model container starts up.
    baseInferenceComponentName String

    The name of an existing inference component that is to contain the inference component that you're creating with your request.

    Specify this parameter only if your request is meant to create an adapter inference component. An adapter inference component contains the path to an adapter model. The purpose of the adapter model is to tailor the inference output of a base foundation model, which is hosted by the base inference component. The adapter inference component uses the compute resources that you assigned to the base inference component.

    When you create an adapter inference component, use the Container parameter to specify the location of the adapter artifacts. In the parameter value, use the ArtifactUrl parameter of the InferenceComponentContainerSpecification data type.

    Before you can create an adapter inference component, you must have an existing inference component that contains the foundation model that you want to adapt.

    computeResourceRequirements Property Map

    The compute resources allocated to run the model, plus any adapter models, that you assign to the inference component.

    Omit this parameter if your request is meant to create an adapter inference component. An adapter inference component is loaded by a base inference component, and it uses the compute resources of the base inference component.

    container Property Map
    Defines a container that provides the runtime environment for a model that you deploy with an inference component.
    currentDataCacheConfig Property Map
    The data caching configuration actually in effect, including a value the service chose rather than the template: SageMaker enables caching automatically on instance types with more than 232 GiB of local NVMe storage, whether or not DataCacheConfig was set. Returned by Describe and not settable; set DataCacheConfig instead.
    dataCacheConfig Property Map
    modelName String
    The name of an existing SageMaker AI model object in your account that you want to deploy with the inference component.
    schedulingConfig Property Map
    startupParameters Property Map
    Settings that take effect while the model container starts up.

    InferenceComponentSpecificationForInstanceType

    InstanceType string
    ComputeResourceRequirements Pulumi.AwsNative.SageMaker.Inputs.InferenceComponentComputeResourceRequirements
    Container Pulumi.AwsNative.SageMaker.Inputs.InferenceComponentContainerSpecificationForInstanceType
    CurrentDataCacheConfig Pulumi.AwsNative.SageMaker.Inputs.InferenceComponentDataCacheConfig
    The data caching configuration actually in effect for this instance type, including a value the service chose rather than the template: SageMaker enables caching automatically on instance types with more than 232 GiB of local NVMe storage, whether or not DataCacheConfig was set. Returned by Describe and not settable; set DataCacheConfig instead.
    DataCacheConfig Pulumi.AwsNative.SageMaker.Inputs.InferenceComponentDataCacheConfig
    ModelName string
    SchedulingConfig Pulumi.AwsNative.SageMaker.Inputs.InferenceComponentSchedulingConfig
    StartupParameters Pulumi.AwsNative.SageMaker.Inputs.InferenceComponentStartupParameters
    InstanceType string
    ComputeResourceRequirements InferenceComponentComputeResourceRequirements
    Container InferenceComponentContainerSpecificationForInstanceType
    CurrentDataCacheConfig InferenceComponentDataCacheConfig
    The data caching configuration actually in effect for this instance type, including a value the service chose rather than the template: SageMaker enables caching automatically on instance types with more than 232 GiB of local NVMe storage, whether or not DataCacheConfig was set. Returned by Describe and not settable; set DataCacheConfig instead.
    DataCacheConfig InferenceComponentDataCacheConfig
    ModelName string
    SchedulingConfig InferenceComponentSchedulingConfig
    StartupParameters InferenceComponentStartupParameters
    instance_type string
    compute_resource_requirements object
    container object
    current_data_cache_config object
    The data caching configuration actually in effect for this instance type, including a value the service chose rather than the template: SageMaker enables caching automatically on instance types with more than 232 GiB of local NVMe storage, whether or not DataCacheConfig was set. Returned by Describe and not settable; set DataCacheConfig instead.
    data_cache_config object
    model_name string
    scheduling_config object
    startup_parameters object
    instanceType String
    computeResourceRequirements InferenceComponentComputeResourceRequirements
    container InferenceComponentContainerSpecificationForInstanceType
    currentDataCacheConfig InferenceComponentDataCacheConfig
    The data caching configuration actually in effect for this instance type, including a value the service chose rather than the template: SageMaker enables caching automatically on instance types with more than 232 GiB of local NVMe storage, whether or not DataCacheConfig was set. Returned by Describe and not settable; set DataCacheConfig instead.
    dataCacheConfig InferenceComponentDataCacheConfig
    modelName String
    schedulingConfig InferenceComponentSchedulingConfig
    startupParameters InferenceComponentStartupParameters
    instanceType string
    computeResourceRequirements InferenceComponentComputeResourceRequirements
    container InferenceComponentContainerSpecificationForInstanceType
    currentDataCacheConfig InferenceComponentDataCacheConfig
    The data caching configuration actually in effect for this instance type, including a value the service chose rather than the template: SageMaker enables caching automatically on instance types with more than 232 GiB of local NVMe storage, whether or not DataCacheConfig was set. Returned by Describe and not settable; set DataCacheConfig instead.
    dataCacheConfig InferenceComponentDataCacheConfig
    modelName string
    schedulingConfig InferenceComponentSchedulingConfig
    startupParameters InferenceComponentStartupParameters
    instance_type str
    compute_resource_requirements InferenceComponentComputeResourceRequirements
    container InferenceComponentContainerSpecificationForInstanceType
    current_data_cache_config InferenceComponentDataCacheConfig
    The data caching configuration actually in effect for this instance type, including a value the service chose rather than the template: SageMaker enables caching automatically on instance types with more than 232 GiB of local NVMe storage, whether or not DataCacheConfig was set. Returned by Describe and not settable; set DataCacheConfig instead.
    data_cache_config InferenceComponentDataCacheConfig
    model_name str
    scheduling_config InferenceComponentSchedulingConfig
    startup_parameters InferenceComponentStartupParameters
    instanceType String
    computeResourceRequirements Property Map
    container Property Map
    currentDataCacheConfig Property Map
    The data caching configuration actually in effect for this instance type, including a value the service chose rather than the template: SageMaker enables caching automatically on instance types with more than 232 GiB of local NVMe storage, whether or not DataCacheConfig was set. Returned by Describe and not settable; set DataCacheConfig instead.
    dataCacheConfig Property Map
    modelName String
    schedulingConfig Property Map
    startupParameters Property Map

    InferenceComponentStartupParameters

    ContainerStartupHealthCheckTimeoutInSeconds int
    The timeout value, in seconds, for your inference container to pass health check by Amazon S3 Hosting. For more information about health check, see How Your Container Should Respond to Health Check (Ping) Requests .
    ModelDataDownloadTimeoutInSeconds int
    The timeout value, in seconds, to download and extract the model that you want to host from Amazon S3 to the individual inference instance associated with this inference component.
    ContainerStartupHealthCheckTimeoutInSeconds int
    The timeout value, in seconds, for your inference container to pass health check by Amazon S3 Hosting. For more information about health check, see How Your Container Should Respond to Health Check (Ping) Requests .
    ModelDataDownloadTimeoutInSeconds int
    The timeout value, in seconds, to download and extract the model that you want to host from Amazon S3 to the individual inference instance associated with this inference component.
    container_startup_health_check_timeout_in_seconds number
    The timeout value, in seconds, for your inference container to pass health check by Amazon S3 Hosting. For more information about health check, see How Your Container Should Respond to Health Check (Ping) Requests .
    model_data_download_timeout_in_seconds number
    The timeout value, in seconds, to download and extract the model that you want to host from Amazon S3 to the individual inference instance associated with this inference component.
    containerStartupHealthCheckTimeoutInSeconds Integer
    The timeout value, in seconds, for your inference container to pass health check by Amazon S3 Hosting. For more information about health check, see How Your Container Should Respond to Health Check (Ping) Requests .
    modelDataDownloadTimeoutInSeconds Integer
    The timeout value, in seconds, to download and extract the model that you want to host from Amazon S3 to the individual inference instance associated with this inference component.
    containerStartupHealthCheckTimeoutInSeconds number
    The timeout value, in seconds, for your inference container to pass health check by Amazon S3 Hosting. For more information about health check, see How Your Container Should Respond to Health Check (Ping) Requests .
    modelDataDownloadTimeoutInSeconds number
    The timeout value, in seconds, to download and extract the model that you want to host from Amazon S3 to the individual inference instance associated with this inference component.
    container_startup_health_check_timeout_in_seconds int
    The timeout value, in seconds, for your inference container to pass health check by Amazon S3 Hosting. For more information about health check, see How Your Container Should Respond to Health Check (Ping) Requests .
    model_data_download_timeout_in_seconds int
    The timeout value, in seconds, to download and extract the model that you want to host from Amazon S3 to the individual inference instance associated with this inference component.
    containerStartupHealthCheckTimeoutInSeconds Number
    The timeout value, in seconds, for your inference container to pass health check by Amazon S3 Hosting. For more information about health check, see How Your Container Should Respond to Health Check (Ping) Requests .
    modelDataDownloadTimeoutInSeconds Number
    The timeout value, in seconds, to download and extract the model that you want to host from Amazon S3 to the individual inference instance associated with this inference component.

    InferenceComponentStatus

    Tag

    Key string
    The key name of the tag
    Value string
    The value of the tag
    Key string
    The key name of the tag
    Value string
    The value of the tag
    key string
    The key name of the tag
    value string
    The value of the tag
    key String
    The key name of the tag
    value String
    The value of the tag
    key string
    The key name of the tag
    value string
    The value of the tag
    key str
    The key name of the tag
    value str
    The value of the tag
    key String
    The key name of the tag
    value String
    The value of the tag

    Package Details

    Repository
    AWS Native pulumi/pulumi-aws-native
    License
    Apache-2.0
    aws-native logo aws-native logo

    We recommend new projects start with resources from the AWS provider.

    Viewing docs for AWS Cloud Control v1.81.0
    published on Monday, Sep 28, 2026 by Pulumi

      Try Pulumi Cloud free.
      Your team will thank you.

      Start free trial