We recommend new projects start with resources from the AWS provider.
published on Monday, Sep 28, 2026 by Pulumi
We recommend new projects start with resources from the AWS provider.
published on Monday, Sep 28, 2026 by Pulumi
Resource Type definition for AWS::SageMaker::EndpointConfig
Create EndpointConfig Resource
Resources are created with functions called constructors. To learn more about declaring and configuring resources, see Resources.
Constructor syntax
new EndpointConfig(name: string, args: EndpointConfigArgs, opts?: CustomResourceOptions);@overload
def EndpointConfig(resource_name: str,
args: EndpointConfigArgs,
opts: Optional[ResourceOptions] = None)
@overload
def EndpointConfig(resource_name: str,
opts: Optional[ResourceOptions] = None,
production_variants: Optional[Sequence[EndpointConfigProductionVariantArgs]] = None,
async_inference_config: Optional[EndpointConfigAsyncInferenceConfigArgs] = None,
data_capture_config: Optional[EndpointConfigDataCaptureConfigArgs] = None,
enable_network_isolation: Optional[bool] = None,
endpoint_config_name: Optional[str] = None,
execution_role_arn: Optional[str] = None,
explainer_config: Optional[EndpointConfigExplainerConfigArgs] = None,
kms_key_id: Optional[str] = None,
metrics_config: Optional[EndpointConfigMetricsConfigArgs] = None,
shadow_production_variants: Optional[Sequence[EndpointConfigProductionVariantArgs]] = None,
tags: Optional[Sequence[_root_inputs.TagArgs]] = None,
vpc_config: Optional[EndpointConfigVpcConfigArgs] = None)func NewEndpointConfig(ctx *Context, name string, args EndpointConfigArgs, opts ...ResourceOption) (*EndpointConfig, error)public EndpointConfig(string name, EndpointConfigArgs args, CustomResourceOptions? opts = null)
public EndpointConfig(String name, EndpointConfigArgs args)
public EndpointConfig(String name, EndpointConfigArgs args, CustomResourceOptions options)
type: aws-native:sagemaker:EndpointConfig
properties: # The arguments to resource properties.
options: # Bag of options to control resource's behavior.
resource "aws-native_sagemaker_endpoint_config" "name" {
# resource properties
}Parameters
- name string
- The unique name of the resource.
- args EndpointConfigArgs
- The arguments to resource properties.
- opts CustomResourceOptions
- Bag of options to control resource's behavior.
- resource_name str
- The unique name of the resource.
- args EndpointConfigArgs
- The arguments to resource properties.
- opts ResourceOptions
- Bag of options to control resource's behavior.
- ctx Context
- Context object for the current deployment.
- name string
- The unique name of the resource.
- args EndpointConfigArgs
- The arguments to resource properties.
- opts ResourceOption
- Bag of options to control resource's behavior.
- name string
- The unique name of the resource.
- args EndpointConfigArgs
- The arguments to resource properties.
- opts CustomResourceOptions
- Bag of options to control resource's behavior.
- name String
- The unique name of the resource.
- args EndpointConfigArgs
- The arguments to resource properties.
- options CustomResourceOptions
- Bag of options to control resource's behavior.
EndpointConfig Resource Properties
To learn more about resource properties and how to use them, see Inputs and Outputs in the Architecture and Concepts docs.
Inputs
In Python, inputs that are objects can be passed either as argument classes or as dictionary literals.
The EndpointConfig resource accepts the following input properties:
- Production
Variants List<Pulumi.Aws Native. Sage Maker. Inputs. Endpoint Config Production Variant> - A list of ProductionVariant objects, one for each model that you want to host at this endpoint.
- Async
Inference Pulumi.Config Aws Native. Sage Maker. Inputs. Endpoint Config Async Inference Config - Specifies configuration for how an endpoint performs asynchronous inference.
- Data
Capture Pulumi.Config Aws Native. Sage Maker. Inputs. Endpoint Config Data Capture Config - Specifies how to capture endpoint data for model monitor. The data capture configuration applies to all production variants hosted at the endpoint.
- Enable
Network boolIsolation - Sets whether all model containers deployed to the endpoint are isolated. If they are, no inbound or outbound network calls can be made to or from the model containers.
- Endpoint
Config stringName - The name of the endpoint configuration.
- Execution
Role stringArn - The Amazon Resource Name (ARN) of an IAM role that Amazon SageMaker AI can assume to perform actions on your behalf.
- Explainer
Config Pulumi.Aws Native. Sage Maker. Inputs. Endpoint Config Explainer Config - A parameter to activate explainers.
- Kms
Key stringId - The Amazon Resource Name (ARN) of an AWS Key Management Service key that Amazon SageMaker uses to encrypt data on the storage volume attached to the ML compute instance that hosts the endpoint.
- Metrics
Config Pulumi.Aws Native. Sage Maker. Inputs. Endpoint Config Metrics Config - Specifies the metrics that the endpoint publishes to Amazon CloudWatch, the frequency of publication, and whether to enable enhanced or detailed observability metrics.
- Shadow
Production List<Pulumi.Variants Aws Native. Sage Maker. Inputs. Endpoint Config Production Variant> - Array of ProductionVariant objects. There is one for each model that you want to host at this endpoint in shadow mode with production traffic replicated from the model specified on ProductionVariants. If you use this field, you can only specify one variant for ProductionVariants and one variant for ShadowProductionVariants.
-
List<Pulumi.
Aws Native. Inputs. Tag> - A list of key-value pairs to apply to this resource.
- Vpc
Config Pulumi.Aws Native. Sage Maker. Inputs. Endpoint Config Vpc Config - Specifies an Amazon Virtual Private Cloud (VPC) that your SageMaker jobs, hosted models, and compute resources have access to. You can control access to and from your resources by configuring a VPC.
- Production
Variants []EndpointConfig Production Variant Args - A list of ProductionVariant objects, one for each model that you want to host at this endpoint.
- Async
Inference EndpointConfig Config Async Inference Config Args - Specifies configuration for how an endpoint performs asynchronous inference.
- Data
Capture EndpointConfig Config Data Capture Config Args - Specifies how to capture endpoint data for model monitor. The data capture configuration applies to all production variants hosted at the endpoint.
- Enable
Network boolIsolation - Sets whether all model containers deployed to the endpoint are isolated. If they are, no inbound or outbound network calls can be made to or from the model containers.
- Endpoint
Config stringName - The name of the endpoint configuration.
- Execution
Role stringArn - The Amazon Resource Name (ARN) of an IAM role that Amazon SageMaker AI can assume to perform actions on your behalf.
- Explainer
Config EndpointConfig Explainer Config Args - A parameter to activate explainers.
- Kms
Key stringId - The Amazon Resource Name (ARN) of an AWS Key Management Service key that Amazon SageMaker uses to encrypt data on the storage volume attached to the ML compute instance that hosts the endpoint.
- Metrics
Config EndpointConfig Metrics Config Args - Specifies the metrics that the endpoint publishes to Amazon CloudWatch, the frequency of publication, and whether to enable enhanced or detailed observability metrics.
- Shadow
Production []EndpointVariants Config Production Variant Args - Array of ProductionVariant objects. There is one for each model that you want to host at this endpoint in shadow mode with production traffic replicated from the model specified on ProductionVariants. If you use this field, you can only specify one variant for ProductionVariants and one variant for ShadowProductionVariants.
-
Tag
Args - A list of key-value pairs to apply to this resource.
- Vpc
Config EndpointConfig Vpc Config Args - Specifies an Amazon Virtual Private Cloud (VPC) that your SageMaker jobs, hosted models, and compute resources have access to. You can control access to and from your resources by configuring a VPC.
- production_
variants list(object) - A list of ProductionVariant objects, one for each model that you want to host at this endpoint.
- async_
inference_ objectconfig - Specifies configuration for how an endpoint performs asynchronous inference.
- data_
capture_ objectconfig - Specifies how to capture endpoint data for model monitor. The data capture configuration applies to all production variants hosted at the endpoint.
- enable_
network_ boolisolation - Sets whether all model containers deployed to the endpoint are isolated. If they are, no inbound or outbound network calls can be made to or from the model containers.
- endpoint_
config_ stringname - The name of the endpoint configuration.
- execution_
role_ stringarn - The Amazon Resource Name (ARN) of an IAM role that Amazon SageMaker AI can assume to perform actions on your behalf.
- explainer_
config object - A parameter to activate explainers.
- kms_
key_ stringid - The Amazon Resource Name (ARN) of an AWS Key Management Service key that Amazon SageMaker uses to encrypt data on the storage volume attached to the ML compute instance that hosts the endpoint.
- metrics_
config object - Specifies the metrics that the endpoint publishes to Amazon CloudWatch, the frequency of publication, and whether to enable enhanced or detailed observability metrics.
- shadow_
production_ list(object)variants - Array of ProductionVariant objects. There is one for each model that you want to host at this endpoint in shadow mode with production traffic replicated from the model specified on ProductionVariants. If you use this field, you can only specify one variant for ProductionVariants and one variant for ShadowProductionVariants.
- list(object)
- A list of key-value pairs to apply to this resource.
- vpc_
config object - Specifies an Amazon Virtual Private Cloud (VPC) that your SageMaker jobs, hosted models, and compute resources have access to. You can control access to and from your resources by configuring a VPC.
- production
Variants List<EndpointConfig Production Variant> - A list of ProductionVariant objects, one for each model that you want to host at this endpoint.
- async
Inference EndpointConfig Config Async Inference Config - Specifies configuration for how an endpoint performs asynchronous inference.
- data
Capture EndpointConfig Config Data Capture Config - Specifies how to capture endpoint data for model monitor. The data capture configuration applies to all production variants hosted at the endpoint.
- enable
Network BooleanIsolation - Sets whether all model containers deployed to the endpoint are isolated. If they are, no inbound or outbound network calls can be made to or from the model containers.
- endpoint
Config StringName - The name of the endpoint configuration.
- execution
Role StringArn - The Amazon Resource Name (ARN) of an IAM role that Amazon SageMaker AI can assume to perform actions on your behalf.
- explainer
Config EndpointConfig Explainer Config - A parameter to activate explainers.
- kms
Key StringId - The Amazon Resource Name (ARN) of an AWS Key Management Service key that Amazon SageMaker uses to encrypt data on the storage volume attached to the ML compute instance that hosts the endpoint.
- metrics
Config EndpointConfig Metrics Config - Specifies the metrics that the endpoint publishes to Amazon CloudWatch, the frequency of publication, and whether to enable enhanced or detailed observability metrics.
- shadow
Production List<EndpointVariants Config Production Variant> - Array of ProductionVariant objects. There is one for each model that you want to host at this endpoint in shadow mode with production traffic replicated from the model specified on ProductionVariants. If you use this field, you can only specify one variant for ProductionVariants and one variant for ShadowProductionVariants.
- List<Tag>
- A list of key-value pairs to apply to this resource.
- vpc
Config EndpointConfig Vpc Config - Specifies an Amazon Virtual Private Cloud (VPC) that your SageMaker jobs, hosted models, and compute resources have access to. You can control access to and from your resources by configuring a VPC.
- production
Variants EndpointConfig Production Variant[] - A list of ProductionVariant objects, one for each model that you want to host at this endpoint.
- async
Inference EndpointConfig Config Async Inference Config - Specifies configuration for how an endpoint performs asynchronous inference.
- data
Capture EndpointConfig Config Data Capture Config - Specifies how to capture endpoint data for model monitor. The data capture configuration applies to all production variants hosted at the endpoint.
- enable
Network booleanIsolation - Sets whether all model containers deployed to the endpoint are isolated. If they are, no inbound or outbound network calls can be made to or from the model containers.
- endpoint
Config stringName - The name of the endpoint configuration.
- execution
Role stringArn - The Amazon Resource Name (ARN) of an IAM role that Amazon SageMaker AI can assume to perform actions on your behalf.
- explainer
Config EndpointConfig Explainer Config - A parameter to activate explainers.
- kms
Key stringId - The Amazon Resource Name (ARN) of an AWS Key Management Service key that Amazon SageMaker uses to encrypt data on the storage volume attached to the ML compute instance that hosts the endpoint.
- metrics
Config EndpointConfig Metrics Config - Specifies the metrics that the endpoint publishes to Amazon CloudWatch, the frequency of publication, and whether to enable enhanced or detailed observability metrics.
- shadow
Production EndpointVariants Config Production Variant[] - Array of ProductionVariant objects. There is one for each model that you want to host at this endpoint in shadow mode with production traffic replicated from the model specified on ProductionVariants. If you use this field, you can only specify one variant for ProductionVariants and one variant for ShadowProductionVariants.
- Tag[]
- A list of key-value pairs to apply to this resource.
- vpc
Config EndpointConfig Vpc Config - Specifies an Amazon Virtual Private Cloud (VPC) that your SageMaker jobs, hosted models, and compute resources have access to. You can control access to and from your resources by configuring a VPC.
- production_
variants Sequence[EndpointConfig Production Variant Args] - A list of ProductionVariant objects, one for each model that you want to host at this endpoint.
- async_
inference_ Endpointconfig Config Async Inference Config Args - Specifies configuration for how an endpoint performs asynchronous inference.
- data_
capture_ Endpointconfig Config Data Capture Config Args - Specifies how to capture endpoint data for model monitor. The data capture configuration applies to all production variants hosted at the endpoint.
- enable_
network_ boolisolation - Sets whether all model containers deployed to the endpoint are isolated. If they are, no inbound or outbound network calls can be made to or from the model containers.
- endpoint_
config_ strname - The name of the endpoint configuration.
- execution_
role_ strarn - The Amazon Resource Name (ARN) of an IAM role that Amazon SageMaker AI can assume to perform actions on your behalf.
- explainer_
config EndpointConfig Explainer Config Args - A parameter to activate explainers.
- kms_
key_ strid - The Amazon Resource Name (ARN) of an AWS Key Management Service key that Amazon SageMaker uses to encrypt data on the storage volume attached to the ML compute instance that hosts the endpoint.
- metrics_
config EndpointConfig Metrics Config Args - Specifies the metrics that the endpoint publishes to Amazon CloudWatch, the frequency of publication, and whether to enable enhanced or detailed observability metrics.
- shadow_
production_ Sequence[Endpointvariants Config Production Variant Args] - Array of ProductionVariant objects. There is one for each model that you want to host at this endpoint in shadow mode with production traffic replicated from the model specified on ProductionVariants. If you use this field, you can only specify one variant for ProductionVariants and one variant for ShadowProductionVariants.
-
Sequence[Tag
Args] - A list of key-value pairs to apply to this resource.
- vpc_
config EndpointConfig Vpc Config Args - Specifies an Amazon Virtual Private Cloud (VPC) that your SageMaker jobs, hosted models, and compute resources have access to. You can control access to and from your resources by configuring a VPC.
- production
Variants List<Property Map> - A list of ProductionVariant objects, one for each model that you want to host at this endpoint.
- async
Inference Property MapConfig - Specifies configuration for how an endpoint performs asynchronous inference.
- data
Capture Property MapConfig - Specifies how to capture endpoint data for model monitor. The data capture configuration applies to all production variants hosted at the endpoint.
- enable
Network BooleanIsolation - Sets whether all model containers deployed to the endpoint are isolated. If they are, no inbound or outbound network calls can be made to or from the model containers.
- endpoint
Config StringName - The name of the endpoint configuration.
- execution
Role StringArn - The Amazon Resource Name (ARN) of an IAM role that Amazon SageMaker AI can assume to perform actions on your behalf.
- explainer
Config Property Map - A parameter to activate explainers.
- kms
Key StringId - The Amazon Resource Name (ARN) of an AWS Key Management Service key that Amazon SageMaker uses to encrypt data on the storage volume attached to the ML compute instance that hosts the endpoint.
- metrics
Config Property Map - Specifies the metrics that the endpoint publishes to Amazon CloudWatch, the frequency of publication, and whether to enable enhanced or detailed observability metrics.
- shadow
Production List<Property Map>Variants - Array of ProductionVariant objects. There is one for each model that you want to host at this endpoint in shadow mode with production traffic replicated from the model specified on ProductionVariants. If you use this field, you can only specify one variant for ProductionVariants and one variant for ShadowProductionVariants.
- List<Property Map>
- A list of key-value pairs to apply to this resource.
- vpc
Config Property Map - Specifies an Amazon Virtual Private Cloud (VPC) that your SageMaker jobs, hosted models, and compute resources have access to. You can control access to and from your resources by configuring a VPC.
Outputs
All input properties are implicitly available as output properties. Additionally, the EndpointConfig resource produces the following output properties:
- Endpoint
Config stringArn - The Amazon Resource Name (ARN) of the endpoint configuration.
- Id string
- The provider-assigned unique ID for this managed resource.
- Endpoint
Config stringArn - The Amazon Resource Name (ARN) of the endpoint configuration.
- Id string
- The provider-assigned unique ID for this managed resource.
- endpoint_
config_ stringarn - The Amazon Resource Name (ARN) of the endpoint configuration.
- id string
- The provider-assigned unique ID for this managed resource.
- endpoint
Config StringArn - The Amazon Resource Name (ARN) of the endpoint configuration.
- id String
- The provider-assigned unique ID for this managed resource.
- endpoint
Config stringArn - The Amazon Resource Name (ARN) of the endpoint configuration.
- id string
- The provider-assigned unique ID for this managed resource.
- endpoint_
config_ strarn - The Amazon Resource Name (ARN) of the endpoint configuration.
- id str
- The provider-assigned unique ID for this managed resource.
- endpoint
Config StringArn - The Amazon Resource Name (ARN) of the endpoint configuration.
- id String
- The provider-assigned unique ID for this managed resource.
Supporting Types
EndpointConfigAsyncInferenceClientConfig, EndpointConfigAsyncInferenceClientConfigArgs
Configures the behavior of the client used by SageMaker to interact with the model container during asynchronous inference.- Max
Concurrent intInvocations Per Instance - The maximum number of concurrent requests sent by the SageMaker client to the model container. If no value is provided, SageMaker will choose an optimal value for you.
- Max
Concurrent intInvocations Per Instance - The maximum number of concurrent requests sent by the SageMaker client to the model container. If no value is provided, SageMaker will choose an optimal value for you.
- max_
concurrent_ numberinvocations_ per_ instance - The maximum number of concurrent requests sent by the SageMaker client to the model container. If no value is provided, SageMaker will choose an optimal value for you.
- max
Concurrent IntegerInvocations Per Instance - The maximum number of concurrent requests sent by the SageMaker client to the model container. If no value is provided, SageMaker will choose an optimal value for you.
- max
Concurrent numberInvocations Per Instance - The maximum number of concurrent requests sent by the SageMaker client to the model container. If no value is provided, SageMaker will choose an optimal value for you.
- max_
concurrent_ intinvocations_ per_ instance - The maximum number of concurrent requests sent by the SageMaker client to the model container. If no value is provided, SageMaker will choose an optimal value for you.
- max
Concurrent NumberInvocations Per Instance - The maximum number of concurrent requests sent by the SageMaker client to the model container. If no value is provided, SageMaker will choose an optimal value for you.
EndpointConfigAsyncInferenceConfig, EndpointConfigAsyncInferenceConfigArgs
Specifies configuration for how an endpoint performs asynchronous inference.- Output
Config Pulumi.Aws Native. Sage Maker. Inputs. Endpoint Config Async Inference Output Config - Specifies the configuration for asynchronous inference invocation outputs.
- Client
Config Pulumi.Aws Native. Sage Maker. Inputs. Endpoint Config Async Inference Client Config - Configures the behavior of the client used by SageMaker to interact with the model container during asynchronous inference.
- Output
Config EndpointConfig Async Inference Output Config - Specifies the configuration for asynchronous inference invocation outputs.
- Client
Config EndpointConfig Async Inference Client Config - Configures the behavior of the client used by SageMaker to interact with the model container during asynchronous inference.
- output_
config object - Specifies the configuration for asynchronous inference invocation outputs.
- client_
config object - Configures the behavior of the client used by SageMaker to interact with the model container during asynchronous inference.
- output
Config EndpointConfig Async Inference Output Config - Specifies the configuration for asynchronous inference invocation outputs.
- client
Config EndpointConfig Async Inference Client Config - Configures the behavior of the client used by SageMaker to interact with the model container during asynchronous inference.
- output
Config EndpointConfig Async Inference Output Config - Specifies the configuration for asynchronous inference invocation outputs.
- client
Config EndpointConfig Async Inference Client Config - Configures the behavior of the client used by SageMaker to interact with the model container during asynchronous inference.
- output_
config EndpointConfig Async Inference Output Config - Specifies the configuration for asynchronous inference invocation outputs.
- client_
config EndpointConfig Async Inference Client Config - Configures the behavior of the client used by SageMaker to interact with the model container during asynchronous inference.
- output
Config Property Map - Specifies the configuration for asynchronous inference invocation outputs.
- client
Config Property Map - Configures the behavior of the client used by SageMaker to interact with the model container during asynchronous inference.
EndpointConfigAsyncInferenceNotificationConfig, EndpointConfigAsyncInferenceNotificationConfigArgs
Specifies the configuration for notifications of inference results for asynchronous inference.- Error
Topic string - Amazon SNS topic to post a notification to when an inference fails. If no topic is provided, no notification is sent on failure.
- Include
Inference List<string>Response In - The Amazon SNS topics where you want the inference response to be included.
- Success
Topic string - Amazon SNS topic to post a notification to when an inference completes successfully. If no topic is provided, no notification is sent on success.
- Error
Topic string - Amazon SNS topic to post a notification to when an inference fails. If no topic is provided, no notification is sent on failure.
- Include
Inference []stringResponse In - The Amazon SNS topics where you want the inference response to be included.
- Success
Topic string - Amazon SNS topic to post a notification to when an inference completes successfully. If no topic is provided, no notification is sent on success.
- error_
topic string - Amazon SNS topic to post a notification to when an inference fails. If no topic is provided, no notification is sent on failure.
- include_
inference_ list(string)response_ in - The Amazon SNS topics where you want the inference response to be included.
- success_
topic string - Amazon SNS topic to post a notification to when an inference completes successfully. If no topic is provided, no notification is sent on success.
- error
Topic String - Amazon SNS topic to post a notification to when an inference fails. If no topic is provided, no notification is sent on failure.
- include
Inference List<String>Response In - The Amazon SNS topics where you want the inference response to be included.
- success
Topic String - Amazon SNS topic to post a notification to when an inference completes successfully. If no topic is provided, no notification is sent on success.
- error
Topic string - Amazon SNS topic to post a notification to when an inference fails. If no topic is provided, no notification is sent on failure.
- include
Inference string[]Response In - The Amazon SNS topics where you want the inference response to be included.
- success
Topic string - Amazon SNS topic to post a notification to when an inference completes successfully. If no topic is provided, no notification is sent on success.
- error_
topic str - Amazon SNS topic to post a notification to when an inference fails. If no topic is provided, no notification is sent on failure.
- include_
inference_ Sequence[str]response_ in - The Amazon SNS topics where you want the inference response to be included.
- success_
topic str - Amazon SNS topic to post a notification to when an inference completes successfully. If no topic is provided, no notification is sent on success.
- error
Topic String - Amazon SNS topic to post a notification to when an inference fails. If no topic is provided, no notification is sent on failure.
- include
Inference List<String>Response In - The Amazon SNS topics where you want the inference response to be included.
- success
Topic String - Amazon SNS topic to post a notification to when an inference completes successfully. If no topic is provided, no notification is sent on success.
EndpointConfigAsyncInferenceOutputConfig, EndpointConfigAsyncInferenceOutputConfigArgs
Specifies the configuration for asynchronous inference invocation outputs.- Kms
Key stringId - The AWS Key Management Service (AWS KMS) key that Amazon SageMaker uses to encrypt the asynchronous inference output in Amazon S3.
- Notification
Config Pulumi.Aws Native. Sage Maker. Inputs. Endpoint Config Async Inference Notification Config - Specifies the configuration for notifications of inference results for asynchronous inference.
- S3Failure
Path string - The Amazon S3 location to upload failure inference responses to.
- S3Output
Path string - The Amazon S3 location to upload inference responses to.
- Kms
Key stringId - The AWS Key Management Service (AWS KMS) key that Amazon SageMaker uses to encrypt the asynchronous inference output in Amazon S3.
- Notification
Config EndpointConfig Async Inference Notification Config - Specifies the configuration for notifications of inference results for asynchronous inference.
- S3Failure
Path string - The Amazon S3 location to upload failure inference responses to.
- S3Output
Path string - The Amazon S3 location to upload inference responses to.
- kms_
key_ stringid - The AWS Key Management Service (AWS KMS) key that Amazon SageMaker uses to encrypt the asynchronous inference output in Amazon S3.
- notification_
config object - Specifies the configuration for notifications of inference results for asynchronous inference.
- s3_
failure_ stringpath - The Amazon S3 location to upload failure inference responses to.
- s3_
output_ stringpath - The Amazon S3 location to upload inference responses to.
- kms
Key StringId - The AWS Key Management Service (AWS KMS) key that Amazon SageMaker uses to encrypt the asynchronous inference output in Amazon S3.
- notification
Config EndpointConfig Async Inference Notification Config - Specifies the configuration for notifications of inference results for asynchronous inference.
- s3Failure
Path String - The Amazon S3 location to upload failure inference responses to.
- s3Output
Path String - The Amazon S3 location to upload inference responses to.
- kms
Key stringId - The AWS Key Management Service (AWS KMS) key that Amazon SageMaker uses to encrypt the asynchronous inference output in Amazon S3.
- notification
Config EndpointConfig Async Inference Notification Config - Specifies the configuration for notifications of inference results for asynchronous inference.
- s3Failure
Path string - The Amazon S3 location to upload failure inference responses to.
- s3Output
Path string - The Amazon S3 location to upload inference responses to.
- kms_
key_ strid - The AWS Key Management Service (AWS KMS) key that Amazon SageMaker uses to encrypt the asynchronous inference output in Amazon S3.
- notification_
config EndpointConfig Async Inference Notification Config - Specifies the configuration for notifications of inference results for asynchronous inference.
- s3_
failure_ strpath - The Amazon S3 location to upload failure inference responses to.
- s3_
output_ strpath - The Amazon S3 location to upload inference responses to.
- kms
Key StringId - The AWS Key Management Service (AWS KMS) key that Amazon SageMaker uses to encrypt the asynchronous inference output in Amazon S3.
- notification
Config Property Map - Specifies the configuration for notifications of inference results for asynchronous inference.
- s3Failure
Path String - The Amazon S3 location to upload failure inference responses to.
- s3Output
Path String - The Amazon S3 location to upload inference responses to.
EndpointConfigCapacityReservationConfig, EndpointConfigCapacityReservationConfigArgs
Settings for the capacity reservation for the compute instances that SageMaker AI reserves for an endpoint.- Capacity
Reservation stringPreference - Options that you can choose for the capacity reservation.
- Ml
Reservation stringArn - The Amazon Resource Name (ARN) that uniquely identifies the ML capacity reservation that SageMaker AI applies when it deploys the endpoint.
- Capacity
Reservation stringPreference - Options that you can choose for the capacity reservation.
- Ml
Reservation stringArn - The Amazon Resource Name (ARN) that uniquely identifies the ML capacity reservation that SageMaker AI applies when it deploys the endpoint.
- capacity_
reservation_ stringpreference - Options that you can choose for the capacity reservation.
- ml_
reservation_ stringarn - The Amazon Resource Name (ARN) that uniquely identifies the ML capacity reservation that SageMaker AI applies when it deploys the endpoint.
- capacity
Reservation StringPreference - Options that you can choose for the capacity reservation.
- ml
Reservation StringArn - The Amazon Resource Name (ARN) that uniquely identifies the ML capacity reservation that SageMaker AI applies when it deploys the endpoint.
- capacity
Reservation stringPreference - Options that you can choose for the capacity reservation.
- ml
Reservation stringArn - The Amazon Resource Name (ARN) that uniquely identifies the ML capacity reservation that SageMaker AI applies when it deploys the endpoint.
- capacity_
reservation_ strpreference - Options that you can choose for the capacity reservation.
- ml_
reservation_ strarn - The Amazon Resource Name (ARN) that uniquely identifies the ML capacity reservation that SageMaker AI applies when it deploys the endpoint.
- capacity
Reservation StringPreference - Options that you can choose for the capacity reservation.
- ml
Reservation StringArn - The Amazon Resource Name (ARN) that uniquely identifies the ML capacity reservation that SageMaker AI applies when it deploys the endpoint.
EndpointConfigCaptureContentTypeHeader, EndpointConfigCaptureContentTypeHeaderArgs
Specifies the JSON and CSV content types of the data that the endpoint captures.- Csv
Content List<string>Types - A list of the CSV content types of the data that the endpoint captures. For the endpoint to capture the data, you must also specify the content type when you invoke the endpoint.
- Json
Content List<string>Types - A list of the JSON content types of the data that the endpoint captures. For the endpoint to capture the data, you must also specify the content type when you invoke the endpoint.
- Csv
Content []stringTypes - A list of the CSV content types of the data that the endpoint captures. For the endpoint to capture the data, you must also specify the content type when you invoke the endpoint.
- Json
Content []stringTypes - A list of the JSON content types of the data that the endpoint captures. For the endpoint to capture the data, you must also specify the content type when you invoke the endpoint.
- csv_
content_ list(string)types - A list of the CSV content types of the data that the endpoint captures. For the endpoint to capture the data, you must also specify the content type when you invoke the endpoint.
- json_
content_ list(string)types - A list of the JSON content types of the data that the endpoint captures. For the endpoint to capture the data, you must also specify the content type when you invoke the endpoint.
- csv
Content List<String>Types - A list of the CSV content types of the data that the endpoint captures. For the endpoint to capture the data, you must also specify the content type when you invoke the endpoint.
- json
Content List<String>Types - A list of the JSON content types of the data that the endpoint captures. For the endpoint to capture the data, you must also specify the content type when you invoke the endpoint.
- csv
Content string[]Types - A list of the CSV content types of the data that the endpoint captures. For the endpoint to capture the data, you must also specify the content type when you invoke the endpoint.
- json
Content string[]Types - A list of the JSON content types of the data that the endpoint captures. For the endpoint to capture the data, you must also specify the content type when you invoke the endpoint.
- csv_
content_ Sequence[str]types - A list of the CSV content types of the data that the endpoint captures. For the endpoint to capture the data, you must also specify the content type when you invoke the endpoint.
- json_
content_ Sequence[str]types - A list of the JSON content types of the data that the endpoint captures. For the endpoint to capture the data, you must also specify the content type when you invoke the endpoint.
- csv
Content List<String>Types - A list of the CSV content types of the data that the endpoint captures. For the endpoint to capture the data, you must also specify the content type when you invoke the endpoint.
- json
Content List<String>Types - A list of the JSON content types of the data that the endpoint captures. For the endpoint to capture the data, you must also specify the content type when you invoke the endpoint.
EndpointConfigCaptureOption, EndpointConfigCaptureOptionArgs
Specifies whether the endpoint captures input data or output data.- Capture
Mode string - Specifies whether the endpoint captures input data or output data.
- Capture
Mode string - Specifies whether the endpoint captures input data or output data.
- capture_
mode string - Specifies whether the endpoint captures input data or output data.
- capture
Mode String - Specifies whether the endpoint captures input data or output data.
- capture
Mode string - Specifies whether the endpoint captures input data or output data.
- capture_
mode str - Specifies whether the endpoint captures input data or output data.
- capture
Mode String - Specifies whether the endpoint captures input data or output data.
EndpointConfigClarifyExplainerConfig, EndpointConfigClarifyExplainerConfigArgs
The configuration parameters for the SageMaker Clarify explainer.- Shap
Config Pulumi.Aws Native. Sage Maker. Inputs. Endpoint Config Clarify Shap Config - The configuration for SHAP analysis.
- Enable
Explanations string - A JMESPath boolean expression used to filter which records to explain. Explanations are activated by default.
- Inference
Config Pulumi.Aws Native. Sage Maker. Inputs. Endpoint Config Clarify Inference Config - The inference configuration parameter for the model container.
- Shap
Config EndpointConfig Clarify Shap Config - The configuration for SHAP analysis.
- Enable
Explanations string - A JMESPath boolean expression used to filter which records to explain. Explanations are activated by default.
- Inference
Config EndpointConfig Clarify Inference Config - The inference configuration parameter for the model container.
- shap_
config object - The configuration for SHAP analysis.
- enable_
explanations string - A JMESPath boolean expression used to filter which records to explain. Explanations are activated by default.
- inference_
config object - The inference configuration parameter for the model container.
- shap
Config EndpointConfig Clarify Shap Config - The configuration for SHAP analysis.
- enable
Explanations String - A JMESPath boolean expression used to filter which records to explain. Explanations are activated by default.
- inference
Config EndpointConfig Clarify Inference Config - The inference configuration parameter for the model container.
- shap
Config EndpointConfig Clarify Shap Config - The configuration for SHAP analysis.
- enable
Explanations string - A JMESPath boolean expression used to filter which records to explain. Explanations are activated by default.
- inference
Config EndpointConfig Clarify Inference Config - The inference configuration parameter for the model container.
- shap_
config EndpointConfig Clarify Shap Config - The configuration for SHAP analysis.
- enable_
explanations str - A JMESPath boolean expression used to filter which records to explain. Explanations are activated by default.
- inference_
config EndpointConfig Clarify Inference Config - The inference configuration parameter for the model container.
- shap
Config Property Map - The configuration for SHAP analysis.
- enable
Explanations String - A JMESPath boolean expression used to filter which records to explain. Explanations are activated by default.
- inference
Config Property Map - The inference configuration parameter for the model container.
EndpointConfigClarifyInferenceConfig, EndpointConfigClarifyInferenceConfigArgs
The inference configuration parameter for the model container.- Content
Template string - A template string used to format a JSON record into an acceptable model container input.
- Feature
Headers List<string> - The names of the features. If provided, these are included in the endpoint response payload to help readability of the InvokeEndpoint output.
- Feature
Types List<string> - A list of data types of the features (optional). Applicable only to NLP explainability. If provided, FeatureTypes must have at least one 'text' string (for example, ['text']). If FeatureTypes is not provided, the explainer infers the feature types based on the baseline data.
- Features
Attribute string - Provides the JMESPath expression to extract the features from a model container input in JSON Lines format.
- Label
Attribute string - A JMESPath expression used to locate the list of label headers in the model container output.
- Label
Headers List<string> - For multiclass classification problems, the label headers are the names of the classes. Otherwise, the label header is the name of the predicted label.
- Label
Index int - A zero-based index used to extract a label header or list of label headers from model container output in CSV format.
- Max
Payload intIn Mb - The maximum payload size (MB) allowed of a request from the explainer to the model container. Defaults to 6 MB.
- Max
Record intCount - The maximum number of records in a request that the model container can process when querying the model container for the predictions of a synthetic dataset. A record is a unit of input data that inference can be made on, for example, a single line in CSV data.
- Probability
Attribute string - A JMESPath expression used to extract the probability (or score) from the model container output if the model container is in JSON Lines format.
- Probability
Index int - A zero-based index used to extract a probability value (score) or list from model container output in CSV format. If this value is not provided, the entire model container output will be treated as a probability value (score) or list.
- Content
Template string - A template string used to format a JSON record into an acceptable model container input.
- Feature
Headers []string - The names of the features. If provided, these are included in the endpoint response payload to help readability of the InvokeEndpoint output.
- Feature
Types []string - A list of data types of the features (optional). Applicable only to NLP explainability. If provided, FeatureTypes must have at least one 'text' string (for example, ['text']). If FeatureTypes is not provided, the explainer infers the feature types based on the baseline data.
- Features
Attribute string - Provides the JMESPath expression to extract the features from a model container input in JSON Lines format.
- Label
Attribute string - A JMESPath expression used to locate the list of label headers in the model container output.
- Label
Headers []string - For multiclass classification problems, the label headers are the names of the classes. Otherwise, the label header is the name of the predicted label.
- Label
Index int - A zero-based index used to extract a label header or list of label headers from model container output in CSV format.
- Max
Payload intIn Mb - The maximum payload size (MB) allowed of a request from the explainer to the model container. Defaults to 6 MB.
- Max
Record intCount - The maximum number of records in a request that the model container can process when querying the model container for the predictions of a synthetic dataset. A record is a unit of input data that inference can be made on, for example, a single line in CSV data.
- Probability
Attribute string - A JMESPath expression used to extract the probability (or score) from the model container output if the model container is in JSON Lines format.
- Probability
Index int - A zero-based index used to extract a probability value (score) or list from model container output in CSV format. If this value is not provided, the entire model container output will be treated as a probability value (score) or list.
- content_
template string - A template string used to format a JSON record into an acceptable model container input.
- feature_
headers list(string) - The names of the features. If provided, these are included in the endpoint response payload to help readability of the InvokeEndpoint output.
- feature_
types list(string) - A list of data types of the features (optional). Applicable only to NLP explainability. If provided, FeatureTypes must have at least one 'text' string (for example, ['text']). If FeatureTypes is not provided, the explainer infers the feature types based on the baseline data.
- features_
attribute string - Provides the JMESPath expression to extract the features from a model container input in JSON Lines format.
- label_
attribute string - A JMESPath expression used to locate the list of label headers in the model container output.
- label_
headers list(string) - For multiclass classification problems, the label headers are the names of the classes. Otherwise, the label header is the name of the predicted label.
- label_
index number - A zero-based index used to extract a label header or list of label headers from model container output in CSV format.
- max_
payload_ numberin_ mb - The maximum payload size (MB) allowed of a request from the explainer to the model container. Defaults to 6 MB.
- max_
record_ numbercount - The maximum number of records in a request that the model container can process when querying the model container for the predictions of a synthetic dataset. A record is a unit of input data that inference can be made on, for example, a single line in CSV data.
- probability_
attribute string - A JMESPath expression used to extract the probability (or score) from the model container output if the model container is in JSON Lines format.
- probability_
index number - A zero-based index used to extract a probability value (score) or list from model container output in CSV format. If this value is not provided, the entire model container output will be treated as a probability value (score) or list.
- content
Template String - A template string used to format a JSON record into an acceptable model container input.
- feature
Headers List<String> - The names of the features. If provided, these are included in the endpoint response payload to help readability of the InvokeEndpoint output.
- feature
Types List<String> - A list of data types of the features (optional). Applicable only to NLP explainability. If provided, FeatureTypes must have at least one 'text' string (for example, ['text']). If FeatureTypes is not provided, the explainer infers the feature types based on the baseline data.
- features
Attribute String - Provides the JMESPath expression to extract the features from a model container input in JSON Lines format.
- label
Attribute String - A JMESPath expression used to locate the list of label headers in the model container output.
- label
Headers List<String> - For multiclass classification problems, the label headers are the names of the classes. Otherwise, the label header is the name of the predicted label.
- label
Index Integer - A zero-based index used to extract a label header or list of label headers from model container output in CSV format.
- max
Payload IntegerIn Mb - The maximum payload size (MB) allowed of a request from the explainer to the model container. Defaults to 6 MB.
- max
Record IntegerCount - The maximum number of records in a request that the model container can process when querying the model container for the predictions of a synthetic dataset. A record is a unit of input data that inference can be made on, for example, a single line in CSV data.
- probability
Attribute String - A JMESPath expression used to extract the probability (or score) from the model container output if the model container is in JSON Lines format.
- probability
Index Integer - A zero-based index used to extract a probability value (score) or list from model container output in CSV format. If this value is not provided, the entire model container output will be treated as a probability value (score) or list.
- content
Template string - A template string used to format a JSON record into an acceptable model container input.
- feature
Headers string[] - The names of the features. If provided, these are included in the endpoint response payload to help readability of the InvokeEndpoint output.
- feature
Types string[] - A list of data types of the features (optional). Applicable only to NLP explainability. If provided, FeatureTypes must have at least one 'text' string (for example, ['text']). If FeatureTypes is not provided, the explainer infers the feature types based on the baseline data.
- features
Attribute string - Provides the JMESPath expression to extract the features from a model container input in JSON Lines format.
- label
Attribute string - A JMESPath expression used to locate the list of label headers in the model container output.
- label
Headers string[] - For multiclass classification problems, the label headers are the names of the classes. Otherwise, the label header is the name of the predicted label.
- label
Index number - A zero-based index used to extract a label header or list of label headers from model container output in CSV format.
- max
Payload numberIn Mb - The maximum payload size (MB) allowed of a request from the explainer to the model container. Defaults to 6 MB.
- max
Record numberCount - The maximum number of records in a request that the model container can process when querying the model container for the predictions of a synthetic dataset. A record is a unit of input data that inference can be made on, for example, a single line in CSV data.
- probability
Attribute string - A JMESPath expression used to extract the probability (or score) from the model container output if the model container is in JSON Lines format.
- probability
Index number - A zero-based index used to extract a probability value (score) or list from model container output in CSV format. If this value is not provided, the entire model container output will be treated as a probability value (score) or list.
- content_
template str - A template string used to format a JSON record into an acceptable model container input.
- feature_
headers Sequence[str] - The names of the features. If provided, these are included in the endpoint response payload to help readability of the InvokeEndpoint output.
- feature_
types Sequence[str] - A list of data types of the features (optional). Applicable only to NLP explainability. If provided, FeatureTypes must have at least one 'text' string (for example, ['text']). If FeatureTypes is not provided, the explainer infers the feature types based on the baseline data.
- features_
attribute str - Provides the JMESPath expression to extract the features from a model container input in JSON Lines format.
- label_
attribute str - A JMESPath expression used to locate the list of label headers in the model container output.
- label_
headers Sequence[str] - For multiclass classification problems, the label headers are the names of the classes. Otherwise, the label header is the name of the predicted label.
- label_
index int - A zero-based index used to extract a label header or list of label headers from model container output in CSV format.
- max_
payload_ intin_ mb - The maximum payload size (MB) allowed of a request from the explainer to the model container. Defaults to 6 MB.
- max_
record_ intcount - The maximum number of records in a request that the model container can process when querying the model container for the predictions of a synthetic dataset. A record is a unit of input data that inference can be made on, for example, a single line in CSV data.
- probability_
attribute str - A JMESPath expression used to extract the probability (or score) from the model container output if the model container is in JSON Lines format.
- probability_
index int - A zero-based index used to extract a probability value (score) or list from model container output in CSV format. If this value is not provided, the entire model container output will be treated as a probability value (score) or list.
- content
Template String - A template string used to format a JSON record into an acceptable model container input.
- feature
Headers List<String> - The names of the features. If provided, these are included in the endpoint response payload to help readability of the InvokeEndpoint output.
- feature
Types List<String> - A list of data types of the features (optional). Applicable only to NLP explainability. If provided, FeatureTypes must have at least one 'text' string (for example, ['text']). If FeatureTypes is not provided, the explainer infers the feature types based on the baseline data.
- features
Attribute String - Provides the JMESPath expression to extract the features from a model container input in JSON Lines format.
- label
Attribute String - A JMESPath expression used to locate the list of label headers in the model container output.
- label
Headers List<String> - For multiclass classification problems, the label headers are the names of the classes. Otherwise, the label header is the name of the predicted label.
- label
Index Number - A zero-based index used to extract a label header or list of label headers from model container output in CSV format.
- max
Payload NumberIn Mb - The maximum payload size (MB) allowed of a request from the explainer to the model container. Defaults to 6 MB.
- max
Record NumberCount - The maximum number of records in a request that the model container can process when querying the model container for the predictions of a synthetic dataset. A record is a unit of input data that inference can be made on, for example, a single line in CSV data.
- probability
Attribute String - A JMESPath expression used to extract the probability (or score) from the model container output if the model container is in JSON Lines format.
- probability
Index Number - A zero-based index used to extract a probability value (score) or list from model container output in CSV format. If this value is not provided, the entire model container output will be treated as a probability value (score) or list.
EndpointConfigClarifyShapBaselineConfig, EndpointConfigClarifyShapBaselineConfigArgs
The configuration for the SHAP baseline (also called the background or reference dataset) of the Kernal SHAP algorithm.- Mime
Type string - The MIME type of the baseline data. Choose from 'text/csv' or 'application/jsonlines'. Defaults to 'text/csv'.
- Shap
Baseline string - The inline SHAP baseline data in string format. ShapBaseline can have one or multiple records to be used as the baseline dataset. The format of the SHAP baseline file should be the same format as the training dataset.
- Shap
Baseline stringUri - The uniform resource identifier (URI) of the S3 bucket where the SHAP baseline file is stored. The format of the SHAP baseline file should be the same format as the format of the training dataset.
- Mime
Type string - The MIME type of the baseline data. Choose from 'text/csv' or 'application/jsonlines'. Defaults to 'text/csv'.
- Shap
Baseline string - The inline SHAP baseline data in string format. ShapBaseline can have one or multiple records to be used as the baseline dataset. The format of the SHAP baseline file should be the same format as the training dataset.
- Shap
Baseline stringUri - The uniform resource identifier (URI) of the S3 bucket where the SHAP baseline file is stored. The format of the SHAP baseline file should be the same format as the format of the training dataset.
- mime_
type string - The MIME type of the baseline data. Choose from 'text/csv' or 'application/jsonlines'. Defaults to 'text/csv'.
- shap_
baseline string - The inline SHAP baseline data in string format. ShapBaseline can have one or multiple records to be used as the baseline dataset. The format of the SHAP baseline file should be the same format as the training dataset.
- shap_
baseline_ stringuri - The uniform resource identifier (URI) of the S3 bucket where the SHAP baseline file is stored. The format of the SHAP baseline file should be the same format as the format of the training dataset.
- mime
Type String - The MIME type of the baseline data. Choose from 'text/csv' or 'application/jsonlines'. Defaults to 'text/csv'.
- shap
Baseline String - The inline SHAP baseline data in string format. ShapBaseline can have one or multiple records to be used as the baseline dataset. The format of the SHAP baseline file should be the same format as the training dataset.
- shap
Baseline StringUri - The uniform resource identifier (URI) of the S3 bucket where the SHAP baseline file is stored. The format of the SHAP baseline file should be the same format as the format of the training dataset.
- mime
Type string - The MIME type of the baseline data. Choose from 'text/csv' or 'application/jsonlines'. Defaults to 'text/csv'.
- shap
Baseline string - The inline SHAP baseline data in string format. ShapBaseline can have one or multiple records to be used as the baseline dataset. The format of the SHAP baseline file should be the same format as the training dataset.
- shap
Baseline stringUri - The uniform resource identifier (URI) of the S3 bucket where the SHAP baseline file is stored. The format of the SHAP baseline file should be the same format as the format of the training dataset.
- mime_
type str - The MIME type of the baseline data. Choose from 'text/csv' or 'application/jsonlines'. Defaults to 'text/csv'.
- shap_
baseline str - The inline SHAP baseline data in string format. ShapBaseline can have one or multiple records to be used as the baseline dataset. The format of the SHAP baseline file should be the same format as the training dataset.
- shap_
baseline_ struri - The uniform resource identifier (URI) of the S3 bucket where the SHAP baseline file is stored. The format of the SHAP baseline file should be the same format as the format of the training dataset.
- mime
Type String - The MIME type of the baseline data. Choose from 'text/csv' or 'application/jsonlines'. Defaults to 'text/csv'.
- shap
Baseline String - The inline SHAP baseline data in string format. ShapBaseline can have one or multiple records to be used as the baseline dataset. The format of the SHAP baseline file should be the same format as the training dataset.
- shap
Baseline StringUri - The uniform resource identifier (URI) of the S3 bucket where the SHAP baseline file is stored. The format of the SHAP baseline file should be the same format as the format of the training dataset.
EndpointConfigClarifyShapConfig, EndpointConfigClarifyShapConfigArgs
The configuration for SHAP analysis using SageMaker Clarify Explainer.- Shap
Baseline Pulumi.Config Aws Native. Sage Maker. Inputs. Endpoint Config Clarify Shap Baseline Config - The configuration for the SHAP baseline of the Kernal SHAP algorithm.
- Number
Of intSamples - The number of samples to be used for analysis by the Kernal SHAP algorithm.
- Seed int
- The starting value used to initialize the random number generator in the explainer. Provide a value for this parameter to obtain a deterministic SHAP result.
- Text
Config Pulumi.Aws Native. Sage Maker. Inputs. Endpoint Config Clarify Text Config - A parameter that indicates if text features are treated as text and explanations are provided for individual units of text. Required for natural language processing (NLP) explainability only.
- Use
Logit bool - A Boolean toggle to indicate if you want to use the logit function (true) or log-odds units (false) for model predictions. Defaults to false.
- Shap
Baseline EndpointConfig Config Clarify Shap Baseline Config - The configuration for the SHAP baseline of the Kernal SHAP algorithm.
- Number
Of intSamples - The number of samples to be used for analysis by the Kernal SHAP algorithm.
- Seed int
- The starting value used to initialize the random number generator in the explainer. Provide a value for this parameter to obtain a deterministic SHAP result.
- Text
Config EndpointConfig Clarify Text Config - A parameter that indicates if text features are treated as text and explanations are provided for individual units of text. Required for natural language processing (NLP) explainability only.
- Use
Logit bool - A Boolean toggle to indicate if you want to use the logit function (true) or log-odds units (false) for model predictions. Defaults to false.
- shap_
baseline_ objectconfig - The configuration for the SHAP baseline of the Kernal SHAP algorithm.
- number_
of_ numbersamples - The number of samples to be used for analysis by the Kernal SHAP algorithm.
- seed number
- The starting value used to initialize the random number generator in the explainer. Provide a value for this parameter to obtain a deterministic SHAP result.
- text_
config object - A parameter that indicates if text features are treated as text and explanations are provided for individual units of text. Required for natural language processing (NLP) explainability only.
- use_
logit bool - A Boolean toggle to indicate if you want to use the logit function (true) or log-odds units (false) for model predictions. Defaults to false.
- shap
Baseline EndpointConfig Config Clarify Shap Baseline Config - The configuration for the SHAP baseline of the Kernal SHAP algorithm.
- number
Of IntegerSamples - The number of samples to be used for analysis by the Kernal SHAP algorithm.
- seed Integer
- The starting value used to initialize the random number generator in the explainer. Provide a value for this parameter to obtain a deterministic SHAP result.
- text
Config EndpointConfig Clarify Text Config - A parameter that indicates if text features are treated as text and explanations are provided for individual units of text. Required for natural language processing (NLP) explainability only.
- use
Logit Boolean - A Boolean toggle to indicate if you want to use the logit function (true) or log-odds units (false) for model predictions. Defaults to false.
- shap
Baseline EndpointConfig Config Clarify Shap Baseline Config - The configuration for the SHAP baseline of the Kernal SHAP algorithm.
- number
Of numberSamples - The number of samples to be used for analysis by the Kernal SHAP algorithm.
- seed number
- The starting value used to initialize the random number generator in the explainer. Provide a value for this parameter to obtain a deterministic SHAP result.
- text
Config EndpointConfig Clarify Text Config - A parameter that indicates if text features are treated as text and explanations are provided for individual units of text. Required for natural language processing (NLP) explainability only.
- use
Logit boolean - A Boolean toggle to indicate if you want to use the logit function (true) or log-odds units (false) for model predictions. Defaults to false.
- shap_
baseline_ Endpointconfig Config Clarify Shap Baseline Config - The configuration for the SHAP baseline of the Kernal SHAP algorithm.
- number_
of_ intsamples - The number of samples to be used for analysis by the Kernal SHAP algorithm.
- seed int
- The starting value used to initialize the random number generator in the explainer. Provide a value for this parameter to obtain a deterministic SHAP result.
- text_
config EndpointConfig Clarify Text Config - A parameter that indicates if text features are treated as text and explanations are provided for individual units of text. Required for natural language processing (NLP) explainability only.
- use_
logit bool - A Boolean toggle to indicate if you want to use the logit function (true) or log-odds units (false) for model predictions. Defaults to false.
- shap
Baseline Property MapConfig - The configuration for the SHAP baseline of the Kernal SHAP algorithm.
- number
Of NumberSamples - The number of samples to be used for analysis by the Kernal SHAP algorithm.
- seed Number
- The starting value used to initialize the random number generator in the explainer. Provide a value for this parameter to obtain a deterministic SHAP result.
- text
Config Property Map - A parameter that indicates if text features are treated as text and explanations are provided for individual units of text. Required for natural language processing (NLP) explainability only.
- use
Logit Boolean - A Boolean toggle to indicate if you want to use the logit function (true) or log-odds units (false) for model predictions. Defaults to false.
EndpointConfigClarifyTextConfig, EndpointConfigClarifyTextConfigArgs
A parameter used to configure the SageMaker Clarify explainer to treat text features as text so that explanations are provided for individual units of text. Required only for natural language processing (NLP) explainability.- Granularity string
- The unit of granularity for the analysis of text features. For example, if the unit is 'token', then each token (like a word in English) of the text is treated as a feature. SHAP values are computed for each unit/feature.
- Language string
- Specifies the language of the text features in ISO 639-1 or ISO 639-3 code of a supported language.
- Granularity string
- The unit of granularity for the analysis of text features. For example, if the unit is 'token', then each token (like a word in English) of the text is treated as a feature. SHAP values are computed for each unit/feature.
- Language string
- Specifies the language of the text features in ISO 639-1 or ISO 639-3 code of a supported language.
- granularity string
- The unit of granularity for the analysis of text features. For example, if the unit is 'token', then each token (like a word in English) of the text is treated as a feature. SHAP values are computed for each unit/feature.
- language string
- Specifies the language of the text features in ISO 639-1 or ISO 639-3 code of a supported language.
- granularity String
- The unit of granularity for the analysis of text features. For example, if the unit is 'token', then each token (like a word in English) of the text is treated as a feature. SHAP values are computed for each unit/feature.
- language String
- Specifies the language of the text features in ISO 639-1 or ISO 639-3 code of a supported language.
- granularity string
- The unit of granularity for the analysis of text features. For example, if the unit is 'token', then each token (like a word in English) of the text is treated as a feature. SHAP values are computed for each unit/feature.
- language string
- Specifies the language of the text features in ISO 639-1 or ISO 639-3 code of a supported language.
- granularity str
- The unit of granularity for the analysis of text features. For example, if the unit is 'token', then each token (like a word in English) of the text is treated as a feature. SHAP values are computed for each unit/feature.
- language str
- Specifies the language of the text features in ISO 639-1 or ISO 639-3 code of a supported language.
- granularity String
- The unit of granularity for the analysis of text features. For example, if the unit is 'token', then each token (like a word in English) of the text is treated as a feature. SHAP values are computed for each unit/feature.
- language String
- Specifies the language of the text features in ISO 639-1 or ISO 639-3 code of a supported language.
EndpointConfigCoreDumpConfig, EndpointConfigCoreDumpConfigArgs
Specifies where SageMaker writes core dumps from the model container when the process crashes, and how it encrypts them.- Destination
S3Uri string - The Amazon S3 bucket to send the core dump to.
- Kms
Key stringId - The AWS Key Management Service (AWS KMS) key that SageMaker uses to encrypt the core dump data at rest using Amazon S3 server-side encryption. If you use a KMS key ID or an alias of your KMS key, the SageMaker execution role must include permissions to call kms:Encrypt.
- Destination
S3Uri string - The Amazon S3 bucket to send the core dump to.
- Kms
Key stringId - The AWS Key Management Service (AWS KMS) key that SageMaker uses to encrypt the core dump data at rest using Amazon S3 server-side encryption. If you use a KMS key ID or an alias of your KMS key, the SageMaker execution role must include permissions to call kms:Encrypt.
- destination_
s3_ stringuri - The Amazon S3 bucket to send the core dump to.
- kms_
key_ stringid - The AWS Key Management Service (AWS KMS) key that SageMaker uses to encrypt the core dump data at rest using Amazon S3 server-side encryption. If you use a KMS key ID or an alias of your KMS key, the SageMaker execution role must include permissions to call kms:Encrypt.
- destination
S3Uri String - The Amazon S3 bucket to send the core dump to.
- kms
Key StringId - The AWS Key Management Service (AWS KMS) key that SageMaker uses to encrypt the core dump data at rest using Amazon S3 server-side encryption. If you use a KMS key ID or an alias of your KMS key, the SageMaker execution role must include permissions to call kms:Encrypt.
- destination
S3Uri string - The Amazon S3 bucket to send the core dump to.
- kms
Key stringId - The AWS Key Management Service (AWS KMS) key that SageMaker uses to encrypt the core dump data at rest using Amazon S3 server-side encryption. If you use a KMS key ID or an alias of your KMS key, the SageMaker execution role must include permissions to call kms:Encrypt.
- destination_
s3_ struri - The Amazon S3 bucket to send the core dump to.
- kms_
key_ strid - The AWS Key Management Service (AWS KMS) key that SageMaker uses to encrypt the core dump data at rest using Amazon S3 server-side encryption. If you use a KMS key ID or an alias of your KMS key, the SageMaker execution role must include permissions to call kms:Encrypt.
- destination
S3Uri String - The Amazon S3 bucket to send the core dump to.
- kms
Key StringId - The AWS Key Management Service (AWS KMS) key that SageMaker uses to encrypt the core dump data at rest using Amazon S3 server-side encryption. If you use a KMS key ID or an alias of your KMS key, the SageMaker execution role must include permissions to call kms:Encrypt.
EndpointConfigDataCaptureConfig, EndpointConfigDataCaptureConfigArgs
Specifies how to capture endpoint data for model monitor. The data capture configuration applies to all production variants hosted at the endpoint.- Capture
Options List<Pulumi.Aws Native. Sage Maker. Inputs. Endpoint Config Capture Option> - Specifies whether the endpoint captures input data to your model, output data from your model, or both.
- Destination
S3Uri string - The S3 bucket where model monitor stores captured data.
- Initial
Sampling intPercentage - The percentage of data to capture.
- Capture
Content Pulumi.Type Header Aws Native. Sage Maker. Inputs. Endpoint Config Capture Content Type Header - A list of the JSON and CSV content type that the endpoint captures.
- Enable
Capture bool - Set to True to enable data capture.
- Kms
Key stringId - The AWS Key Management Service (AWS KMS) key that Amazon SageMaker uses to encrypt the captured data at rest using Amazon S3 server-side encryption.
- Capture
Options []EndpointConfig Capture Option - Specifies whether the endpoint captures input data to your model, output data from your model, or both.
- Destination
S3Uri string - The S3 bucket where model monitor stores captured data.
- Initial
Sampling intPercentage - The percentage of data to capture.
- Capture
Content EndpointType Header Config Capture Content Type Header - A list of the JSON and CSV content type that the endpoint captures.
- Enable
Capture bool - Set to True to enable data capture.
- Kms
Key stringId - The AWS Key Management Service (AWS KMS) key that Amazon SageMaker uses to encrypt the captured data at rest using Amazon S3 server-side encryption.
- capture_
options list(object) - Specifies whether the endpoint captures input data to your model, output data from your model, or both.
- destination_
s3_ stringuri - The S3 bucket where model monitor stores captured data.
- initial_
sampling_ numberpercentage - The percentage of data to capture.
- capture_
content_ objecttype_ header - A list of the JSON and CSV content type that the endpoint captures.
- enable_
capture bool - Set to True to enable data capture.
- kms_
key_ stringid - The AWS Key Management Service (AWS KMS) key that Amazon SageMaker uses to encrypt the captured data at rest using Amazon S3 server-side encryption.
- capture
Options List<EndpointConfig Capture Option> - Specifies whether the endpoint captures input data to your model, output data from your model, or both.
- destination
S3Uri String - The S3 bucket where model monitor stores captured data.
- initial
Sampling IntegerPercentage - The percentage of data to capture.
- capture
Content EndpointType Header Config Capture Content Type Header - A list of the JSON and CSV content type that the endpoint captures.
- enable
Capture Boolean - Set to True to enable data capture.
- kms
Key StringId - The AWS Key Management Service (AWS KMS) key that Amazon SageMaker uses to encrypt the captured data at rest using Amazon S3 server-side encryption.
- capture
Options EndpointConfig Capture Option[] - Specifies whether the endpoint captures input data to your model, output data from your model, or both.
- destination
S3Uri string - The S3 bucket where model monitor stores captured data.
- initial
Sampling numberPercentage - The percentage of data to capture.
- capture
Content EndpointType Header Config Capture Content Type Header - A list of the JSON and CSV content type that the endpoint captures.
- enable
Capture boolean - Set to True to enable data capture.
- kms
Key stringId - The AWS Key Management Service (AWS KMS) key that Amazon SageMaker uses to encrypt the captured data at rest using Amazon S3 server-side encryption.
- capture_
options Sequence[EndpointConfig Capture Option] - Specifies whether the endpoint captures input data to your model, output data from your model, or both.
- destination_
s3_ struri - The S3 bucket where model monitor stores captured data.
- initial_
sampling_ intpercentage - The percentage of data to capture.
- capture_
content_ Endpointtype_ header Config Capture Content Type Header - A list of the JSON and CSV content type that the endpoint captures.
- enable_
capture bool - Set to True to enable data capture.
- kms_
key_ strid - The AWS Key Management Service (AWS KMS) key that Amazon SageMaker uses to encrypt the captured data at rest using Amazon S3 server-side encryption.
- capture
Options List<Property Map> - Specifies whether the endpoint captures input data to your model, output data from your model, or both.
- destination
S3Uri String - The S3 bucket where model monitor stores captured data.
- initial
Sampling NumberPercentage - The percentage of data to capture.
- capture
Content Property MapType Header - A list of the JSON and CSV content type that the endpoint captures.
- enable
Capture Boolean - Set to True to enable data capture.
- kms
Key StringId - The AWS Key Management Service (AWS KMS) key that Amazon SageMaker uses to encrypt the captured data at rest using Amazon S3 server-side encryption.
EndpointConfigExplainerConfig, EndpointConfigExplainerConfigArgs
A parameter to activate explainers.- Clarify
Explainer Pulumi.Config Aws Native. Sage Maker. Inputs. Endpoint Config Clarify Explainer Config - A member of ExplainerConfig that contains configuration parameters for the SageMaker Clarify explainer.
- Clarify
Explainer EndpointConfig Config Clarify Explainer Config - A member of ExplainerConfig that contains configuration parameters for the SageMaker Clarify explainer.
- clarify_
explainer_ objectconfig - A member of ExplainerConfig that contains configuration parameters for the SageMaker Clarify explainer.
- clarify
Explainer EndpointConfig Config Clarify Explainer Config - A member of ExplainerConfig that contains configuration parameters for the SageMaker Clarify explainer.
- clarify
Explainer EndpointConfig Config Clarify Explainer Config - A member of ExplainerConfig that contains configuration parameters for the SageMaker Clarify explainer.
- clarify_
explainer_ Endpointconfig Config Clarify Explainer Config - A member of ExplainerConfig that contains configuration parameters for the SageMaker Clarify explainer.
- clarify
Explainer Property MapConfig - A member of ExplainerConfig that contains configuration parameters for the SageMaker Clarify explainer.
EndpointConfigInstancePool, EndpointConfigInstancePoolArgs
Specifies an instance type and its priority for a heterogeneous endpoint. Use instance pools to configure a production variant with multiple instance types, enabling the endpoint to provision instances across different types based on priority.- Instance
Type string - The ML compute instance type for the instance pool.
- Priority int
- The priority for the instance pool. SageMaker attempts to provision instances in order of priority, starting with the lowest value. If instances for a higher-priority pool are unavailable, SageMaker attempts to provision from the next pool. Valid values: 1 to 5, where 1 is the highest priority.
- Model
Name stringOverride - The name of a SageMaker model to use for this instance pool instead of the model specified for the production variant. Use this to deploy a different model optimized for the instance type in this pool.
- Instance
Type string - The ML compute instance type for the instance pool.
- Priority int
- The priority for the instance pool. SageMaker attempts to provision instances in order of priority, starting with the lowest value. If instances for a higher-priority pool are unavailable, SageMaker attempts to provision from the next pool. Valid values: 1 to 5, where 1 is the highest priority.
- Model
Name stringOverride - The name of a SageMaker model to use for this instance pool instead of the model specified for the production variant. Use this to deploy a different model optimized for the instance type in this pool.
- instance_
type string - The ML compute instance type for the instance pool.
- priority number
- The priority for the instance pool. SageMaker attempts to provision instances in order of priority, starting with the lowest value. If instances for a higher-priority pool are unavailable, SageMaker attempts to provision from the next pool. Valid values: 1 to 5, where 1 is the highest priority.
- model_
name_ stringoverride - The name of a SageMaker model to use for this instance pool instead of the model specified for the production variant. Use this to deploy a different model optimized for the instance type in this pool.
- instance
Type String - The ML compute instance type for the instance pool.
- priority Integer
- The priority for the instance pool. SageMaker attempts to provision instances in order of priority, starting with the lowest value. If instances for a higher-priority pool are unavailable, SageMaker attempts to provision from the next pool. Valid values: 1 to 5, where 1 is the highest priority.
- model
Name StringOverride - The name of a SageMaker model to use for this instance pool instead of the model specified for the production variant. Use this to deploy a different model optimized for the instance type in this pool.
- instance
Type string - The ML compute instance type for the instance pool.
- priority number
- The priority for the instance pool. SageMaker attempts to provision instances in order of priority, starting with the lowest value. If instances for a higher-priority pool are unavailable, SageMaker attempts to provision from the next pool. Valid values: 1 to 5, where 1 is the highest priority.
- model
Name stringOverride - The name of a SageMaker model to use for this instance pool instead of the model specified for the production variant. Use this to deploy a different model optimized for the instance type in this pool.
- instance_
type str - The ML compute instance type for the instance pool.
- priority int
- The priority for the instance pool. SageMaker attempts to provision instances in order of priority, starting with the lowest value. If instances for a higher-priority pool are unavailable, SageMaker attempts to provision from the next pool. Valid values: 1 to 5, where 1 is the highest priority.
- model_
name_ stroverride - The name of a SageMaker model to use for this instance pool instead of the model specified for the production variant. Use this to deploy a different model optimized for the instance type in this pool.
- instance
Type String - The ML compute instance type for the instance pool.
- priority Number
- The priority for the instance pool. SageMaker attempts to provision instances in order of priority, starting with the lowest value. If instances for a higher-priority pool are unavailable, SageMaker attempts to provision from the next pool. Valid values: 1 to 5, where 1 is the highest priority.
- model
Name StringOverride - The name of a SageMaker model to use for this instance pool instead of the model specified for the production variant. Use this to deploy a different model optimized for the instance type in this pool.
EndpointConfigManagedInstanceScaling, EndpointConfigManagedInstanceScalingArgs
Settings that control the range in the number of instances that the endpoint provisions as it scales up or down to accommodate traffic.- Max
Instance intCount - The maximum number of instances that the endpoint can provision when it scales up to accommodate an increase in traffic.
- Min
Instance intCount - The minimum number of instances that the endpoint must retain when it scales down to accommodate a decrease in traffic.
- Scale
In Pulumi.Policy Aws Native. Sage Maker. Inputs. Endpoint Config Scale In Policy - Configures the scale-in behavior for managed instance scaling.
- Status string
- Indicates whether managed instance scaling is enabled.
- Max
Instance intCount - The maximum number of instances that the endpoint can provision when it scales up to accommodate an increase in traffic.
- Min
Instance intCount - The minimum number of instances that the endpoint must retain when it scales down to accommodate a decrease in traffic.
- Scale
In EndpointPolicy Config Scale In Policy - Configures the scale-in behavior for managed instance scaling.
- Status string
- Indicates whether managed instance scaling is enabled.
- max_
instance_ numbercount - The maximum number of instances that the endpoint can provision when it scales up to accommodate an increase in traffic.
- min_
instance_ numbercount - The minimum number of instances that the endpoint must retain when it scales down to accommodate a decrease in traffic.
- scale_
in_ objectpolicy - Configures the scale-in behavior for managed instance scaling.
- status string
- Indicates whether managed instance scaling is enabled.
- max
Instance IntegerCount - The maximum number of instances that the endpoint can provision when it scales up to accommodate an increase in traffic.
- min
Instance IntegerCount - The minimum number of instances that the endpoint must retain when it scales down to accommodate a decrease in traffic.
- scale
In EndpointPolicy Config Scale In Policy - Configures the scale-in behavior for managed instance scaling.
- status String
- Indicates whether managed instance scaling is enabled.
- max
Instance numberCount - The maximum number of instances that the endpoint can provision when it scales up to accommodate an increase in traffic.
- min
Instance numberCount - The minimum number of instances that the endpoint must retain when it scales down to accommodate a decrease in traffic.
- scale
In EndpointPolicy Config Scale In Policy - Configures the scale-in behavior for managed instance scaling.
- status string
- Indicates whether managed instance scaling is enabled.
- max_
instance_ intcount - The maximum number of instances that the endpoint can provision when it scales up to accommodate an increase in traffic.
- min_
instance_ intcount - The minimum number of instances that the endpoint must retain when it scales down to accommodate a decrease in traffic.
- scale_
in_ Endpointpolicy Config Scale In Policy - Configures the scale-in behavior for managed instance scaling.
- status str
- Indicates whether managed instance scaling is enabled.
- max
Instance NumberCount - The maximum number of instances that the endpoint can provision when it scales up to accommodate an increase in traffic.
- min
Instance NumberCount - The minimum number of instances that the endpoint must retain when it scales down to accommodate a decrease in traffic.
- scale
In Property MapPolicy - Configures the scale-in behavior for managed instance scaling.
- status String
- Indicates whether managed instance scaling is enabled.
EndpointConfigMetricsConfig, EndpointConfigMetricsConfigArgs
Specifies the metrics that the endpoint publishes to Amazon CloudWatch, the frequency of publication, and whether to enable enhanced or detailed observability metrics.- Enable
Detailed boolObservability - Specifies whether to enable detailed observability for the endpoint. When set to true, the endpoint publishes container-level inference metrics, per-GPU metrics, per-instance host metrics, and inference component placement metrics.
- Enable
Enhanced boolMetrics - Specifies whether to enable enhanced metrics for the endpoint. Enhanced metrics provide utilization and invocation data at instance and container granularity.
- Metric
Publish intFrequency In Seconds - The interval, in seconds, at which the endpoint publishes metrics to Amazon CloudWatch. Valid values are 10, 30, 60, 120, 180, 240, and 300. The default is 60.
- Enable
Detailed boolObservability - Specifies whether to enable detailed observability for the endpoint. When set to true, the endpoint publishes container-level inference metrics, per-GPU metrics, per-instance host metrics, and inference component placement metrics.
- Enable
Enhanced boolMetrics - Specifies whether to enable enhanced metrics for the endpoint. Enhanced metrics provide utilization and invocation data at instance and container granularity.
- Metric
Publish intFrequency In Seconds - The interval, in seconds, at which the endpoint publishes metrics to Amazon CloudWatch. Valid values are 10, 30, 60, 120, 180, 240, and 300. The default is 60.
- enable_
detailed_ boolobservability - Specifies whether to enable detailed observability for the endpoint. When set to true, the endpoint publishes container-level inference metrics, per-GPU metrics, per-instance host metrics, and inference component placement metrics.
- enable_
enhanced_ boolmetrics - Specifies whether to enable enhanced metrics for the endpoint. Enhanced metrics provide utilization and invocation data at instance and container granularity.
- metric_
publish_ numberfrequency_ in_ seconds - The interval, in seconds, at which the endpoint publishes metrics to Amazon CloudWatch. Valid values are 10, 30, 60, 120, 180, 240, and 300. The default is 60.
- enable
Detailed BooleanObservability - Specifies whether to enable detailed observability for the endpoint. When set to true, the endpoint publishes container-level inference metrics, per-GPU metrics, per-instance host metrics, and inference component placement metrics.
- enable
Enhanced BooleanMetrics - Specifies whether to enable enhanced metrics for the endpoint. Enhanced metrics provide utilization and invocation data at instance and container granularity.
- metric
Publish IntegerFrequency In Seconds - The interval, in seconds, at which the endpoint publishes metrics to Amazon CloudWatch. Valid values are 10, 30, 60, 120, 180, 240, and 300. The default is 60.
- enable
Detailed booleanObservability - Specifies whether to enable detailed observability for the endpoint. When set to true, the endpoint publishes container-level inference metrics, per-GPU metrics, per-instance host metrics, and inference component placement metrics.
- enable
Enhanced booleanMetrics - Specifies whether to enable enhanced metrics for the endpoint. Enhanced metrics provide utilization and invocation data at instance and container granularity.
- metric
Publish numberFrequency In Seconds - The interval, in seconds, at which the endpoint publishes metrics to Amazon CloudWatch. Valid values are 10, 30, 60, 120, 180, 240, and 300. The default is 60.
- enable_
detailed_ boolobservability - Specifies whether to enable detailed observability for the endpoint. When set to true, the endpoint publishes container-level inference metrics, per-GPU metrics, per-instance host metrics, and inference component placement metrics.
- enable_
enhanced_ boolmetrics - Specifies whether to enable enhanced metrics for the endpoint. Enhanced metrics provide utilization and invocation data at instance and container granularity.
- metric_
publish_ intfrequency_ in_ seconds - The interval, in seconds, at which the endpoint publishes metrics to Amazon CloudWatch. Valid values are 10, 30, 60, 120, 180, 240, and 300. The default is 60.
- enable
Detailed BooleanObservability - Specifies whether to enable detailed observability for the endpoint. When set to true, the endpoint publishes container-level inference metrics, per-GPU metrics, per-instance host metrics, and inference component placement metrics.
- enable
Enhanced BooleanMetrics - Specifies whether to enable enhanced metrics for the endpoint. Enhanced metrics provide utilization and invocation data at instance and container granularity.
- metric
Publish NumberFrequency In Seconds - The interval, in seconds, at which the endpoint publishes metrics to Amazon CloudWatch. Valid values are 10, 30, 60, 120, 180, 240, and 300. The default is 60.
EndpointConfigPrefixAwareRoutingConfig, EndpointConfigPrefixAwareRoutingConfigArgs
The configuration for prefix-aware routing on a SageMaker real-time inference endpoint. Specify PrefixLength and ConcurrencyThreshold to control routing behavior.- Concurrency
Threshold int - The maximum number of in-flight requests on the target instance before the endpoint routes to another instance. Required when RoutingStrategy is PREFIX_AWARE. Valid values are 1 to 1024.
- Prefix
Length int - The maximum length of the prefix used for routing decisions. Required when RoutingStrategy is PREFIX_AWARE. Valid values are 1024 to 65536.
- Concurrency
Threshold int - The maximum number of in-flight requests on the target instance before the endpoint routes to another instance. Required when RoutingStrategy is PREFIX_AWARE. Valid values are 1 to 1024.
- Prefix
Length int - The maximum length of the prefix used for routing decisions. Required when RoutingStrategy is PREFIX_AWARE. Valid values are 1024 to 65536.
- concurrency_
threshold number - The maximum number of in-flight requests on the target instance before the endpoint routes to another instance. Required when RoutingStrategy is PREFIX_AWARE. Valid values are 1 to 1024.
- prefix_
length number - The maximum length of the prefix used for routing decisions. Required when RoutingStrategy is PREFIX_AWARE. Valid values are 1024 to 65536.
- concurrency
Threshold Integer - The maximum number of in-flight requests on the target instance before the endpoint routes to another instance. Required when RoutingStrategy is PREFIX_AWARE. Valid values are 1 to 1024.
- prefix
Length Integer - The maximum length of the prefix used for routing decisions. Required when RoutingStrategy is PREFIX_AWARE. Valid values are 1024 to 65536.
- concurrency
Threshold number - The maximum number of in-flight requests on the target instance before the endpoint routes to another instance. Required when RoutingStrategy is PREFIX_AWARE. Valid values are 1 to 1024.
- prefix
Length number - The maximum length of the prefix used for routing decisions. Required when RoutingStrategy is PREFIX_AWARE. Valid values are 1024 to 65536.
- concurrency_
threshold int - The maximum number of in-flight requests on the target instance before the endpoint routes to another instance. Required when RoutingStrategy is PREFIX_AWARE. Valid values are 1 to 1024.
- prefix_
length int - The maximum length of the prefix used for routing decisions. Required when RoutingStrategy is PREFIX_AWARE. Valid values are 1024 to 65536.
- concurrency
Threshold Number - The maximum number of in-flight requests on the target instance before the endpoint routes to another instance. Required when RoutingStrategy is PREFIX_AWARE. Valid values are 1 to 1024.
- prefix
Length Number - The maximum length of the prefix used for routing decisions. Required when RoutingStrategy is PREFIX_AWARE. Valid values are 1024 to 65536.
EndpointConfigProductionVariant, EndpointConfigProductionVariantArgs
Specifies a model that you want to host and the resources to deploy for hosting it.- Variant
Name string - The name of the production variant.
- Capacity
Reservation Pulumi.Config Aws Native. Sage Maker. Inputs. Endpoint Config Capacity Reservation Config - Settings for the capacity reservation for the compute instances that SageMaker AI reserves for an endpoint.
- Container
Startup intHealth Check Timeout In Seconds - The timeout value, in seconds, for your inference container to pass health check by SageMaker Hosting.
- Core
Dump Pulumi.Config Aws Native. Sage Maker. Inputs. Endpoint Config Core Dump Config - Specifies configuration for a core dump from the model container when the process crashes.
- Enable
Ssm boolAccess - You can use this parameter to turn on native AWS Systems Manager (SSM) access for a production variant behind an endpoint. By default, SSM access is disabled for all production variants behind an endpoint.
- Inference
Ami stringVersion - Specifies an option from a collection of preconfigured Amazon Machine Image (AMI) images. Each image is configured by AWS with a set of software and driver versions. AWS optimizes these configurations for different machine learning workloads. By selecting an AMI version, you can ensure that your inference environment is compatible with specific software requirements, such as CUDA driver versions, Linux kernel versions, or AWS Neuron driver versions
- Initial
Instance intCount - Number of instances to launch initially.
- Initial
Variant doubleWeight - Determines initial traffic distribution among all of the models that you specify in the endpoint configuration.
- Instance
Pools List<Pulumi.Aws Native. Sage Maker. Inputs. Endpoint Config Instance Pool> - A list of instance pools for the production variant. Each instance pool specifies an instance type and its priority for provisioning. Use instance pools to configure heterogeneous endpoints that deploy models across multiple instance types.
- Instance
Type string - The ML compute instance type.
- Managed
Instance Pulumi.Scaling Aws Native. Sage Maker. Inputs. Endpoint Config Managed Instance Scaling - Model
Data intDownload Timeout In Seconds - The timeout value, in seconds, to download and extract the model that you want to host from Amazon S3 to the individual inference instance associated with this production variant.
- Model
Name string - The name of the model that you want to host. This is the name that you specified when creating the model.
- Routing
Config Pulumi.Aws Native. Sage Maker. Inputs. Endpoint Config Routing Config - Settings that control how the endpoint routes incoming traffic to the instances that the endpoint hosts.
- Serverless
Config Pulumi.Aws Native. Sage Maker. Inputs. Endpoint Config Serverless Config - The serverless configuration for an endpoint. Specifies a serverless endpoint configuration instead of an instance-based endpoint configuration.
- Variant
Instance intProvision Timeout In Seconds - The timeout value, in seconds, for provisioning instances for the production variant. When SageMaker encounters an insufficient capacity error while provisioning instances, it retries with the next instance pool (if configured) or waits until the timeout expires. This timeout applies only to capacity provisioning and does not include the time for model download or container startup.
- Volume
Size intIn Gb - The size, in GB, of the ML storage volume attached to individual inference instance associated with the production variant. Currently only Amazon EBS gp2 storage volumes are supported.
- Variant
Name string - The name of the production variant.
- Capacity
Reservation EndpointConfig Config Capacity Reservation Config - Settings for the capacity reservation for the compute instances that SageMaker AI reserves for an endpoint.
- Container
Startup intHealth Check Timeout In Seconds - The timeout value, in seconds, for your inference container to pass health check by SageMaker Hosting.
- Core
Dump EndpointConfig Config Core Dump Config - Specifies configuration for a core dump from the model container when the process crashes.
- Enable
Ssm boolAccess - You can use this parameter to turn on native AWS Systems Manager (SSM) access for a production variant behind an endpoint. By default, SSM access is disabled for all production variants behind an endpoint.
- Inference
Ami stringVersion - Specifies an option from a collection of preconfigured Amazon Machine Image (AMI) images. Each image is configured by AWS with a set of software and driver versions. AWS optimizes these configurations for different machine learning workloads. By selecting an AMI version, you can ensure that your inference environment is compatible with specific software requirements, such as CUDA driver versions, Linux kernel versions, or AWS Neuron driver versions
- Initial
Instance intCount - Number of instances to launch initially.
- Initial
Variant float64Weight - Determines initial traffic distribution among all of the models that you specify in the endpoint configuration.
- Instance
Pools []EndpointConfig Instance Pool - A list of instance pools for the production variant. Each instance pool specifies an instance type and its priority for provisioning. Use instance pools to configure heterogeneous endpoints that deploy models across multiple instance types.
- Instance
Type string - The ML compute instance type.
- Managed
Instance EndpointScaling Config Managed Instance Scaling - Model
Data intDownload Timeout In Seconds - The timeout value, in seconds, to download and extract the model that you want to host from Amazon S3 to the individual inference instance associated with this production variant.
- Model
Name string - The name of the model that you want to host. This is the name that you specified when creating the model.
- Routing
Config EndpointConfig Routing Config - Settings that control how the endpoint routes incoming traffic to the instances that the endpoint hosts.
- Serverless
Config EndpointConfig Serverless Config - The serverless configuration for an endpoint. Specifies a serverless endpoint configuration instead of an instance-based endpoint configuration.
- Variant
Instance intProvision Timeout In Seconds - The timeout value, in seconds, for provisioning instances for the production variant. When SageMaker encounters an insufficient capacity error while provisioning instances, it retries with the next instance pool (if configured) or waits until the timeout expires. This timeout applies only to capacity provisioning and does not include the time for model download or container startup.
- Volume
Size intIn Gb - The size, in GB, of the ML storage volume attached to individual inference instance associated with the production variant. Currently only Amazon EBS gp2 storage volumes are supported.
- variant_
name string - The name of the production variant.
- capacity_
reservation_ objectconfig - Settings for the capacity reservation for the compute instances that SageMaker AI reserves for an endpoint.
- container_
startup_ numberhealth_ check_ timeout_ in_ seconds - The timeout value, in seconds, for your inference container to pass health check by SageMaker Hosting.
- core_
dump_ objectconfig - Specifies configuration for a core dump from the model container when the process crashes.
- enable_
ssm_ boolaccess - You can use this parameter to turn on native AWS Systems Manager (SSM) access for a production variant behind an endpoint. By default, SSM access is disabled for all production variants behind an endpoint.
- inference_
ami_ stringversion - Specifies an option from a collection of preconfigured Amazon Machine Image (AMI) images. Each image is configured by AWS with a set of software and driver versions. AWS optimizes these configurations for different machine learning workloads. By selecting an AMI version, you can ensure that your inference environment is compatible with specific software requirements, such as CUDA driver versions, Linux kernel versions, or AWS Neuron driver versions
- initial_
instance_ numbercount - Number of instances to launch initially.
- initial_
variant_ numberweight - Determines initial traffic distribution among all of the models that you specify in the endpoint configuration.
- instance_
pools list(object) - A list of instance pools for the production variant. Each instance pool specifies an instance type and its priority for provisioning. Use instance pools to configure heterogeneous endpoints that deploy models across multiple instance types.
- instance_
type string - The ML compute instance type.
- managed_
instance_ objectscaling - model_
data_ numberdownload_ timeout_ in_ seconds - The timeout value, in seconds, to download and extract the model that you want to host from Amazon S3 to the individual inference instance associated with this production variant.
- model_
name string - The name of the model that you want to host. This is the name that you specified when creating the model.
- routing_
config object - Settings that control how the endpoint routes incoming traffic to the instances that the endpoint hosts.
- serverless_
config object - The serverless configuration for an endpoint. Specifies a serverless endpoint configuration instead of an instance-based endpoint configuration.
- variant_
instance_ numberprovision_ timeout_ in_ seconds - The timeout value, in seconds, for provisioning instances for the production variant. When SageMaker encounters an insufficient capacity error while provisioning instances, it retries with the next instance pool (if configured) or waits until the timeout expires. This timeout applies only to capacity provisioning and does not include the time for model download or container startup.
- volume_
size_ numberin_ gb - The size, in GB, of the ML storage volume attached to individual inference instance associated with the production variant. Currently only Amazon EBS gp2 storage volumes are supported.
- variant
Name String - The name of the production variant.
- capacity
Reservation EndpointConfig Config Capacity Reservation Config - Settings for the capacity reservation for the compute instances that SageMaker AI reserves for an endpoint.
- container
Startup IntegerHealth Check Timeout In Seconds - The timeout value, in seconds, for your inference container to pass health check by SageMaker Hosting.
- core
Dump EndpointConfig Config Core Dump Config - Specifies configuration for a core dump from the model container when the process crashes.
- enable
Ssm BooleanAccess - You can use this parameter to turn on native AWS Systems Manager (SSM) access for a production variant behind an endpoint. By default, SSM access is disabled for all production variants behind an endpoint.
- inference
Ami StringVersion - Specifies an option from a collection of preconfigured Amazon Machine Image (AMI) images. Each image is configured by AWS with a set of software and driver versions. AWS optimizes these configurations for different machine learning workloads. By selecting an AMI version, you can ensure that your inference environment is compatible with specific software requirements, such as CUDA driver versions, Linux kernel versions, or AWS Neuron driver versions
- initial
Instance IntegerCount - Number of instances to launch initially.
- initial
Variant DoubleWeight - Determines initial traffic distribution among all of the models that you specify in the endpoint configuration.
- instance
Pools List<EndpointConfig Instance Pool> - A list of instance pools for the production variant. Each instance pool specifies an instance type and its priority for provisioning. Use instance pools to configure heterogeneous endpoints that deploy models across multiple instance types.
- instance
Type String - The ML compute instance type.
- managed
Instance EndpointScaling Config Managed Instance Scaling - model
Data IntegerDownload Timeout In Seconds - The timeout value, in seconds, to download and extract the model that you want to host from Amazon S3 to the individual inference instance associated with this production variant.
- model
Name String - The name of the model that you want to host. This is the name that you specified when creating the model.
- routing
Config EndpointConfig Routing Config - Settings that control how the endpoint routes incoming traffic to the instances that the endpoint hosts.
- serverless
Config EndpointConfig Serverless Config - The serverless configuration for an endpoint. Specifies a serverless endpoint configuration instead of an instance-based endpoint configuration.
- variant
Instance IntegerProvision Timeout In Seconds - The timeout value, in seconds, for provisioning instances for the production variant. When SageMaker encounters an insufficient capacity error while provisioning instances, it retries with the next instance pool (if configured) or waits until the timeout expires. This timeout applies only to capacity provisioning and does not include the time for model download or container startup.
- volume
Size IntegerIn Gb - The size, in GB, of the ML storage volume attached to individual inference instance associated with the production variant. Currently only Amazon EBS gp2 storage volumes are supported.
- variant
Name string - The name of the production variant.
- capacity
Reservation EndpointConfig Config Capacity Reservation Config - Settings for the capacity reservation for the compute instances that SageMaker AI reserves for an endpoint.
- container
Startup numberHealth Check Timeout In Seconds - The timeout value, in seconds, for your inference container to pass health check by SageMaker Hosting.
- core
Dump EndpointConfig Config Core Dump Config - Specifies configuration for a core dump from the model container when the process crashes.
- enable
Ssm booleanAccess - You can use this parameter to turn on native AWS Systems Manager (SSM) access for a production variant behind an endpoint. By default, SSM access is disabled for all production variants behind an endpoint.
- inference
Ami stringVersion - Specifies an option from a collection of preconfigured Amazon Machine Image (AMI) images. Each image is configured by AWS with a set of software and driver versions. AWS optimizes these configurations for different machine learning workloads. By selecting an AMI version, you can ensure that your inference environment is compatible with specific software requirements, such as CUDA driver versions, Linux kernel versions, or AWS Neuron driver versions
- initial
Instance numberCount - Number of instances to launch initially.
- initial
Variant numberWeight - Determines initial traffic distribution among all of the models that you specify in the endpoint configuration.
- instance
Pools EndpointConfig Instance Pool[] - A list of instance pools for the production variant. Each instance pool specifies an instance type and its priority for provisioning. Use instance pools to configure heterogeneous endpoints that deploy models across multiple instance types.
- instance
Type string - The ML compute instance type.
- managed
Instance EndpointScaling Config Managed Instance Scaling - model
Data numberDownload Timeout In Seconds - The timeout value, in seconds, to download and extract the model that you want to host from Amazon S3 to the individual inference instance associated with this production variant.
- model
Name string - The name of the model that you want to host. This is the name that you specified when creating the model.
- routing
Config EndpointConfig Routing Config - Settings that control how the endpoint routes incoming traffic to the instances that the endpoint hosts.
- serverless
Config EndpointConfig Serverless Config - The serverless configuration for an endpoint. Specifies a serverless endpoint configuration instead of an instance-based endpoint configuration.
- variant
Instance numberProvision Timeout In Seconds - The timeout value, in seconds, for provisioning instances for the production variant. When SageMaker encounters an insufficient capacity error while provisioning instances, it retries with the next instance pool (if configured) or waits until the timeout expires. This timeout applies only to capacity provisioning and does not include the time for model download or container startup.
- volume
Size numberIn Gb - The size, in GB, of the ML storage volume attached to individual inference instance associated with the production variant. Currently only Amazon EBS gp2 storage volumes are supported.
- variant_
name str - The name of the production variant.
- capacity_
reservation_ Endpointconfig Config Capacity Reservation Config - Settings for the capacity reservation for the compute instances that SageMaker AI reserves for an endpoint.
- container_
startup_ inthealth_ check_ timeout_ in_ seconds - The timeout value, in seconds, for your inference container to pass health check by SageMaker Hosting.
- core_
dump_ Endpointconfig Config Core Dump Config - Specifies configuration for a core dump from the model container when the process crashes.
- enable_
ssm_ boolaccess - You can use this parameter to turn on native AWS Systems Manager (SSM) access for a production variant behind an endpoint. By default, SSM access is disabled for all production variants behind an endpoint.
- inference_
ami_ strversion - Specifies an option from a collection of preconfigured Amazon Machine Image (AMI) images. Each image is configured by AWS with a set of software and driver versions. AWS optimizes these configurations for different machine learning workloads. By selecting an AMI version, you can ensure that your inference environment is compatible with specific software requirements, such as CUDA driver versions, Linux kernel versions, or AWS Neuron driver versions
- initial_
instance_ intcount - Number of instances to launch initially.
- initial_
variant_ floatweight - Determines initial traffic distribution among all of the models that you specify in the endpoint configuration.
- instance_
pools Sequence[EndpointConfig Instance Pool] - A list of instance pools for the production variant. Each instance pool specifies an instance type and its priority for provisioning. Use instance pools to configure heterogeneous endpoints that deploy models across multiple instance types.
- instance_
type str - The ML compute instance type.
- managed_
instance_ Endpointscaling Config Managed Instance Scaling - model_
data_ intdownload_ timeout_ in_ seconds - The timeout value, in seconds, to download and extract the model that you want to host from Amazon S3 to the individual inference instance associated with this production variant.
- model_
name str - The name of the model that you want to host. This is the name that you specified when creating the model.
- routing_
config EndpointConfig Routing Config - Settings that control how the endpoint routes incoming traffic to the instances that the endpoint hosts.
- serverless_
config EndpointConfig Serverless Config - The serverless configuration for an endpoint. Specifies a serverless endpoint configuration instead of an instance-based endpoint configuration.
- variant_
instance_ intprovision_ timeout_ in_ seconds - The timeout value, in seconds, for provisioning instances for the production variant. When SageMaker encounters an insufficient capacity error while provisioning instances, it retries with the next instance pool (if configured) or waits until the timeout expires. This timeout applies only to capacity provisioning and does not include the time for model download or container startup.
- volume_
size_ intin_ gb - The size, in GB, of the ML storage volume attached to individual inference instance associated with the production variant. Currently only Amazon EBS gp2 storage volumes are supported.
- variant
Name String - The name of the production variant.
- capacity
Reservation Property MapConfig - Settings for the capacity reservation for the compute instances that SageMaker AI reserves for an endpoint.
- container
Startup NumberHealth Check Timeout In Seconds - The timeout value, in seconds, for your inference container to pass health check by SageMaker Hosting.
- core
Dump Property MapConfig - Specifies configuration for a core dump from the model container when the process crashes.
- enable
Ssm BooleanAccess - You can use this parameter to turn on native AWS Systems Manager (SSM) access for a production variant behind an endpoint. By default, SSM access is disabled for all production variants behind an endpoint.
- inference
Ami StringVersion - Specifies an option from a collection of preconfigured Amazon Machine Image (AMI) images. Each image is configured by AWS with a set of software and driver versions. AWS optimizes these configurations for different machine learning workloads. By selecting an AMI version, you can ensure that your inference environment is compatible with specific software requirements, such as CUDA driver versions, Linux kernel versions, or AWS Neuron driver versions
- initial
Instance NumberCount - Number of instances to launch initially.
- initial
Variant NumberWeight - Determines initial traffic distribution among all of the models that you specify in the endpoint configuration.
- instance
Pools List<Property Map> - A list of instance pools for the production variant. Each instance pool specifies an instance type and its priority for provisioning. Use instance pools to configure heterogeneous endpoints that deploy models across multiple instance types.
- instance
Type String - The ML compute instance type.
- managed
Instance Property MapScaling - model
Data NumberDownload Timeout In Seconds - The timeout value, in seconds, to download and extract the model that you want to host from Amazon S3 to the individual inference instance associated with this production variant.
- model
Name String - The name of the model that you want to host. This is the name that you specified when creating the model.
- routing
Config Property Map - Settings that control how the endpoint routes incoming traffic to the instances that the endpoint hosts.
- serverless
Config Property Map - The serverless configuration for an endpoint. Specifies a serverless endpoint configuration instead of an instance-based endpoint configuration.
- variant
Instance NumberProvision Timeout In Seconds - The timeout value, in seconds, for provisioning instances for the production variant. When SageMaker encounters an insufficient capacity error while provisioning instances, it retries with the next instance pool (if configured) or waits until the timeout expires. This timeout applies only to capacity provisioning and does not include the time for model download or container startup.
- volume
Size NumberIn Gb - The size, in GB, of the ML storage volume attached to individual inference instance associated with the production variant. Currently only Amazon EBS gp2 storage volumes are supported.
EndpointConfigRoutingConfig, EndpointConfigRoutingConfigArgs
Settings that control how the endpoint routes incoming traffic to the instances that the endpoint hosts.- Prefix
Aware Pulumi.Routing Config Aws Native. Sage Maker. Inputs. Endpoint Config Prefix Aware Routing Config - The configuration for prefix-aware routing. Specify this property only when you set RoutingStrategy to PREFIX_AWARE.
- Routing
Strategy string - Sets how the endpoint routes incoming traffic.
- Prefix
Aware EndpointRouting Config Config Prefix Aware Routing Config - The configuration for prefix-aware routing. Specify this property only when you set RoutingStrategy to PREFIX_AWARE.
- Routing
Strategy string - Sets how the endpoint routes incoming traffic.
- prefix_
aware_ objectrouting_ config - The configuration for prefix-aware routing. Specify this property only when you set RoutingStrategy to PREFIX_AWARE.
- routing_
strategy string - Sets how the endpoint routes incoming traffic.
- prefix
Aware EndpointRouting Config Config Prefix Aware Routing Config - The configuration for prefix-aware routing. Specify this property only when you set RoutingStrategy to PREFIX_AWARE.
- routing
Strategy String - Sets how the endpoint routes incoming traffic.
- prefix
Aware EndpointRouting Config Config Prefix Aware Routing Config - The configuration for prefix-aware routing. Specify this property only when you set RoutingStrategy to PREFIX_AWARE.
- routing
Strategy string - Sets how the endpoint routes incoming traffic.
- prefix_
aware_ Endpointrouting_ config Config Prefix Aware Routing Config - The configuration for prefix-aware routing. Specify this property only when you set RoutingStrategy to PREFIX_AWARE.
- routing_
strategy str - Sets how the endpoint routes incoming traffic.
- prefix
Aware Property MapRouting Config - The configuration for prefix-aware routing. Specify this property only when you set RoutingStrategy to PREFIX_AWARE.
- routing
Strategy String - Sets how the endpoint routes incoming traffic.
EndpointConfigScaleInPolicy, EndpointConfigScaleInPolicyArgs
Specifies how the endpoint releases instances when managed instance scaling scales in.- Strategy string
- The strategy for scaling in instances. IDLE_RELEASE releases instances that have no hosted inference component copies. CONSOLIDATION consolidates inference component copies onto fewer instances to release more instances.
- Cooldown
In intMinutes - The cooldown period, in minutes, after the last endpoint operation before the endpoint evaluates consolidation scale-in opportunities. Valid values are 5 to 1440. The default is 20.
- Maximum
Step intSize - The maximum number of instances that the endpoint can terminate at a time during a consolidation scale-in operation. Valid values are 1 to 100. The default is 1.
- Strategy string
- The strategy for scaling in instances. IDLE_RELEASE releases instances that have no hosted inference component copies. CONSOLIDATION consolidates inference component copies onto fewer instances to release more instances.
- Cooldown
In intMinutes - The cooldown period, in minutes, after the last endpoint operation before the endpoint evaluates consolidation scale-in opportunities. Valid values are 5 to 1440. The default is 20.
- Maximum
Step intSize - The maximum number of instances that the endpoint can terminate at a time during a consolidation scale-in operation. Valid values are 1 to 100. The default is 1.
- strategy string
- The strategy for scaling in instances. IDLE_RELEASE releases instances that have no hosted inference component copies. CONSOLIDATION consolidates inference component copies onto fewer instances to release more instances.
- cooldown_
in_ numberminutes - The cooldown period, in minutes, after the last endpoint operation before the endpoint evaluates consolidation scale-in opportunities. Valid values are 5 to 1440. The default is 20.
- maximum_
step_ numbersize - The maximum number of instances that the endpoint can terminate at a time during a consolidation scale-in operation. Valid values are 1 to 100. The default is 1.
- strategy String
- The strategy for scaling in instances. IDLE_RELEASE releases instances that have no hosted inference component copies. CONSOLIDATION consolidates inference component copies onto fewer instances to release more instances.
- cooldown
In IntegerMinutes - The cooldown period, in minutes, after the last endpoint operation before the endpoint evaluates consolidation scale-in opportunities. Valid values are 5 to 1440. The default is 20.
- maximum
Step IntegerSize - The maximum number of instances that the endpoint can terminate at a time during a consolidation scale-in operation. Valid values are 1 to 100. The default is 1.
- strategy string
- The strategy for scaling in instances. IDLE_RELEASE releases instances that have no hosted inference component copies. CONSOLIDATION consolidates inference component copies onto fewer instances to release more instances.
- cooldown
In numberMinutes - The cooldown period, in minutes, after the last endpoint operation before the endpoint evaluates consolidation scale-in opportunities. Valid values are 5 to 1440. The default is 20.
- maximum
Step numberSize - The maximum number of instances that the endpoint can terminate at a time during a consolidation scale-in operation. Valid values are 1 to 100. The default is 1.
- strategy str
- The strategy for scaling in instances. IDLE_RELEASE releases instances that have no hosted inference component copies. CONSOLIDATION consolidates inference component copies onto fewer instances to release more instances.
- cooldown_
in_ intminutes - The cooldown period, in minutes, after the last endpoint operation before the endpoint evaluates consolidation scale-in opportunities. Valid values are 5 to 1440. The default is 20.
- maximum_
step_ intsize - The maximum number of instances that the endpoint can terminate at a time during a consolidation scale-in operation. Valid values are 1 to 100. The default is 1.
- strategy String
- The strategy for scaling in instances. IDLE_RELEASE releases instances that have no hosted inference component copies. CONSOLIDATION consolidates inference component copies onto fewer instances to release more instances.
- cooldown
In NumberMinutes - The cooldown period, in minutes, after the last endpoint operation before the endpoint evaluates consolidation scale-in opportunities. Valid values are 5 to 1440. The default is 20.
- maximum
Step NumberSize - The maximum number of instances that the endpoint can terminate at a time during a consolidation scale-in operation. Valid values are 1 to 100. The default is 1.
EndpointConfigServerlessConfig, EndpointConfigServerlessConfigArgs
Specifies the serverless configuration for an endpoint variant.- Max
Concurrency int - The maximum number of concurrent invocations your serverless endpoint can process.
- Memory
Size intIn Mb - The memory size of your serverless endpoint. Valid values are in 1 GB increments: 1024 MB, 2048 MB, 3072 MB, 4096 MB, 5120 MB, or 6144 MB.
- Provisioned
Concurrency int - The amount of provisioned concurrency to allocate for the serverless endpoint. Should be less than or equal to MaxConcurrency.
- Max
Concurrency int - The maximum number of concurrent invocations your serverless endpoint can process.
- Memory
Size intIn Mb - The memory size of your serverless endpoint. Valid values are in 1 GB increments: 1024 MB, 2048 MB, 3072 MB, 4096 MB, 5120 MB, or 6144 MB.
- Provisioned
Concurrency int - The amount of provisioned concurrency to allocate for the serverless endpoint. Should be less than or equal to MaxConcurrency.
- max_
concurrency number - The maximum number of concurrent invocations your serverless endpoint can process.
- memory_
size_ numberin_ mb - The memory size of your serverless endpoint. Valid values are in 1 GB increments: 1024 MB, 2048 MB, 3072 MB, 4096 MB, 5120 MB, or 6144 MB.
- provisioned_
concurrency number - The amount of provisioned concurrency to allocate for the serverless endpoint. Should be less than or equal to MaxConcurrency.
- max
Concurrency Integer - The maximum number of concurrent invocations your serverless endpoint can process.
- memory
Size IntegerIn Mb - The memory size of your serverless endpoint. Valid values are in 1 GB increments: 1024 MB, 2048 MB, 3072 MB, 4096 MB, 5120 MB, or 6144 MB.
- provisioned
Concurrency Integer - The amount of provisioned concurrency to allocate for the serverless endpoint. Should be less than or equal to MaxConcurrency.
- max
Concurrency number - The maximum number of concurrent invocations your serverless endpoint can process.
- memory
Size numberIn Mb - The memory size of your serverless endpoint. Valid values are in 1 GB increments: 1024 MB, 2048 MB, 3072 MB, 4096 MB, 5120 MB, or 6144 MB.
- provisioned
Concurrency number - The amount of provisioned concurrency to allocate for the serverless endpoint. Should be less than or equal to MaxConcurrency.
- max_
concurrency int - The maximum number of concurrent invocations your serverless endpoint can process.
- memory_
size_ intin_ mb - The memory size of your serverless endpoint. Valid values are in 1 GB increments: 1024 MB, 2048 MB, 3072 MB, 4096 MB, 5120 MB, or 6144 MB.
- provisioned_
concurrency int - The amount of provisioned concurrency to allocate for the serverless endpoint. Should be less than or equal to MaxConcurrency.
- max
Concurrency Number - The maximum number of concurrent invocations your serverless endpoint can process.
- memory
Size NumberIn Mb - The memory size of your serverless endpoint. Valid values are in 1 GB increments: 1024 MB, 2048 MB, 3072 MB, 4096 MB, 5120 MB, or 6144 MB.
- provisioned
Concurrency Number - The amount of provisioned concurrency to allocate for the serverless endpoint. Should be less than or equal to MaxConcurrency.
EndpointConfigVpcConfig, EndpointConfigVpcConfigArgs
Specifies an Amazon Virtual Private Cloud (VPC) that your SageMaker jobs, hosted models, and compute resources have access to. You can control access to and from your resources by configuring a VPC.- Security
Group List<string>Ids - The VPC security group IDs, in the form sg-xxxxxxxx. Specify the security groups for the VPC that is specified in the Subnets field.
- Subnets List<string>
- The ID of the subnets in the VPC to which you want to connect your training job or model.
- Security
Group []stringIds - The VPC security group IDs, in the form sg-xxxxxxxx. Specify the security groups for the VPC that is specified in the Subnets field.
- Subnets []string
- The ID of the subnets in the VPC to which you want to connect your training job or model.
- security_
group_ list(string)ids - The VPC security group IDs, in the form sg-xxxxxxxx. Specify the security groups for the VPC that is specified in the Subnets field.
- subnets list(string)
- The ID of the subnets in the VPC to which you want to connect your training job or model.
- security
Group List<String>Ids - The VPC security group IDs, in the form sg-xxxxxxxx. Specify the security groups for the VPC that is specified in the Subnets field.
- subnets List<String>
- The ID of the subnets in the VPC to which you want to connect your training job or model.
- security
Group string[]Ids - The VPC security group IDs, in the form sg-xxxxxxxx. Specify the security groups for the VPC that is specified in the Subnets field.
- subnets string[]
- The ID of the subnets in the VPC to which you want to connect your training job or model.
- security_
group_ Sequence[str]ids - The VPC security group IDs, in the form sg-xxxxxxxx. Specify the security groups for the VPC that is specified in the Subnets field.
- subnets Sequence[str]
- The ID of the subnets in the VPC to which you want to connect your training job or model.
- security
Group List<String>Ids - The VPC security group IDs, in the form sg-xxxxxxxx. Specify the security groups for the VPC that is specified in the Subnets field.
- subnets List<String>
- The ID of the subnets in the VPC to which you want to connect your training job or model.
Tag, TagArgs
A set of tags to apply to the resource.Package Details
- Repository
- AWS Native pulumi/pulumi-aws-native
- License
- Apache-2.0
We recommend new projects start with resources from the AWS provider.
published on Monday, Sep 28, 2026 by Pulumi