| Column Name | Column Data Type |
inference_recommendations_job_name Required Input Column
The name of a previously completed Inference Recommender job. | VARCHAR |
endpoint_name Input Column
The name of an endpoint benchmarked during a previously completed Inference Recommender job. | VARCHAR |
recommendation_id Input Column
The recommendation ID of a previously completed inference recommendation. | VARCHAR |
scaling_policy_objective Input Column
An object representing the anticipated traffic pattern for an endpoint that you specified in the request. | STRUCT( "min_invocations_per_minute" BIGINT, "max_invocations_per_minute" BIGINT ) |
Show child fields- scaling_policy_objective.max_invocations_per_minute
The maximum number of expected requests to your endpoint per minute.
- scaling_policy_objective.min_invocations_per_minute
The minimum number of expected requests to your endpoint per minute.
|
target_cpu_utilization_per_core Input Column
The percentage of how much utilization you want an instance to use before autoscaling, which you specified in the request. The default value is 50%. | BIGINT |
_aws_profile Input Column
The AWS profile defines the AWS identity used. It can be defined via credentials or by assuming a IAM role. | STRUCT( "type" VARCHAR, "name" VARCHAR, "account_id" VARCHAR, "via_profile_name" VARCHAR, "assumed_role_arn" VARCHAR, "organization" STRUCT( "account_name" VARCHAR, "id" VARCHAR, "tags" STRUCT( "key" VARCHAR, "value" VARCHAR )[], "master_account" STRUCT( "id" VARCHAR, "email" VARCHAR ), "parents" STRUCT( "type" VARCHAR, "id" VARCHAR, "name" VARCHAR, "tags" STRUCT( "key" VARCHAR, "value" VARCHAR )[] )[] ) ) |
Show child fields- _aws_profile.account_id
The AWS account id
- _aws_profile.assumed_role_arn
The ARN of the assumed role
- _aws_profile.name
The unique name of the profile.
- _aws_profile.organization
Information about this profile's membership in the AWS organization. Show child fields- _aws_profile.organization.account_name
The name of account speciifed by the organization
- _aws_profile.organization.id
The organization id
- _aws_profile.organization.master_account
Show child fields- _aws_profile.organization.master_account.email
The organization master account email address
- _aws_profile.organization.master_account.id
The organization master account id
- _aws_profile.organization.parents[]
Show child fields- _aws_profile.organization.parents[].id
The id of the parent
- _aws_profile.organization.parents[].name
The name of the parent
- _aws_profile.organization.parents[].tags[]
Show child fields- _aws_profile.organization.parents[].tags[].key
- _aws_profile.organization.parents[].tags[].value
- _aws_profile.organization.parents[].type
The type of parent can be an organization unit or a root
- _aws_profile.organization.tags[]
Show child fields- _aws_profile.organization.tags[].key
- _aws_profile.organization.tags[].value
- _aws_profile.type
The type of profile, either 'credentials' or 'assumed_role'
- _aws_profile.via_profile_name
This IAM role for this profile is assumed by first utilizing another profile with this name to obtain credentials.
|
dynamic_scaling_configuration
An object with the recommended values for you to specify when creating an autoscaling policy. | STRUCT( "min_capacity" BIGINT, "max_capacity" BIGINT, "scale_in_cooldown" BIGINT, "scale_out_cooldown" BIGINT, "scaling_policies" STRUCT( "target_tracking" STRUCT( "metric_specification" STRUCT( "predefined" STRUCT( "predefined_metric_type" VARCHAR ), "customized" STRUCT( "metric_name" VARCHAR, "namespace" VARCHAR, "statistic" VARCHAR ) ), "target_value" DOUBLE ) )[] ) |
Show child fields- dynamic_scaling_configuration.max_capacity
The recommended maximum capacity to specify for your autoscaling policy.
- dynamic_scaling_configuration.min_capacity
The recommended minimum capacity to specify for your autoscaling policy.
- dynamic_scaling_configuration.scale_in_cooldown
The recommended scale in cooldown time for your autoscaling policy.
- dynamic_scaling_configuration.scale_out_cooldown
The recommended scale out cooldown time for your autoscaling policy.
- dynamic_scaling_configuration.scaling_policies[]
Show child fields- dynamic_scaling_configuration.scaling_policies[].target_tracking
A target tracking scaling policy. Includes support for predefined or customized metrics. Show child fields- dynamic_scaling_configuration.scaling_policies[].target_tracking.metric_specification
An object containing information about a metric. Show child fields- dynamic_scaling_configuration.scaling_policies[].target_tracking.metric_specification.customized
Information about a customized metric. Show child fields- dynamic_scaling_configuration.scaling_policies[].target_tracking.metric_specification.customized.metric_name
The name of the customized metric.
- dynamic_scaling_configuration.scaling_policies[].target_tracking.metric_specification.customized.namespace
The namespace of the customized metric.
- dynamic_scaling_configuration.scaling_policies[].target_tracking.metric_specification.customized.statistic
The statistic of the customized metric.
- dynamic_scaling_configuration.scaling_policies[].target_tracking.metric_specification.predefined
Information about a predefined metric. Show child fields- dynamic_scaling_configuration.scaling_policies[].target_tracking.metric_specification.predefined.predefined_metric_type
The metric type. You can only apply SageMaker metric types to SageMaker endpoints.
- dynamic_scaling_configuration.scaling_policies[].target_tracking.target_value
The recommended target value to specify for the metric when creating a scaling policy.
|
metric
An object with a list of metrics that were benchmarked during the previously completed Inference Recommender job. | STRUCT( "invocations_per_instance" BIGINT, "model_latency" BIGINT ) |
Show child fields- metric.invocations_per_instance
The number of invocations sent to a model, normalized by InstanceCount in each ProductionVariant. 1/numberOfInstances is sent as the value on each request, where numberOfInstances is the number of active instances for the ProductionVariant behind the endpoint at the time of the request.
- metric.model_latency
The interval of time taken by a model to respond as viewed from SageMaker. This interval includes the local communication times taken to send the request and to fetch the response from the container of a model and the time taken to complete the inference in the container.
|