Skip to content

aws.sagemaker.get_scaling_configuration_recommendation

Example SQL Queries

SELECT * FROM
aws.sagemaker.get_scaling_configuration_recommendation
WHERE
"inference_recommendations_job_name" = 'VALUE';

Description

Starts an Amazon SageMaker Inference Recommender autoscaling recommendation job. Returns recommendations for autoscaling policies that you can apply to your SageMaker endpoint.

Table Definition

Column NameColumn Data Type
inference_recommendations_job_name Required Input Column

The name of a previously completed Inference Recommender job.

VARCHAR
endpoint_name Input Column

The name of an endpoint benchmarked during a previously completed Inference Recommender job.

VARCHAR
recommendation_id Input Column

The recommendation ID of a previously completed inference recommendation.

VARCHAR
scaling_policy_objective Input Column

An object representing the anticipated traffic pattern for an endpoint that you specified in the request.

STRUCT(
"min_invocations_per_minute" BIGINT,
"max_invocations_per_minute" BIGINT
)
Show child fields
scaling_policy_objective.max_invocations_per_minute

The maximum number of expected requests to your endpoint per minute.

scaling_policy_objective.min_invocations_per_minute

The minimum number of expected requests to your endpoint per minute.

target_cpu_utilization_per_core Input Column

The percentage of how much utilization you want an instance to use before autoscaling, which you specified in the request. The default value is 50%.

BIGINT
_aws_profile Input Column

The AWS profile defines the AWS identity used. It can be defined via credentials or by assuming a IAM role.

STRUCT(
"type" VARCHAR,
"name" VARCHAR,
"account_id" VARCHAR,
"via_profile_name" VARCHAR,
"assumed_role_arn" VARCHAR,
"organization" STRUCT(
"account_name" VARCHAR,
"id" VARCHAR,
"tags" STRUCT(
"key" VARCHAR,
"value" VARCHAR
)[],
"master_account" STRUCT(
"id" VARCHAR,
"email" VARCHAR
),
"parents" STRUCT(
"type" VARCHAR,
"id" VARCHAR,
"name" VARCHAR,
"tags" STRUCT(
"key" VARCHAR,
"value" VARCHAR
)[]
)[]
)
)
Show child fields
_aws_profile.account_id

The AWS account id

_aws_profile.assumed_role_arn

The ARN of the assumed role

_aws_profile.name

The unique name of the profile.

_aws_profile.organization

Information about this profile's membership in the AWS organization.

Show child fields
_aws_profile.organization.account_name

The name of account speciifed by the organization

_aws_profile.organization.id

The organization id

_aws_profile.organization.master_account
Show child fields
_aws_profile.organization.master_account.email

The organization master account email address

_aws_profile.organization.master_account.id

The organization master account id

_aws_profile.organization.parents[]
Show child fields
_aws_profile.organization.parents[].id

The id of the parent

_aws_profile.organization.parents[].name

The name of the parent

_aws_profile.organization.parents[].tags[]
Show child fields
_aws_profile.organization.parents[].tags[].key
_aws_profile.organization.parents[].tags[].value
_aws_profile.organization.parents[].type

The type of parent can be an organization unit or a root

_aws_profile.organization.tags[]
Show child fields
_aws_profile.organization.tags[].key
_aws_profile.organization.tags[].value
_aws_profile.type

The type of profile, either 'credentials' or 'assumed_role'

_aws_profile.via_profile_name

This IAM role for this profile is assumed by first utilizing another profile with this name to obtain credentials.

dynamic_scaling_configuration

An object with the recommended values for you to specify when creating an autoscaling policy.

STRUCT(
"min_capacity" BIGINT,
"max_capacity" BIGINT,
"scale_in_cooldown" BIGINT,
"scale_out_cooldown" BIGINT,
"scaling_policies" STRUCT(
"target_tracking" STRUCT(
"metric_specification" STRUCT(
"predefined" STRUCT(
"predefined_metric_type" VARCHAR
),
"customized" STRUCT(
"metric_name" VARCHAR,
"namespace" VARCHAR,
"statistic" VARCHAR
)
),
"target_value" DOUBLE
)
)[]
)
Show child fields
dynamic_scaling_configuration.max_capacity

The recommended maximum capacity to specify for your autoscaling policy.

dynamic_scaling_configuration.min_capacity

The recommended minimum capacity to specify for your autoscaling policy.

dynamic_scaling_configuration.scale_in_cooldown

The recommended scale in cooldown time for your autoscaling policy.

dynamic_scaling_configuration.scale_out_cooldown

The recommended scale out cooldown time for your autoscaling policy.

dynamic_scaling_configuration.scaling_policies[]
Show child fields
dynamic_scaling_configuration.scaling_policies[].target_tracking

A target tracking scaling policy. Includes support for predefined or customized metrics.

Show child fields
dynamic_scaling_configuration.scaling_policies[].target_tracking.metric_specification

An object containing information about a metric.

Show child fields
dynamic_scaling_configuration.scaling_policies[].target_tracking.metric_specification.customized

Information about a customized metric.

Show child fields
dynamic_scaling_configuration.scaling_policies[].target_tracking.metric_specification.customized.metric_name

The name of the customized metric.

dynamic_scaling_configuration.scaling_policies[].target_tracking.metric_specification.customized.namespace

The namespace of the customized metric.

dynamic_scaling_configuration.scaling_policies[].target_tracking.metric_specification.customized.statistic

The statistic of the customized metric.

dynamic_scaling_configuration.scaling_policies[].target_tracking.metric_specification.predefined

Information about a predefined metric.

Show child fields
dynamic_scaling_configuration.scaling_policies[].target_tracking.metric_specification.predefined.predefined_metric_type

The metric type. You can only apply SageMaker metric types to SageMaker endpoints.

dynamic_scaling_configuration.scaling_policies[].target_tracking.target_value

The recommended target value to specify for the metric when creating a scaling policy.

metric

An object with a list of metrics that were benchmarked during the previously completed Inference Recommender job.

STRUCT(
"invocations_per_instance" BIGINT,
"model_latency" BIGINT
)
Show child fields
metric.invocations_per_instance

The number of invocations sent to a model, normalized by InstanceCount in each ProductionVariant. 1/numberOfInstances is sent as the value on each request, where numberOfInstances is the number of active instances for the ProductionVariant behind the endpoint at the time of the request.

metric.model_latency

The interval of time taken by a model to respond as viewed from SageMaker. This interval includes the local communication times taken to send the request and to fetch the response from the container of a model and the time taken to complete the inference in the container.