<?xml version="1.0" encoding="UTF-8" standalone="no"?><rss xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:slash="http://purl.org/rss/1.0/modules/slash/" xmlns:sy="http://purl.org/rss/1.0/modules/syndication/" xmlns:wfw="http://wellformedweb.org/CommentAPI/" version="2.0">

<channel>
	<title>AWS Compute Blog</title>
	<atom:link href="https://aws.amazon.com/blogs/compute/feed/" rel="self" type="application/rss+xml"/>
	<link>https://aws.amazon.com/blogs/compute/</link>
	<description/>
	<lastBuildDate>Tue, 29 Sep 2026 17:14:00 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	
	<item>
		<title>Implementing customer managed keys for AWS Lambda durable functions with Terraform</title>
		<link>https://aws.amazon.com/blogs/compute/implementing-customer-managed-keys-for-aws-lambda-durable-functions-with-terraform/</link>
		
		<dc:creator><![CDATA[Rajdeep Banerjee]]></dc:creator>
		<pubDate>Tue, 29 Sep 2026 17:14:00 +0000</pubDate>
				<category><![CDATA[Advanced (300)]]></category>
		<category><![CDATA[AWS Key Management Service]]></category>
		<category><![CDATA[AWS Lambda]]></category>
		<category><![CDATA[Technical How-to]]></category>
		<guid isPermaLink="false">db91eabc9cfca26ee3fbbed72ac7160e425845ca</guid>

					<description>Lambda durable functions checkpoint execution state to durable storage, and for regulated payment workloads that data is sensitive. This post shows how to configure a customer managed key in AWS KMS to encrypt durable execution data, define a least-privilege key policy, and verify encryption through AWS CloudTrail, all deployed with Terraform.</description>
										<content:encoded>&lt;p&gt;If you run regulated workloads, you must control how persisted data is encrypted and who can access it. You need to manage encryption key rotation schedules, restrict decryption to authorized principals, and produce audit evidence that proves encryption controls are operating as designed.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/durable-functions.html" target="_blank" rel="noopener"&gt;AWS Lambda durable functions&lt;/a&gt; build resilient, multi-step workflows that survive failures through automatic checkpointing. The checkpoint mechanism persists execution state, including step results, payloads, and callback responses, to durable storage. For payment processing workloads, this persisted data is sensitive. AWS Lambda durable functions support &lt;a href="https://docs.aws.amazon.com/kms/latest/cryptographic-details/basic-concepts.html" target="_blank" rel="noopener"&gt;customer managed keys&lt;/a&gt; from &lt;a href="https://aws.amazon.com/kms/" target="_blank" rel="noopener"&gt;AWS Key Management Service (AWS KMS)&lt;/a&gt;. A customer managed key gives you three controls: you set the key rotation schedule, you restrict decryption access through the key policy, and you generate per-function audit trails in &lt;a href="https://aws.amazon.com/cloudtrail/" target="_blank" rel="noopener"&gt;AWS CloudTrail&lt;/a&gt;. A durable execution uses the same encryption key it started with for its entire lifetime. Changing or removing the key affects only executions that start after the change.&lt;/p&gt;
&lt;p&gt;Updating the customer managed key policy to remove decrypt permissions, or disabling the key, stops the Lambda service from accessing previously checkpointed state. Customer managed key deletion is a permanent action, and all durable executions encrypted with that key become unrecoverable because the Lambda service has no mechanism to restore the data. Before scheduling key deletion, use the AWS KMS waiting period (7 to 30 days) and monitor AWS CloudTrail for &lt;code&gt;Decrypt&lt;/code&gt; calls to confirm that the key is no longer in active use.&lt;/p&gt;
&lt;p&gt;In this post, you learn to configure a customer managed key to encrypt durable execution data in an event-driven payment processing workflow. You create a &lt;a href="https://docs.aws.amazon.com/kms/latest/developerguide/create-symmetric-cmk.html" target="_blank" rel="noopener"&gt;symmetric encryption key in AWS KMS&lt;/a&gt; and define a key policy that grants the Lambda service, the function’s execution role, the function author, and durable execution operators only the AWS KMS actions each principal requires. You then configure the function to use the key for durable execution encryption and verify encryption operations through AWS CloudTrail logs. By the end, you have a deployable reference architecture you can adapt for regulated workloads running on Lambda durable functions.&lt;/p&gt;
&lt;p&gt;To learn more about how AWS Lambda encrypts durable execution data, see &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/durable-encryption.html" target="_blank" rel="noopener"&gt;Encrypting AWS Lambda durable execution data&lt;/a&gt; in the AWS Lambda Developer Guide.&lt;/p&gt;
&lt;h2 id="solution-overview"&gt;Solution overview&lt;/h2&gt;
&lt;p&gt;The sample application implements an event-driven payment processing pipeline using &lt;a href="https://aws.amazon.com/dynamodb/" target="_blank" rel="noopener"&gt;Amazon DynamoDB&lt;/a&gt;, &lt;a href="https://aws.amazon.com/eventbridge/" target="_blank" rel="noopener"&gt;Amazon EventBridge&lt;/a&gt;, &lt;a href="https://docs.aws.amazon.com/eventbridge/latest/userguide/eb-pipes.html" target="_blank" rel="noopener"&gt;Amazon EventBridge Pipes&lt;/a&gt;, &lt;a href="https://aws.amazon.com/lambda/" target="_blank" rel="noopener"&gt;AWS Lambda&lt;/a&gt;, and &lt;a href="https://aws.amazon.com/sqs/" target="_blank" rel="noopener"&gt;Amazon SQS&lt;/a&gt;. The pipeline receives authorized payment transactions, validates and enriches them. A Lambda durable function applies business rules to the enriched transactions. The approved transactions are sent to a downstream settlement system for posting.&lt;/p&gt;
&lt;p&gt;The following section covers the key architectural steps.&lt;/p&gt;
&lt;h3 id="architecture-steps"&gt;Architecture steps&lt;/h3&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;The upstream authorization system writes authorized payment records to a DynamoDB table.&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/Streams.html" target="_blank" rel="noopener"&gt;DynamoDB Streams&lt;/a&gt; captures each new record as an ordered change event.&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://aws.amazon.com/eventbridge/pipes/" target="_blank" rel="noopener"&gt;Amazon EventBridge Pipes&lt;/a&gt; polls the record from the DynamoDB stream. The pipe triggers a Lambda function as part of &lt;a href="https://docs.aws.amazon.com/eventbridge/latest/userguide/eb-pipes-logs-execution-steps.html" target="_blank" rel="noopener"&gt;enrichment&lt;/a&gt; step for duplicate checking.&lt;/li&gt;
 &lt;li&gt;The deduplication Lambda uses a DynamoDB table with &lt;a href="https://aws.amazon.com/blogs/database/building-distributed-locks-with-the-dynamodb-lock-client/" target="_blank" rel="noopener"&gt;conditional writes&lt;/a&gt; to identify duplicate inbound transactions based on transaction properties and time window.&lt;/li&gt;
 &lt;li&gt;When the deduplication is successful, the pipe publishes an event to the Amazon EventBridge custom event bus.&lt;/li&gt;
 &lt;li&gt;An Amazon EventBridge rule invokes a Lambda function for matching events. The function adds business context such as account type, bank routing details, and merchant category codes. The function publishes a new enriched event to the custom event bus.&lt;/li&gt;
 &lt;li&gt;Another Amazon EventBridge rule matches the enriched events to a Lambda durable function. The durable function applies business rules to the incoming event. When the event passes all business rules, the function publishes a new event to the event bus.&lt;/li&gt;
 &lt;li&gt;An Amazon EventBridge rule routes the approved event to an Amazon SQS queue preserving ordering for settlement and buffering against downstream throughput limits.&lt;/li&gt;
 &lt;li&gt;The Posting Lambda function reads from the Amazon SQS and invokes the downstream posting subsystem to post the transaction. Finally, the function publishes a completion event to the event bus completing the transaction lifecycle.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;With customer managed keys configured on DynamoDB, Amazon EventBridge, SQS, and the AWS Lambda durable function, every piece of persisted data in this pipeline is encrypted with keys you own and control. The walkthrough that follows shows you how to deploy this configuration with Terraform.&lt;/p&gt;
&lt;p&gt;Figure 1 shows the reference architecture for this solution.&lt;/p&gt;
&lt;h3 id="reference-architecture"&gt;Reference architecture&lt;/h3&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/28/ComputeBlog-2578-1.jpg" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/28/ComputeBlog-2578-1.jpg" alt="Reference architecture for payment processing using Lambda durable functions" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 1: Payment processing using Lambda durable functions&lt;/p&gt;
&lt;/div&gt;
&lt;h2 id="prerequisites"&gt;Prerequisites&lt;/h2&gt;
&lt;p&gt;To deploy this solution, you need the following prerequisites:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;&lt;strong&gt;AWS account and CLI&lt;/strong&gt;: An active AWS account with the &lt;a href="https://docs.aws.amazon.com/cli/latest/userguide/getting-started-install.html" target="_blank" rel="noopener"&gt;AWS CLI&lt;/a&gt; installed and configured with appropriate credentials.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Terraform&lt;/strong&gt;: &lt;a href="https://developer.hashicorp.com/terraform/install" target="_blank" rel="noopener"&gt;Terraform&lt;/a&gt; installed (version 1.0 or later) for infrastructure provisioning.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Python environment&lt;/strong&gt;: &lt;a href="https://www.python.org/downloads/release/python-3110/" target="_blank" rel="noopener"&gt;Python 3.11&lt;/a&gt; or later, with &lt;a href="https://docs.pytest.org/en/stable/" target="_blank" rel="noopener"&gt;pytest&lt;/a&gt; for running unit tests. The &lt;code&gt;aws-durable-execution-sdk-python&lt;/code&gt; package requires Python 3.11 or later.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;&lt;a href="https://aws.amazon.com/iam/" target="_blank" rel="noopener"&gt;AWS Identity and Access Management (IAM)&lt;/a&gt; permissions&lt;/strong&gt;: The IAM permissions to create the resources. Follow the &lt;a href="https://github.com/aws-samples/sample-payment-processing-with-lambda-durable-functions" target="_blank" rel="noopener"&gt;sample repository&lt;/a&gt; for the sample policy.&lt;/li&gt;
 &lt;li&gt;Basic understanding and familiarity with &lt;a href="https://aws.amazon.com/serverless/" target="_blank" rel="noopener"&gt;AWS Serverless services&lt;/a&gt;.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="solution-walkthrough"&gt;Solution walkthrough&lt;/h2&gt;
&lt;p&gt;The following is a step-by-step guide to deploy and test the payment processing solution.&lt;/p&gt;
&lt;h3 id="step-1-clone-the-repository"&gt;Step 1: Clone the repository&lt;/h3&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;git clone https://github.com/aws-samples/sample-payment-processing-with-lambda-durable-functions.git
cd sample-payment-processing-with-lambda-durable-functions/source&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h3 id="step-2-run-unit-tests"&gt;Step 2: Run unit tests&lt;/h3&gt;
&lt;p&gt;Validate the payment processing logic locally before deploying:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;cd lambda-src/business_rules
pip3 install -r requirements-test.txt
pytest test_app.py -v&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;This runs unit tests that cover transaction validation, business rule checks (foreign transaction detection, currency conversion, merchant type), event schema validation, and misconfiguration handling. The tests use the &lt;a href="https://docs.aws.amazon.com/durable-execution/testing/" target="_blank" rel="noopener"&gt;AWS Durable Execution Testing SDK&lt;/a&gt; to run the handler locally without deploying AWS resources.&lt;/p&gt;
&lt;p&gt;Figure 2 shows an example of test results running locally.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/28/ComputeBlog-2578-2.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/28/ComputeBlog-2578-2.png" alt="Test run results of the business rules" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 2: Test run results of the business rules&lt;/p&gt;
&lt;/div&gt;
&lt;h3 id="step-3-inspect-the-lambda-durable-functions-construct"&gt;Step 3: Inspect the Lambda durable functions construct&lt;/h3&gt;
&lt;p&gt;Open the payments-business-rules Lambda function in &lt;code&gt;source/lambda-src/business_rules/business-rules-app.py&lt;/code&gt; for a sample Lambda durable function. Refer to Figure 3 for the code walkthrough.&lt;/p&gt;
&lt;h4 id="key-features-used"&gt;Key features used&lt;/h4&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;&lt;code&gt;@durable_execution&lt;/code&gt; decorator: Transforms a standard Lambda handler into a durable function handler. The durable execution SDK manages checkpointing automatically. No infrastructure changes are required.&lt;/li&gt;
 &lt;li&gt;&lt;code&gt;context.step("validate-transaction")&lt;/code&gt;: Validates that the transaction has a non-empty &lt;code&gt;issuingCountryCode&lt;/code&gt;. The durable execution checkpoints the result (&lt;code&gt;True&lt;/code&gt; or &lt;code&gt;False&lt;/code&gt;) to durable storage. The durable execution restores checkpoint results instead of re-executing steps during the replay phase. This phase occurs whenever the function is re-invoked after an interruption such as a wait period completing, a failure, or a suspension. This checkpointed result is part of the durable execution data encrypted by your customer managed key.&lt;/li&gt;
 &lt;li&gt;&lt;code&gt;context.step("publish-posting-failure")&lt;/code&gt;: Publishes the full Amazon EventBridge envelope to &lt;a href="https://aws.amazon.com/sns/" target="_blank" rel="noopener"&gt;Amazon SNS&lt;/a&gt; when validation fails. This step only runs on the failure path. The runtime checkpoints the Amazon SNS publish response to durable storage.&lt;/li&gt;
 &lt;li&gt;&lt;code&gt;context.parallel("run-business-rules")&lt;/code&gt;: Runs three independent rule checks concurrently: foreign transaction detection, currency conversion, and merchant type validation. Each branch checkpoints independently. If one branch fails, the others are not replayed on resume. Each branch result is persisted to durable storage and encrypted by the customer managed key.&lt;/li&gt;
 &lt;li&gt;&lt;code&gt;ctx.step("trigger-foreign-transaction-rule")&lt;/code&gt; (inside parallel): Compares &lt;code&gt;billingAmount&lt;/code&gt; against &lt;code&gt;transactionAmount&lt;/code&gt;. If they differ, it emits a &lt;code&gt;ForeignTransactionFound&lt;/code&gt; event to Amazon EventBridge. This step is checkpointed independently within the parallel group.&lt;/li&gt;
 &lt;li&gt;&lt;code&gt;ctx.step("trigger-conversion-rate-rule")&lt;/code&gt; (inside parallel): Checks whether &lt;code&gt;conversionRate&lt;/code&gt; equals &lt;code&gt;1&lt;/code&gt;. If so, it emits a &lt;code&gt;CurrencyConversionTransactionFound&lt;/code&gt; event to Amazon EventBridge. This step is checkpointed independently within the parallel group.&lt;/li&gt;
 &lt;li&gt;&lt;code&gt;ctx.step("trigger-merchant-rule")&lt;/code&gt; (inside parallel): Checks whether &lt;code&gt;merchantType&lt;/code&gt; equals &lt;code&gt;AAFF&lt;/code&gt;. If so, it emits a &lt;code&gt;WarningMerchantTypeTransactionFound&lt;/code&gt; event to Amazon EventBridge. This step is checkpointed independently within the parallel group.&lt;/li&gt;
 &lt;li&gt;&lt;code&gt;context.step("post-transaction-processed")&lt;/code&gt;: Emits the final &lt;code&gt;TransactionPostingApproved&lt;/code&gt; event to Amazon EventBridge. This step is only reached when validation passes and all business rules complete. The runtime checkpoints the Amazon EventBridge response. On replay, if this step already succeeded, the event is not re-published, which guarantees exactly-once approval semantics.&lt;/li&gt;
 &lt;li&gt;&lt;code&gt;context.logger&lt;/code&gt;: Provides replay-aware logging throughout the handler. During replay of previously completed steps, log statements are suppressed to prevent duplicate log entries in &lt;a href="https://aws.amazon.com/cloudwatch/" target="_blank" rel="noopener"&gt;Amazon CloudWatch&lt;/a&gt;.&lt;/li&gt;
&lt;/ol&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/28/ComputeBlog-2578-3.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/28/ComputeBlog-2578-3.png" alt="Lambda durable functions code walkthrough" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 3: Sample Lambda durable functions code&lt;/p&gt;
&lt;/div&gt;
&lt;h3 id="step-4-deploy-infrastructure-with-terraform"&gt;Step 4: Deploy infrastructure with Terraform&lt;/h3&gt;
&lt;p&gt;Terraform currently doesn’t support attaching a customer managed key directly to the durable function. You create the symmetric key in Terraform and then associate the key with the durable function on the AWS Management Console. Refer to &lt;code&gt;source/durable_kms.tf&lt;/code&gt; for the key configuration.&lt;/p&gt;
&lt;p&gt;Initialize and deploy the AWS resources that make up the solution:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;cd ../../
terraform init
terraform plan -var="region=us-east-2"&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Review the plan output, then apply:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;terraform apply -var="region=us-east-2" --auto-approve&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;&lt;em&gt;Note: Replace &lt;code&gt;us-east-2&lt;/code&gt; with your preferred AWS Region.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;On successful completion, Terraform outputs the AWS KMS key alias, key ARN, and DynamoDB Streams ARN used by the event-driven pipeline:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-plaintext"&gt;Apply complete! Resources: N added, 0 changed, 0 destroyed.

Outputs:

durable_kms_key_alias = "durable-function-encryption"
durable_kms_key_arn = "arn:aws:kms:us-east-2:xxxxxxxxxxxx:key/4e87d4c2-1190-4db4-8b97-46657f83ee00"
stream_arn = "arn:aws:dynamodb:us-east-2:xxxxxxxxxxxx:table/visa/stream/2026-03-16T14:25:47.847"&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h3 id="step-5-verify-lambda-durable-functions-configuration"&gt;Step 5: Verify Lambda durable functions configuration&lt;/h3&gt;
&lt;p&gt;In the AWS Lambda console, navigate to the &lt;code&gt;payments-business-rules&lt;/code&gt; function. Confirm that the function &lt;strong&gt;Type&lt;/strong&gt; displays &lt;strong&gt;Durable&lt;/strong&gt;, which indicates that the checkpoint-and-replay mechanism is active. Figure 4 shows the expected function configuration.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/28/ComputeBlog-2578-4.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/28/ComputeBlog-2578-4.png" alt="The durable function in the AWS Lambda console" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 4: The Lambda durable function in the AWS Lambda console&lt;/p&gt;
&lt;/div&gt;
&lt;h3 id="step-6-add-the-aws-kms-key-to-the-lambda-durable-function"&gt;Step 6: Add the AWS KMS key to the Lambda durable function&lt;/h3&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;The durable function is not encrypted with a customer managed key. Figure 5 shows the function’s encryption configuration as empty.
  &lt;p&gt;&lt;/p&gt;
  &lt;div style="width: 810px" class="wp-caption alignnone"&gt;
   &lt;a href="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/28/ComputeBlog-2578-5.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/28/ComputeBlog-2578-5.png" alt="The Lambda durable function missing an AWS KMS key in the AWS Lambda console" width="800"&gt;&lt;/a&gt;
   &lt;p class="wp-caption-text"&gt;Figure 5: The Lambda durable function missing a customer managed key in the AWS Lambda console&lt;/p&gt;
  &lt;/div&gt;&lt;/li&gt;
 &lt;li&gt;Choose &lt;strong&gt;Edit&lt;/strong&gt;, then turn on &lt;strong&gt;Customize encryption settings&lt;/strong&gt; as shown in Figure 6.
  &lt;p&gt;&lt;/p&gt;
  &lt;div style="width: 810px" class="wp-caption alignnone"&gt;
   &lt;a href="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/28/ComputeBlog-2578-6.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/28/ComputeBlog-2578-6.png" alt="The Lambda durable function encryption settings in the AWS Lambda console" width="800"&gt;&lt;/a&gt;
   &lt;p class="wp-caption-text"&gt;Figure 6: The Lambda durable function check encryption in the AWS Lambda console&lt;/p&gt;
  &lt;/div&gt;&lt;/li&gt;
 &lt;li&gt;Select the AWS KMS key ARN created for the durable function. The key ARN is available in the Terraform output from Step 4. Figure 7 shows the key selection.
  &lt;p&gt;&lt;/p&gt;
  &lt;div style="width: 810px" class="wp-caption alignnone"&gt;
   &lt;a href="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/28/ComputeBlog-2578-7.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/28/ComputeBlog-2578-7.png" alt="Selecting the AWS KMS key for the durable function in the AWS Lambda console" width="800"&gt;&lt;/a&gt;
   &lt;p class="wp-caption-text"&gt;Figure 7: Select the AWS KMS key ARN for the durable function in the AWS Lambda console&lt;/p&gt;
  &lt;/div&gt;&lt;/li&gt;
 &lt;li&gt;Choose &lt;strong&gt;Save&lt;/strong&gt; and confirm that the durable function is now encrypted with a customer managed key, as shown in Figure 8.
  &lt;p&gt;&lt;/p&gt;
  &lt;div style="width: 810px" class="wp-caption alignnone"&gt;
   &lt;a href="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/28/ComputeBlog-2578-8.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/28/ComputeBlog-2578-8.png" alt="The Lambda durable function with the AWS KMS key in the AWS Lambda console" width="800"&gt;&lt;/a&gt;
   &lt;p class="wp-caption-text"&gt;Figure 8: AWS Lambda durable function with the customer managed key in the AWS Lambda console&lt;/p&gt;
  &lt;/div&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id="step-7-execute-a-test-payment"&gt;Step 7: Execute a test payment&lt;/h3&gt;
&lt;p&gt;Invoke the &lt;code&gt;payments-visa-mock&lt;/code&gt; Lambda function to simulate an end-to-end authorization flow. The mock function reads sample Visa authorization messages from a CSV file and writes them to DynamoDB, which triggers the event-driven pipeline. Figure 9 shows a sample test invocation.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/28/ComputeBlog-2578-9.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/28/ComputeBlog-2578-9.png" alt="Invoking the payments-visa-mock function to trigger the workflow" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 9: Invoke the payments-visa-mock function to trigger workflow&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;Figure 10 shows a sample response after invocation.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/28/ComputeBlog-2578-10.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/28/ComputeBlog-2578-10.png" alt="Test results from the payments-visa-mock function" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 10: Test results from the payments-visa-mock function to trigger workflow&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;The mock Lambda invocation creates records that follow the process described in the preceding architecture steps.&lt;/p&gt;
&lt;h3 id="step-8-verify-results"&gt;Step 8: Verify results&lt;/h3&gt;
&lt;p&gt;Open Amazon CloudWatch Logs and inspect the log group &lt;code&gt;/aws/lambda/payments-business_rules&lt;/code&gt;. This log group belongs to the Lambda durable function for this use case. Figure 11 shows the CloudWatch log group on the console.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/28/ComputeBlog-2578-11.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/28/ComputeBlog-2578-11.png" alt="Search in CloudWatch Logs for the durable function log group" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 11: Search in CloudWatch&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;You see the complete business rules lifecycle for each transaction, as shown in Figure 12. The highlighted sections show all the business rules performed by the durable function. Each step is checkpointed by the runtime and encrypted by the customer managed key.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/29/computeblog-2578-fig-12.png" target="_blank" rel="noopener"&gt;&lt;img loading="lazy" src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/29/computeblog-2578-fig-12.png" alt="Search Results in lambda durable functions console" width="800" height="1143"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 12: Search Results in lambda durable functions console&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;You can also check the other log groups to trace the full pipeline:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;code&gt;/aws/lambda/payments-enrich&lt;/code&gt;: Transaction enrichment logs.&lt;/li&gt;
 &lt;li&gt;&lt;code&gt;/aws/lambda/payments-posting&lt;/code&gt;: Settlement posting logs.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="step-9-verify-the-customer-managed-key-configuration"&gt;Step 9: Verify the customer managed key configuration&lt;/h3&gt;
&lt;p&gt;You can verify the key configuration by using the AWS CLI:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws lambda get-function-configuration \
    --function-name payments-business-rules \
    --query "DurableConfig" \
    --region us-east-2&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Expected response:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-json"&gt;{
    "KMSKeyArn": "arn:aws:kms:us-east-2:xxxxxxxxxxxx:key/4e87d4c2-1190-4db4-8b97-46657f83ee00",
    "RetentionPeriodInDays": 7,
    "ExecutionTimeout": 180
}&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;You can search in AWS CloudTrail to track the AWS KMS calls. When you configure or update the customer managed key on a durable function, Lambda validates the key policy with dry-run &lt;code&gt;GenerateDataKey&lt;/code&gt; and &lt;code&gt;Decrypt&lt;/code&gt; calls. These appear in CloudTrail with a &lt;code&gt;DryRunOperationException&lt;/code&gt; error code, which confirms that the key policy permissions are correct and does not indicate an actual error. For more details, see &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/durable-encryption.html" target="_blank" rel="noopener"&gt;Encrypting AWS Lambda durable execution data&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="clean-up"&gt;Clean up&lt;/h2&gt;
&lt;p&gt;To avoid ongoing charges, destroy all deployed resources using the following command:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;terraform destroy -var="region=us-east-2" --auto-approve&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Expected output:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-plaintext"&gt;Destroy complete! Resources: N destroyed.&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;In this post, you configured a customer managed key to encrypt durable execution data in a Lambda durable function. With a customer managed key, you control the key rotation schedule, restrict decryption access through the key policy, and generate per-function audit trails in AWS CloudTrail. You can revoke access to durable execution data at any time by updating the key policy, giving you full control over who can read execution state. In-flight executions stop at the next checkpoint call and new executions must be started after restoring access. For details, see &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/durable-encryption.html#durable-encryption-key-unavailable" target="_blank" rel="noopener"&gt;When the customer managed key is unavailable&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;For payment processors and financial institutions, encrypting durable execution data with a customer managed key satisfies compliance obligations for data-at-rest encryption, key governance, and access auditability across multi-step transaction workflows.&lt;/p&gt;
&lt;p&gt;To get started, clone the &lt;a href="https://github.com/aws-samples/sample-payment-processing-with-lambda-durable-functions" target="_blank" rel="noopener"&gt;sample repository&lt;/a&gt; and follow the preceding walkthrough. To learn more about Lambda durable functions, see the &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/durable-functions.html" target="_blank" rel="noopener"&gt;AWS Lambda Developer Guide&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="related-resources"&gt;Related resources&lt;/h2&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/durable-encryption.html" target="_blank" rel="noopener"&gt;AWS Lambda durable functions encryption documentation&lt;/a&gt;&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/durable-functions.html" target="_blank" rel="noopener"&gt;AWS Lambda durable functions Developer Guide&lt;/a&gt;&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/eventbridge/" target="_blank" rel="noopener"&gt;Amazon EventBridge Documentation&lt;/a&gt;&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/step-functions/" target="_blank" rel="noopener"&gt;AWS Step Functions Comparison Guide&lt;/a&gt;&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/Streams.html" target="_blank" rel="noopener"&gt;Amazon DynamoDB Streams Documentation&lt;/a&gt;&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://aws.amazon.com/solutions/guidance/payment-systems-using-event-driven-architecture-on-aws/" target="_blank" rel="noopener"&gt;AWS Guidance for Payment Systems using Event-Driven Architecture&lt;/a&gt;&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://workshops.aws/" target="_blank" rel="noopener"&gt;AWS Serverless Workshops&lt;/a&gt; – Search for Lambda and serverless services workshops.&lt;/li&gt;
&lt;/ul&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>Simplify AMI discovery with Amazon EC2 and SSM Parameter Store</title>
		<link>https://aws.amazon.com/blogs/compute/simplify-ami-discovery-with-amazon-ec2-and-ssm-parameter-store/</link>
		
		<dc:creator><![CDATA[Ashwani Tyagi]]></dc:creator>
		<pubDate>Tue, 29 Sep 2026 16:14:17 +0000</pubDate>
				<category><![CDATA[Amazon EC2]]></category>
		<category><![CDATA[Intermediate (200)]]></category>
		<category><![CDATA[Technical How-to]]></category>
		<guid isPermaLink="false">a24666aa6abc73d7f0d54d34b2ea4d898192c856</guid>

					<description>Managing Amazon EC2 AMIs at scale means constantly mapping AMI IDs to their AWS Systems Manager parameter paths by hand. A recent enhancement to the DescribeImages API returns the associated SSM parameter directly. This post shows how to use it across the AWS CLI, AWS CloudFormation, Terraform, and Auto Scaling launch templates.</description>
										<content:encoded>&lt;p&gt;If you manage &lt;a href="https://aws.amazon.com/ec2/" target="_blank" rel="noopener"&gt;Amazon Elastic Compute Cloud (Amazon EC2)&lt;/a&gt; infrastructure at scale, you have likely encountered the following situation. You release an infrastructure change with the correct Region, the correct instance type, and a launch template that has operated reliably for months. The deployment nevertheless comes up on an &lt;a href="https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/AMIs.html" target="_blank" rel="noopener"&gt;Amazon Machine Image (AMI)&lt;/a&gt; that is several patch cycles out of date, because the AMI ID hardcoded in the template had become stale weeks earlier. The condition goes unnoticed until a security scan flags the instance, at which point you must reconcile AMI IDs across Regions rather than close out the week.&lt;/p&gt;
&lt;p&gt;That scenario is rarely a one-time event. It is one example of a broader pattern that quietly taxes teams running Amazon EC2 at scale: stale AMI IDs, manual parameter lookups, inconsistent Region mappings, and pipelines that silently fail to update. The following section examines four variations of this pattern in detail.&lt;/p&gt;
&lt;p&gt;The common thread across all of these is the same. Locating the correct image is not the hard part. The difficulty lies in wiring that image into your infrastructure as code (IaC) in a manner that remains current. You identify the appropriate AMI on the console, then search &lt;a href="https://aws.amazon.com/systems-manager/features/#Parameter_Store" target="_blank" rel="noopener"&gt;AWS Systems Manager (SSM) Parameter Store&lt;/a&gt; paths to obtain the dynamic reference that maps to it. The workflow spans two tools and two mental models, with a gap in between where errors accumulate. Because the authoritative link between an AMI and its SSM parameter lived outside the API, teams had to reconstruct it by hand, and hands make mistakes.&lt;/p&gt;
&lt;p&gt;A recent enhancement to the Amazon EC2 DescribeImages API closes that gap. When you call DescribeImages on a public AMI, the response now contains a PublicSsmParameterName field: the SSM parameter that resolves to the latest AMI in that lineage. A single API call replaces manual correlation.&lt;/p&gt;
&lt;p&gt;In this post, we examine the operational friction that makes AMI management harder than it should be and show how this enhancement addresses it. We walk through practical examples using the &lt;a href="https://aws.amazon.com/cli/" target="_blank" rel="noopener"&gt;AWS Command Line Interface (AWS CLI)&lt;/a&gt;, &lt;a href="https://aws.amazon.com/cloudformation/" target="_blank" rel="noopener"&gt;AWS CloudFormation&lt;/a&gt;, &lt;a href="https://www.terraform.io/" target="_blank" rel="noopener"&gt;Terraform&lt;/a&gt;, and &lt;a href="https://aws.amazon.com/ec2/autoscaling/" target="_blank" rel="noopener"&gt;Amazon EC2 Auto Scaling&lt;/a&gt; launch templates. We conclude with best practices for golden AMI pipelines, including operational considerations to review before adopting the feature in production.&lt;/p&gt;
&lt;h2 id="prerequisites"&gt;Prerequisites&lt;/h2&gt;
&lt;p&gt;To follow the examples in this post, you will need the following:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;An AWS account.&lt;/li&gt;
 &lt;li&gt;The AWS CLI v2 installed and configured with appropriate permissions (ec2:DescribeImages, ssm:GetParameters).&lt;/li&gt;
 &lt;li&gt;Basic familiarity with AMIs, SSM Parameter Store, and at least one IaC tool (CloudFormation or Terraform).&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="understanding-the-operational-challenges"&gt;Understanding the operational challenges&lt;/h2&gt;
&lt;p&gt;Before addressing the solution, it is worth examining the problem in detail, because the problem seldom manifests as a single, dramatic failure. It is instead a gradual accumulation of minor frictions that, in aggregate, impose a measurable cost on teams responsible for compute.&lt;/p&gt;
&lt;p&gt;AMI IDs are Region-specific, version-specific, and change frequently. The workflow of finding an AMI, locating its SSM parameter, and referencing it in templates spans multiple tools, and the boundaries between steps are where errors accumulate.&lt;/p&gt;
&lt;h3 id="challenge-1-silent-image-aging"&gt;Challenge 1: Silent image aging&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Scenario:&lt;/strong&gt; An engineer copies an AMI ID into a Terraform module as an interim measure. Several months later, that identifier is embedded across four environments. New instances launch on an image that predates numerous patches. There is no error and no alert, only drift that remains invisible until an audit or a review brings it to light.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Impact:&lt;/strong&gt; Hardcoded AMI IDs do not fail conspicuously. They fail quietly, by launching a prior image at a later date. The distance between “this was correct when written” and “this remains correct” widens continuously, and no owner is assigned to monitor it.&lt;/p&gt;
&lt;h3 id="challenge-2-the-multi-region-maintenance-burden"&gt;Challenge 2: The multi-region maintenance burden&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Scenario:&lt;/strong&gt; An application operates across three Regions. The same logical image (for example, the latest &lt;a href="https://aws.amazon.com/linux/amazon-linux-2023/" target="_blank" rel="noopener"&gt;Amazon Linux 2023&lt;/a&gt;) carries a different AMI ID in each Region. Templates therefore accrue region-to-AMI mapping blocks, lookup logic, or both. Each additional Region introduces another entry to maintain, and each AMI refresh requires updating all of them.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Impact:&lt;/strong&gt; The team ends up maintaining a translation table that AWS already maintains on its behalf. The mapping logic becomes load-bearing infrastructure in its own right, and a single stale entry in one Region produces inconsistent fleets that are difficult to diagnose.&lt;/p&gt;
&lt;h3 id="challenge-3-barriers-to-onboarding"&gt;Challenge 3: Barriers to onboarding&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Scenario:&lt;/strong&gt; A new engineer joins the team and poses a reasonable question: which SSM parameter corresponds to a given AMI? The answer resides in an internal knowledge-base page that was accurate eighteen months earlier. The engineer copies a path that appears correct, deploys, and inadvertently references the wrong lineage.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Impact:&lt;/strong&gt; When the relationship between an AMI and its parameter is not discoverable from the API, it must be documented manually. Manually maintained mappings degrade over time. Each new team member re-learns the same institutional knowledge, and each instance of degradation introduces an opportunity to reference an incorrect value.&lt;/p&gt;
&lt;h3 id="challenge-4-uncertainty-about-update-success"&gt;Challenge 4: Uncertainty about update success&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Scenario:&lt;/strong&gt; A golden AMI pipeline completes a build and updates a parameter. The command returns a success response, and the team assumes the new image is in effect. However, for certain parameter data types, a success response does not always indicate that the value was accepted. This specific behavior is examined in the best-practices section, as it is particularly relevant to golden AMI pipelines.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Impact:&lt;/strong&gt; Confidence without confirmation carries substantial risk. A pipeline that presumes success can propagate a stale image across a fleet before the discrepancy is identified.&lt;/p&gt;
&lt;p&gt;Considered individually, none of these situations constitutes a crisis. Considered collectively, they explain why “launch the latest image” is never, in fact, a single step. The common root cause is consistent across all four: the authoritative link between an AMI and its SSM parameter existed outside the API, requiring teams to reconstruct it manually, a process inherently prone to error.&lt;/p&gt;
&lt;h2 id="whats-new-describeimages-returns-the-associated-ssm-parameter"&gt;What’s new: DescribeImages returns the associated SSM parameter&lt;/h2&gt;
&lt;p&gt;The new feature addresses precisely this boundary.&lt;/p&gt;
&lt;p&gt;As of July 16, 2026, the Amazon EC2 DescribeImages API response includes a new field, PublicSsmParameterName, for public AMIs that have an associated SSM parameter. This capability is available at no additional cost in supported AWS Regions, including AWS GovCloud (US) Regions and the China Regions.&lt;/p&gt;
&lt;p&gt;In place of the previous three-step correlation exercise, the workflow reduces to a single call:&lt;/p&gt;
&lt;table border="1px" width="100%" cellpadding="10px"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;Before&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;After&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Find AMI → manually search SSM paths → confirm the correct match&lt;/td&gt;
   &lt;td&gt;Find AMI → PublicSsmParameterName returns the SSM path immediately&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Two separate API calls or console workflows&lt;/td&gt;
   &lt;td&gt;A single DescribeImages call provides the complete mapping&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Prone to mapping an incorrect parameter to an AMI&lt;/td&gt;
   &lt;td&gt;Authoritative mapping obtained directly from the API&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The change introduces neither a new service nor a new pricing dimension. It relocates information that previously lived in knowledge bases into the API response.&lt;/p&gt;
&lt;p&gt;In addition, you can now use the public-ssm-parameter-name filter in DescribeImages to identify all AMIs associated with a specific SSM parameter, making the relationship queryable in either direction.&lt;/p&gt;
&lt;h2 id="how-it-works-api-response-walkthrough"&gt;How it works: API response walkthrough&lt;/h2&gt;
&lt;p&gt;Call DescribeImages on a public AMI with an associated SSM parameter. The response includes PublicSsmParameterName:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-json"&gt;{
  "Images": [
    {
      "ImageId": "ami-0abcdef1234567890",
      "Name": "al2023-ami-2023.7.20260601.0-kernel-6.1-arm64",
      "Description": "Amazon Linux 2023 AMI 2023.7.20260601.0 arm64 HVM kernel-6.1",
      "Architecture": "arm64",
      "PlatformDetails": "Linux/UNIX",
      "State": "available",
      "Public": true,
      "OwnerId": "111122223333",
      "ImageType": "machine",
      "RootDeviceType": "ebs",
      "VirtualizationType": "hvm",
      "EnaSupport": true,
      "PublicSsmParameterName": "aws/service/ami-amazon-linux-latest/al2023-ami-kernel-default-arm64"
    }
  ]
}&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;The &lt;code&gt;PublicSsmParameterName&lt;/code&gt; value (in this case, &lt;code&gt;aws/service/ami-amazon-linux-latest/al2023-ami-kernel-default-arm64&lt;/code&gt;) identifies the SSM parameter associated with this AMI lineage.&lt;/p&gt;
&lt;blockquote&gt;
 &lt;p&gt;&lt;em&gt;&lt;strong&gt;Tip:&lt;/strong&gt; The field is returned under the aws/service/ namespace without a leading slash. When using this value in SSM API calls, resolve:ssm: references, or CloudFormation dynamic references, prepend a forward slash. For example, use &lt;code&gt;/aws/service/ami-amazon-linux-latest/al2023-ami-kernel-default-arm64&lt;/code&gt;. The SSM parameter is intended to resolve to the latest AMI in the lineage, which can help you keep infrastructure current.&lt;/em&gt;&lt;/p&gt;
 &lt;p&gt;&lt;em&gt;&lt;strong&gt;Note:&lt;/strong&gt; Not every public AMI has an associated parameter. The field is present only for lineages for which AWS publishes parameters. The field is also populated only for public AMIs. If you query one of your own private AMIs and observe an empty field, this is expected behavior rather than a defect.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="practical-examples"&gt;Practical examples&lt;/h2&gt;
&lt;p&gt;The following four examples show how to use the new PublicSsmParameterName field across common IaC tools.&lt;/p&gt;
&lt;h3 id="example-1-discover-the-ssm-parameter-for-an-ami-using-the-aws-cli"&gt;Example 1: Discover the SSM parameter for an AMI using the AWS CLI&lt;/h3&gt;
&lt;p&gt;Suppose you have identified an AMI on the console and wish to determine its SSM parameter path for use in your templates:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws ec2 describe-images \
  --image-ids ami-0abcdef1234567890 \
  --query "Images[0].PublicSsmParameterName" \
  --output text&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Output:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-plaintext"&gt;aws/service/ami-amazon-linux-latest/al2023-ami-kernel-default-arm64&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;You can also perform the inverse operation and determine which AMI a given SSM parameter currently references:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;# Linux example
aws ssm get-parameter \
  --name "/aws/service/ami-amazon-linux-latest/al2023-ami-kernel-default-arm64" \
  --query "Parameter.Value" \
  --output text

# Windows example
aws ssm get-parameter \
  --name "/aws/service/ami-windows-latest/Windows_Server-2022-English-Full-Base" \
  --query "Parameter.Value" \
  --output text&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;blockquote&gt;
 &lt;p&gt;&lt;em&gt;&lt;strong&gt;Tip:&lt;/strong&gt; Public SSM parameters are available for both Linux (/aws/service/ami-amazon-linux-latest) and Windows (/aws/service/ami-windows-latest) AMIs. You can list all available parameters under these paths using aws ssm get-parameters-by-path –path .&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Alternatively, you can use the new filter to identify AMIs by their SSM parameter name:&lt;/p&gt;
&lt;blockquote&gt;
 &lt;p&gt;&lt;em&gt;&lt;strong&gt;Note:&lt;/strong&gt; The public-ssm-parameter-name filter returns all AMIs that have ever been associated with the specified parameter, including previous versions. Use sorting or additional filters (such as –query with CreationDate) to identify the most recent AMI.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws ec2 describe-images \
  --filters "Name=public-ssm-parameter-name,Values=/aws/service/ami-amazon-linux-latest/al2023-ami-kernel-default-arm64" \
  --query "Images[].{ImageId:ImageId,Name:Name,CreationDate:CreationDate}" \
  --output table&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h3 id="example-2-cloudformation-with-dynamic-ssm-references"&gt;Example 2: CloudFormation with dynamic SSM references&lt;/h3&gt;
&lt;p&gt;Once the SSM parameter path is known from DescribeImages, you can use CloudFormation dynamic references to resolve to the latest AMI at deployment time:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-yaml"&gt;AWSTemplateFormatVersion: '2010-09-09'
Description: EC2 instance using SSM parameter for latest Amazon Linux 2023 AMI

Parameters:
  InstanceType:
    Type: String
    Default: t4g.micro

Resources:
  MyInstance:
    Type: AWS::EC2::Instance
    Properties:
      InstanceType: !Ref InstanceType
      ImageId: '{{resolve:ssm:/aws/service/ami-amazon-linux-latest/al2023-ami-kernel-default-arm64}}'
      Tags:
        - Key: Name
          Value: MyLatestAL2023Instance&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Alternatively, you can use the AWS::SSM::Parameter::Value parameter type to permit users to override the SSM path at stack creation time:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-yaml"&gt;AWSTemplateFormatVersion: '2010-09-09'
Description: EC2 instance with configurable SSM-based AMI lookup

Parameters:
  AmiSsmParameter:
    Type: 'AWS::SSM::Parameter::Value&amp;lt;AWS::EC2::Image::Id&amp;gt;'
    Default: '/aws/service/ami-amazon-linux-latest/al2023-ami-kernel-default-arm64'
    Description: SSM parameter path for the AMI (discovered via DescribeImages)
  InstanceType:
    Type: String
    Default: t4g.micro

Resources:
  MyInstance:
    Type: AWS::EC2::Instance
    Properties:
      InstanceType: !Ref InstanceType
      ImageId: !Ref AmiSsmParameter
      Tags:
        - Key: Name
          Value: DynamicAMIInstance&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;CloudFormation resolves the AMI ID at deployment time, so the template never contains a hardcoded AMI ID, and any stack update adopts the latest AMI automatically. Two considerations warrant attention before relying on this approach. First, running instances are not affected. A stack update is required to roll out a newer AMI. Second, CloudFormation does not support drift detection on dynamic references, so if the underlying SSM parameter value changes between deployments, CloudFormation will not report it as drift. For ssm dynamic references in which a version has not been pinned, AWS recommends performing a stack update whenever the parameter changes, so that the stack retrieves the current value.&lt;/p&gt;
&lt;h3 id="example-3-terraform-with-ssm-parameter-data-source"&gt;Example 3: Terraform with SSM parameter data source&lt;/h3&gt;
&lt;p&gt;Use the aws_ssm_parameter data source to resolve the SSM path to the latest AMI ID:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-hcl"&gt;# Use the SSM parameter path discovered from DescribeImages
data "aws_ssm_parameter" "latest_al2023" {
  name = "/aws/service/ami-amazon-linux-latest/al2023-ami-kernel-default-arm64"
}

resource "aws_instance" "web" {
  ami           = data.aws_ssm_parameter.latest_al2023.value
  instance_type = "t4g.micro"

  tags = {
    Name = "LatestAL2023-Instance"
  }
}&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;blockquote&gt;
 &lt;p&gt;&lt;em&gt;&lt;strong&gt;Important:&lt;/strong&gt; In Terraform, ami is a replacement-forcing argument on aws_instance. When the SSM parameter changes, Terraform proposes to destroy and recreate the instance. For stateful workloads, add lifecycle { ignore_changes = [ami] } or use launch templates with Auto Scaling (Example 4) instead.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3 id="example-4-auto-scaling-launch-templates-with-ssm-parameters"&gt;Example 4: Auto Scaling launch templates with SSM parameters&lt;/h3&gt;
&lt;p&gt;For Auto Scaling groups, you can reference the SSM parameter directly in the launch template using the resolve:ssm: prefix:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;# Create a launch template that uses the SSM parameter
aws ec2 create-launch-template \
  --launch-template-name al2023-auto-scaling \
  --launch-template-data '{
  "ImageId": "resolve:ssm:/aws/service/ami-amazon-linux-latest/al2023-ami-kernel-default-arm64",
  "InstanceType": "t4g.micro"
}'&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;When EC2 Auto Scaling launches a new instance, it resolves the SSM parameter at launch time to obtain the current AMI ID. You can verify the AMI ID to which a launch template resolves:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws ec2 describe-launch-template-versions \
  --launch-template-name al2023-auto-scaling \
  --versions '$Latest' \
  --resolve-alias&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;The response shows the resolved ImageId:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-json"&gt;{
  "LaunchTemplateVersions": [
    {
      "LaunchTemplateId": "lt-089c023a30example",
      "LaunchTemplateName": "al2023-auto-scaling",
      "VersionNumber": 1,
      "LaunchTemplateData": {
        "ImageId": "ami-0ac394d6a3example",
        "InstanceType": "t4g.micro"
      }
    }
  ]
}&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;The parameter is stored in the launch template. When the Auto Scaling group scales out or replaces an instance, it uses the launch template to resolve the SSM parameter and determine the AMI to launch. This is the most direct of the four patterns: the parameter serves as the single source of truth, and the Auto Scaling group’s normal instance lifecycle effects the rollout.&lt;/p&gt;
&lt;h2 id="before-and-after-workflow-comparison"&gt;Before-and-after workflow comparison&lt;/h2&gt;
&lt;p&gt;The following table summarizes how this feature improves common workflows, and relates each entry to the challenges described earlier.&lt;/p&gt;
&lt;table border="1px" width="100%" cellpadding="10px"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;Workflow&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Before&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;After&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Discover the SSM path for a known AMI&lt;/td&gt;
   &lt;td&gt;Search SSM parameter namespaces manually. Test multiple paths. Confirm a correct match&lt;/td&gt;
   &lt;td&gt;A single DescribeImages call returns PublicSsmParameterName&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Validate that an SSM parameter maps to the expected AMI&lt;/td&gt;
   &lt;td&gt;Call GetParameter, then call DescribeImages on the returned ID to verify&lt;/td&gt;
   &lt;td&gt;Use the public-ssm-parameter-name filter to view all associated AMIs directly&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Set up IaC templates&lt;/td&gt;
   &lt;td&gt;Find AMI → search for SSM path → copy path to template → verify correctness over time&lt;/td&gt;
   &lt;td&gt;Find AMI → read PublicSsmParameterName from the response → use directly in the template&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Onboard new team members&lt;/td&gt;
   &lt;td&gt;Document AMI-to-parameter mappings in knowledge bases, which become stale&lt;/td&gt;
   &lt;td&gt;New members self-discover using standard API calls&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Audit AMI usage across teams&lt;/td&gt;
   &lt;td&gt;Cross-reference AMI IDs with SSM parameters in separate calls&lt;/td&gt;
   &lt;td&gt;A single API call provides the complete picture&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="best-practices-using-ssm-parameters-for-golden-ami-pipelines"&gt;Best practices: Using SSM parameters for golden AMI pipelines&lt;/h2&gt;
&lt;p&gt;The following recommendations describe how to derive the greatest benefit from this feature, with operational considerations identified where they are material.&lt;/p&gt;
&lt;h3 id="discontinue-hardcoding-ami-ids"&gt;1. Discontinue hardcoding AMI IDs&lt;/h3&gt;
&lt;p&gt;With PublicSsmParameterName removing the discovery barrier, switch all templates to SSM parameter references. Use {{resolve:ssm:}} in CloudFormation, the aws_ssm_parameter data source in Terraform (note the replacement behavior in Example 3), or the resolve:ssm: prefix in launch templates.&lt;/p&gt;
&lt;h3 id="create-custom-ssm-parameters-for-your-golden-amis"&gt;2. Create custom SSM parameters for your golden AMIs&lt;/h3&gt;
&lt;p&gt;For internally built golden AMIs, create your own SSM parameters using the aws:ec2:image data type:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws ssm put-parameter \
  --name "/my-org/golden-ami/amazon-linux-hardened" \
  --type "String" \
  --data-type "aws:ec2:image" \
  --value "ami-0abcdef1234567890"&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;When the pipeline produces a new golden AMI, update the parameter:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws ssm put-parameter \
  --name "/my-org/golden-ami/amazon-linux-hardened" \
  --type "String" \
  --data-type "aws:ec2:image" \
  --value "ami-0fedcba9876543210" \
  --overwrite&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Stacks, launch templates, or Terraform configurations that reference this parameter can adopt the new AMI on their next deployment, with no template edits required in most cases.&lt;/p&gt;
&lt;blockquote&gt;
 &lt;p&gt;&lt;em&gt;&lt;strong&gt;Operational consideration:&lt;/strong&gt; Because PutParameter validates aws:ec2:image values asynchronously, an HTTP 200 does not confirm the value was accepted. Subscribe to Parameter Store change events in &lt;a href="https://docs.aws.amazon.com/eventbridge/latest/userguide/eb-what-is.html" target="_blank" rel="noopener"&gt;Amazon EventBridge&lt;/a&gt; and confirm the operation succeeded before considering the rollout complete.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3 id="use-parameter-versions-and-labels-for-controlled-rollouts"&gt;3. Use parameter versions and labels for controlled rollouts&lt;/h3&gt;
&lt;p&gt;SSM Parameter Store supports versioning and labels, which provide control over rollouts:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;# Label the current production version
aws ssm label-parameter-version \
  --name "/my-org/golden-ami/amazon-linux-hardened" \
  --parameter-version 5 \
  --labels "prod"&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Production launch templates reference the labeled version:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-yaml"&gt;ImageId: resolve:ssm:/my-org/golden-ami/amazon-linux-hardened:prod&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;With this approach, you can update the parameter with a new AMI without immediately affecting production. Promotion to production is accomplished by moving the prod label, a deliberate, auditable action rather than an automatic side effect.&lt;/p&gt;
&lt;h3 id="combine-with-amazon-ec2-image-builder-for-end-to-end-automation"&gt;4. Combine with Amazon EC2 Image Builder for end-to-end automation&lt;/h3&gt;
&lt;p&gt;Use &lt;a href="https://docs.aws.amazon.com/imagebuilder/latest/userguide/what-is-image-builder.html" target="_blank" rel="noopener"&gt;Amazon EC2 Image Builder&lt;/a&gt; to automate AMI creation, then configure the distribution settings to update your SSM parameter automatically when a new AMI is built. Combined with the new discovery feature, this establishes a closed loop:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;Image Builder creates a new AMI on a schedule.&lt;/li&gt;
 &lt;li&gt;Distribution settings update the SSM parameter to point to the new AMI.&lt;/li&gt;
 &lt;li&gt;Auto Scaling and IaC resolve the parameter to the latest AMI at launch time.&lt;/li&gt;
 &lt;li&gt;With DescribeImages, any authorized party can determine which SSM parameter an AMI maps to.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="scope-iam-permissions-appropriately"&gt;5. Scope IAM permissions appropriately&lt;/h3&gt;
&lt;p&gt;Two permission requirements apply.&lt;/p&gt;
&lt;p&gt;To launch instances by using SSM-referenced AMIs, the launching principal requires ssm:GetParameters on the relevant parameter paths:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-json"&gt;{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Action": "ssm:GetParameters",
      "Resource": "arn:aws:ssm:*:*:parameter/aws/service/ami-amazon-linux-latest/*"
    }
  ]
}&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Scope the Resource element to the paths actually in use. If you reference Windows parameters (/aws/service/ami-windows-latest/&lt;em&gt;) or your own golden AMI paths (/my-org/golden-ami/&lt;/em&gt;), include those ARNs as well. Otherwise, launches will fail with an AccessDenied error.&lt;/p&gt;
&lt;p&gt;To create a custom aws:ec2:image parameter, the pipeline principal also requires ssm:PutParameter and ec2:DescribeImages:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-json"&gt;{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Action": "ssm:PutParameter",
      "Resource": "arn:aws:ssm:*:*:parameter/my-org/golden-ami/*"
    },
    {
      "Effect": "Allow",
      "Action": "ec2:DescribeImages",
      "Resource": "*"
    }
  ]
}&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;For broader guidance on keeping infrastructure current and automating operational processes, see the &lt;a href="https://docs.aws.amazon.com/wellarchitected/latest/operational-excellence-pillar/welcome.html" target="_blank" rel="noopener"&gt;Operational Excellence Pillar&lt;/a&gt; of the AWS Well-Architected Framework.&lt;/p&gt;
&lt;h2 id="clean-up"&gt;Clean up&lt;/h2&gt;
&lt;p&gt;The examples in this post use read-only API calls (DescribeImages, GetParameter) and do not create billable resources. If you created a launch template while following Example 4, you can delete it as follows:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws ec2 delete-launch-template --launch-template-name al2023-auto-scaling&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;The difficulty of AMI management was never attributable to any single failure. It arose from the steady accumulation of stale identifiers, region-mapping tables, stale documentation, and pipelines that presumed success, all of which are minor frictions that together produced significant operational effort and risk. The common thread was that the authoritative link between an AMI and its SSM parameter existed outside the API, requiring teams to reconstruct it manually.&lt;/p&gt;
&lt;p&gt;The new PublicSsmParameterName field in the Amazon EC2 DescribeImages API relocates that link into the response, where it appropriately belongs. With a single API call, you can determine the SSM parameter for any public AMI. You can then reference it directly in CloudFormation templates, Terraform configurations, or Auto Scaling launch templates for automatic AMI updates.&lt;/p&gt;
&lt;p&gt;To begin, call DescribeImages on any public AMI and examine the &lt;code&gt;PublicSsmParameterName&lt;/code&gt; field. For further detail, see &lt;a href="https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/finding-an-ami-parameter-store.html" target="_blank" rel="noopener"&gt;Reference the latest AMIs using Systems Manager public parameters&lt;/a&gt; in the Amazon EC2 User Guide.&lt;/p&gt;
&lt;p&gt;For additional learning resources on AMI management and IaC on AWS, explore &lt;a href="https://aws.amazon.com/ec2/" target="_blank" rel="noopener"&gt;Amazon EC2&lt;/a&gt;, &lt;a href="https://docs.aws.amazon.com/systems-manager/latest/userguide/systems-manager-parameter-store.html" target="_blank" rel="noopener"&gt;AWS Systems Manager Parameter Store&lt;/a&gt;, and &lt;a href="https://aws.amazon.com/image-builder/" target="_blank" rel="noopener"&gt;Amazon EC2 Image Builder&lt;/a&gt;.&lt;/p&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>Frozen package management for air-gapped RHEL-family AMIs</title>
		<link>https://aws.amazon.com/blogs/compute/frozen-package-management-for-air-gapped-rhel-family-amis/</link>
		
		<dc:creator><![CDATA[Anand Krishna Varanasi]]></dc:creator>
		<pubDate>Tue, 29 Sep 2026 15:55:07 +0000</pubDate>
				<category><![CDATA[Advanced (300)]]></category>
		<category><![CDATA[Amazon EC2]]></category>
		<category><![CDATA[Technical How-to]]></category>
		<guid isPermaLink="false">34d380583160feea2d919a68f65541555032776f</guid>

					<description>Learn a serverless, two-account pattern for running regulated, air-gapped RHEL-family fleets on AWS. It separates connected package ingestion from the air-gapped workload, adds explicit human approval for package changes, and keeps EC2 Image Builder AMIs and Patch Manager on one frozen Amazon S3 repository snapshot.</description>
										<content:encoded>&lt;p&gt;If you run a regulated, air-gapped compute fleet on RHEL-family instances, you have probably felt three requirements pulling against each other. Your organization must configure the network to remove internet access from the instances. Your team must review and approve new packages or version upgrades before you adopt them. Your team removes public repository definitions, restricts network paths, and configures instances to use only the internal repository your team has approved. Teams in chip design, finance, healthcare, defense, and the public sector often face this combination while still needing operating system updates.&lt;/p&gt;
&lt;p&gt;This post describes a two-account pattern that separates the connected package-ingestion path from the air-gapped fleet. You create an authorized initial baseline and approve later changes to form a versioned package snapshot in &lt;a href="https://aws.amazon.com/s3/" target="_blank" rel="noopener"&gt;Amazon Simple Storage Service (Amazon S3)&lt;/a&gt;. &lt;a href="https://aws.amazon.com/image-builder/" target="_blank" rel="noopener"&gt;EC2 Image Builder&lt;/a&gt; uses the frozen snapshot to build &lt;a href="https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/AMIs.html" target="_blank" rel="noopener"&gt;Amazon Machine Images (AMIs)&lt;/a&gt;. Your team configures &lt;a href="https://docs.aws.amazon.com/systems-manager/latest/userguide/patch-manager.html" target="_blank" rel="noopener"&gt;AWS Systems Manager Patch Manager&lt;/a&gt; to patch the instances your organization runs from the same internal package source.&lt;/p&gt;
&lt;p&gt;The accompanying reference implementation demonstrates the pattern for RPM-based RHEL-family systems (AlmaLinux for example). It is a reference, not a substitute for distribution of licensing, vulnerability analysis, testing, or an organization’s change-management process.&lt;/p&gt;
&lt;h2 id="the-challenge-getting-packages-into-an-air-gapped-approval-gated-fleet"&gt;The challenge: Getting packages into an air-gapped approval-gated fleet&lt;/h2&gt;
&lt;p&gt;Common delivery models each assume something an air-gapped fleet might not provide:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;Red Hat Update Infrastructure (RHUI)&lt;/strong&gt; expects each instance to reach the service. A fleet with no internet egress needs a different content path.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Red Hat Satellite&lt;/strong&gt; supports disconnected content management, but it is a separate product and operational footprint. Teams that need a custom package-level approval workflow must integrate that workflow with their content-management process.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;The Red Hat CDN&lt;/strong&gt; requires a connected, entitled content-management path. Centralizing that path changes the network architecture, not the customer’s Red Hat subscription obligations.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The objective is not to replace these products universally. It is to show a serverless AWS pattern for teams that need an authorized repository baseline, explicit approval for later package changes, and a fleet with no public package source.&lt;/p&gt;
&lt;h2 id="how-the-pattern-works"&gt;How the pattern works&lt;/h2&gt;
&lt;p&gt;The pattern combines three controls:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;&lt;strong&gt;A frozen package repository on Amazon S3:&lt;/strong&gt; The pattern stores a deployment-authorized baseline and subsequent approved package changes in versioned, per-OS repository prefixes. The repository manifest records the package inventory for each state.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;EC2 Image Builder Orchestration builds AMIs from that repository:&lt;/strong&gt; The build helps remove upstream repository definitions and configures the internal frozen mirror to be used for all package operations.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;The launched fleet has no internet egress:&lt;/strong&gt; The &lt;code&gt;dnf&lt;/code&gt; operations are configured to resolve the internal mirror. Patch Manager uses the same repository source, so image builds and in-place patching draw from one frozen snapshot.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="choosing-the-upstream-source"&gt;Choosing the upstream source&lt;/h2&gt;
&lt;p&gt;Choose one package lineage end to end. The parent AMI, repository content, and trusted signing keys must belong to that same lineage.&lt;/p&gt;
&lt;p&gt;The reference implementation defaults to AlmaLinux vault content plus EPEL and an AlmaLinux parent AMI. The AlmaLinux OS Foundation states that AlmaLinux aims for binary and application binary interface (ABI) compatibility with RHEL. This is an AlmaLinux compatibility goal, not a Red Hat certification, and it does not make repository mixing a supported practice.&lt;/p&gt;
&lt;p&gt;For genuine RHEL systems, use a Red Hat parent AMI, entitled Red Hat repositories, and Red Hat signing keys. A connected content-management host can retrieve content for the isolated environment. This centralizes the network path but does not reduce or change the customer’s Red Hat subscription obligations. Confirm those obligations against the applicable Red Hat agreement.&lt;/p&gt;
&lt;p&gt;Do not pair AlmaLinux repositories with genuine RHEL hosts, or Red Hat repositories with AlmaLinux hosts. Mixed-vendor package lineages can create support, stability, and maintainability problems even when the RPMs appear mechanically compatible.&lt;/p&gt;
&lt;h2 id="architecture-and-workflow"&gt;Architecture and workflow&lt;/h2&gt;
&lt;p&gt;The account boundary provides a primary security boundary for this architecture. The following diagram shows the connected Distribution account, the read-only Workload account, and an example cross-Region layout.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/25/ComputeBlog-2688-1.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/25/ComputeBlog-2688-1.png" alt="Two-account architecture showing the connected Distribution account with the control plane and internet path, and the air-gapped read-only Workload account, spanning two Regions" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 1: Two-account, cross-Region architecture separating the connected Distribution account from the air-gapped Workload account&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;The &lt;strong&gt;Distribution account&lt;/strong&gt; owns the writable control plane and the only internet path. It runs &lt;a href="https://aws.amazon.com/eventbridge/" target="_blank" rel="noopener"&gt;Amazon EventBridge&lt;/a&gt;, three &lt;a href="https://aws.amazon.com/lambda/lambda-functions/" target="_blank" rel="noopener"&gt;AWS Lambda functions&lt;/a&gt;, &lt;a href="https://aws.amazon.com/dynamodb/" target="_blank" rel="noopener"&gt;Amazon DynamoDB&lt;/a&gt;, &lt;a href="https://aws.amazon.com/sns/" target="_blank" rel="noopener"&gt;Amazon Simple Notification Service (Amazon SNS)&lt;/a&gt;, and the &lt;a href="https://aws.amazon.com/fargate/" target="_blank" rel="noopener"&gt;AWS Fargate&lt;/a&gt; sync task. It also owns the frozen S3 repository and its &lt;a href="https://aws.amazon.com/kms/" target="_blank" rel="noopener"&gt;AWS Key Management Service (AWS KMS) key&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The &lt;strong&gt;Workload account&lt;/strong&gt; is air-gapped and read-only with respect to the repository. It runs the internal HTTPS mirror, EC2 Image Builder, Patch Manager, and the compute fleet. Its mirror task role can read and decrypt frozen content but cannot write it.&lt;/p&gt;
&lt;p&gt;The sample repository places the Distribution control plane in US East (N. Virginia), the frozen store in US West (Oregon), and the Workload resources in US West (Oregon) to demonstrate API-only cross-account and cross-Region operation. This Region split is not required. In most deployments, place the Distribution control plane and frozen store in the same Region unless data residency, disaster recovery, or an existing regional footprint justifies the additional latency, transfer cost, and KMS policy complexity.&lt;/p&gt;
&lt;p&gt;The two accounts do not need &lt;a href="https://aws.amazon.com/vpc/" target="_blank" rel="noopener"&gt;Amazon Virtual Private Cloud (VPC)&lt;/a&gt; peering or a transit gateway. Cross-account access uses S3, KMS, and IAM policies. The VPC address ranges can overlap because no &lt;code&gt;VPC-to-VPC&lt;/code&gt; route is required.&lt;/p&gt;
&lt;h2 id="package-baseline-and-scheduled-upgrade-workflow"&gt;Package baseline and scheduled upgrade workflow&lt;/h2&gt;
&lt;p&gt;Before the scheduled workflow begins, your organization must authorize and run a full sync to establish the initial repository baseline. This bootstrap does not provide package-by-package approval. If your organization requires individual approval for every initial RPM, your team should generate and review the baseline manifest before promotion instead of relying solely on deployment authorization.&lt;/p&gt;
&lt;p&gt;After the baseline, the detector runs on a customer-defined schedule. The reference implementation defaults to monthly. The following diagram shows the bootstrap distinction and the selective approval flow.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/25/ComputeBlog-2688-2.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/25/ComputeBlog-2688-2.png" alt="Workflow diagram distinguishing the initial baseline bootstrap sync from the recurring detect, request approval, review, record, selective sync, and manifest update steps" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 2: Package baseline bootstrap and the scheduled selective approval workflow&lt;/p&gt;
&lt;/div&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;&lt;strong&gt;Detect.&lt;/strong&gt; Amazon EventBridge invokes the detector Lambda function on the configured schedule. The detector compares upstream repository metadata with &lt;code&gt;manifest.json&lt;/code&gt;, which records the current frozen inventory. It classifies a newer version as an upgrade and an absent package as new.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Request approval.&lt;/strong&gt; The detector writes candidates to S3, creates a KMS-protected review token carrying the request ID and expiry, and is designed to send a review link through SNS. The detector can use &lt;code&gt;kms:Encrypt&lt;/code&gt; but not &lt;code&gt;kms:Decrypt&lt;/code&gt;.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Review.&lt;/strong&gt; A human opens the review page through an &lt;a href="https://aws.amazon.com/api-gateway/" target="_blank" rel="noopener"&gt;Amazon API Gateway&lt;/a&gt; HTTP API, reviews the proposed package versions, and chooses which changes to approve. The approver can use &lt;code&gt;kms:Decrypt&lt;/code&gt; but not &lt;code&gt;kms:Encrypt&lt;/code&gt;.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Record and start.&lt;/strong&gt; A conditional DynamoDB update changes a request from &lt;code&gt;pending&lt;/code&gt; to &lt;code&gt;approved&lt;/code&gt; only once. The approver then starts the Fargate sync task and passes the request ID.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Selective sync.&lt;/strong&gt; The task reads the approved package list, downloads those package versions, is designed to perform verification checks, and regenerates repository metadata.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Update the manifest.&lt;/strong&gt; When the task stops, Amazon EventBridge invokes the manifest-updater Lambda function. It archives the outgoing manifest and records the resulting repository inventory.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The approval decision controls adoption. It does not prove that package code is safe. Advisory review, vulnerability scanning, testing, and staged rollout remain in separate controls.&lt;/p&gt;
&lt;h2 id="evidence-from-the-approval-workflow"&gt;Evidence from the approval workflow&lt;/h2&gt;
&lt;p&gt;The token ties a review action to a specific request and expiry. Separating &lt;code&gt;kms:Encrypt&lt;/code&gt; from &lt;code&gt;kms:Decrypt&lt;/code&gt; prevents either Lambda function from performing both token roles. The conditional DynamoDB write makes the approval transition single-use.&lt;/p&gt;
&lt;p&gt;DynamoDB records request state, &lt;a href="https://aws.amazon.com/cloudtrail/" target="_blank" rel="noopener"&gt;AWS CloudTrail&lt;/a&gt; records control-plane API activity, and manifest history records repository inventory changes. These service records can feed the organization’s existing audit and evidence-management workflow. Object-level S3 access auditing requires CloudTrail S3 data events. KMS activity alone is not a substitute for those events.&lt;/p&gt;
&lt;h2 id="the-frozen-package-repository-on-amazon-s3"&gt;The frozen package repository on Amazon S3&lt;/h2&gt;
&lt;p&gt;The following diagram shows the per-OS, per-component prefix layout, and manifest objects.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/25/ComputeBlog-2688-3.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/25/ComputeBlog-2688-3.png" alt="Amazon S3 prefix layout with one prefix per operating system version, each holding BaseOS, AppStream, and EPEL components with Packages and repodata trees plus manifest objects" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 3: Per-OS, per-component prefix layout of the frozen repository on Amazon S3&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;Each pinned operating system version receives its own prefix. Repository components such as &lt;strong&gt;BaseOS&lt;/strong&gt;, &lt;strong&gt;AppStream&lt;/strong&gt;, and &lt;strong&gt;EPEL&lt;/strong&gt; contain &lt;code&gt;Packages/&lt;/code&gt; and &lt;code&gt;repodata/&lt;/code&gt; trees. &lt;code&gt;manifest.json&lt;/code&gt; records the active inventory, and archived manifests preserve historical evidence and comparison points.&lt;/p&gt;
&lt;p&gt;Your organization configures the bucket with versioning and SSE-KMS. Public RPM content does not require a customer-managed KMS key for confidentiality, so your organization could instead configure SSE-S3 for encryption at rest. However, SSE-S3 would remove the separate cross-account authorization control provided by the customer-managed KMS key policy. The customer managed key is used here for explicit cross-account key-policy control and revocation, and the manifests reveal the fleet’s exact software inventory. S3 Bucket Keys reduce KMS request volume. If object-level access evidence is required, enable CloudTrail S3 data events.&lt;/p&gt;
&lt;p&gt;A rollback must restore a coherent repository state, including metadata and any required object versions. Restoring only &lt;code&gt;manifest.json&lt;/code&gt; does not roll back repository contents.&lt;/p&gt;
&lt;h2 id="building-patching-and-running-the-fleet"&gt;Building, patching, and running the fleet&lt;/h2&gt;
&lt;p&gt;At AMI build time, an Image Builder component installs the configured repository keys, moves existing repository definitions aside, and writes one frozen repository definition per component. It locks the package manager to the frozen repository directory, fetches metadata through the internal mirror, and fails the build if the mirror validation step fails. An optional curated package list demonstrates that the AMI can install real packages through the frozen path.&lt;/p&gt;
&lt;p&gt;Patch Manager uses the same mirror for the running fleet. A host created from an older AMI and a newly built host are therefore patched toward the same frozen snapshot. Instances run without an &lt;a href="https://docs.aws.amazon.com/vpc/latest/userguide/vpc-nat-gateway.html" target="_blank" rel="noopener"&gt;Amazon VPC NAT gateway&lt;/a&gt;, public IP, or an &lt;a href="https://docs.aws.amazon.com/vpc/latest/userguide/VPC_Internet_Gateway.html" target="_blank" rel="noopener"&gt;Amazon VPC internet gateway&lt;/a&gt; route in the Workload VPC, and their repository configuration contains no public fallback.&lt;/p&gt;
&lt;p&gt;The intended verification model is defense in depth: the sync task helps verify a vendor’s signature before content enters the trusted repository, and the system verifies it again at installation through &lt;code&gt;dnf&lt;/code&gt;. The ingestion gate helps reject digest-only results and can be configured to help confirm that only valid package signatures from a trusted lineage key are accepted.&lt;/p&gt;
&lt;h2 id="the-package-mirror"&gt;The package mirror&lt;/h2&gt;
&lt;p&gt;Nginx fronts &lt;code&gt;aws-sigv4-proxy&lt;/code&gt;, which signs cross-account S3 GET requests using the mirror task role. To a client, the service appears as a standard HTTPS package repository behind an internal &lt;a href="https://aws.amazon.com/elasticloadbalancing/application-load-balancer/" target="_blank" rel="noopener"&gt;Application Load Balancer&lt;/a&gt; and private DNS name.&lt;/p&gt;
&lt;p&gt;Use the latest version of &lt;code&gt;aws-sigv4-proxy&lt;/code&gt; (current latest is v1.12). This version 1.12 contains the fix for signing S3 paths (or the OS package names) with special characters such as &lt;code&gt;+&lt;/code&gt;. Earlier versions can return &lt;code&gt;SignatureDoesNotMatch&lt;/code&gt;. The reference implementation pins the reviewed v1.12 release commit immutably. Keep it current through dependency-update reviews.&lt;/p&gt;
&lt;h2 id="security-boundaries-and-limits"&gt;Security boundaries and limits&lt;/h2&gt;
&lt;p&gt;The design provides the following controls:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;No automatic public-repository adoption:&lt;/strong&gt; A new upstream version enters the selective path only after an explicit, recorded decision by the user.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Repository ingestion and installation checks:&lt;/strong&gt; You configure strict sync-time signature validation to help validate content before it enters the trusted store. &lt;code&gt;dnf&lt;/code&gt; verifies again during installation.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;No package-channel egress:&lt;/strong&gt; Workload instances are configured to prevent access to public package sources.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;A read-only workload boundary:&lt;/strong&gt; A Workload-account principal cannot modify the frozen repository.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Human approval is not a malware detection. A reviewer cannot reliably identify a backdoor in a legitimately signed package merely by seeing its name, version, or changelog. Use vulnerability intelligence, scanning, pre-production tests, and staged deployment as additional controls.&lt;/p&gt;
&lt;p&gt;The approval state, manifests, and CloudTrail records can help support evidence for control frameworks such as SOC 2 change management, ISO 27001 patch-management controls, and FDA 21 CFR Part 11 electronic records. Applicability depends on the organization’s environment, audit scope, and assessor. Confirm it with the compliance team under the AWS shared responsibility model.&lt;/p&gt;
&lt;h2 id="cost-and-operations"&gt;Cost and operations&lt;/h2&gt;
&lt;p&gt;Cost depends on the amount of repository content and the chosen networking and availability design. Components can include S3 storage and requests, KMS requests, Lambda invocations, DynamoDB, SNS, Fargate tasks, the internal load balancer, Distribution-account internet egress, &lt;a href="https://docs.aws.amazon.com/vpc/latest/privatelink/create-interface-endpoint.html" target="_blank" rel="noopener"&gt;Amazon VPC endpoints&lt;/a&gt;, and AMI snapshots. A three-task always-on mirror costs more than an S3 bucket alone. Estimate the target topology with current AWS pricing rather than applying a fixed monthly figure from the sample.&lt;/p&gt;
&lt;p&gt;Run detection and review at an interval defined by patch policy and risk tolerance. The supplied default is monthly, but the Terraform input is configurable. If a package change must be reversed, restore a tested, coherent repository version and rebuild or patch affected hosts as appropriate.&lt;/p&gt;
&lt;h2 id="prerequisites"&gt;Prerequisites&lt;/h2&gt;
&lt;p&gt;To set up the reference implementation, work through these in order:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;&lt;strong&gt;Two AWS accounts:&lt;/strong&gt; a connected Distribution account and an air-gapped Workload account.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Deployment tools:&lt;/strong&gt; Terraform 1.5 or later, Terragrunt, Finch or Docker, and Python with &lt;code&gt;pip&lt;/code&gt;.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;AWS Command Line Interface (AWS CLI):&lt;/strong&gt; one named profile per account.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Distribution networking:&lt;/strong&gt; private subnets with internet egress that works without public IPs, security-group egress on port 443, and DNS resolution for public names.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Workload networking:&lt;/strong&gt; VPC interface endpoints for &lt;code&gt;ssm&lt;/code&gt;, &lt;code&gt;ssmmessages&lt;/code&gt;, &lt;code&gt;ec2messages&lt;/code&gt;, &lt;code&gt;logs&lt;/code&gt;, &lt;code&gt;kms&lt;/code&gt;, and &lt;code&gt;imagebuilder&lt;/code&gt;, plus an S3 gateway endpoint.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Internal mirror identity:&lt;/strong&gt; an AWS Certificate Manager (ACM) certificate and a private hosted zone. If the parent AMI does not trust the issuing CA, configure the CA file so the build installs the trust anchor.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Package lineage:&lt;/strong&gt; a parent AMI, repositories, and signing keys from the same distribution lineage. A RHEL subscription is required when retrieving genuine entitled Red Hat content.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Optional deployment roles:&lt;/strong&gt; otherwise, the stack uses each profile’s credentials.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The deployment creates state backend and &lt;a href="https://aws.amazon.com/ecr/" target="_blank" rel="noopener"&gt;Amazon Elastic Container Registry (ECR)&lt;/a&gt; repositories. Do not create those ECR repositories separately before applying their own Terraform units.&lt;/p&gt;
&lt;h2 id="reference-implementation"&gt;Reference implementation&lt;/h2&gt;
&lt;p&gt;The &lt;a href="https://github.com/aws-samples/sample-frozen-repo-image-build" target="_blank" rel="noopener"&gt;companion repository&lt;/a&gt; provides Terraform modules, Lambda handlers, two container images, a Terragrunt two-account layout, and Makefile targets for deployment and verification. The shipped &lt;code&gt;alma810&lt;/code&gt; example defaults to the AlmaLinux lineage (RHEL family).&lt;/p&gt;
&lt;p&gt;Choose one distribution lineage before deployment:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;AlmaLinux Parent Image default:&lt;/strong&gt; use an AlmaLinux parent AMI, AlmaLinux vault repositories, EPEL, and the included AlmaLinux and EPEL signing keys. No Red Hat subscription is required.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Genuine RHEL Parent Image:&lt;/strong&gt; use a Red Hat parent AMI, an entitled Red Hat content source, and Red Hat signing keys. The customer supplies the Red Hat subscription and content-access integration.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; The reference implementation only provides AlmaLinux lineage setup, not genuine RHEL. If you choose to use a genuine RHEL parent image lineage, only the reference implementation code needs to be updated to fetch the Red Hat credentials or subscription access, and the rest of the workflow remains the same.&lt;/p&gt;
&lt;p&gt;Please follow the &lt;a href="https://github.com/aws-samples/sample-frozen-repo-image-build/blob/main/README.md" target="_blank" rel="noopener"&gt;repository README&lt;/a&gt; for detailed setup instructions:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;Fill in &lt;code&gt;environments/config.hcl&lt;/code&gt; and both account files with the two accounts, networking, mirror certificate, package lineage, and parent AMI.&lt;/li&gt;
 &lt;li&gt;Create the Terraform state backend with &lt;code&gt;make bootstrap DIST_PROFILE=&amp;lt;dist&amp;gt; WORK_PROFILE=&amp;lt;work&amp;gt;&lt;/code&gt;.&lt;/li&gt;
 &lt;li&gt;Review both account plans with &lt;code&gt;make plan DIST_PROFILE=&amp;lt;dist&amp;gt; WORK_PROFILE=&amp;lt;work&amp;gt;&lt;/code&gt;.&lt;/li&gt;
 &lt;li&gt;Deploy in dependency order with &lt;code&gt;make all BASELINE_APPROVED=true DIST_PROFILE=&amp;lt;dist&amp;gt; WORK_PROFILE=&amp;lt;work&amp;gt;&lt;/code&gt;. The baseline sync can take several hours.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="validate-the-internal-package-management-workflow"&gt;Validate the internal package management workflow&lt;/h2&gt;
&lt;p&gt;Validation begins by running the AlmaLinux EC2 Image Builder pipeline. A successful build produces private, encrypted AMIs in the configured AWS Regions. To validate the configuration, launch a test instance from the generated AMI in a no-egress Workload subnet and verify that the instance is connected to the internal frozen repository for all package management workflow.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/25/ComputeBlog-2688-4.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/25/ComputeBlog-2688-4.png" alt="Terminal output listing only the internal frozen-baseos, frozen-appstream, and frozen-epel repositories enabled, with dnf makecache downloading metadata through the private mirror" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 4: Instance showing only the internal frozen repositories enabled, with no public repository configured&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;The output confirms that only the internal frozen AlmaLinux repositories are enabled: &lt;code&gt;frozen-baseos&lt;/code&gt;, &lt;code&gt;frozen-appstream&lt;/code&gt;, and &lt;code&gt;frozen-epel&lt;/code&gt;. The &lt;code&gt;dnf makecache&lt;/code&gt; command successfully downloads metadata for all three repositories through the private mirror. No public repository is configured or used.&lt;/p&gt;
&lt;p&gt;Now try installing, upgrading, or installing a new package to test that package operations are served by the internal frozen repository.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/25/ComputeBlog-2688-5.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/25/ComputeBlog-2688-5.png" alt="Terminal output showing a package install and upgrade completing successfully through the internal frozen repository mirror" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 5: Package install and upgrade served by the internal frozen repository&lt;/p&gt;
&lt;/div&gt;
&lt;h2 id="recommended-practices"&gt;Recommended practices&lt;/h2&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;Fail closed on approval data.&lt;/strong&gt; If the approved package list cannot be retrieved, stop the sync task and avoid substituting a full sync.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Govern the initial baseline.&lt;/strong&gt; Record who authorized the bootstrap full sync, or require explicit review of its manifest before promotion.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Require vendor signatures at ingestion.&lt;/strong&gt; Do not treat a valid package digest as equivalent to a trusted signature.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Keep package lineages consistent.&lt;/strong&gt; Parent AMI, repository content, and signing keys must come from the same distribution lineage.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Use a customer-defined review cadence.&lt;/strong&gt; Monthly is only the sample default.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Use immutable dependency references with active updates.&lt;/strong&gt; Require &lt;code&gt;aws-sigv4-proxy&lt;/code&gt; v1.12 or later, pin the reviewed artifact by digest or full SHA, and automate update proposals.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Test rollback as a repository operation.&lt;/strong&gt; Restore metadata and objects together, then validate the mirror before using the restored state.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="clean-up"&gt;Clean up&lt;/h2&gt;
&lt;p&gt;The walkthrough deploys billable resources in both accounts. Please follow the &lt;a href="https://github.com/aws-samples/sample-frozen-repo-image-build/blob/main/README.md" target="_blank" rel="noopener"&gt;repository README&lt;/a&gt; for detailed setup cleanup instructions.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;This pattern separates connected package ingestion from an air-gapped fleet, provides an authorized package baseline and an explicit decision point for later package changes or upgrades, and keeps image builds and running hosts on one frozen repository snapshot. To learn more, visit the EC2 Image Builder service page, the EC2 Image Builder documentation, the Patch Manager documentation, and the Amazon S3 user guide. The reference implementation is available in &lt;a href="https://github.com/aws-samples/sample-frozen-repo-image-build" target="_blank" rel="noopener"&gt;aws-samples&lt;/a&gt;.&lt;/p&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>Building event-driven applications at scale with Amazon EventBridge</title>
		<link>https://aws.amazon.com/blogs/compute/building-event-driven-applications-at-scale-with-amazon-eventbridge/</link>
		
		<dc:creator><![CDATA[Nahid Karimaghalou]]></dc:creator>
		<pubDate>Mon, 28 Sep 2026 22:38:57 +0000</pubDate>
				<category><![CDATA[Advanced (300)]]></category>
		<category><![CDATA[Amazon EventBridge]]></category>
		<category><![CDATA[Announcements]]></category>
		<guid isPermaLink="false">752597509722bb2a043058e1df06ec9ac2c7e4cd</guid>

					<description>Amazon EventBridge relaunched the Custom event bus so a platform team can share one governed bus across the organization while application teams publish and subscribe from their own accounts. See how retention, ordering, open formats, transformation, and direct target delivery change what one bus can carry.</description>
										<content:encoded>&lt;p&gt;Event-driven applications on &lt;a href="https://aws.amazon.com/eventbridge/" target="_blank" rel="noopener"&gt;Amazon EventBridge&lt;/a&gt; usually start small and then spread. One team creates a Custom event bus, adds a few rules, and ships. Another team needs some of those events, so a rule forwards them to a bus in a second account. A third team needs a subset of what the second team receives, so another rule forwards again. A year later the organization runs dozens of Custom event buses joined by forwarding rules, and that topology has become a thing to operate in its own right.&lt;/p&gt;
&lt;p&gt;That shape has a price, and the smallest part of it is the bill. Every forwarding hop is a separate ingestion, so cost tracks the topology rather than the number of consumers that needed the event. The harder problem is that nobody can see the whole picture. Governance spreads across the accounts it was meant to cover. Answering who publishes to a bus, who consumes a given event type, or what breaks when a team stops publishing means visiting each account and reading its rule configuration. Tracing one event is harder still: its path crosses several buses in several accounts, each with its own metrics and logs, and no single view follows it from publication to the consumer that never received it.&lt;/p&gt;
&lt;p&gt;Application teams also wait. Publishing to a bus in another account, or consuming from one, needs a resource policy, a role, and a forwarding rule owned by a central team. The team that wants to build opens a ticket, and the platform team becomes a queue. Both the missing visibility and the waiting grow with every team onboarded.&lt;/p&gt;
&lt;p&gt;Amazon EventBridge recently relaunched the Custom event bus, which tackles these challenges directly. A platform team creates one bus, shares it across the organization, and keeps control of who can publish and who can subscribe. Every consumer of those events is listed on the one bus rather than inferred from configuration spread across accounts. Application teams create their own Subscribers in their own accounts. The bus stores events for a retention period you choose, preserves order within a key the publisher sets, accepts Avro and Protocol Buffers (Protobuf) alongside JSON (including CloudEvents), and delivers to targets without a function in the path to translate a call. It runs alongside the Custom event bus – classic, so adoption is incremental.&lt;/p&gt;
&lt;p&gt;In this post, you see how a platform team stands up a shared bus and governs access to it, how application teams onboard themselves with a single Subscriber resource, and how retention, ordering, open formats, transformation, and direct target integrations change what one bus can carry.&lt;/p&gt;
&lt;h2 id="one-bus-shared-with-the-organization"&gt;One bus, shared with the organization&lt;/h2&gt;
&lt;p&gt;The platform team’s job on a shared bus is narrower than it was on a fleet of them. It owns the bus and sets the boundaries: which principals can publish and what their events can declare, which principals can subscribe, and, where it matters, what those principals are allowed to filter on. Application teams then manage their own configuration within those boundaries, such as filters, targets, delivery roles, retry policies, and failure destinations, none of which the platform team needs to write or review. That division is the point of the design. The platform team keeps governance of the bus and stops owning everyone else’s configuration, which is what takes it out of the provisioning path without giving up control of who is on the bus.&lt;/p&gt;
&lt;p&gt;Creating the bus is a single call in a platform account.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;BUS_ARN=$(aws eventsv2 create-event-bus \
    --name company-events \
    --storage-configuration '{"RetentionPeriodInDays":7}' \
    --query EventBusArn --output text)&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Retention is the one setting worth deciding deliberately here rather than revisiting after an incident. It runs from 1 to 365 days and can be modified later, but a change only applies going forward. Raising it widens the window for events published from that point on, and does not make older events readable again. Seven days covers a working week of history, which is usually enough to onboard a consumer or reprocess after a bug without paying to store a year of events nobody will read.&lt;/p&gt;
&lt;p&gt;Sharing the bus is the second decision. &lt;a href="https://aws.amazon.com/ram/" target="_blank" rel="noopener"&gt;AWS Resource Access Manager&lt;/a&gt; is the route to reach for first: it associates automatically for accounts in the same organization and reaches accounts outside it by invitation the consumer accepts. A resource policy written on the bus directly is the alternative, and can also name accounts inside or outside the organization.&lt;/p&gt;
&lt;p&gt;Access is granted per principal, and publishing and subscribing are separate permissions. A team that produces order events gains no ability to read payment events from the same bus. One grant is not enough for a cross-account caller, as usual on AWS: the role that publishes or subscribes also needs its own IAM policy allowing those actions. The platform team decides which accounts can reach the bus, and each consuming team decides which of its own principals can use that access.&lt;/p&gt;
&lt;p&gt;Taken together, those decisions produce the architecture in the following diagram. One bus lives in a platform account, and application teams publish to it and subscribe from their own accounts. An AWS Lambda function in Team A’s account calls &lt;code&gt;PutRawEvents&lt;/code&gt; to publish events onto the Amazon EventBridge bus in the platform account. Team B and Team C each attach their own Subscriber: Team B’s delivers to a Lambda function, Team C’s to an &lt;a href="https://aws.amazon.com/dynamodb/" target="_blank" rel="noopener"&gt;Amazon DynamoDB&lt;/a&gt; table.&lt;/p&gt;
&lt;figure&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/28/ComputeBlog-2709-1.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/28/ComputeBlog-2709-1.png" alt="Architecture diagram of one Custom event bus in a platform account. A Lambda function in Team A’s account calls PutRawEvents to publish events onto the Amazon EventBridge bus in the platform account. Team B and Team C each attach their own Subscriber in their own accounts: Team B’s Subscriber delivers to a Lambda function, and Team C’s Subscriber delivers to an Amazon DynamoDB table." width="800"&gt;&lt;/a&gt;
&lt;/figure&gt;
&lt;p&gt;Figure 1: Multi-account sharing&lt;/p&gt;
&lt;p&gt;Cost follows team boundaries because charges separate ingestion from delivery. The account that publishes an event pays to put it on the bus, and the account that owns a Subscriber pays for what that Subscriber consumes. Each team’s usage appears on its own bill, which is what makes a shared bus something a platform team can charge back rather than a shared cost center nobody can decompose. Removing the forwarding hops also removes the duplicated ingestion and delivery those hops created: the same event reaching the same three consumers is ingested once instead of three times.&lt;/p&gt;
&lt;h2 id="publishing-in-the-format-teams-already-use"&gt;Publishing in the format teams already use&lt;/h2&gt;
&lt;p&gt;Not every producer speaks JSON. Teams that standardize event exchange across an organization often register schemas and publish compact binary payloads, because the schema is the contract between teams that deploy on their own timetables. Accepting the formats those producers already emit is simpler than changing each one to convert to JSON first.&lt;/p&gt;
&lt;p&gt;With the new Custom event bus, application teams can publish events in Avro, Protobuf, and CloudEvents (JSON) formats. For the binary formats, a schema registry named on the request is used to deserialize the events.&lt;/p&gt;
&lt;p&gt;There are two publish APIs, and the payload decides which one to call. &lt;code&gt;PutEvents&lt;/code&gt; takes structured JSON with the familiar &lt;code&gt;Detail&lt;/code&gt;, &lt;code&gt;Source&lt;/code&gt;, and &lt;code&gt;DetailType&lt;/code&gt; fields. &lt;code&gt;PutRawEvents&lt;/code&gt; takes a binary payload plus metadata you define, and is the one to use for Avro, Protobuf, CloudEvents, or bytes the bus should not interpret.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-python"&gt;import boto3

events = boto3.client("eventbridgev2")
events.put_raw_events(
    EventBusArn=BUS_ARN,
    SchemaRegistryConfiguration={"RegistryUri": GLUE_REGISTRY_ARN},
    Entries=[
        {
            "Data": avro_encoded_order,  # bytes, straight from your existing producer
            "SystemMetadata": {"ContentType": "application/avro"},
            "Metadata": {"eventType": "OrderPlaced"},
        }
    ],
)&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;The schema registry can be either the &lt;a href="https://docs.aws.amazon.com/glue/latest/dg/schema-registry.html" target="_blank" rel="noopener"&gt;AWS Glue Schema Registry&lt;/a&gt; or the Confluent Cloud Schema Registry.&lt;/p&gt;
&lt;p&gt;Because the bus decodes the event before filters and transformations run, a consumer subscribing to Avro events written by another team needs no schema, no decoder, and no access to the registry. It writes the same filter it would write against JSON. Producers and consumers stay decoupled, and no deserialization code has to be repeated in each consuming team.&lt;/p&gt;
&lt;p&gt;Publishers get one more setting on the same request: deduplication. A retry that already succeeded would otherwise leave a duplicate for every consumer to handle. It works one of two ways: the bus hashes the content of each event, or it uses a deduplication ID you supply. Content-based hashing suits producers with no natural key, since two identical events hash the same. A deduplication ID fits when you already have one, such as an order ID combined with a state transition. It keeps matching even when parts of the payload differ in ways that should not count as a new event.&lt;/p&gt;
&lt;h2 id="self-service-onboarding-for-application-teams"&gt;Self-service onboarding for application teams&lt;/h2&gt;
&lt;p&gt;The new Custom event bus introduces a new resource called a Subscriber. Application teams create and configure their own Subscribers in their own accounts, provided they have been granted subscribe access to the bus. A Subscriber is the one place a consumer’s behavior is defined: which events it receives, where they are delivered, how delivery is retried, and where events go when delivery does not succeed. Reviewing or changing a consumer is one thing to read and one thing to update.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;SUBSCRIBER_ARN=$(aws eventsv2 create-subscriber \
    --name orders-to-fulfilment \
    --event-bus-arn "$BUS_ARN" \
    --filter-configuration '{"Filters":[{"Scope":"METADATA","Pattern":"{\"eventType\":[\"OrderPlaced\"]}"}]}' \
    --invoke-configuration '{"TargetArn":"'"$QUEUE_ARN"'","RoleArn":"'"$ROLE_ARN"'"}' \
    --retry-policy '{"MaxRetryAttempts":10,"MaxEventAgeInSeconds":3600}' \
    --on-failure-configuration '{"Arn":"'"$DLQ_ARN"'"}' \
    --query SubscriberArn --output text)&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;A filter’s scope decides which part of the event the pattern is matched against. &lt;code&gt;DATA&lt;/code&gt; matches the payload, &lt;code&gt;METADATA&lt;/code&gt; matches the key-value pairs the publisher attached to the event, and &lt;code&gt;SYSTEM_METADATA&lt;/code&gt; matches the event’s system fields: the content type and ordering key a publisher declares, plus the fields Amazon EventBridge adds itself. Because Avro and Protobuf payloads are decoded as they are published, a &lt;code&gt;DATA&lt;/code&gt; filter reads their fields directly, the same as it would for JSON.&lt;/p&gt;
&lt;p&gt;The retry policy says how the bus should behave when a target is failing. &lt;code&gt;MaxRetryAttempts&lt;/code&gt; sets how many times a delivery is retried, and &lt;code&gt;MaxEventAgeInSeconds&lt;/code&gt; sets how long an event stays eligible for retry, measured from when it was published. Retries stop as soon as either limit is reached, so both bound the same delivery.&lt;/p&gt;
&lt;p&gt;When deliveries do fail, the reason shows up in the Subscriber’s own logs, which application teams can turn on themselves. They record the error from each delivery attempt alongside the exact input sent to the target, which makes a problem quick to place. Seeing what the target actually received separates a transformation that produced the wrong shape from a target that rejected a correct one.&lt;/p&gt;
&lt;figure&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/28/ComputeBlog-2709-2.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/28/ComputeBlog-2709-2.png" alt="Screenshot of the Amazon EventBridge console showing the Create subscriber form, with fields for the subscriber name, event bus, filter configuration, target (invoke configuration), retry policy, and on-failure destination." width="800"&gt;&lt;/a&gt;
 &lt;figcaption aria-hidden="true"&gt;Figure 2: Creating a Subscriber on console&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h3 id="history-for-consumers-that-did-not-exist-yet"&gt;History for consumers that did not exist yet&lt;/h3&gt;
&lt;p&gt;A Subscriber sometimes needs events that were published before it existed. For example, a new analytics service needs hydrating with recent history, or a target processed a window of events incorrectly and needs that window replayed. Because the bus retains events for the period configured on it, a Subscriber can be created with a starting position in the past, so it reads history, catches up, and continues with live traffic:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws eventsv2 create-subscriber \
    --name analytics-backfill \
    --event-bus-arn "$BUS_ARN" \
    --starting-position POINT_IN_TIME \
    --point-in-time-configuration '{"PointType":"TIMESTAMP","StartingPoint":"2026-09-14T06:00:00Z"}' \
    --filter-configuration '{"Filters":[{"Scope":"METADATA","Pattern":"{\"eventType\":[\"OrderPlaced\"]}"}]}' \
    --invoke-configuration '{"TargetArn":"'"$ANALYTICS_ARN"'","RoleArn":"'"$ROLE_ARN"'"}'&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;A starting position is either &lt;code&gt;LATEST&lt;/code&gt; or &lt;code&gt;POINT_IN_TIME&lt;/code&gt;. Choosing &lt;code&gt;POINT_IN_TIME&lt;/code&gt; then needs a point-in-time configuration: a &lt;code&gt;PointType&lt;/code&gt; of &lt;code&gt;TIMESTAMP&lt;/code&gt; with a starting point, or &lt;code&gt;HORIZON&lt;/code&gt; to begin at the earliest event still retained. An optional end point stops the read at a chosen time, which is what you want when reprocessing a known-bad window rather than catching up to live traffic.&lt;/p&gt;
&lt;p&gt;Two things to keep in mind. The starting position is fixed when the Subscriber is created, so reading a different window means a new Subscriber. Treat the starting position as part of a Subscriber’s identity rather than a dial to turn later. And retention cannot reach back beyond the retention window, so the read starts at the earliest retained event however far back the timestamp asks for.&lt;/p&gt;
&lt;h3 id="order-where-order-matters"&gt;Order, where order matters&lt;/h3&gt;
&lt;p&gt;In event-driven architectures, where components are built to work asynchronously, the order events arrive in usually does not matter. There are still use cases where a consumer relies on ordered delivery, and the new Custom event bus offers it as an option on individual Subscribers.&lt;/p&gt;
&lt;p&gt;Ordering is scoped by a key the publisher sets. A publisher includes an event group ID (a customer ID, an order ID, a driver ID), and a Subscriber created with FIFO delivery type receives the events for each group in the order they were published. A FIFO Subscriber reading events published without a group ID has nothing to sequence by, so the two sides work together. Creating one takes the same call as an unordered Subscriber, with the delivery type set to FIFO:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws eventsv2 create-subscriber \
    --name inventory-ordered \
    --event-bus-arn "$BUS_ARN" \
    --type FIFO \
    --filter-configuration '{"Filters":[{"Scope":"METADATA","Pattern":"{\"eventType\":[\"OrderPlaced\"]}"}]}' \
    --invoke-configuration '{"TargetArn":"'"$FIFO_QUEUE_ARN"'","RoleArn":"'"$ROLE_ARN"'","SqsParameters":{"MessageGroupId":"{% $events.SystemMetadata.EventGroupId %}","MessageDeduplicationId":"{% $events.SystemMetadata.DeduplicationId %}"}}'&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Ordering is per group, so throughput scales with the number of groups. If an event cannot be delivered, it holds up the rest of its own group while other groups keep moving. Choosing the key therefore matters: one that maps to a business entity, such as an order or a customer, gives sequencing where it is needed and independence everywhere else. A key so broad that most events share it puts them all in a single sequence, and a key so specific that every event has its own leaves nothing to order.&lt;/p&gt;
&lt;p&gt;Because ordering is set on each Subscriber, consumers of the same events do not need to agree on it. An inventory service can receive a group’s events in sequence while an analytics service subscribing to those same events takes them as they arrive.&lt;/p&gt;
&lt;h2 id="reshaping-events-and-delivering-directly-to-a-target"&gt;Reshaping events, and delivering directly to a target&lt;/h2&gt;
&lt;p&gt;A consumer’s business logic expects events in a particular shape, and the events on the bus are not always in that shape. Where the two get reconciled is an ownership decision: inside the consumer, where it becomes part of that team’s code, or on the Subscriber, ahead of it.&lt;/p&gt;
&lt;p&gt;The first case is reformatting. A downstream system, often owned by another domain or outside the organization entirely, expects a different structure from the one the publisher emits. A JSONata transformer on the Subscriber produces that structure before delivery, so the consumer receives what it already expects. The business logic stays where it belongs, and when the published shape changes upstream, or another event type needs deriving into the same input, it is the transformer that changes rather than the consumer:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;--transformer '{
    "Type":"JSONATA",
    "JsonataConfiguration":{
        "Expression":"{% {\"orderRef\": $events.Data.detail.orderId, \"total\": $events.Data.detail.amount} %}"
    }
}'&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;The transformer type determines the shape of what gets delivered. &lt;code&gt;RAW&lt;/code&gt; delivers the event payload as is and is the default, so a Subscriber with no transformer configuration receives only the payload. &lt;code&gt;WITH_METADATA&lt;/code&gt; adds the event envelope alongside it, and &lt;code&gt;JSONATA&lt;/code&gt; reshapes it with an expression wrapped in &lt;code&gt;{% %}&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;The transformation reshapes events only for the Subscriber that owns it and does not affect what other Subscribers of the same bus receive. That also makes it a data minimization control: a partner can receive only the fields it needs rather than a whole internal event. Defining it at the Subscriber means it holds for every event without anyone remembering to strip fields.&lt;/p&gt;
&lt;p&gt;The second case is calling an AWS service API. A Subscriber delivers directly to targets including &lt;a href="https://aws.amazon.com/sqs/" target="_blank" rel="noopener"&gt;Amazon Simple Queue Service (Amazon SQS)&lt;/a&gt;, &lt;a href="https://aws.amazon.com/sns/" target="_blank" rel="noopener"&gt;Amazon Simple Notification Service (Amazon SNS)&lt;/a&gt;, &lt;a href="https://aws.amazon.com/lambda/" target="_blank" rel="noopener"&gt;AWS Lambda&lt;/a&gt;, and &lt;a href="https://aws.amazon.com/kinesis/data-streams/" target="_blank" rel="noopener"&gt;Amazon Kinesis Data Streams&lt;/a&gt;. For other services it has been common practice to add a proxy step whose only job is to make the call. With universal targets, the new Custom event bus can call a supported AWS service API directly, with the request body built by a JSONata expression.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-yaml"&gt;TargetArn: arn:aws:events:::aws-sdk:dynamodb:putItem
UniversalTargetParameters.Input:
{% { "TableName": "orders", "Item": { "pk": { "S": $events.Data.detail.orderId } } } %}&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Note that a universal target shapes its input through that parameter rather than through the preceding transformer, and setting a transformer on one is rejected when the Subscriber is created. The two mechanisms do the same kind of work on different targets.&lt;/p&gt;
&lt;p&gt;That removes the proxy processing that existed only to make the call. The delivery role still needs the action the target requires and getting that wrong is the most common cause of a Subscriber that looks healthy and delivers nothing.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;Running an event-driven application across many accounts no longer means running many event buses and the forwarding between them. A platform team creates one new Custom event bus, shares it across the organization through AWS Resource Access Manager or a resource policy on the bus, and keeps one place to decide who publishes and who consumes. Application teams create and own their Subscribers without waiting for provisioning. Ingestion and delivery are charged separately, so each team’s usage appears on its own bill, and the duplicated ingestion that forwarding hops created disappears with the hops.&lt;/p&gt;
&lt;p&gt;The capabilities that used to send individual teams elsewhere now sit on the same bus. Ordering is per Subscriber and scoped by a publisher-supplied key, so one team’s sequencing requirement no longer fragments an architecture. Retention makes it possible to onboard a consumer that needs history it was never subscribed to. Avro and Protobuf are decoded by the bus, so producers keep their binary contracts. Transformation and universal targets keep business logic where it belongs, removing the proxy steps that existed only to reshape an event or make an API call.&lt;/p&gt;
&lt;p&gt;Because the new Custom event bus runs alongside the Custom event bus – classic, adoption is incremental. Point one new consumer at a shared bus or forward a slice of an existing bus into it and move the rest as teams are ready.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Next steps.&lt;/strong&gt; Create a bus, add a Subscriber, and publish an event, starting from the &lt;a href="https://docs.aws.amazon.com/eventbridge/" target="_blank" rel="noopener"&gt;Amazon EventBridge documentation&lt;/a&gt; for the resource model and the &lt;a href="https://aws.amazon.com/cli/" target="_blank" rel="noopener"&gt;AWS Command Line Interface (AWS CLI)&lt;/a&gt; reference. If you already run Custom event buses, the migration guidance covers routing existing events into a new Custom event bus without changing producers. From there, look at the Subscriber logging and metrics options for tracing an event from publication to delivery, and at &lt;a href="https://aws.amazon.com/ram/" target="_blank" rel="noopener"&gt;AWS Resource Access Manager&lt;/a&gt; for how sharing and permissions work across an organization. If you have questions or feedback about the new Custom event bus, leave a comment on this post. We’d like to hear how you’re using it.&lt;/p&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>Improving Lambda function latency with scalable network bandwidth</title>
		<link>https://aws.amazon.com/blogs/compute/improving-lambda-function-latency-with-scalable-network-bandwidth/</link>
		
		<dc:creator><![CDATA[Rahul Shandilya]]></dc:creator>
		<pubDate>Mon, 28 Sep 2026 21:17:09 +0000</pubDate>
				<category><![CDATA[Announcements]]></category>
		<category><![CDATA[AWS Lambda]]></category>
		<category><![CDATA[Intermediate (200)]]></category>
		<guid isPermaLink="false">716ba562f3f1d920bd3aa3997c7f825fd79c5e7f</guid>

					<description>AWS Lambda now supports scalable network bandwidth for functions with 2,048 MB of memory or more running outside a VPC. Sustained throughput scales from 625 Mbps up to 3,000 Mbps at 10,240 MB, reducing execution times and per-invocation costs for latency-sensitive data processing workloads.</description>
										<content:encoded>&lt;p&gt;&lt;a href="https://aws.amazon.com/lambda/" target="_blank" rel="noopener"&gt;AWS Lambda&lt;/a&gt; now supports scalable network bandwidth for functions configured with 2,048 MB of memory or more, running outside of a virtual private cloud (VPC). Previously, sustained network throughput was capped at 625 Mbps regardless of your function’s memory configuration. Now, sustained throughput scales proportionally from 625 Mbps at configurations below 2,048 MB up to 3,000 Mbps at 10,240 MB, increasing the rate at which data moves to and from your execution environment.&lt;/p&gt;
&lt;p&gt;In this post, you learn how to apply this new capability to latency-sensitive data processing workloads, helping reduce function execution times and per-invocation costs while improving the end-user experience through reduced latency. You also walk through a deployable implementation that demonstrates the performance improvements this capability unlocks.&lt;/p&gt;
&lt;h2&gt;Latency-sensitive data processing&lt;/h2&gt;
&lt;p&gt;Latency-sensitive data processing applications are data processing workloads that must be completed in a defined period of time. They often experience bursty, ad hoc traffic patterns while being required to download gigabytes or even terabytes of data from a data store, process it in a compute environment, and return a result to a waiting end user.&lt;/p&gt;
&lt;p&gt;Latency-sensitive data processing is often highly parallelizable. Data can be divided into smaller pieces with each piece being individually processed before combining them together to obtain a result.&lt;/p&gt;
&lt;p&gt;These workloads can be found in multiple industries and verticals. Examples include:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;Log querying engines&lt;/strong&gt; – An end user initiates an on-demand search across terabytes of log data and expects results within seconds.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Insurance underwriting&lt;/strong&gt; – A prospective customer submits an application, triggering real-time evaluation of historical claims and risk data. The underwriting process determines what coverage and premiums to offer to the prospective customer.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Financial ETL pipelines&lt;/strong&gt; – An economic announcement triggers an unexpected burst of market data that must be ingested, transformed, and made available to downstream trading systems before the next market tick.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Genomics platforms&lt;/strong&gt; – A clinician orders a diagnostic test, requiring gigabytes of DNA or RNA sequencing data to pass through a bioinformatics pipeline and be compared against a reference genome while the patient awaits results.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These workloads are challenging to build on traditional compute clusters. Their spiky and unpredictable nature forces you to choose between under-provisioning compute to optimize costs (and risk missing your SLA) or over-provisioning and paying for idle capacity.&lt;/p&gt;
&lt;h2&gt;Why Lambda fits latency-sensitive data processing&lt;/h2&gt;
&lt;p&gt;Lambda eliminates this tradeoff. Instead of pre-provisioning a compute cluster, Lambda scales compute capacity in response to incoming requests, matching processing power to unpredictable traffic patterns. Because latency-sensitive data processing is highly parallelizable, the ability of Lambda to rapidly scale out execution environments makes it a natural fit. You can fan out across thousands of concurrent functions to process data in parallel, paying only for the compute you use.&lt;/p&gt;
&lt;p&gt;However, as data volume and performance requirements grow, network bandwidth to and from the compute environment can become the limiting factor in minimizing workload latency.&lt;/p&gt;
&lt;p&gt;Scalable network bandwidth directly addresses this limitation by raising the per-environment network throughput ceiling, improving the rate at which data can be transferred to and from the execution environment. Each execution environment can now drive up to 3,000 Mbps of sustained throughput when configured with 10,240 MB of memory, a 4.8x increase from the previous ceiling of 625 Mbps. Combined with the ability of Lambda to scale out at a rate of 1,000 execution environments every 10 seconds, you can download more than 3 TB of data in under 10 seconds.&lt;/p&gt;
&lt;h2&gt;New network throughput behavior for Lambda functions&lt;/h2&gt;
&lt;p&gt;Scalable network bandwidth applies to both data ingress to and egress from an execution environment for functions outside of a VPC. For functions configured with 2 GB of memory or more, network bandwidth scales by approximately 280 Mbps increments for every 1 GB of additional memory allocated.&lt;/p&gt;
&lt;p&gt;The following table shows the maximum sustained bandwidth available to each execution environment at each memory configuration.&lt;/p&gt;
&lt;table border="1px" width="100%" cellpadding="10px"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;Memory Configuration&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Max Sustained Bandwidth&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Less than 2,048 MB&lt;/td&gt;
   &lt;td&gt;625 Mbps&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;2,048 MB&lt;/td&gt;
   &lt;td&gt;765 Mbps&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;3,072 MB&lt;/td&gt;
   &lt;td&gt;1,044 Mbps&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;4,096 MB&lt;/td&gt;
   &lt;td&gt;1,324 Mbps&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;5,120 MB&lt;/td&gt;
   &lt;td&gt;1,603 Mbps&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;6,144 MB&lt;/td&gt;
   &lt;td&gt;1,883 Mbps&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;7,168 MB&lt;/td&gt;
   &lt;td&gt;2,162 Mbps&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;8,192 MB&lt;/td&gt;
   &lt;td&gt;2,441 Mbps&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;9,216 MB&lt;/td&gt;
   &lt;td&gt;2,721 Mbps&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;10,240 MB&lt;/td&gt;
   &lt;td&gt;3,000 Mbps (4.8x increase)&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;em&gt;Table 1. Lambda sustained network bandwidth by memory configuration. Bandwidth scales at ~280 Mbps per additional GB of memory above 2 GB.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;In the following section, you learn how scalable network bandwidth improves end-user latency by building an ETL pipeline that demonstrates it. You can find the source code in the &lt;a href="https://github.com/aws-samples/sample-lambda-enhanced-bandwidth" target="_blank" rel="noopener"&gt;GitHub repository&lt;/a&gt;.&lt;/p&gt;
&lt;h2&gt;Solution overview&lt;/h2&gt;
&lt;p&gt;Consider a SaaS analytics platform where users submit ad hoc queries against a data store. The application must extract the relevant data, apply a filter or transformation, and return an aggregate result while the user waits. In this example, the result needs to be returned in 8 seconds or less.&lt;/p&gt;
&lt;p&gt;The following diagram illustrates the architecture of the solution.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/18/ComputeBlog-2609-1.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/18/ComputeBlog-2609-1.png" alt="ETL fan-out architecture: a client calls an orchestrator Lambda function, which fans out to multiple worker Lambda functions that read data in parallel from Amazon S3, with bandwidth scaling callouts for each memory tier." width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 1. ETL fan-out pattern: an orchestrator Lambda function distributes work to multiple worker Lambda functions that read from Amazon S3 in parallel, with bandwidth scaling callouts per memory tier.&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;A client initiates an ad hoc query by calling the orchestrator Lambda function through the Lambda API. The orchestrator function determines how to split the work. To process the data in parallel, the orchestrator function uses a &lt;code&gt;ThreadPoolExecutor&lt;/code&gt; to issue synchronous invoke requests to the Lambda worker function, fanning out the worker across multiple execution environments at the same time.&lt;/p&gt;
&lt;p&gt;Each Lambda worker function is configured with 10,240 MB of memory, so it has access to up to 3,000 Mbps of sustained network throughput. After the data is processed, the aggregated result is returned to the client.&lt;/p&gt;
&lt;h2 id="prerequisites"&gt;Prerequisites&lt;/h2&gt;
&lt;p&gt;Before you start the deployment process, make sure that you have completed the following steps:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;Install the AWS SAM CLI on your computer and confirm that you are running Python 3.12 or later.&lt;/li&gt;
 &lt;li&gt;Have your AWS account credentials ready.&lt;/li&gt;
 &lt;li&gt;Submit a request to &lt;a href="https://console.aws.amazon.com/servicequotas/home/services/lambda/quotas" target="_blank" rel="noopener"&gt;AWS Service Quotas&lt;/a&gt; to turn on scalable network bandwidth for your Lambda functions. This quota is listed under Network bandwidth per execution environment.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Clone the source code from the GitHub repo and deploy the application within your AWS account. Creating the 10 GB test dataset and running the benchmark can incur charges to your AWS account.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;git clone https://github.com/aws-samples/sample-lambda-enhanced-bandwidth
cd sample-lambda-enhanced-bandwidth
sam build
sam deploy --guided&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;After the AWS CloudFormation stack is deployed, record the &lt;code&gt;DataBucketName&lt;/code&gt; and orchestrator function name from the stack outputs to use in subsequent commands.&lt;/p&gt;
&lt;p&gt;To simulate data for the end user to query, the GitHub repo has a script that creates 10 GB of synthetic data and uploads it to your S3 bucket.&lt;/p&gt;
&lt;h2 id="mode-1-processing-pre-partitioned-data"&gt;Mode 1: Processing pre-partitioned data&lt;/h2&gt;
&lt;p&gt;In Mode 1, the 10 GB of synthetic data is pre-partitioned. Pre-partitioned data is typically produced incrementally by many sources over a period of time, which can be the case with IoT data or access logs. The following command creates 10 GB of data divided into 20 partitions that are 512 MB each.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;python scripts/generate_data.py \
    --bucket &amp;lt;DATA_BUCKET_NAME&amp;gt; \
    --total-gb 10 \
    --chunk-mb 512&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Turning on scalable network bandwidth does not, on its own, make your downloads faster. A single download request only opens one connection to Amazon S3, and one connection does not move data fast enough to fill all the bandwidth now available to your Lambda function. To actually use your full allotment of network bandwidth, the execution environment has to pull the data over several connections at once. It does this by preferring the AWS Common Runtime (CRT) transfer client, a high-performance download engine built into Boto3. When the worker calls &lt;code&gt;download_fileobj&lt;/code&gt;, the CRT client automatically breaks the 512 MB object into smaller parts and downloads them in parallel across multiple Amazon S3 requests. Those parallel downloads are what let a single Lambda worker take advantage of its full network bandwidth.&lt;/p&gt;
&lt;p&gt;The following command runs the benchmark on the pre-partitioned data.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;python scripts/run_fanout_benchmark.py \
    --orchestrator-name &amp;lt;STACK_NAME&amp;gt;-orchestrator \
    --bucket &amp;lt;DATA_BUCKET_NAME&amp;gt; \
    --iterations 5&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h2 id="mode-2-processing-single-large-objects"&gt;Mode 2: Processing single large objects&lt;/h2&gt;
&lt;p&gt;Mode 2 generates 10 GB of data in one large object. This arrangement is more common when data is produced or delivered as one complete unit, such as database backups or genomic datasets. The following command creates 10 GB of data in a single large object.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;python scripts/generate_data.py \
    --bucket &amp;lt;DATA_BUCKET_NAME&amp;gt; \
    --single-object-gb 10 \
    --key large-object/large-file.bin&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;In Mode 1, the CRT preference applies to Boto3 managed transfer methods such as &lt;code&gt;download_file&lt;/code&gt; and &lt;code&gt;download_fileobj&lt;/code&gt;. Mode 2 takes a different approach. Each worker reads a specific byte range of a single large object using &lt;code&gt;get_object&lt;/code&gt;. The CRT preference setting has no effect on these calls. Instead, you can control concurrency by explicitly tuning the number of Lambda workers and using a bounded &lt;code&gt;ThreadPoolExecutor&lt;/code&gt; to issue multiple byte-range requests at the same time.&lt;/p&gt;
&lt;p&gt;When you run the following command, the orchestrator takes the single large object and divides it into consecutive byte ranges of 512 MB each. Each of the individual ranges is then processed by a Lambda worker execution environment in parallel.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;python scripts/run_fanout_benchmark.py \
    --orchestrator-name &amp;lt;STACK_NAME&amp;gt;-orchestrator \
    --bucket &amp;lt;DATA_BUCKET_NAME&amp;gt; \
    --key large-object/large-file.bin \
    --slice-size-mb 512 \
    --iterations 5&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;The benchmark reports wall-clock duration, client-observed duration, aggregate throughput across workers, worker completion counts, and target compliance. When comparing memory configurations, keep the code, dataset, AWS Region, partition count, warm-up policy, and measurement count identical. You should run the benchmark multiple times in your account because placement, cold starts, concurrency, S3 behavior, and execution-environment reuse could affect results.&lt;/p&gt;
&lt;h2 id="results"&gt;Results&lt;/h2&gt;
&lt;p&gt;To compare results, we ran the benchmark using a baseline configuration where the worker Lambda function is configured with only 1,024 MB of memory, well below the 2,048 MB threshold required for scalable network bandwidth to take effect. The 1,024 MB configuration limits network throughput to the previous &lt;em&gt;sustained&lt;/em&gt; ceiling of 625 Mbps.&lt;/p&gt;
&lt;p&gt;In our baseline test run, a worker downloaded and processed a single 512 MB partition with a 6.61-second download time at a 649.8 Mbps throughput (at p50). The 649.8 Mbps throughput exceeds the 625 Mbps ceiling because Lambda is capable of bursts in network throughput over a short period of time. Across twenty measured fan-out queries, the complete 10 GB query was completed with a 7.113-second wall-clock at p50. This fits within the 8-second SLA but leaves very little headroom.&lt;/p&gt;
&lt;p&gt;To run our scalable network bandwidth benchmark, we re-deployed our worker Lambda function with a 10,240 MB memory configuration and re-ran the application. At a 10,240 MB memory configuration, each execution environment can now access up to 3,000 Mbps in sustained throughput. Direct 512 MB downloads achieved a 1.70-second download time and 2,521.3 Mbps throughput (both at p50). The complete 10 GB query was completed with a 2.640-second wall-clock at p50. That is 2.69 times faster, or 62.9% lower median latency, than the 1,024 MB configuration.&lt;/p&gt;
&lt;p&gt;Table 2 summarizes the direct worker and end-to-end fan-out measurements for the same 10 GB dataset and 20 × 512 MB orchestration pattern. Aggregate throughput is the total data transfer rate across all twenty execution environments spun up to run the benchmark.&lt;/p&gt;
&lt;table border="1px" width="100%" cellpadding="10px"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;Memory Configuration&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Single 512 MB partition download time and throughput (p50)&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;10 GB fan-out wall time (p50)&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Aggregate throughput p50&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;1,024 MB baseline tier (sustained 625 Mbps)&lt;/td&gt;
   &lt;td&gt;6.61s / 649.8 Mbps&lt;/td&gt;
   &lt;td&gt;7.113s&lt;/td&gt;
   &lt;td&gt;12.08 Gbps&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;10,240 MB scalable tier (up to 3,000 Mbps)&lt;/td&gt;
   &lt;td&gt;1.70s / 2,521.3 Mbps&lt;/td&gt;
   &lt;td&gt;2.640s&lt;/td&gt;
   &lt;td&gt;32.54 Gbps&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;em&gt;Table 2. Measured 1,024 MB baseline tier and 10,240 MB scalable bandwidth performance for a 10 GB fan-out ETL query.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Using scalable network bandwidth, the customer’s SLA headroom has improved by nearly 5 seconds. The Lambda function can now handle larger partitions within the same SLA window, reducing costs while still remaining comfortably within the customer’s SLA.&lt;/p&gt;
&lt;h2 id="clean-up"&gt;Clean up&lt;/h2&gt;
&lt;p&gt;To clean up the resources you created for the benchmark test, run the following commands:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws s3 rm s3://&amp;lt;DATA_BUCKET_NAME&amp;gt; --recursive
sam delete --stack-name &amp;lt;STACK_NAME&amp;gt;&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h2&gt;Best practices&lt;/h2&gt;
&lt;p&gt;After scalable network bandwidth is turned on for your AWS account, the following practices help you get the most out of it.&lt;/p&gt;
&lt;h2 id="profiling-and-planning"&gt;Profiling and planning&lt;/h2&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;Test before you tune.&lt;/strong&gt; Not every function is network-bound. Before increasing memory, profile your function to confirm that network I/O is the primary contributor to invocation duration and not CPU or application logic. Use Amazon CloudWatch Lambda Insights to inspect &lt;code&gt;rx_bytes&lt;/code&gt;, &lt;code&gt;tx_bytes&lt;/code&gt;, and duration. Functions where network I/O dominates invocation time are prime candidates for tuning.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Design for parallelism.&lt;/strong&gt; Break your data into parallelizable chunks that can be processed independently in a fan-out pattern across multiple execution environments. You can use &lt;a href="https://aws.amazon.com/s3/" target="_blank" rel="noopener"&gt;Amazon S3&lt;/a&gt; byte-range reads to split large files into independently downloadable partitions. For implementation details, see &lt;a href="https://docs.aws.amazon.com/AmazonS3/latest/userguide/downloading-an-object.html" target="_blank" rel="noopener"&gt;Downloading an object with part numbers&lt;/a&gt; in the Amazon S3 User Guide.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Run AWS Lambda Power Tuning.&lt;/strong&gt; &lt;a href="https://github.com/alexcasalboni/aws-lambda-power-tuning" target="_blank" rel="noopener"&gt;Lambda Power Tuning&lt;/a&gt; is a state machine that helps you optimize your Lambda functions for cost and performance. Use Power Tuning to sweep memory configurations from 1,024 MB to 10,240 MB and identify the optimal cost-vs-latency point for your workload.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="implementation"&gt;Implementation&lt;/h2&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;Check upstream and downstream limits.&lt;/strong&gt; Check the throughput limits of your data sources. For example, a Lambda function running at 3,000 Mbps can exceed the throughput capacity of a single S3 prefix, which supports up to 5,500 GET requests per second. When this happens, you will see HTTP 503 (Slow Down) errors in your application logs. Distribute your S3 objects across multiple prefixes to parallelize reads and avoid per-prefix throttling.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Balance bandwidth and CPU.&lt;/strong&gt; Lambda allocates CPU proportionally to memory. For example, at a 1.7 GB memory configuration you are allocated 1 vCPU while a 10 GB memory configuration is allocated up to 6 vCPU. If your function processes data in parallel threads, the higher memory tiers give you both more network bandwidth and more CPU to process it. Use the &lt;a href="https://docs.python.org/3/library/concurrent.futures.html" target="_blank" rel="noopener"&gt;concurrent.futures&lt;/a&gt; module in Python or &lt;a href="https://nodejs.org/api/worker_threads.html" target="_blank" rel="noopener"&gt;worker_threads&lt;/a&gt; in &lt;code&gt;Node.js&lt;/code&gt; to process data across parallel threads and maximize both CPU and network utilization.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Turn on Amazon S3 CRT for Boto3.&lt;/strong&gt; If your function uses the Python runtime, initialize your Amazon S3 client with &lt;code&gt;preferred_transfer_client: 'crt'&lt;/code&gt; to maximize single-connection throughput. The AWS Common Runtime automatically parallelizes requests across multiple TCP connections, which matters because individual TCP connections have a throughput ceiling.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Use SnapStart for JVM workloads.&lt;/strong&gt; If you use &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/snapstart.html" target="_blank" rel="noopener"&gt;Lambda SnapStart&lt;/a&gt; for Java functions, scalable network bandwidth reduces &lt;code&gt;afterRestore&lt;/code&gt; hook latency. Network activity that occurs during function restore, such as pre-warming connections or pre-fetching configuration data, can complete faster.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;Scalable network bandwidth raises the per-environment sustained throughput ceiling of AWS Lambda from 625 Mbps to 3,000 Mbps, directly reducing end-to-end latency for data-intensive workloads. Combined with the Lambda scaling rate, you can now move terabytes of data in seconds, without provisioning or managing infrastructure.&lt;/p&gt;
&lt;p&gt;To get started, request the Network bandwidth per execution environment quota increase through &lt;a href="https://console.aws.amazon.com/servicequotas/home/services/lambda/quotas" target="_blank" rel="noopener"&gt;AWS Service Quotas&lt;/a&gt; and deploy the sample application from the &lt;a href="https://github.com/aws-samples/sample-lambda-enhanced-bandwidth" target="_blank" rel="noopener"&gt;GitHub repository&lt;/a&gt; to see the improvement firsthand.&lt;/p&gt;
&lt;p style="clear: both"&gt;&lt;/p&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>Adding custom domains to AWS Lambda MicroVMs with Application Load Balancer</title>
		<link>https://aws.amazon.com/blogs/compute/adding-custom-domains-to-aws-lambda-microvms-with-application-load-balancer/</link>
		
		<dc:creator><![CDATA[Frank Scarfo]]></dc:creator>
		<pubDate>Tue, 22 Sep 2026 16:48:32 +0000</pubDate>
				<category><![CDATA[Advanced (300)]]></category>
		<category><![CDATA[AWS Lambda]]></category>
		<category><![CDATA[Technical How-to]]></category>
		<guid isPermaLink="false">eed40814f635759befd9a9e6f671e1e93dc9f934</guid>

					<description>Many teams want to expose their AWS Lambda MicroVMs under a custom domain they own, and satisfy CORS for browser clients, without changing the application. This post shows how, using an Application Load Balancer that rewrites the Host header and forwards over AWS PrivateLink, deployed with the AWS CDK.</description>
										<content:encoded>&lt;p&gt;AWS Lambda MicroVMs is a serverless compute building block that provides VM-level isolation, near-instant startup performance, and state retention. You can now give each user or job their own execution environment to securely run just-in-time code, whether user or AI-generated. You do this without managing virtualization infrastructure or choosing between isolation, speed, and state retention. Lambda MicroVMs are powered by Firecracker virtualization, the technology underpinning AWS Lambda.&lt;/p&gt;
&lt;p&gt;When you run a workload on &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/lambda-microvms-guide.html" target="_blank" rel="noopener"&gt;AWS Lambda MicroVMs&lt;/a&gt;, each MicroVM is reachable at a service-generated endpoint that looks like &lt;code&gt;92cfc7f9-….lambda-microvm-….on.aws&lt;/code&gt;. That works, but many teams want to expose their MicroVMs under a domain they own, such as &lt;code&gt;92cfc7f9-….microvms.example.com&lt;/code&gt;. When a browser is the client, they also want to satisfy &lt;a href="https://developer.mozilla.org/en-US/docs/Web/HTTP/CORS" target="_blank" rel="noopener"&gt;cross-origin resource sharing (CORS)&lt;/a&gt; without changing the application inside the MicroVM.&lt;/p&gt;
&lt;p&gt;Both are achievable today, entirely from load-balancing and networking primitives. There is no &lt;a href="https://aws.amazon.com/cloudfront/" target="_blank" rel="noopener"&gt;Amazon CloudFront&lt;/a&gt; distribution and no compute in the request path. All you need is an &lt;a href="https://aws.amazon.com/elasticloadbalancing/application-load-balancer/" target="_blank" rel="noopener"&gt;Application Load Balancer (ALB)&lt;/a&gt; that terminates TLS with your &lt;a href="https://docs.aws.amazon.com/acm/latest/userguide/acm-overview.html" target="_blank" rel="noopener"&gt;AWS Certificate Manager (ACM)&lt;/a&gt; certificate, rewrites the &lt;code&gt;Host&lt;/code&gt; header, and forwards the request over &lt;a href="https://aws.amazon.com/privatelink/" target="_blank" rel="noopener"&gt;AWS PrivateLink&lt;/a&gt;. In this post you’ll deploy that pattern with the &lt;a href="https://aws.amazon.com/cdk/" target="_blank" rel="noopener"&gt;AWS Cloud Development Kit (AWS CDK)&lt;/a&gt;, map a wildcard of custom domains onto your MicroVMs, and let the ALB handle CORS for you.&lt;/p&gt;
&lt;p&gt;The complete, deployable example is available as a pattern on &lt;a href="https://serverlessland.com/patterns/lambda-microvm-custom-domain-cdk" target="_blank" rel="noopener"&gt;Serverless Land&lt;/a&gt;. This walkthrough centers on the reusable networking pattern. The sample also includes a small demo application that &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/microvms-launching.html" target="_blank" rel="noopener"&gt;provisions&lt;/a&gt; a MicroVM and mints an &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/microvms-security.html#microvms-security-auth-tokens" target="_blank" rel="noopener"&gt;access token&lt;/a&gt;, which we reference but do not detail here.&lt;/p&gt;
&lt;h2 id="what-youll-build"&gt;What you’ll build&lt;/h2&gt;
&lt;p&gt;By the end you’ll have:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;A wildcard custom domain like &lt;code&gt;*.microvms.example.com&lt;/code&gt;, where each &lt;code&gt;&amp;lt;uuid&amp;gt;.microvms.example.com&lt;/code&gt; maps transparently to the corresponding MicroVM.&lt;/li&gt;
 &lt;li&gt;An internet-facing ALB that rewrites the incoming request’s &lt;code&gt;Host&lt;/code&gt; header to the real MicroVM endpoint and forwards requests to it privately over PrivateLink.&lt;/li&gt;
 &lt;li&gt;CORS preflight and response headers handled at the ALB, with no change to the code running in the MicroVM.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Calling &lt;code&gt;https://&amp;lt;uuid&amp;gt;.microvms.example.com/&amp;lt;path&amp;gt;&lt;/code&gt; (with the MicroVM access headers described later) reaches the right MicroVM, with your domain intact end to end.&lt;/p&gt;
&lt;h2 id="solution-overview"&gt;Solution overview&lt;/h2&gt;
&lt;p&gt;The request flow looks like this:&lt;/p&gt;
&lt;figure&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/17/ComputeBlog-2746-1.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/17/ComputeBlog-2746-1.png" alt="Request flow from a browser through the Application Load Balancer, which terminates TLS and rewrites the Host header, then forwards over AWS PrivateLink to the Lambda MicroVM service." width="661"&gt;&lt;/a&gt;
 &lt;figcaption aria-hidden="true"&gt;Request flow from a browser through the Application Load Balancer, which terminates TLS and rewrites the Host header, then forwards over AWS PrivateLink to the Lambda MicroVM service.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The key component is the ALB &lt;strong&gt;host header rewrite&lt;/strong&gt;, introduced in &lt;a href="https://aws.amazon.com/blogs/networking-and-content-delivery/introducing-url-and-host-header-rewrite-with-aws-application-load-balancers/" target="_blank" rel="noopener"&gt;URL and host header rewrite for Application Load Balancers&lt;/a&gt;. A listener rule matches the incoming custom host with a &lt;strong&gt;regex condition&lt;/strong&gt;, captures the MicroVM ID from the left-most label, and a &lt;code&gt;host-header-rewrite&lt;/code&gt; &lt;strong&gt;transform&lt;/strong&gt; rewrites the &lt;code&gt;Host&lt;/code&gt; header to &lt;code&gt;&amp;lt;uuid&amp;gt;.lambda-microvm.&amp;lt;region&amp;gt;.on.aws&lt;/code&gt; before forwarding. Because the MicroVM service front-end routes on the &lt;code&gt;Host&lt;/code&gt; header, the request lands on the correct MicroVM, while the customer’s domain stays in the browser’s address bar the whole time.&lt;/p&gt;
&lt;h3 id="why-not-cloudfront-why-not-an-alb-redirect"&gt;Why not CloudFront? Why not an ALB redirect?&lt;/h3&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;CloudFront&lt;/strong&gt; can also rewrite &lt;code&gt;Host&lt;/code&gt;/SNI toward the origin, but a single distribution has static origins. Mapping a &lt;em&gt;wildcard&lt;/em&gt; of MicroVM IDs through one distribution would require a CloudFront Function to compute the origin per request. The ALB transform performs the same rewrite for the entire wildcard with zero code.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;An ALB redirect action&lt;/strong&gt; only issues an &lt;code&gt;HTTP 301 Moved Permanently&lt;/code&gt;/&lt;code&gt;302 Found&lt;/code&gt; response. The browser would follow it, and the address bar would then show the &lt;code&gt;.on.aws&lt;/code&gt; URL, which breaks our design as it is not a real custom domain. The transform (not a redirect) is what makes the custom domain transparent.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="walkthrough"&gt;Walkthrough&lt;/h2&gt;
&lt;p&gt;The example is an AWS CDK application. Configuration lives under the &lt;code&gt;microvm-custom-domains&lt;/code&gt; key in &lt;code&gt;cdk.json&lt;/code&gt; (hosted zone, wildcard base, the endpoint base to rewrite to, the PrivateLink service name, and the CORS origin). Set those values, then deploy. The sections below explain what the stack creates and why.&lt;/p&gt;
&lt;h3 id="prerequisites"&gt;Prerequisites&lt;/h3&gt;
&lt;ul&gt;
 &lt;li&gt;An AWS account.&lt;/li&gt;
 &lt;li&gt;An existing &lt;a href="https://docs.aws.amazon.com/Route53/latest/DeveloperGuide/AboutHZWorkingWith.html" target="_blank" rel="noopener"&gt;Amazon Route 53 public hosted zone&lt;/a&gt; for your domain (for example, &lt;code&gt;example.com&lt;/code&gt;). The stack imports it and adds a wildcard record.&lt;/li&gt;
 &lt;li&gt;&lt;code&gt;Node.js&lt;/code&gt; v22 or higher and the &lt;a href="https://docs.aws.amazon.com/cdk/v2/guide/getting_started.html" target="_blank" rel="noopener"&gt;AWS CDK v2&lt;/a&gt;.&lt;/li&gt;
 &lt;li&gt;Credentials for the target account with permission to create Lambda, &lt;a href="https://aws.amazon.com/vpc/" target="_blank" rel="noopener"&gt;Amazon Virtual Private Cloud (VPC)&lt;/a&gt;, Elastic Load Balancing, &lt;a href="https://aws.amazon.com/certificate-manager/" target="_blank" rel="noopener"&gt;AWS Certificate Manager (ACM)&lt;/a&gt;, and Route 53 resources through &lt;a href="https://aws.amazon.com/cloudformation/" target="_blank" rel="noopener"&gt;AWS CloudFormation&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="create-the-network-and-the-privatelink-endpoint"&gt;1. Create the network and the PrivateLink endpoint&lt;/h3&gt;
&lt;p&gt;A small VPC (two Availability Zones, which is the minimum for an internet-facing ALB) hosts the ALB and an interface VPC endpoint to the AWS managed MicroVM service. There are no NAT gateways, because nothing here needs egress, which keeps the footprint lean.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-typescript"&gt;// Interface (PrivateLink) endpoint to the AWS managed MicroVM service.
const endpoint = new ec2.InterfaceVpcEndpoint(this, 'MicroVmEndpoint', {
  vpc,
  service: new ec2.InterfaceVpcEndpointService(cfg.microvmVpceServiceName, 443),
  subnets: { subnetType: ec2.SubnetType.PRIVATE_ISOLATED },
});&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h3 id="discover-the-endpoints-private-ip-addresses-at-deploy-time"&gt;2. Discover the endpoint’s private IP addresses at deploy time&lt;/h3&gt;
&lt;p&gt;An ALB IP target group needs the private ENI IP addresses of the interface endpoint (one per Availability Zone). CloudFormation does not expose those IPs as a usable attribute, so the stack resolves them during deployment with an &lt;a href="https://docs.aws.amazon.com/cdk/api/v2/docs/aws-cdk-lib.custom_resources.AwsCustomResource.html" target="_blank" rel="noopener"&gt;AwsCustomResource&lt;/a&gt; that reads the endpoint’s own ENIs by ID (&lt;code&gt;DescribeNetworkInterfaces&lt;/code&gt; on &lt;code&gt;vpcEndpointNetworkInterfaceIds&lt;/code&gt;).&lt;/p&gt;
&lt;p&gt;This is the &lt;strong&gt;only&lt;/strong&gt; compute the package deploys, it runs &lt;strong&gt;only during&lt;/strong&gt; &lt;code&gt;cdk deploy&lt;/code&gt;, and it is &lt;strong&gt;never in the request path&lt;/strong&gt;.&lt;/p&gt;
&lt;h3 id="request-a-wildcard-tls-certificate"&gt;3. Request a wildcard TLS certificate&lt;/h3&gt;
&lt;p&gt;ACM issues a DNS-validated wildcard certificate for &lt;code&gt;*.microvms.example.com&lt;/code&gt;, validated through the hosted zone you imported. The ALB presents this certificate for every custom domain under the wildcard.&lt;/p&gt;
&lt;h3 id="create-the-alb-and-the-microvm-target-group"&gt;4. Create the ALB and the MicroVM target group&lt;/h3&gt;
&lt;p&gt;The internet-facing ALB has an HTTPS:443 listener using the wildcard certificate. The target group holds the endpoint ENI IPs as IP targets, reached over HTTPS:443.&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;Encrypted in transit.&lt;/strong&gt; A customer-provided AWS Certificate Manager (ACM) certificate is used to securely terminate encryption between the client and the ALB. The ALB re-originates TLS to the MicroVM service so traffic stays encrypted through the network.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;IP-based targets.&lt;/strong&gt; The target group uses IP-based targets with the local IP addresses of the VPC endpoints.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Health check matcher &lt;code&gt;200,403,404&lt;/code&gt;.&lt;/strong&gt; The load balancer’s health probes are unauthenticated, so the MicroVM endpoint answers them with &lt;code&gt;403&lt;/code&gt;. A &lt;code&gt;403&lt;/code&gt; here means “endpoint is reachable,” not “auth is broken,” so the matcher treats it as healthy.&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-typescript"&gt;const targetGroup = new elbv2.ApplicationTargetGroup(this, 'MicroVmTargets', {
  vpc,
  protocol: elbv2.ApplicationProtocol.HTTPS,
  port: 443,
  targetType: elbv2.TargetType.IP,
  targets: targetIps.map((ip) =&amp;gt; new elbv2t.IpTarget(ip, 443)),
  healthCheck: {
    protocol: elbv2.Protocol.HTTPS,
    path: '/',
    healthyHttpCodes: '200,403,404',
  },
});

const listener = alb.addListener('Https', {
  port: 443,
  protocol: elbv2.ApplicationProtocol.HTTPS,
  certificates: [certificate],
  // Default action for anything that doesn't match our host regex.
  defaultAction: elbv2.ListenerAction.fixedResponse(404, {
    contentType: 'text/plain',
    messageBody: 'Unknown custom domain',
  }),
});&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h3 id="add-the-host-header-rewrite-rule"&gt;5. Add the host-header rewrite rule&lt;/h3&gt;
&lt;p&gt;A listener rule matches &lt;code&gt;&amp;lt;uuid&amp;gt;.microvms.example.com&lt;/code&gt; with a regex condition and rewrites the &lt;code&gt;Host&lt;/code&gt; header to &lt;code&gt;&amp;lt;uuid&amp;gt;.lambda-microvm.&amp;lt;region&amp;gt;.on.aws&lt;/code&gt; with a &lt;code&gt;host-header-rewrite&lt;/code&gt; transform. The regex captures the left-most label (the MicroVM ID) and reuses it in the replacement.&lt;/p&gt;
&lt;p&gt;At the time of writing, the CDK L2 constructs don’t yet model regex host conditions or transforms, so the example reaches the underlying &lt;code&gt;CfnListenerRule&lt;/code&gt; to set them:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-typescript"&gt;const escapedBase = customDomainBase.replace(/[.]/g, '\\.');
const matchRegex = `^(.+)\\.${escapedBase}$`;     // capture &amp;lt;uuid&amp;gt;
const replaceWith = `$1.${microvmEndpointBase}`;   // &amp;lt;uuid&amp;gt;.lambda-microvm.&amp;lt;region&amp;gt;.on.aws

const cfnRule = forwardingRule.node.defaultChild as elbv2.CfnListenerRule;

cfnRule.conditions = [{ field: 'host-header', regexValues: [matchRegex] }];

cfnRule.addPropertyOverride('Transforms', [
  {
    Type: 'host-header-rewrite',
    HostHeaderRewriteConfig: { Rewrites: [{ Regex: matchRegex, Replace: replaceWith }] },
  },
]);&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h3 id="point-route-53-at-the-alb"&gt;6. Point Route 53 at the ALB&lt;/h3&gt;
&lt;p&gt;Wildcard A and AAAA alias records (&lt;code&gt;*.microvms.example.com&lt;/code&gt;) target the ALB, so every MicroVM custom subdomain resolves to it.&lt;/p&gt;
&lt;h3 id="deploy"&gt;7. Deploy&lt;/h3&gt;
&lt;p&gt;Run the following commands to install the dependencies and then deploy the application.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;npm install
npx cdk deploy&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h2 id="handling-cors-at-the-alb"&gt;Handling CORS at the ALB&lt;/h2&gt;
&lt;p&gt;If your clients are browsers calling the MicroVM from another origin, CORS is handled entirely at the ALB, with no change to the application inside the MicroVM.&lt;/p&gt;
&lt;p&gt;The listener uses &lt;a href="https://docs.aws.amazon.com/elasticloadbalancing/latest/application/header-modification.html#insert-header" target="_blank" rel="noopener"&gt;ALB header-modification attributes&lt;/a&gt; to insert the &lt;code&gt;Access-Control-Allow-*&lt;/code&gt; headers on &lt;strong&gt;every&lt;/strong&gt; response. A higher-priority rule answers &lt;code&gt;OPTIONS&lt;/code&gt; preflight requests at the edge with a fast &lt;code&gt;204&lt;/code&gt; response. Otherwise, preflight requests would reach the origin and be rejected without an access token.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-typescript"&gt;// Insert CORS headers on every response on this listener.
const cfnListener = listener.node.defaultChild as elbv2.CfnListener;
cfnListener.addPropertyOverride('ListenerAttributes', [
  { Key: 'routing.http.response.access_control_allow_origin.header_value',  Value: cfg.corsAllowOrigin },
  { Key: 'routing.http.response.access_control_allow_methods.header_value', Value: 'GET,POST,PUT,DELETE,OPTIONS,PATCH,HEAD' },
  { Key: 'routing.http.response.access_control_allow_headers.header_value', Value: 'x-aws-proxy-auth,x-aws-proxy-port,content-type,authorization' },
  { Key: 'routing.http.response.access_control_expose_headers.header_value', Value: 'content-type,content-length' },
  { Key: 'routing.http.response.access_control_max_age.header_value',        Value: '86400' },
]);

// Answer OPTIONS preflights at the ALB.
new elbv2.ApplicationListenerRule(this, 'CorsPreflightRule', {
  listener,
  priority: 10,
  conditions: [elbv2.ListenerCondition.httpRequestMethods(['OPTIONS'])],
  action: elbv2.ListenerAction.fixedResponse(204, { contentType: 'text/plain', messageBody: '' }),
});&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Because the ALB adds those headers to both the preflight &lt;code&gt;204&lt;/code&gt; and the forwarded MicroVM response, a browser’s cross-origin call succeeds without any application change. Set &lt;code&gt;corsAllowOrigin&lt;/code&gt; to &lt;code&gt;*&lt;/code&gt; for quick testing, and pin it to your own site for anything beyond a demo.&lt;/p&gt;
&lt;h2 id="test-it-end-to-end"&gt;Test it end to end&lt;/h2&gt;
&lt;p&gt;First, launch a Lambda MicroVM and mint an access token (follow &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/microvms-getting-started.html" target="_blank" rel="noopener"&gt;Create your first Lambda MicroVM&lt;/a&gt;). When it’s running, the service gives you a generated endpoint that looks like:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-text"&gt;012345678-9abc-defg.lambda-microvm.us-east-2.on.aws&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;To get the custom-domain equivalent, replace the endpoint suffix (&lt;code&gt;.lambda-microvm.&amp;lt;region&amp;gt;.on.aws&lt;/code&gt;) with your wildcard base: &lt;code&gt;.microvms.example.com&lt;/code&gt;. Everything ahead of that suffix is preserved exactly:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-text"&gt;012345678-9abc-defg.microvms.example.com&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;The ALB’s rewrite rule captures whatever precedes the suffix and re-attaches it to the real endpoint base, so the mapping holds for the entire wildcard. You never register anything per-MicroVM.&lt;/p&gt;
&lt;p&gt;With your token in hand, call the custom domain you derived:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;curl "https://012345678-9abc-defg.microvms.example.com/&amp;lt;path&amp;gt;" \
  -H "X-aws-proxy-auth: &amp;lt;token&amp;gt;" \
  -H "X-aws-proxy-port: 8080"&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;The request travels to the ALB, which terminates TLS, rewrites the host header, and forwards over PrivateLink to the MicroVM. The response comes back under your domain.&lt;/p&gt;
&lt;p&gt;The reference architecture also includes a single-page demo and a &lt;code&gt;POST /api/provision&lt;/code&gt; endpoint that runs or reuses a MicroVM and mints a short-lived token. With it, you can try the flow without wiring up token creation yourself. It even performs this suffix swap for you and hands back a ready-to-click custom-domain URL. See the repository for that piece.&lt;/p&gt;
&lt;h2 id="important-considerations"&gt;Important considerations&lt;/h2&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;Authentication is still the client’s job.&lt;/strong&gt; This pattern only rewrites &lt;code&gt;Host&lt;/code&gt;. The client must still supply a valid, unexpired access token in &lt;code&gt;X-aws-proxy-auth&lt;/code&gt;. This is deliberate. MicroVM tokens are per-MicroVM and short-lived, so baking them into infrastructure would be fragile and insecure.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Region pinning.&lt;/strong&gt; PrivateLink is regional, so the ALB, the endpoint, and the MicroVM service must all be in the same Region.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Production hardening.&lt;/strong&gt; If you adapt the sample’s provisioning endpoint, put &lt;a href="https://docs.aws.amazon.com/elasticloadbalancing/latest/application/listener-authenticate-users.html" target="_blank" rel="noopener"&gt;authentication&lt;/a&gt; and &lt;a href="https://aws.amazon.com/waf/" target="_blank" rel="noopener"&gt;rate limiting&lt;/a&gt; in front of it, pin CORS to your origin, and scope IAM to the minimum. The sample’s provisioning path is intentionally open for demonstration and is not production-safe as written.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Cost.&lt;/strong&gt; You pay for the ALB and the interface endpoint (hourly plus data processing) in addition to the Lambda MicroVM usage. There is no CloudFront distribution and no per-request compute in the data path.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="clean-up"&gt;Clean up&lt;/h2&gt;
&lt;p&gt;Run the following command in the same directory where you deployed the application from.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;npx cdk destroy&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;This removes the ALB, target groups, endpoint, certificate, VPC, and Route 53 records created by the stack.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;You can front AWS Lambda MicroVMs with customer-owned wildcard custom domains using an Application Load Balancer and AWS PrivateLink. The key is the ALB’s host-header rewrite. Because the MicroVM service routes requests based on the &lt;code&gt;Host&lt;/code&gt; header, a single rewrite rule can transparently map an entire wildcard of custom domains onto your MicroVMs. CORS is handled at the edge as well. The whole setup relies only on networking primitives, with no CloudFront distribution and no compute in the request path.&lt;/p&gt;
&lt;p&gt;To try it yourself, deploy the &lt;a href="https://serverlessland.com/patterns/lambda-microvm-custom-domain-cdk" target="_blank" rel="noopener"&gt;reference architecture&lt;/a&gt; and review the &lt;a href="https://aws.amazon.com/blogs/networking-and-content-delivery/introducing-url-and-host-header-rewrite-with-aws-application-load-balancers/" target="_blank" rel="noopener"&gt;ALB URL and host header rewrite launch post&lt;/a&gt; for more on the transform feature.&lt;/p&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>Running self-hosted AI agent sandboxes with AWS Lambda MicroVMs</title>
		<link>https://aws.amazon.com/blogs/compute/running-self-hosted-ai-agent-sandboxes-with-aws-lambda-microvms/</link>
		
		<dc:creator><![CDATA[Brian Krygsman]]></dc:creator>
		<pubDate>Fri, 18 Sep 2026 15:10:33 +0000</pubDate>
				<category><![CDATA[Advanced (300)]]></category>
		<category><![CDATA[Announcements]]></category>
		<category><![CDATA[AWS Lambda]]></category>
		<guid isPermaLink="false">944ab8480d283c186edab190849be500bbcf4b16</guid>

					<description>Learn how to run AI agent tool calls in secure, isolated sandboxes using AWS Lambda MicroVMs. This post shows how to architect a self-hosted control plane that launches a fresh, VM-isolated MicroVM for each agent session, keeping credentials, networking, and governance entirely within your own AWS account.</description>
										<content:encoded>&lt;p&gt;Organizations are building AI agents that autonomously write code, query databases, and interact with internal systems on behalf of their teams. These agents handle use cases such as automated code review, data pipeline optimization, and infrastructure troubleshooting. When your AI agent generates a shell command, queries a database, or writes to a file system, that code needs a secure environment to run in. Without isolation, one session’s tool calls can contaminate another session’s state, inadvertently expose sensitive data across tenants, or unintentionally allow untrusted code to reach production resources. Self-hosted sandboxes solve this by keeping agent execution within your own AWS account, giving you full control over networking, secrets, and governance.&lt;/p&gt;
&lt;p&gt;Say you’re building an internal AI agent that optimizes database queries for your engineering team. A developer asks it to find the ten slowest queries in your analytics database, rewrite them with better indexing, and test the results. That’s three tool calls in a single session. One hits a live database with real credentials. One generates code. One executes it. Now multiply that by fifty developers using the assistant at the same time. Each session needs its own credentials, its own filesystem, its own network boundary. If credentials or state cross session boundaries, you have inadvertent data exposure.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://docs.aws.amazon.com/lambda/" target="_blank" rel="noopener"&gt;AWS Lambda MicroVMs&lt;/a&gt; is a serverless compute environment that provides general-purpose runtimes with the strong isolation of virtual machines and the rapid scaling of &lt;a href="https://aws.amazon.com/lambda/" target="_blank" rel="noopener"&gt;AWS Lambda&lt;/a&gt;. Powered by Firecracker virtualization, each MicroVM runs Amazon Linux with full OS access for up to 8 hours. You launch, suspend, resume, and terminate MicroVMs programmatically. You get the serverless benefits of managed infrastructure, responsive scaling, and pay-per-use pricing. Three capabilities make Lambda MicroVMs a strong fit for agent sandboxes:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;VM-level isolation per environment&lt;/strong&gt;: Each MicroVM runs in its own Firecracker virtual machine, providing hardware-virtualization-based isolation between sessions without the resource overhead and startup time required of full VMs. One developer cannot see a teammate’s session, even when both run at the same time.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Launch from snapshot:&lt;/strong&gt; Like &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/snapstart.html" target="_blank" rel="noopener"&gt;Lambda SnapStart&lt;/a&gt;, MicroVMs boot from a pre-captured memory and disk snapshot, skipping application initialization entirely. Your agent gets a near-instant ready-to-use environment.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;4x vertical scaling without re-provisioning&lt;/strong&gt;: A running MicroVM can scale CPU and memory up to 4x its initial allocation, which can range from 0.25 vCPU/0.5 GB to 4 vCPU/8 GB, without terminating or re-creating the environment. If the agent needs to run a heavy data transformation mid-session, it can get more resources without starting over.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In this post, we show you how to architect and build a self-hosted AI agent that uses Lambda MicroVMs as secure, isolated sandboxes for tool-call execution. Lambda MicroVMs can handle the compute isolation for running tool calls, while the host for production AI agents, such as Amazon Bedrock AgentCore, manages the agent logic, model routing, and session state. A complete reference solution is available in &lt;a href="https://github.com/aws-samples/sample-lambda-microvm-claude-managed-agents" target="_blank" rel="noopener"&gt;aws-samples&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="how-self-hosted-sandboxes-work"&gt;How self-hosted sandboxes work&lt;/h2&gt;
&lt;p&gt;A developer asks the agent to “find the ten slowest queries in our analytics database and suggest index improvements.” The agent orchestration system starts a session then breaks the objective into tool calls and distributes them. A worker needs to pick up that session, run the queries, and return results.&lt;/p&gt;
&lt;p&gt;Most AI agent orchestration services and frameworks use a work queue model to distribute tool-call execution. The orchestration service enqueues sessions representing tool-call work. A worker, the process that claims a session and executes its tool calls, runs inside a compute environment, posts results, and exits. In this architecture, each Lambda MicroVM is the compute environment, and the worker is the process running inside it. Claude Managed Agents self-hosted sandboxes run those workers inside your own infrastructure rather than on a shared, multi-tenant compute pool. Your database credentials stay in your virtual private cloud (VPC). Your network, introspection, and governance rules apply.&lt;/p&gt;
&lt;p&gt;You can trigger workers in two ways:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;Webhook-triggered:&lt;/strong&gt; The orchestration application sends a notification when a session is ready. Your control plane launches a worker on demand.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Always-on:&lt;/strong&gt; A long-running process continuously polls the work queue for new sessions.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The Lambda MicroVMs lifecycle aligns with the webhook-triggered pattern, where each session produces one inbound event that launches a fresh MicroVM. Lambda MicroVMs support configurable &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/microvms-launching.html#microvms-launching-idle-policy" target="_blank" rel="noopener"&gt;idle policies&lt;/a&gt;. After a configurable idle period, a MicroVM suspends automatically, preserving disk and memory state. It resumes when inbound traffic arrives or when you call the resume API. The MicroVM runs for the duration of the session, the worker exits, and the idle policy suspends then finally terminates the VM. &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/microvms-launching.html#microvms-launching-lifecycle-hooks" target="_blank" rel="noopener"&gt;Lifecycle hooks&lt;/a&gt; allow you to run custom logic at key steps in the MicroVM lifecycle.&lt;/p&gt;
&lt;p&gt;In contrast, the always-on pattern risks breaking the polling loop by suspending the MicroVM when idle, since there’s no inbound traffic between sessions. You could disable the configurable idle period, but then you pay for empty polling. Use the webhook-triggered approach for self-hosted sandboxes on Lambda MicroVMs.&lt;/p&gt;
&lt;h2 id="architecture"&gt;Architecture&lt;/h2&gt;
&lt;p&gt;The following figure shows the reference solution’s architecture, with the Anthropic agent orchestration service control plane on the left interacting with a self-hosted sandbox environment in AWS on the right.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/16/ComputeBlog-2668-1.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/16/ComputeBlog-2668-1.png" alt="Reference architecture showing the Anthropic orchestration control plane sending a webhook through API Gateway to a launcher Lambda function that starts a MicroVM worker in your AWS account" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 1: Reference architecture for self-hosted AI agent sandboxes on Lambda MicroVMs&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;The sample architecture is event-driven. The only inbound traffic is the webhook call. When the event arrives, the handler launches a MicroVM. Once launched, the MicroVM pulls its assigned session from the orchestration system’s work queue and runs the task. In our example, the developer’s “find slow queries” request has been queued as a session. The agent now needs to reach your infrastructure, spin up an isolated environment, and hand off the work. The following sequence shows how each component interacts to fulfill a single session.&lt;/p&gt;
&lt;p&gt;The orchestration service queues work as &lt;strong&gt;sessions&lt;/strong&gt;. A &lt;strong&gt;MicroVM&lt;/strong&gt; launches to service each session, and the &lt;strong&gt;worker&lt;/strong&gt; is the process running inside that MicroVM that claims the session, executes tool calls, and returns results.&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;Once the orchestration service marks a session as ready to run, it sends a &lt;code&gt;session.status_run_started&lt;/code&gt; webhook to an &lt;a href="https://aws.amazon.com/api-gateway/" target="_blank" rel="noopener"&gt;Amazon API Gateway&lt;/a&gt; endpoint, triggering a MicroVM launch.&lt;/li&gt;
 &lt;li&gt;The launcher verifies the webhook signature using a signing secret from &lt;a href="https://aws.amazon.com/systems-manager/" target="_blank" rel="noopener"&gt;AWS Systems Manager Parameter Store&lt;/a&gt;, rejecting invalid or stale deliveries before spending compute.&lt;/li&gt;
 &lt;li&gt;The launcher calls &lt;code&gt;RunMicrovm&lt;/code&gt;, passing the session ID and a secret reference through &lt;code&gt;runHookPayload&lt;/code&gt;. It deduplicates on the webhook event ID (backed by &lt;a href="https://aws.amazon.com/dynamodb/" target="_blank" rel="noopener"&gt;Amazon DynamoDB&lt;/a&gt;) so retries do not launch duplicate VMs.&lt;/li&gt;
 &lt;li&gt;The MicroVM boots from a pre-captured Firecracker snapshot and receives the dispatch on its &lt;code&gt;/run&lt;/code&gt; lifecycle hook. The worker fetches the environment key from &lt;a href="https://aws.amazon.com/systems-manager/" target="_blank" rel="noopener"&gt;Parameter Store&lt;/a&gt; using its execution role. It pulls the matching session from the work queue, claims it, and executes tool calls in an isolated &lt;code&gt;/workspace&lt;/code&gt; directory. When finished, it posts results and exits. The idle policy suspends then terminates the VM.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;strong&gt;Deduplication.&lt;/strong&gt; The webhook event ID serves as the idempotency key. The launcher uses &lt;a href="https://docs.powertools.aws.dev/lambda/python/latest/utilities/idempotency/" target="_blank" rel="noopener"&gt;Powertools for AWS Lambda (Python)&lt;/a&gt; with a DynamoDB persistence layer to verify exactly-once processing. If the orchestration application retries a delivery with the same event ID, Powertools protects the system from launching extra MicroVMs and doing extra work.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Credential boundaries&lt;/strong&gt;. Each component accesses only the single secret it needs. The launcher reads only the webhook signing secret to verify inbound events. It passes only an ARN reference to the environment key into the MicroVM payload. The MicroVM’s execution role retrieves only that environment key at runtime. No single component holds both secrets.&lt;/p&gt;
&lt;table border="1px" width="100%" cellpadding="10px"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;Component&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Has access to&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Launcher Lambda&lt;/td&gt;
   &lt;td&gt;Webhook signing secret (verify inbound events)&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;MicroVM worker&lt;/td&gt;
   &lt;td&gt;Environment key (through the execution role, to poll and claim sessions)&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;strong&gt;Cost model&lt;/strong&gt;. &lt;a href="https://aws.amazon.com/lambda/pricing/#Lambda_MicroVMs_Pricing" target="_blank" rel="noopener"&gt;You pay for MicroVM run time per session&lt;/a&gt;, plus standard charges for API Gateway requests, Parameter Store API calls, and Lambda invocations for the launcher. When no sessions are active, no MicroVMs run. Cost scales with concurrent sessions and their duration, avoiding idle compute charges.&lt;/p&gt;
&lt;h2 id="implementation"&gt;Implementation&lt;/h2&gt;
&lt;p&gt;The following sections explore the &lt;a href="https://github.com/aws-samples/sample-lambda-microvm-claude-managed-agents" target="_blank" rel="noopener"&gt;reference architecture&lt;/a&gt; in more depth.&lt;/p&gt;
&lt;h3 id="project-structure"&gt;Project structure&lt;/h3&gt;
&lt;p&gt;The reference solution uses &lt;a href="https://aws.amazon.com/serverless/sam/" target="_blank" rel="noopener"&gt;AWS Serverless Application Model (AWS SAM)&lt;/a&gt; for infrastructure-as-code. Alternatively, if you use an AI coding agent such as Claude Code, Kiro, or Cursor, the &lt;a href="https://aws.amazon.com/products/developer-tools/agent-toolkit-for-aws/" target="_blank" rel="noopener"&gt;Agent Toolkit for AWS&lt;/a&gt; includes a &lt;a href="https://github.com/aws/agent-toolkit-for-aws/tree/main/skills/specialized-skills/serverless-skills/aws-lambda-microvms" target="_blank" rel="noopener"&gt;Lambda MicroVMs skill&lt;/a&gt; that gives your agent the procedures to provision, configure, and deploy MicroVM-based sandbox environments on your behalf.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-plaintext"&gt;├── template.yaml                    # SAM: launcher, API, WAF, secrets, roles
├── src/
│   ├── functions/launcher.py        # Verify signature, RunMicrovm
│   ├── microvm-image/
│   │   ├── Dockerfile               # AL2023 + Node.js worker
│   │   └── worker/worker.mjs        # Lifecycle hook server
│   └── scripts/build-image.sh       # Package + create MicroVM image&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h3 id="launcher-verify-the-webhook-before-spinning-up-compute"&gt;Launcher: verify the webhook before spinning up compute&lt;/h3&gt;
&lt;p&gt;When the webhook arrives saying a developer’s session is ready, the launcher’s first action is signature verification. If it fails, the function returns 401 immediately. No MicroVM launches. No DynamoDB writes. You don’t pay for fraudulent or replayed requests.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-python"&gt;signing_secret = [REDACTED_PASSWORD]  # Verify webhook before spending compute
if not verify_signature(raw_body, headers, signing_secret):
    return {"statusCode": 401, "body": "invalid signature"}&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;After verification, the launcher builds a dispatch payload containing the session ID, environment ID, region, and an ARN reference to the environment key secret. It passes this to &lt;code&gt;RunMicrovm&lt;/code&gt; through &lt;code&gt;runHookPayload&lt;/code&gt;:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-python"&gt;launched = microvm_client.run_microvm(
    image_identifier="arn:aws:lambda:us-east-1:123456789012:microvm-image:worker",
    run_hook_payload=json.dumps({"session": dispatch}),
    execution_role_arn=config.execution_role_arn,
    maximum_duration_in_seconds=28800,
    ingress_network_connectors=["arn:aws:lambda:::network-connector:aws-network-connector:ALL_INGRESS"],
    egress_network_connectors=["arn:aws:lambda:::network-connector:aws-network-connector:INTERNET_EGRESS"],
)&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h3 id="microvm-worker-claim-one-session-execute-exit"&gt;MicroVM worker: claim one session, execute, exit&lt;/h3&gt;
&lt;p&gt;The MicroVM image is built from a &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/microvms-images-snapshots.html" target="_blank" rel="noopener"&gt;Firecracker snapshot&lt;/a&gt;. The worker process starts during image creation and is captured in the snapshot, so there is no application startup at run time. The /run &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/microvms-launching.html#microvms-launching-lifecycle-hooks" target="_blank" rel="noopener"&gt;lifecycle hook&lt;/a&gt; delivers the dispatch payload:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-javascript"&gt;// POST /aws/lambda-microvms/runtime/v1/run
case "run": {
    const envelope = JSON.parse(rawBody);
    const dispatch = JSON.parse(envelope.runHookPayload);
    res.writeHead(200); // Acknowledge hook immediately
    res.end();
    const key = await fetchParameter(dispatch.session.ENVIRONMENT_KEY_PARAM_NAME);
    await pollAndHandleSession(dispatch.session.ANTHROPIC_SESSION_ID, key);
    // Session complete; terminate this MicroVM to release all resources
    await terminateMicroVm(envelope.microvmId);
}&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;The worker acknowledges the hook within its timeout, fetches the environment key, and claims the session. This is where the requested work begins. The worker connects to the analytics database, runs &lt;code&gt;EXPLAIN ANALYZE&lt;/code&gt; on the flagged queries, writes optimized alternatives to &lt;code&gt;/workspace/suggestions.sql&lt;/code&gt;, and posts the results back to the developer. All of that happens inside this single VM. When the session completes, the worker calls &lt;code&gt;terminate-microvm&lt;/code&gt; to release all compute resources.&lt;/p&gt;
&lt;h3 id="deployment"&gt;Deployment&lt;/h3&gt;
&lt;p&gt;For full deployment instructions, see the &lt;a href="https://github.com/aws-samples/sample-lambda-microvm-claude-managed-agents" target="_blank" rel="noopener"&gt;reference solution README&lt;/a&gt;. Before deploying, make sure you have these prerequisites.&lt;/p&gt;
&lt;h4 id="prerequisites"&gt;Prerequisites&lt;/h4&gt;
&lt;ul&gt;
 &lt;li&gt;An AWS account with permissions for &lt;a href="https://aws.amazon.com/s3/" target="_blank" rel="noopener"&gt;Amazon Simple Storage Service (Amazon S3)&lt;/a&gt;, &lt;a href="https://aws.amazon.com/iam/" target="_blank" rel="noopener"&gt;AWS Identity and Access Management (IAM)&lt;/a&gt;, &lt;a href="https://aws.amazon.com/systems-manager/" target="_blank" rel="noopener"&gt;AWS Systems Manager Parameter Store&lt;/a&gt;, &lt;a href="https://aws.amazon.com/api-gateway/" target="_blank" rel="noopener"&gt;Amazon API Gateway&lt;/a&gt;, &lt;a href="https://aws.amazon.com/lambda/" target="_blank" rel="noopener"&gt;AWS Lambda&lt;/a&gt;, &lt;a href="https://aws.amazon.com/waf/" target="_blank" rel="noopener"&gt;AWS WAF&lt;/a&gt;, &lt;a href="https://aws.amazon.com/cloudwatch/" target="_blank" rel="noopener"&gt;Amazon CloudWatch Logs&lt;/a&gt;, and &lt;a href="https://aws.amazon.com/lambda/lambda-microvms/" target="_blank" rel="noopener"&gt;AWS Lambda MicroVMs&lt;/a&gt;.&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/cli/latest/userguide/install-cliv2.html" target="_blank" rel="noopener"&gt;AWS Command Line Interface (AWS CLI) v2+&lt;/a&gt;.&lt;/li&gt;
 &lt;li&gt;The &lt;a href="https://docs.aws.amazon.com/serverless-application-model/latest/developerguide/install-sam-cli.html" target="_blank" rel="noopener"&gt;AWS SAM CLI&lt;/a&gt;.&lt;/li&gt;
 &lt;li&gt;An existing &lt;a href="https://platform.claude.com/docs/en/managed-agents/self-hosted-sandboxes" target="_blank" rel="noopener"&gt;Anthropic Claude Managed Agents&lt;/a&gt; agent configured with a &lt;code&gt;self_hosted&lt;/code&gt; environment (note the agent ID and environment ID).&lt;/li&gt;
 &lt;li&gt;A webhook signing secret and environment key, both generated in the &lt;a href="https://platform.claude.com/" target="_blank" rel="noopener"&gt;Anthropic Claude Console&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 id="four-steps"&gt;Four steps&lt;/h4&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;&lt;strong&gt;Deploy the control plane.&lt;/strong&gt; Build and deploy the SAM stack, which creates the launcher Lambda, API Gateway endpoint, WAF WebACL, DynamoDB idempotency table, Parameter Store entries, and MicroVM execution role.
  &lt;div class="hide-language"&gt;
   &lt;pre&gt;&lt;code class="language-bash"&gt;sam build
sam deploy --guided --capabilities CAPABILITY_NAMED_IAM&lt;/code&gt;&lt;/pre&gt;
  &lt;/div&gt;&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Register the webhook and populate secrets.&lt;/strong&gt; In the &lt;a href="https://platform.claude.com/" target="_blank" rel="noopener"&gt;Claude Console&lt;/a&gt;, register the stack’s &lt;code&gt;WebhookUrl&lt;/code&gt; output as a webhook endpoint subscribed to &lt;code&gt;session.status_run_started&lt;/code&gt;. Store the signing secret and environment key in the Parameter Store resources created by the stack.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Build the MicroVM image.&lt;/strong&gt; Package the Dockerfile and worker code, upload to Amazon S3, and create the image. The service runs your Dockerfile, launches the worker, and captures a Firecracker snapshot. Monitor build progress in Amazon CloudWatch under &lt;code&gt;/aws/lambda/microvms/&amp;lt;image-name&amp;gt;&lt;/code&gt;.
  &lt;div class="hide-language"&gt;
   &lt;pre&gt;&lt;code class="language-bash"&gt;./src/scripts/build-image.sh&lt;/code&gt;&lt;/pre&gt;
  &lt;/div&gt;&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Verify.&lt;/strong&gt; Create a test session and confirm a MicroVM launches and completes end-to-end. The reference solution includes a verification script that creates a session, triggers the webhook, and validates the full flow.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="using-claude-platform-on-aws-cpoa"&gt;Using Claude Platform on AWS (CPOA)&lt;/h2&gt;
&lt;p&gt;The preceding architecture works similarly when you access Claude through &lt;a href="https://aws.amazon.com/claude-platform/" target="_blank" rel="noopener"&gt;Claude Platform on AWS&lt;/a&gt; rather than the first-party API. Three things change in the worker:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;&lt;strong&gt;Client initialization.&lt;/strong&gt; Replace the first-party client with the AWS client and supply your workspace ID:
  &lt;div class="hide-language"&gt;
   &lt;pre&gt;&lt;code class="language-python"&gt;from anthropic import AnthropicAWS

client = AnthropicAWS(aws_region="us-east-1")

# Workspace ID is required on every request
# Set via ANTHROPIC_AWS_WORKSPACE_ID env var or pass per-call&lt;/code&gt;&lt;/pre&gt;
  &lt;/div&gt;&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Authentication options.&lt;/strong&gt; CPOA supports two modes:
  &lt;ol type="a"&gt;
   &lt;li&gt;&lt;strong&gt;CPOA API key&lt;/strong&gt; (&lt;code&gt;aws-external-anthropic-api-key-...&lt;/code&gt;): Store it in Parameter Store the same way as the first-party environment key. These keys are short-lived (12-hour STS tokens) and must be regenerated when they expire.&lt;/li&gt;
   &lt;li&gt;&lt;strong&gt;SigV4 (IAM)&lt;/strong&gt;: The MicroVM execution role can sign requests directly, so there is no secret to store or rotate. Set the environment key secret to a placeholder value (for example, &lt;code&gt;use-sigv4&lt;/code&gt;) and the SDK falls through to IAM credentials automatically. This is the recommended path for production.&lt;/li&gt;
  &lt;/ol&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;In both authentication modes, attach the AWS managed policy &lt;code&gt;AnthropicSelfHostedEnvironmentAccess&lt;/code&gt; to the MicroVM execution role. This policy grants the &lt;code&gt;aws-external-anthropic&lt;/code&gt; actions needed to poll the work queue, claim sessions, and post results. See &lt;a href="https://platform.claude.com/docs/en/api/claude-platform-on-aws-iam-actions" target="_blank" rel="noopener"&gt;IAM actions for Claude Platform on AWS&lt;/a&gt; for the full reference.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Prerequisite:&lt;/strong&gt; &lt;a href="https://aws.amazon.com/blogs/aws/simplify-access-to-external-services-using-aws-iam-outbound-identity-federation/" target="_blank" rel="noopener"&gt;Enable outbound web identity federation&lt;/a&gt; once per AWS account:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws iam enable-outbound-web-identity-federation&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Everything else, including webhook verification, deduplication, credential separation, and idle policy remains the same.&lt;/p&gt;
&lt;h2 id="security"&gt;Security&lt;/h2&gt;
&lt;p&gt;Earlier we talked about what goes wrong without isolation. Credentials exposed between sessions. Scripts unintentionally reaching production. Agents escaping their sandbox. This architecture implements defense in depth to help prevent these.&lt;/p&gt;
&lt;p&gt;Each component accesses a single, scoped secret. The launcher passes only an ARN reference to the worker credential into the MicroVM. The MicroVM’s execution role retrieves only that credential at runtime. The analytics database connection string does not touch the launcher and does not leave your environment.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://aws.amazon.com/waf/" target="_blank" rel="noopener"&gt;AWS WAF&lt;/a&gt; applies managed rule sets (OWASP, known bad inputs, IP reputation) and per-IP rate limiting. &lt;a href="https://aws.amazon.com/api-gateway/" target="_blank" rel="noopener"&gt;Amazon API Gateway&lt;/a&gt; request validation rejects malformed bodies. The launcher performs HMAC signature verification as the true authentication boundary.&lt;/p&gt;
&lt;p&gt;Each session runs in its own MicroVM. Sessions do not share memory, disk, or network namespaces. Firecracker provides hardware-virtualization-based isolation. The launcher IAM role reads only the signing secret. The MicroVM execution role reads only the worker credential. Both are scoped to specific Parameter Store ARNs. The &lt;a href="https://aws.amazon.com/s3/" target="_blank" rel="noopener"&gt;Amazon S3&lt;/a&gt; artifact bucket blocks public access, enables versioning, and uses server-side encryption.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;This post walked through how to give your internal AI agent a safe place to run database queries, generate code, and execute scripts on behalf of fifty developers without leaking data between sessions or reaching resources it shouldn’t.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://docs.aws.amazon.com/lambda/" target="_blank" rel="noopener"&gt;AWS Lambda MicroVMs&lt;/a&gt; provide ephemeral, VM-isolated compute environments that align with the per-session execution model of AI agent sandboxes. Snapshot-based launch avoids application startup latency. Idle policies terminate VMs once sessions complete. Firecracker isolation verifies that sessions do not share state. You pay only for active execution time and maintain full control over credentials, networking, and governance within your AWS boundary.&lt;/p&gt;
&lt;p&gt;You build and operate a serverless control plane. You get per-session VM isolation with no idle compute cost and no shared tenancy.&lt;/p&gt;
&lt;p&gt;To get started, explore these resources:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;Read more about Lambda MicroVMs in the &lt;a href="https://aws.amazon.com/blogs/aws/run-isolated-sandboxes-with-full-lifecycle-control-aws-lambda-introduces-microvms/" target="_blank" rel="noopener"&gt;launch announcement post&lt;/a&gt;.&lt;/li&gt;
 &lt;li&gt;Clone the reference solution in the &lt;a href="https://github.com/aws-samples/sample-lambda-microvm-claude-managed-agents" target="_blank" rel="noopener"&gt;aws-samples repository&lt;/a&gt;.&lt;/li&gt;
 &lt;li&gt;Add the &lt;a href="https://github.com/aws/agent-toolkit-for-aws/tree/main/skills/specialized-skills/serverless-skills/aws-lambda-microvms" target="_blank" rel="noopener"&gt;MicroVMs skill&lt;/a&gt; to your coding agent.&lt;/li&gt;
 &lt;li&gt;Learn more about AWS Lambda MicroVMs in the &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/lambda-microvms-guide.html" target="_blank" rel="noopener"&gt;Developer Guide&lt;/a&gt;.&lt;/li&gt;
 &lt;li&gt;Learn about Anthropic self-hosted sandbox configuration in the &lt;a href="https://platform.claude.com/docs/en/managed-agents/self-hosted-sandboxes" target="_blank" rel="noopener"&gt;self-hosted sandbox documentation&lt;/a&gt;.&lt;/li&gt;
 &lt;li&gt;For more serverless learning resources, visit &lt;a href="https://serverlessland.com/" target="_blank" rel="noopener"&gt;Serverless Land&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p style="clear: both"&gt;&lt;/p&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>Multi-modal autoscaling with Amazon EC2 Auto Scaling: adding signals for faster, more reliable scaling</title>
		<link>https://aws.amazon.com/blogs/compute/multi-modal-autoscaling-with-amazon-ec2-auto-scaling-adding-signals-for-faster-more-reliable-scaling/</link>
		
		<dc:creator><![CDATA[Shubhendu Dubey]]></dc:creator>
		<pubDate>Thu, 17 Sep 2026 18:18:38 +0000</pubDate>
				<category><![CDATA[Advanced (300)]]></category>
		<category><![CDATA[Amazon EC2]]></category>
		<category><![CDATA[Auto Scaling]]></category>
		<category><![CDATA[Technical How-to]]></category>
		<guid isPermaLink="false">f0cd20d867fe47384347b6529408872a9b3b0ef7</guid>

					<description>Multi-modal autoscaling with Amazon EC2 Auto Scaling combines infrastructure metrics like CPU with application-level signals, so a group scales on the demand its users create. In this post, we show you how to implement it, with code samples and results from a controlled test.</description>
										<content:encoded>&lt;p&gt;How do you handle unpredictable workload patterns that spike during promotional events or seasonal peaks? Multi-modal autoscaling with &lt;a href="https://aws.amazon.com/ec2/autoscaling/" target="_blank" rel="noopener"&gt;Amazon EC2 Auto Scaling&lt;/a&gt; combines infrastructure metrics like CPU with application-level signals, so a group scales on the demand its users create and not only on how busy the servers look. Those signals track the load that drives your business outcomes, such as sales or sign-ups.&lt;/p&gt;
&lt;p&gt;CPU-based autoscaling works well for many workloads, but some demand does not register as CPU right away. Adding signals such as request counts and application metrics lets a group respond to the load its users create. By publishing &lt;a href="https://aws.amazon.com/cloudwatch/" target="_blank" rel="noopener"&gt;Amazon CloudWatch&lt;/a&gt; custom metrics and application-driven triggers, you give Auto Scaling more information to act on.&lt;/p&gt;
&lt;p&gt;In our testing, a group that scaled only on CPU rejected about 7,000 checkout sessions during a demand spike, and a stronger baseline that added Application Load Balancer request count still rejected about 6,900. A group that added an application metric rejected none, and it held p99 latency to about 0.43 seconds against 2.25 seconds for the CPU-only group. Predictive scaling can add a forecasting layer for cyclical demand, but it needs days of history to be useful, so we treat it as a complement. In this post, we show you how to implement multi-modal autoscaling on EC2 Auto Scaling, with code samples and results from a controlled test.&lt;/p&gt;
&lt;h2 id="prerequisites"&gt;Prerequisites&lt;/h2&gt;
&lt;p&gt;To follow along, you need access to the following AWS services with appropriate permissions:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;
  &lt;p&gt;EC2 Auto Scaling, for scaling policies and group management.&lt;/p&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;CloudWatch, for metrics, alarms, and dashboards.&lt;/p&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;&lt;a href="https://aws.amazon.com/cloudformation/" target="_blank" rel="noopener"&gt;AWS CloudFormation&lt;/a&gt;, for infrastructure deployment.&lt;/p&gt;
 &lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="expanding-beyond-single-metric-scaling"&gt;Expanding beyond single-metric scaling&lt;/h2&gt;
&lt;p&gt;The default target tracking policy in EC2 Auto Scaling uses average CPU utilization, a practical starting point because CPU usage is a universal characteristic of compute workloads. Adding complementary signals, such as application-level metrics or predictive forecasting, gives Auto Scaling more information to make timely capacity decisions.&lt;/p&gt;
&lt;p&gt;For workloads that need a faster response from target tracking alone, see &lt;a href="https://aws.amazon.com/blogs/compute/faster-scaling-with-amazon-ec2-auto-scaling-target-tracking/" target="_blank" rel="noopener"&gt;Faster scaling with Amazon EC2 Auto Scaling target tracking&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;In distributed architectures, different components can have distinct scaling characteristics. An API gateway might correlate well with request rate, while a background processor scales better on queue depth. With multi-modal scaling, you can match each component’s policy to its actual workload pattern. For containerized workloads, consider &lt;a href="https://aws.amazon.com/solutions/guidance/event-driven-application-autoscaling-with-keda-on-amazon-eks/" target="_blank" rel="noopener"&gt;event-driven autoscaling with KEDA on Amazon Elastic Kubernetes Service (Amazon EKS)&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="multi-modal-autoscaling-architecture"&gt;Multi-modal autoscaling architecture&lt;/h2&gt;
&lt;p&gt;Multi-modal autoscaling combines three approaches to capacity management. Reactive scaling responds to current CloudWatch metrics, such as CPU utilization, memory, network throughput, response times, and custom application indicators. Application-metric scaling brings workload-specific signals into the decision, using custom CloudWatch metrics like active user sessions, queue depth, or transaction volume. These application metrics are often the closest measurable proxy for business activity such as orders or sign-ups. Predictive scaling uses machine learning in EC2 Auto Scaling to forecast capacity needs from historical patterns, so infrastructure scales before demand increases.&lt;/p&gt;
&lt;p&gt;With application-metric scaling, applications can scale on signals that infrastructure metrics miss. An ecommerce platform might scale on active checkout sessions, while a streaming service scales on concurrent stream counts. In the test later in this post, we use active checkout sessions as the custom metric.&lt;/p&gt;
&lt;h2 id="implementing-multi-modal-autoscaling"&gt;Implementing multi-modal autoscaling&lt;/h2&gt;
&lt;p&gt;This section builds the configuration in layers. Start with CPU target tracking as a baseline that every group keeps, then add a custom application metric that reflects real user load. The test later in this post compares these signals against a request-count baseline. Predictive scaling is an optional forecasting layer described at the end.&lt;/p&gt;
&lt;h3 id="step-1-cpu-target-tracking"&gt;Step 1: CPU target tracking&lt;/h3&gt;
&lt;p&gt;Start with the foundation that most workloads already use: a target tracking policy on average CPU utilization. Target tracking is a managed policy that adjusts capacity to keep a metric at or near a target value. It supports predefined metrics, including CPU utilization and request count per target, and custom CloudWatch metrics. When multiple target tracking policies are active, Auto Scaling coordinates them: it scales out if any policy requires it, but scales in only when all policies agree, which helps prevent oscillation.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-yaml"&gt;# CPU target tracking scaling policy (ASG A)
CPUTargetTrackingPolicy:
  Type: AWS::AutoScaling::ScalingPolicy
  Properties:
    AutoScalingGroupName: !Ref AutoScalingGroupName
    PolicyType: TargetTrackingScaling
    TargetTrackingConfiguration:
      PredefinedMetricSpecification:
        PredefinedMetricType: ASGAverageCPUUtilization
      TargetValue: 70&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Our test also included a second infrastructure baseline, a target tracking policy on the load balancer’s request count per target. It uses the same structure with a predefined metric:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-yaml"&gt;# Request count target tracking (ASG B)
RequestCountPTTargetTrackingPolicy:
  Type: AWS::AutoScaling::ScalingPolicy
  Properties:
    AutoScalingGroupName: !Ref AutoScalingGroupName
    PolicyType: TargetTrackingScaling
    TargetTrackingConfiguration:
      PredefinedMetricSpecification:
        PredefinedMetricType: ALBRequestCountPerTarget
        ResourceLabel: !Sub "${Alb.LoadBalancerFullName}/${TargetGroupB.TargetGroupFullName}"
      TargetValue: 300
      DisableScaleIn: false&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Step scaling is another option for spike handling. With step scaling, you can define different capacity increments for different alarm thresholds. It keeps evaluating the alarm during scaling activities, which can make it react faster than target tracking’s default evaluation window. Step scaling policies do not coordinate with each other.&lt;/p&gt;
&lt;h3 id="step-2-add-a-custom-application-metric"&gt;Step 2: Add a custom application metric&lt;/h3&gt;
&lt;p&gt;Next, add a second target tracking policy on a custom CloudWatch metric that reflects application load. In our test, instances publish an active checkout sessions metric at a 10-second resolution. To act on that resolution, set a Period of 10 seconds on the policy. Without it, the policy waits for three 1-minute datapoints like any other and the high-resolution metric only adds publishing cost. With it, a scale-out can begin in about 30 seconds. The policy includes the Auto Scaling group dimension so it tracks the metric for the right group. We set the target to 100 active sessions per instance, about 75 percent of the measured per-instance capacity of 135. This leaves headroom to absorb a spike while new instances boot.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-yaml"&gt;# Custom application metric target tracking (ASG C)
CustomMetricTargetTrackingPolicy:
  Type: AWS::AutoScaling::ScalingPolicy
  Properties:
    AutoScalingGroupName: !Ref AutoScalingGroupName
    PolicyType: TargetTrackingScaling
    TargetTrackingConfiguration:
      CustomizedMetricSpecification:
        MetricName: ActiveCheckoutSessions
        Namespace: ECommerce/CheckoutMetrics
        Dimensions:
          - Name: AutoScalingGroupName
            Value: !Ref AutoScalingGroupName
        Statistic: Average
        Period: 10
      TargetValue: 100
      DisableScaleIn: false&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h3 id="step-3-add-predictive-scaling"&gt;Step 3: Add predictive scaling&lt;/h3&gt;
&lt;p&gt;Predictive scaling is an optional forecasting layer. It uses machine learning in EC2 Auto Scaling to analyze historical load and scale ahead of recurring, cyclical demand, using customized metric specifications in ForecastAndScale mode. You need to provide several days of history for it to forecast well, so it complements reactive signals rather than replacing them. Start in ForecastOnly mode to watch the forecast before it drives any scaling.&lt;/p&gt;
&lt;h3 id="monitoring"&gt;Monitoring&lt;/h3&gt;
&lt;p&gt;Use CloudWatch dashboards to track how each policy contributes to scaling decisions, and set alarms on the metrics that matter for your workload, such as per-instance load or latency. Enable detailed monitoring on the launch template, with Monitoring set to true, so that the system publishes CPU metrics every minute. Without it, you cannot complete the CPU policy’s scale-in evaluation and your group will stop scaling in. Watching the policies side by side is what surfaced this scale-in behavior.&lt;/p&gt;
&lt;h2 id="performance-results"&gt;Performance results&lt;/h2&gt;
&lt;p&gt;We compared three Auto Scaling groups under an identical load profile in a single 75-minute test in the us-east-1 Region. Each group used c8g.large instances with a minimum of 6 and a maximum of 40 instances, and every group carried the same CPU target tracking policy at 70 percent as a fallback:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;
  &lt;p&gt;&lt;strong&gt;ASG A&lt;/strong&gt;: CPU target tracking only. This is the single-signal infrastructure baseline.&lt;/p&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;&lt;strong&gt;ASG B&lt;/strong&gt;: CPU target tracking plus an Application Load Balancer request-count policy. Request rate is a stronger infrastructure baseline than CPU alone.&lt;/p&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;&lt;strong&gt;ASG C&lt;/strong&gt;: CPU target tracking plus the custom checkout-sessions metric at 10-second resolution, published with a Period of 10 seconds.&lt;/p&gt;
 &lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;All the groups received the same load at the same time. During the shared ramp, arrival rate rose and every group scaled correctly, which makes the comparison fair. CPU crossed 70 percent on the CPU group, request count crossed its target of 300 on the request-count group, and all three converged to a similar size.&lt;/p&gt;
&lt;table border="1px" cellpadding="10px" width="100%"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;Ramp phase (arrivals 30% → 85%)&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;A: CPU only&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;B: A + ALB requests&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;C: A + app sessions&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Instances (start → peak)&lt;/td&gt;
   &lt;td&gt;6 → 9&lt;/td&gt;
   &lt;td&gt;6 → 10&lt;/td&gt;
   &lt;td&gt;6 → 10&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;CPU&lt;/td&gt;
   &lt;td&gt;73.6%&lt;/td&gt;
   &lt;td&gt;74.0%&lt;/td&gt;
   &lt;td&gt;69.0%&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Requests per target (target 300)&lt;/td&gt;
   &lt;td&gt;319&lt;/td&gt;
   &lt;td&gt;323&lt;/td&gt;
   &lt;td&gt;297&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Then arrival rate was held flat while the number of concurrent checkout sessions kept rising, a shape that infrastructure signals cannot see. The next table reports that divergence phase, measured directly from CloudWatch and the load balancer.&lt;/p&gt;
&lt;table border="1px" cellpadding="10px" width="100%"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;Measured metric (divergence)&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;A: CPU only&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;B: A + ALB requests&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;C: A + app sessions&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Instances (start → end)&lt;/td&gt;
   &lt;td&gt;11 → 11&lt;/td&gt;
   &lt;td&gt;11 → 11&lt;/td&gt;
   &lt;td&gt;11 → 22&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Rejected checkouts&lt;/td&gt;
   &lt;td&gt;7,064&lt;/td&gt;
   &lt;td&gt;6,886&lt;/td&gt;
   &lt;td&gt;0&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;CPU (start → end)&lt;/td&gt;
   &lt;td&gt;67.2% → 45.1%&lt;/td&gt;
   &lt;td&gt;66.5% → 45.5%&lt;/td&gt;
   &lt;td&gt;66.2% → 36.8%&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Peak sessions per instance&lt;/td&gt;
   &lt;td&gt;135&lt;/td&gt;
   &lt;td&gt;135&lt;/td&gt;
   &lt;td&gt;116&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Requests per target&lt;/td&gt;
   &lt;td&gt;285 → 279&lt;/td&gt;
   &lt;td&gt;282 → 279&lt;/td&gt;
   &lt;td&gt;277 → 154&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The difference is what each group could see. Arrival rate was held flat while the number of concurrent sessions rose, so CPU and request count stayed in range while the application saturated. The CPU-only and request-count groups held at 11 instances and rejected 7,064 and 6,886 checkouts. Their CPU even fell, from about 67 percent to about 45 percent, because a rejected request never reaches the work it would have done, so a policy targeting 70 percent saw spare capacity at the moment the application was failing users. The application-metric group read the rising sessions directly and scaled from 11 to 22 instances, rejecting none.&lt;/p&gt;
&lt;h3 id="effect-on-latency-and-errors"&gt;Effect on latency and errors&lt;/h3&gt;
&lt;p&gt;We measured latency and rejected checkouts on the load balancer during the test. At rest, all groups were identical. The gap opened only in the divergence phase, when concurrency rose without a matching change in arrival rate. Session slots are the scarce resource here, so sessions per instance is the causal driver of latency. The application group scales on sessions and we report latency as the outcome, rather than scaling on latency directly, which is not recommended for target tracking. The latency figures come from the load balancer’s TargetResponseTime at the end of the divergence phase. A client-side number measured over the internet would reflect network round-trip rather than the service.&lt;/p&gt;
&lt;table border="1px" cellpadding="10px" width="100%"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;Measured metric&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;A: CPU only&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;B: A + ALB requests&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;C: A + app sessions&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;TargetResponseTime (average), end of divergence&lt;/td&gt;
   &lt;td&gt;1.122 s&lt;/td&gt;
   &lt;td&gt;1.120 s&lt;/td&gt;
   &lt;td&gt;0.284 s&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;TargetResponseTime (p99), end of divergence&lt;/td&gt;
   &lt;td&gt;2.254 s&lt;/td&gt;
   &lt;td&gt;2.235 s&lt;/td&gt;
   &lt;td&gt;0.431 s&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;TargetResponseTime (average) at warm-up&lt;/td&gt;
   &lt;td&gt;0.283 s&lt;/td&gt;
   &lt;td&gt;0.283 s&lt;/td&gt;
   &lt;td&gt;0.284 s&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Rejected checkouts, drain phase&lt;/td&gt;
   &lt;td&gt;4,115&lt;/td&gt;
   &lt;td&gt;3,631&lt;/td&gt;
   &lt;td&gt;0&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The application-metric group, ASG C, kept per-instance load near its target and rejected no checkouts. Its average latency at the end of the divergence phase was 0.284 seconds against 1.122 for the CPU-only group, and its p99 was 0.431 seconds against 2.254. The request-count group, ASG B, tracked its own signal within range the whole time, which is exactly why it could not react: request rate was flat while concurrency climbed.&lt;/p&gt;
&lt;p&gt;Once every group has enough capacity, they perform the same. The value of the application signal is in the transition, the gap between when demand arrives and when the fleet is ready, which the infrastructure signals here never detected.&lt;/p&gt;
&lt;h3 id="handling-known-high-traffic-events"&gt;Handling known high-traffic events&lt;/h3&gt;
&lt;p&gt;For planned events like flash sales, &lt;a href="https://docs.aws.amazon.com/autoscaling/ec2/userguide/ec2-auto-scaling-scheduled-scaling.html" target="_blank" rel="noopener"&gt;scheduled scaling&lt;/a&gt; can pre-scale capacity ahead of time. Multi-modal scaling complements scheduled scaling by handling unplanned spikes and organic traffic that does not follow a fixed schedule.&lt;/p&gt;
&lt;h3 id="understanding-cost-implications"&gt;Understanding cost implications&lt;/h3&gt;
&lt;p&gt;Running the application signal requires more instances. During the spike it held about 22 instances, against 11 on the infrastructure-only groups. That extra capacity is what kept sessions per instance near the target and stopped the group from turning checkouts away. For your own workload, the question is whether a spike’s worth of extra instances costs less than the checkouts you would otherwise reject.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;Multi-modal autoscaling combines infrastructure metrics with application-level signals so a group scales on the demand its users create, not only on how busy its servers look. In our test, a group that scaled only on CPU rejected about 7,000 checkout sessions during a demand spike, and a stronger baseline that added load balancer request count still rejected about 6,900. A group that added a custom application metric rejected none, and held p99 latency near 0.43 seconds against 2.25 seconds for the CPU-only group. Its CPU even fell while the infrastructure groups were failing requests, which shows why an infrastructure signal alone can miss the demand that matters.&lt;/p&gt;
&lt;p&gt;Start with CPU target tracking as a fallback. Add a signal that reflects the load your users create, and pick the one that tracks closest to a business outcome like orders or active users. Set a Period on a high-resolution custom metric so the policy can act on it, and enable detailed monitoring so scale-in works. Predictive scaling is worth adding for demand you can forecast, once the group has days of history to learn from.&lt;/p&gt;
&lt;p&gt;To implement multi-modal autoscaling, you can open &lt;a href="https://console.aws.amazon.com/ec2autoscaling/" target="_blank" rel="noopener"&gt;Amazon EC2 Auto Scaling&lt;/a&gt; in the AWS Management Console and add a second scaling signal to one of your existing groups, following the configuration steps in this post. For a deeper look at target tracking behavior, see &lt;a href="https://aws.amazon.com/blogs/compute/faster-scaling-with-amazon-ec2-auto-scaling-target-tracking/" target="_blank" rel="noopener"&gt;Faster scaling with Amazon EC2 Auto Scaling target tracking&lt;/a&gt;. The &lt;a href="https://docs.aws.amazon.com/autoscaling/ec2/userguide/" target="_blank" rel="noopener"&gt;Amazon EC2 Auto Scaling User Guide&lt;/a&gt; covers predictive scaling policies, custom metrics, and scaling cooldowns in detail.&lt;/p&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>Deploying regulated workloads on AWS Local Zones and AWS Outposts</title>
		<link>https://aws.amazon.com/blogs/compute/deploying-regulated-workloads-on-aws-local-zones-and-aws-outposts/</link>
		
		<dc:creator><![CDATA[Brianna Rosentrater]]></dc:creator>
		<pubDate>Thu, 17 Sep 2026 17:29:40 +0000</pubDate>
				<category><![CDATA[Advanced (300)]]></category>
		<category><![CDATA[AWS Local Zones]]></category>
		<category><![CDATA[AWS Outposts]]></category>
		<category><![CDATA[Technical How-to]]></category>
		<guid isPermaLink="false">f530496389817aa7460be89a3363546afa37926a</guid>

					<description>AWS Local Zones and AWS Outposts help you keep regulated workloads within specific geographic boundaries. This post presents a framework of technologies, including the AWS Nitro System, AWS Organizations SCPs with AWS Control Tower, and Amazon VPC Traffic Mirroring, to help you build auditable, secure architectures for data residency.</description>
										<content:encoded>&lt;p&gt;Customers in many industries and geographic locations have specific data sovereignty and residency objectives. &lt;a href="https://aws.amazon.com/about-aws/global-infrastructure/localzones/" target="_blank" rel="noopener"&gt;AWS Local Zones&lt;/a&gt; and &lt;a href="https://aws.amazon.com/outposts/" target="_blank" rel="noopener"&gt;AWS Outposts&lt;/a&gt; are fully managed infrastructure solutions for customers that need to keep data within specific geographic boundaries and also want the scalability and innovation of cloud services. The challenge lies not only in where data resides, but in how to architect, secure, and audit these deployments effectively. Whether you’re architecting a new solution or migrating existing regulated workloads to AWS hybrid edge infrastructure, this post provides an overview of key technologies to help you build auditable and secure architectures that support your organization’s data residency objectives.&lt;/p&gt;
&lt;h2 id="solution-framework"&gt;Solution framework&lt;/h2&gt;
&lt;p&gt;Building a solution for data residency deployments on AWS hybrid infrastructure requires a thoughtful, layered approach. Rather than a prescriptive solution, this post presents a flexible framework that you can adapt to your specific operational requirements.&lt;/p&gt;
&lt;p&gt;The AWS Shared Responsibility Model clearly delineates where the responsibilities of AWS end and yours begin. This model provides a critical separation: AWS controls the management infrastructure, while your data remains inaccessible to AWS operators, as enforced by the &lt;a href="https://docs.aws.amazon.com/whitepapers/latest/security-design-of-aws-nitro-system/security-design-of-aws-nitro-system.html" target="_blank" rel="noopener"&gt;hardware-based isolation&lt;/a&gt; of the Nitro System. There is no operator access to the instances, applications, or data. This architectural separation provides the foundation for implementing stringent data residency controls.&lt;/p&gt;
&lt;p&gt;To build upon this foundation, you can implement security best practices by following the guidance in the &lt;a href="https://docs.aws.amazon.com/wellarchitected/latest/security-pillar/welcome.html" target="_blank" rel="noopener"&gt;AWS Well-Architected security pillar&lt;/a&gt;, which helps you strengthen application-level protections and data security controls. For deeper guidance, see the &lt;a href="https://docs.aws.amazon.com/wellarchitected/latest/data-residency-hybrid-cloud-services-lens/data-residency-with-hybrid-cloud-services-lens.html" target="_blank" rel="noopener"&gt;Data Residency with Hybrid Cloud Services Lens&lt;/a&gt;, which covers considerations for operations, security, cost, performance, and reliability for regulated workloads.&lt;/p&gt;
&lt;p&gt;When implementing data residency controls, you might need auditable evidence of traffic patterns for your internal governance processes. By using third-party monitoring tools combined with port mirroring capabilities, you can generate reports that show all traffic between your applications and databases remains within your Outpost environment. This visibility provides auditable evidence that traffic remains within your designated boundaries. You can also use &lt;a href="https://aws.amazon.com/artifact/" target="_blank" rel="noopener"&gt;AWS Artifact&lt;/a&gt; to access audit reports for your hybrid infrastructure.&lt;/p&gt;
&lt;p&gt;Governance tools form the final layer of this regulatory framework, establishing guardrails around your deployment. These tools continuously monitor and enforce configuration policies, verifying that your environment stays aligned with your security and governance policies, operates within required parameters, and alerts you proactively when issues arise. This shift from reactive to proactive management helps you maintain consistent governance of your environment at scale.&lt;/p&gt;
&lt;p&gt;Together, these layered technologies create a framework for deploying regulated workloads designed to support your data residency objectives while benefiting from the innovation and scalability of AWS services.&lt;/p&gt;
&lt;h3 id="shared-responsibility-model"&gt;Shared responsibility model&lt;/h3&gt;
&lt;p&gt;When extending workloads to Local Zones and Outposts, the shared responsibility model adapts to these hybrid cloud environments while maintaining the same core principles. AWS continues to manage the underlying infrastructure and services, while you retain control over your data, applications, and configurations. This supports consistent security postures whether workloads run in &lt;a href="https://aws.amazon.com/about-aws/global-infrastructure/regions_az/" target="_blank" rel="noopener"&gt;AWS Regions&lt;/a&gt;, Local Zones, or on Outposts infrastructure. You deploy Outposts in a data center or colocation facility of your choice. Under the shared responsibility model, you are responsible for meeting site requirements for power, cooling, on-premises networking, and the &lt;a href="https://docs.aws.amazon.com/outposts/latest/network-userguide/service-links.html" target="_blank" rel="noopener"&gt;Outpost service link&lt;/a&gt; connection to the Region. All traffic between the Outpost and the parent Region traverses an encrypted set of VPN connections over the service link, protecting communications in transit without requiring additional configuration. AWS continues to be responsible for maintaining the Outposts hardware as a managed service.&lt;/p&gt;
&lt;p&gt;This partnership approach to security means you can build auditable solutions with data residency controls without compromising on the innovation and scalability that AWS provides.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/08/28/ComputeBlog-2620-1.png" alt="AWS Shared Responsibility Model showing AWS responsibility for infrastructure and customer responsibility for data and configurations" width="800"&gt;
 &lt;p class="wp-caption-text"&gt;Figure 1: The AWS Shared Responsibility Model in a hybrid edge deployment&lt;/p&gt;
&lt;/div&gt;
&lt;h3 id="aws-nitro-system"&gt;AWS Nitro System&lt;/h3&gt;
&lt;p&gt;The &lt;a href="https://aws.amazon.com/ec2/nitro/" target="_blank" rel="noopener"&gt;AWS Nitro System&lt;/a&gt; is the virtualization platform that powers &lt;a href="https://aws.amazon.com/ec2/" target="_blank" rel="noopener"&gt;Amazon Elastic Compute Cloud (Amazon EC2)&lt;/a&gt; instances. It uses dedicated hardware and software to offload virtualization functions from the server CPU and delivers near-bare-metal performance. Both Outposts and Local Zones also use the Nitro System. By design, the Nitro System has &lt;a href="https://docs.aws.amazon.com/whitepapers/latest/security-design-of-aws-nitro-system/no-aws-operator-access.html" target="_blank" rel="noopener"&gt;no operator access&lt;/a&gt;. There is no way for AWS or any entity to log into the EC2 Nitro hosts, access compute resources, or reach encrypted customer data remotely. The following diagram shows the purpose-built hardware components of the Nitro System.&lt;/p&gt;
&lt;div style="width: 756px" class="wp-caption alignnone"&gt;
 &lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/08/28/ComputeBlog-2620-2.png" alt="AWS Nitro System stack showing the Nitro Card, Nitro Security Chip, and Nitro Hypervisor components" width="746"&gt;
 &lt;p class="wp-caption-text"&gt;Figure 2: The AWS Nitro System hardware and software stack&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;The Nitro System combines purpose-built hardware consisting of the following key security components:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;The &lt;a href="https://docs.aws.amazon.com/whitepapers/latest/security-design-of-aws-nitro-system/the-components-of-the-nitro-system.html#the-nitro-cards" target="_blank" rel="noopener"&gt;Nitro Card&lt;/a&gt; – provides I/O interfaces used for &lt;a href="https://aws.amazon.com/vpc/" target="_blank" rel="noopener"&gt;Amazon Virtual Private Cloud&lt;/a&gt; (Amazon VPC) network virtualization, &lt;a href="https://aws.amazon.com/ebs/" target="_blank" rel="noopener"&gt;Amazon Elastic Block Store&lt;/a&gt; (Amazon EBS), and instance storage, freeing up host CPU resources. Nitro Cards are logically isolated from the system main board that runs customer workloads and can be live-updated, reducing the need for maintenance windows and workload disruption.&lt;/li&gt;
 &lt;li&gt;The &lt;a href="https://docs.aws.amazon.com/whitepapers/latest/security-design-of-aws-nitro-system/the-components-of-the-nitro-system.html#the-nitro-security-chip" target="_blank" rel="noopener"&gt;Nitro Security Chip&lt;/a&gt; – provides the link between the Nitro Controller (used for orchestration) and the system main board. It intercepts and controls all firmware updates, preventing the main CPUs from being used to modify system firmware. This is particularly important when running bare metal EC2 instances. This chip is also used for boot control to validate system firmware integrity.&lt;/li&gt;
 &lt;li&gt;The &lt;a href="https://docs.aws.amazon.com/whitepapers/latest/security-design-of-aws-nitro-system/the-components-of-the-nitro-system.html#the-nitro-hypervisor" target="_blank" rel="noopener"&gt;Nitro Hypervisor&lt;/a&gt; – designed to receive EC2 instance management commands sent by the Nitro Controller, provide compute virtualization and logical instance isolation, and assign &lt;a href="https://en.wikipedia.org/wiki/Single-root_input/output_virtualization" target="_blank" rel="noopener"&gt;SR-IOV&lt;/a&gt; virtual functions as needed. It includes no general-purpose operating system features, only the features absolutely necessary for its function, and works with other purpose-built Nitro components to maintain its small size and bare-metal-like performance. This simple design reduces the risk for remote networking attacks and driver-based privilege escalations.&lt;/li&gt;
 &lt;li&gt;The &lt;a href="https://docs.aws.amazon.com/outposts/latest/userguide/data-protection.html#encryption-rest" target="_blank" rel="noopener"&gt;Nitro Security Key&lt;/a&gt; (Outposts only) – a removable device that stores the external key required to decrypt all data at rest on your Outpost. At the end of your Outposts commitment, after migrating your data off the Outpost, you can destroy this key to cryptographically shred any remaining data on the Outpost.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These components work together to provide a layered security approach that doesn’t compromise performance. By designing each component to have a specific function decoupled from the main system board, the Nitro System provides non-disruptive firmware updates and reduces classes of security issues often found in other hypervisor systems.&lt;/p&gt;
&lt;h3 id="aws-organizations-service-control-policies"&gt;AWS Organizations Service Control Policies&lt;/h3&gt;
&lt;p&gt;&lt;a href="https://docs.aws.amazon.com/organizations/latest/userguide/orgs_manage_policies_scps.html" target="_blank" rel="noopener"&gt;AWS Organizations Service Control Policies (SCPs)&lt;/a&gt; are a governance tool that helps you enforce data residency requirements by controlling where resources can be created and where data can be stored or processed. SCPs function as permission guardrails that define the maximum available permissions for IAM users and roles across your organization’s accounts. By implementing deny guardrails through SCPs, you can prevent resource provisioning in unwanted locations by restricting access to AWS APIs at the infrastructure level.&lt;/p&gt;
&lt;p&gt;When deploying regulated workloads on Local Zones and Outposts, SCPs work in conjunction with &lt;a href="https://aws.amazon.com/controltower/" target="_blank" rel="noopener"&gt;AWS Control Tower&lt;/a&gt; landing zones to create custom guardrails that control data movement, processing, and storage. These policies can be designed with either preventative rules (blocking actions before they occur) or detective rules (identifying compliance violations after the fact). SCPs can restrict data transfer, saving, or snapshot creation outside a specified AWS location, and they can isolate workloads to a specific location. You can apply SCPs across accounts and organizational units (OUs) within your organization. For more information, see &lt;a href="https://aws.amazon.com/blogs/compute/best-practices-for-managing-data-residency-in-aws-local-zones-using-landing-zone-controls/" target="_blank" rel="noopener"&gt;Best practices for managing data residency in AWS Local Zones using landing zone controls&lt;/a&gt; and &lt;a href="https://aws.amazon.com/blogs/compute/architecting-for-data-residency-with-aws-outposts-rack-and-landing-zone-guardrails/" target="_blank" rel="noopener"&gt;Architecting for data residency with AWS Outposts rack and landing zone guardrails&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Here’s an example SCP that restricts EC2 instance launches and network interface creation to only specified AWS Local Zone subnets:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-json"&gt;{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Sid": "DenyNotLocalZonesSubnet",
            "Effect": "Deny",
            "Action": [
                "ec2:RunInstances",
                "ec2:CreateNetworkInterface"
            ],
            "Resource": [
                "arn:aws:ec2:*:*:network-interface/*"
            ],
            "Condition": {
                "ForAllValues:ArnNotEquals": {
                    "ec2:Subnet": [
                        "arn:aws:ec2:us-west-2:123456789012:subnet/subnet-localzone1",
                        "arn:aws:ec2:us-west-2:123456789012:subnet/subnet-localzone2"
                    ]
                }
            }
        }
    ]
}&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h2 id="compliance-monitoring"&gt;Compliance monitoring&lt;/h2&gt;
&lt;p&gt;After you implement the security and governance best practices described in the preceding sections, you can demonstrate that traffic remains within your designated boundaries by using &lt;a href="https://docs.aws.amazon.com/vpc/latest/mirroring/what-is-traffic-mirroring.html" target="_blank" rel="noopener"&gt;Amazon VPC Traffic Mirroring&lt;/a&gt; (also called port mirroring outside of AWS). This mirrors traffic between your application servers and databases. You can use a mirror target report to show that the traffic does not transit the AWS Region. For step-by-step instructions, see &lt;a href="https://docs.aws.amazon.com/vpc/latest/mirroring/traffic-mirroring-getting-started.html" target="_blank" rel="noopener"&gt;Get started using Traffic Mirroring to monitor network traffic&lt;/a&gt;. The key configuration steps include the following:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;&lt;strong&gt;Configure security groups&lt;/strong&gt; – Allow inbound UDP port 4789 only from the security group of the source instances being mirrored, or from specific private CIDR ranges within the VPC. Do not open this port to &lt;code&gt;0.0.0.0/0&lt;/code&gt;.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Create a traffic mirror target&lt;/strong&gt; – Use the elastic network interface (ENI) of your monitoring instance.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Create a traffic mirror filter&lt;/strong&gt; – Define which traffic to capture, either all traffic or specific traffic.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Create mirror sessions&lt;/strong&gt; – Create one for each source instance you want to monitor. Lower session numbers are evaluated first when multiple sessions exist.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Capture traffic&lt;/strong&gt; – Use &lt;code&gt;tcpdump&lt;/code&gt; on the target instance to analyze mirrored packets.&lt;/li&gt;
&lt;/ol&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/08/28/ComputeBlog-2620-3.png" alt="Amazon VPC Traffic Mirroring architecture on Outposts, mirroring traffic between application servers and databases to a monitoring instance" width="800"&gt;
 &lt;p class="wp-caption-text"&gt;Figure 3: Amazon VPC Traffic Mirroring architecture on an Outpost&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;All instances must be in the same VPC or connected through &lt;a href="https://docs.aws.amazon.com/vpc/latest/peering/what-is-vpc-peering.html" target="_blank" rel="noopener"&gt;VPC peering&lt;/a&gt;. Traffic Mirroring encapsulates the mirrored traffic using VXLAN on UDP port 4789. Traffic Mirroring might impact network performance on source instances, so test in a development environment before deploying to production. The following image shows a sample traffic mirroring report that uses NetFlow Analyzer. For this post, all network traffic shown is simulated.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/08/28/ComputeBlog-2620-4.png" alt="NetFlow Analyzer sample report showing traffic captured from an environment with VPC Traffic Mirroring configured" width="800"&gt;
 &lt;p class="wp-caption-text"&gt;Figure 4: Sample traffic mirroring report in NetFlow Analyzer&lt;/p&gt;
&lt;/div&gt;
&lt;h2 id="clean-up"&gt;Clean up&lt;/h2&gt;
&lt;p&gt;If you tested the VPC Traffic Mirroring architecture described in the preceding section, terminate any unnecessary resources to avoid ongoing costs. Remove the resources in the following order to avoid dependency errors:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;&lt;strong&gt;Delete the traffic mirror sessions&lt;/strong&gt; – In the Amazon VPC console, navigate to &lt;strong&gt;Traffic Mirroring&lt;/strong&gt;, &lt;strong&gt;Mirror Sessions&lt;/strong&gt;. Select each mirror session you created and choose &lt;strong&gt;Actions&lt;/strong&gt;, &lt;strong&gt;Delete&lt;/strong&gt;. Repeat for all sessions associated with your source instances.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Delete the traffic mirror filter&lt;/strong&gt; – Navigate to &lt;strong&gt;Traffic Mirroring&lt;/strong&gt;, &lt;strong&gt;Mirror Filters&lt;/strong&gt;. Select the filter you created and choose &lt;strong&gt;Actions&lt;/strong&gt;, &lt;strong&gt;Delete&lt;/strong&gt;. You must delete all associated mirror sessions before you can delete the filter.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Delete the traffic mirror target&lt;/strong&gt; – Navigate to &lt;strong&gt;Traffic Mirroring&lt;/strong&gt;, &lt;strong&gt;Mirror Targets&lt;/strong&gt;. Select the target pointing to the ENI of your monitoring instance and choose &lt;strong&gt;Actions&lt;/strong&gt;, &lt;strong&gt;Delete&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Revoke security group rules&lt;/strong&gt; – Navigate to &lt;strong&gt;Security Groups&lt;/strong&gt; and select the security group attached to your monitoring instance. Remove the inbound rule that allows UDP port 4789 from the security group or CIDR range of the source instances.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Terminate the monitoring instance (optional)&lt;/strong&gt; – If you launched a dedicated EC2 instance solely for traffic capture and analysis, navigate to the EC2 console and terminate the instance. This also releases the associated ENI used as the mirror target.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Delete any stored packet captures (optional)&lt;/strong&gt; – If you saved &lt;code&gt;tcpdump&lt;/code&gt; output to &lt;a href="https://aws.amazon.com/s3/" target="_blank" rel="noopener"&gt;Amazon Simple Storage Service&lt;/a&gt; (Amazon S3) or local storage on the instance, delete those files if they are no longer needed for audit reporting.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;You can verify that all Traffic Mirroring resources have been removed by running the following AWS Command Line Interface (AWS CLI) commands:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws ec2 describe-traffic-mirror-sessions
aws ec2 describe-traffic-mirror-targets
aws ec2 describe-traffic-mirror-filters&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Each command should return an empty list, confirming that no mirroring resources remain active in your account.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;In this post, we covered how the AWS Nitro System, AWS Organizations SCPs with an AWS Control Tower landing zone, and VPC Traffic Mirroring provide capabilities for governing workloads with data residency requirements. Apply the SCP example in this post to test restricting instance launches and network interface creation to specific subnets. To learn more about Outposts for hybrid deployments, review the &lt;a href="https://docs.aws.amazon.com/outposts/latest/network-userguide/get-started-outposts.html" target="_blank" rel="noopener"&gt;Getting started with AWS Outposts&lt;/a&gt; guide and submit the &lt;a href="https://pages.awscloud.com/GLOBAL_PM_LN_outposts-features_2020084_7010z000001Lpcl_01.LandingPage.html" target="_blank" rel="noopener"&gt;AWS Outposts contact form&lt;/a&gt;. To get started with Local Zones, review the &lt;a href="https://docs.aws.amazon.com/local-zones/latest/ug/getting-started.html" target="_blank" rel="noopener"&gt;Getting started with AWS Local Zones&lt;/a&gt; guide, opt in to a Local Zone, and begin trying some of the architecture patterns described in this post.&lt;/p&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>Planning for disaster recovery using AWS Local Zones and AWS Outposts racks</title>
		<link>https://aws.amazon.com/blogs/compute/planning-for-disaster-recovery-using-aws-local-zones-and-aws-outposts-racks/</link>
		
		<dc:creator><![CDATA[Brianna Rosentrater]]></dc:creator>
		<pubDate>Thu, 17 Sep 2026 17:24:21 +0000</pubDate>
				<category><![CDATA[Advanced (300)]]></category>
		<category><![CDATA[AWS Local Zones]]></category>
		<category><![CDATA[AWS Outposts rack]]></category>
		<category><![CDATA[Technical How-to]]></category>
		<guid isPermaLink="false">80404857c8d91d4c3cb08acc28e09cef85c9de05</guid>

					<description>Build highly available architectures that span two AWS Outposts racks, or an Outpost rack and an AWS Local Zone, without single points of failure. This post covers three disaster recovery approaches: DNS-based failover, active/active load balancing, and hybrid database replication, with the RTO/RPO trade-offs of each.</description>
										<content:encoded>&lt;p&gt;AWS customers with data residency, low latency, or local data processing requirements can use &lt;a href="https://aws.amazon.com/hybrid/" target="_blank" rel="noopener"&gt;AWS Hybrid Cloud services&lt;/a&gt; to run their workloads either on-premises or within their regulatory boundary. Many of these workloads might be critical to their business, with minimal thresholds for downtime.&lt;/p&gt;
&lt;p&gt;This post provides practical design guidance for building highly available architectures that span either two &lt;a href="https://aws.amazon.com/outposts/" target="_blank" rel="noopener"&gt;AWS Outposts racks&lt;/a&gt; or an Outpost rack and an &lt;a href="https://aws.amazon.com/about-aws/global-infrastructure/localzones/" target="_blank" rel="noopener"&gt;AWS Local Zone&lt;/a&gt;, which are physically designed without single points of failure. By distributing workloads across two geographically and logically independent edge locations, you can achieve high availability while still benefiting from the low-latency, data-residency, and on-premises integration advantages that edge infrastructure provides. To maintain high availability, we recommend that you put a disaster recovery (DR) plan in place and conduct regular DR drills with your applications.&lt;/p&gt;
&lt;p&gt;The architectures presented here cover a range of approaches to failure detection and site switching. Each approach offers a different balance between &lt;a href="https://aws.amazon.com/blogs/mt/establishing-rpo-and-rto-targets-for-cloud-applications/" target="_blank" rel="noopener"&gt;Recovery Time Objective (RTO) and Recovery Point Objective (RPO)&lt;/a&gt;, operational complexity, and cost. By understanding these trade-offs, you can select the architecture that best aligns to your RPO/RTO targets, data protection and residency requirements, and budget. This helps you achieve the resilience your business requires without over-engineering or over-spending.&lt;/p&gt;
&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;Outposts and Local Zones function as extensions of a single Availability Zone (AZ) within the &lt;a href="https://aws.amazon.com/about-aws/global-infrastructure/regions_az/" target="_blank" rel="noopener"&gt;AWS Region&lt;/a&gt; they’re anchored to. For high availability when planning for failover between the two platforms, anchor each to a different parent Region or, at minimum, a different AZ within the same Region. This geographic separation supports the low RPO and RTO targets required for mission-critical workloads. The architectures in this post follow these principles:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;Shared responsibility&lt;/strong&gt;: AWS manages the Outposts and Local Zone infrastructure. You provide resilient power, cooling, and network connectivity for Outpost sites, and implement application-level failover logic.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Independent failure domains&lt;/strong&gt;: Treat each site as an independent failure domain. Anchoring each to a different parent AZ (or Region) ensures a failure in one AZ doesn’t affect both sites.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Resilient network connectivity&lt;/strong&gt;: Local Zones connect to their parent Region through the &lt;a href="https://aws.amazon.com/about-aws/global-infrastructure/global-network/" target="_blank" rel="noopener"&gt;AWS Global Network&lt;/a&gt;, designed for maximum resilience. Outpost racks include redundant Outpost Networking Devices (ONDs) with eBGP peering for multipath load balancing and failover.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Capacity planning for N+1&lt;/strong&gt;: Provision additional capacity beyond your expected workload so surviving instances can absorb the load during host failures without degradation.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="building-blocks-of-a-disaster-recovery-strategy"&gt;Building blocks of a disaster recovery strategy&lt;/h2&gt;
&lt;p&gt;A key design consideration is how quickly the architecture can detect a site failure and redirect traffic, and what layers of your workload need protection. Your RPO and RTO needs govern this requirement. This post covers three approaches to disaster recovery at different layers of your application, each offering a different balance between time-to-recovery and operational complexity:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;Active/passive DNS-based failover with &lt;a href="https://aws.amazon.com/route53/" target="_blank" rel="noopener"&gt;Amazon Route 53&lt;/a&gt; health checks.&lt;/li&gt;
 &lt;li&gt;Active/active architecture using physical or virtual load balancers deployed at each site.&lt;/li&gt;
 &lt;li&gt;Hybrid database recovery using native database engine replication with &lt;a href="https://aws.amazon.com/rds/" target="_blank" rel="noopener"&gt;Amazon Relational Database Service&lt;/a&gt; (Amazon RDS).&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Depending on your workload, you can implement a combination of these strategies to support the various layers of compute and storage of your application.&lt;/p&gt;
&lt;h3 id="activepassive-dns-based-failover-with-amazon-route-53-health-checks"&gt;Active/passive DNS-based failover with Amazon Route 53 health checks&lt;/h3&gt;
&lt;p&gt;If your workload consists of on-premises web servers accessible from the internet or internal network, you can use a DNS-based failover approach to reroute traffic to a healthy web server in the event of a hardware failure or site outage. Although this method supports any DNS service, the following architecture example uses Amazon Route 53.&lt;/p&gt;
&lt;p&gt;DNS-based failover supports two primary approaches. The first is health check routing, where DNS resolves requests to the IP address of a known good service endpoint. The second is multi-value routing, where the DNS service returns multiple IP addresses. Clients attempt connection to the first address and automatically fail over to subsequent addresses if the connection times out. &lt;a href="https://docs.aws.amazon.com/Route53/latest/DeveloperGuide/dns-failover.html" target="_blank" rel="noopener"&gt;Route 53 health checks&lt;/a&gt; continuously monitor endpoint availability. When a site becomes unreachable, Route 53 automatically updates DNS responses to route traffic to the surviving site. This approach is globally available and works across both Outposts and Local Zones.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/08/28/ComputeBlog-2576-1.png" alt="DNS-based failover architecture showing Route 53 health check monitoring, automatically routes to alternate health endpoint if primary endpoint fails health checks. This is an active/passive architecture." width="800"&gt;
 &lt;p class="wp-caption-text"&gt;Figure 1: Active/passive DNS-based failover architecture&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;When a specific application server fails and Route 53 determines it is unreachable, it is dynamically removed from future DNS responses. DNS systems typically have a Time to Live (TTL) of 300 seconds or longer, during which the DNS resolution is cached locally in the client. During this window, the client uses the cached IP address. New requests are automatically directed to active servers. The total recovery time is governed by the combination of the DNS TTL and health check timeout settings, typically resulting in a recovery time of 5 minutes or the TTL setting.&lt;/p&gt;
&lt;p&gt;This design pattern works between Outposts, between an Outpost and a third-party provider, between an Outpost and a Local Zone, or between Local Zones. Route 53 can also distribute traffic across these sites, supporting blue/green deployments where you gradually shift traffic from one environment to another.&lt;/p&gt;
&lt;p&gt;For AWS Outposts, you can configure Route 53 to monitor an endpoint in the Region. If the &lt;a href="https://docs.aws.amazon.com/outposts/latest/network-userguide/service-links.html" target="_blank" rel="noopener"&gt;Outpost service link&lt;/a&gt; disconnects for more than 5 minutes, DNS failover routes traffic to the secondary site. The Outpost and Local Zone can be anchored to the same or different Regions for added resiliency.&lt;/p&gt;
&lt;p&gt;As with all architectures using the public internet for replication traffic, configure &lt;a href="https://docs.aws.amazon.com/whitepapers/latest/logical-separation/encrypting-data-at-rest-and--in-transit.html" target="_blank" rel="noopener"&gt;Transport Layer Security (TLS) encryption in transit&lt;/a&gt;, &lt;a href="https://docs.aws.amazon.com/vpc/latest/userguide/vpc-security-groups.html" target="_blank" rel="noopener"&gt;security groups&lt;/a&gt;, and &lt;a href="https://docs.aws.amazon.com/vpc/latest/userguide/vpc-network-acls.html" target="_blank" rel="noopener"&gt;network access control lists (NACLs)&lt;/a&gt; to secure your data and control access to your subnet resources.&lt;/p&gt;
&lt;h3 id="activeactive-architecture-using-physical-or-virtual-load-balancers"&gt;Active/active architecture using physical or virtual load balancers&lt;/h3&gt;
&lt;p&gt;For Outposts-to-Outposts high availability when your workload must remain on-premises, an alternative to DNS-based failover is an active/active architecture using physical or virtual load balancers deployed at each site. Outposts racks support &lt;a href="https://aws.amazon.com/elasticloadbalancing/application-load-balancer/" target="_blank" rel="noopener"&gt;Application Load Balancer&lt;/a&gt; (ALB) as well as third-party L4 and L7 virtual or physical load balancers. Like the DNS-based architecture pattern, you can use this strategy to support workloads that consist of on-premises web servers with low latency, data residency, or continued operations requirements.&lt;/p&gt;
&lt;p&gt;In this model, both Outposts can simultaneously serve application traffic, with load balancers continuously monitoring the health of instances. When a failure is detected, the load balancer automatically shifts all traffic to the available Outpost without manual intervention or DNS propagation delays. Typically, the load balancers present a single IP address to service consumers and switch traffic when an endpoint is unavailable. Some load balancers can monitor load and switch traffic based on utilization to maintain response time. This design pattern is specific to Outposts, which support third-party devices connected on premises. It does not work with Local Zones, which are hosted in AWS datacenters.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;img src="//d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/14/ComputeBlog-2576-Figures-1-2-1.png" alt="Active/active architecture using load balancers at each site with data being replicated between sites." width="800"&gt;
 &lt;p class="wp-caption-text"&gt;Figure 2: Active/active architecture using physical or virtual load balancers&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;When you deploy this architecture, make sure the load balancer tier itself does not become a single point of failure. Deploy redundant load balancer instances at each Outpost, with failover between them, so the traffic management layer stays available even if one load balancer instance fails. We also recommend that you configure session persistence and connection draining on your load balancers to minimize disruption to in-flight requests during failover. With this approach, load balancer instances route traffic to your Outpost instances over the &lt;a href="https://docs.aws.amazon.com/outposts/latest/network-userguide/outposts-local-gateways.html" target="_blank" rel="noopener"&gt;local gateway&lt;/a&gt; of each Outpost. Traffic continues to be balanced between instances on each Outpost even if one of the Outposts loses its service link connection. You can anchor the Outposts to the same or different Availability Zones or Regions for added resiliency. This approach does require 2N infrastructure and an external load balancer, making it the most resource-intensive to implement.&lt;/p&gt;
&lt;p&gt;Some load balancers also support multiple endpoint monitoring. The load balancer monitors both the regional instance and the local application. If the service link fails, based on the administrator’s policy, it can drain connections and route traffic to the other Outpost. This keeps service status and logging fully available on the connected Local Zone or Outpost.&lt;/p&gt;
&lt;h3 id="hybrid-database-recovery-using-native-database-engine-replication"&gt;Hybrid database recovery using native database engine replication&lt;/h3&gt;
&lt;p&gt;If you have two or more logical Outpost racks, you can deploy &lt;a href="https://aws.amazon.com/blogs/database/deploy-amazon-rds-on-aws-outposts-with-multi-az-high-availability/" target="_blank" rel="noopener"&gt;Amazon RDS on AWS Outposts with Multi-AZ high availability&lt;/a&gt;. However, depending on your workload criticality, number of sites, and site locations, a more cost-effective disaster recovery option using one Outpost, one Local Zone, or both might be appropriate. For applications that require a database, you can use your chosen database engine’s native replication features or third-party tooling to create hybrid database architectures across an Outpost and a Local Zone, an Outpost and the Region, or a Local Zone and the Region. If using the Region for failover, this can be the same Region your Outpost or Local Zone is anchored to, or a different Region of your choosing. Limitations based on your chosen database engine and licensing terms apply. In this post, all architecture patterns use a PostgreSQL database. The following three hybrid database strategies expand on the &lt;a href="https://aws.amazon.com/blogs/database/understand-and-build-a-hybrid-database-with-amazon-rds-and-aws-outposts/" target="_blank" rel="noopener"&gt;hybrid database with Amazon RDS and AWS Outposts&lt;/a&gt; architecture to show how this design pattern supports disaster recovery across Outposts, Local Zones, and AWS Regions.&lt;/p&gt;
&lt;p&gt;These architectures use a bring-your-own-license (BYOL) model. The replica instance used for high availability and disaster recovery (HA/DR) is customer-managed, running on &lt;a href="https://aws.amazon.com/ec2/" target="_blank" rel="noopener"&gt;Amazon Elastic Compute Cloud&lt;/a&gt; (Amazon EC2) and &lt;a href="https://aws.amazon.com/ebs/" target="_blank" rel="noopener"&gt;Amazon Elastic Block Store&lt;/a&gt; (Amazon EBS). The primary database instance can also be customer-managed, or it can be an RDS-managed database instance so you can use a managed service as your primary operating model. Promoting a replica to primary after a failure is a manual process, but you can automate it with infrastructure as code. Promotion requires updating your DNS entry for the database instance.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/09/compute-2576-figure-3.png" alt="Architecture diagram showing database failover from an Outpost rack to a Local Zone. RDS can only run on 1 platform (Outposts, Local Zones, or in Region) and does not support RDS-native read replicas across platforms. Some database engines support native replication features, and customers can implement a self-managed replica using EC2 and EBS." width="800"&gt;
 &lt;p class="wp-caption-text"&gt;Figure 3: Database failover from an Outpost rack to a Local Zone&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;In the preceding diagram (Figure 3), the Outpost and the Local Zone can be in the same or different Regions for added resiliency.&lt;/p&gt;
&lt;p&gt;In the following diagram (Figure 4), replication traffic can use either the service link or the local gateway of the Outpost as its network path. Replication continues through the local gateway even if the service link fails. If using the service link, the EC2 replica database instance must be in the Outpost anchor Region. If using the local gateway, the EC2 replica database instance can be in the same Region as or a different Region from the Outpost anchor Region for added resiliency. You need to configure a &lt;a href="https://docs.aws.amazon.com/vpn/latest/s2svpn/how_it_works.html" target="_blank" rel="noopener"&gt;Virtual Private Gateway&lt;/a&gt;, &lt;a href="https://docs.aws.amazon.com/vpn/latest/s2svpn/how_it_works.html#Transit-Gateway" target="_blank" rel="noopener"&gt;Transit Gateway&lt;/a&gt;, or &lt;a href="https://docs.aws.amazon.com/vpc/latest/userguide/VPC_Internet_Gateway.html" target="_blank" rel="noopener"&gt;Internet Gateway&lt;/a&gt; in the Region to receive the replication traffic from the Outpost.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/08/28/ComputeBlog-2576-4.png" alt="Architecture showing database failover from an Outpost rack to an AWS Region. RDS can only run on 1 platform (Outposts, Local Zones, or in Region) and does not support RDS-native read replicas across platforms. Some database engines support native replication features, and customers can implement a self-managed replica using EC2 and EBS." width="800"&gt;
 &lt;p class="wp-caption-text"&gt;Figure 4: Database failover from an Outpost rack to an AWS Region&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;In the following diagram (Figure 5), both the primary database instance and the replica are self-hosted on EC2 and EBS. Check your specific &lt;a href="https://aws.amazon.com/about-aws/global-infrastructure/localzones/features/?nc=sn&amp;amp;loc=2" target="_blank" rel="noopener"&gt;Local Zone location&lt;/a&gt; for currently supported services to see if the primary database instance can use RDS. The Region used for the EC2 replica DB instance can be the same Region the Local Zone is a part of, or a different Region for added resiliency. If using a different Region, additional networking such as an &lt;a href="https://docs.aws.amazon.com/local-zones/latest/ug/local-zones-connectivity-igw.html" target="_blank" rel="noopener"&gt;Internet Gateway&lt;/a&gt; is required.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/08/28/ComputeBlog-2576-5.png" alt="Architecture showing database failover from a Local Zone to an AWS Region. Both the primary and Region database and replica instances are customer-managed using EC2 with EBS." width="800"&gt;
 &lt;p class="wp-caption-text"&gt;Figure 5: Database failover from a Local Zone to an AWS Region&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;In all three architectures, you need to update your DNS records and routing to complete failover to the secondary location. If your workload requires data residency, consider whether you can use an AWS Region as a failover destination.&lt;/p&gt;
&lt;h2 id="disaster-recovery-overview"&gt;Disaster recovery overview&lt;/h2&gt;
&lt;p&gt;The strategies discussed in this post support different RTO/RPO objectives. Recovery time depends on the amount of effort to redeploy or reroute to an alternate environment, and whether this process is manual or automated. Recovery point depends on whether the workload has persistent data that needs to be replicated, whether that replication happens synchronously or asynchronously, and whether you use a backup and restore approach. The following table is a high-level overview of the RTO/RPO you can expect for each approach based on these factors:&lt;/p&gt;
&lt;table border="1px" width="100%" cellpadding="10px"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;Architecture&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;RTO&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;RPO&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Active/passive DNS-based failover&lt;/td&gt;
   &lt;td&gt;Total failover time = DNS TTL + (health check interval x failure threshold)&lt;/td&gt;
   &lt;td&gt;Equal to replication schedule, or backup interval&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Active/active with load balancers&lt;/td&gt;
   &lt;td&gt;Seconds, traffic is already being routed to both environments&lt;/td&gt;
   &lt;td&gt;Seconds, data is already being synchronously replicated between sites&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Hybrid database (same anchor Region)&lt;/td&gt;
   &lt;td&gt;Minutes, time needed to reroute to replica instance&lt;/td&gt;
   &lt;td&gt;Equal to replication schedule, faster replication window expected for data traveling less distance&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Hybrid database (different anchor Region)&lt;/td&gt;
   &lt;td&gt;&amp;lt;1 hour, time needed to reroute to replica instance&lt;/td&gt;
   &lt;td&gt;Equal to replication schedule, longer replication window expected for data traveling a greater distance&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;em&gt;Table 1: RTO/RPO disaster recovery overview for each architecture&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;For the active/passive DNS-based failover architecture, DNS TTL, Route 53 health check interval, and failure threshold are all settings you configure to your preferences. The default Route 53 health check interval is 30 seconds, but can be set as low as 10 seconds. The default Route 53 failure threshold is 3 failed checks, but can be set to any number between 1 to 10. Generally, active/active architectures provide the lowest RTO/RPO for your workloads, whereas active/passive architectures incur some downtime during a disaster when rerouting user traffic to your passive standby environment. Review your workload RTO/RPO objectives to determine which approach is right for you. You might require different strategies for different tiers of workload based on your threshold for downtime at each tier.&lt;/p&gt;
&lt;h2 id="considerations"&gt;Considerations&lt;/h2&gt;
&lt;p&gt;When choosing a disaster recovery strategy, consider:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;Latency impact based on the location of your failover site and where your application users are.&lt;/li&gt;
 &lt;li&gt;Resilient network connectivity between your primary and secondary failover locations, or between your on-premises site and the AWS Region. Architecture-specific guidance is included in each section.&lt;/li&gt;
 &lt;li&gt;If your workload requires data residency, evaluate if a particular disaster recovery approach can be used.&lt;/li&gt;
 &lt;li&gt;Promoting a replica (either RDS-managed or customer-managed) is a manual process that you can automate with infrastructure as code, and it requires updating your DNS entry for the database instance.&lt;/li&gt;
 &lt;li&gt;Database replicas might support synchronous or asynchronous replication depending on the database engine. Consider your RPO objectives when evaluating the hybrid database architectures.&lt;/li&gt;
 &lt;li&gt;Limitations based on your chosen database engine and licensing terms apply. Consult your licensing terms and conduct failover drills to test these architecture patterns with your workloads before implementing into production.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;This post showed different architecture patterns for disaster recovery using both Outposts and Local Zones. See &lt;a href="https://aws.amazon.com/blogs/compute/building-highly-resilient-applications-with-on-premises-interdependencies-using-aws-local-zones/" target="_blank" rel="noopener"&gt;Building highly resilient applications with on-premises interdependencies using AWS Local Zones&lt;/a&gt; for additional guidance. Reach out to your AWS account team to learn more about the hybrid edge architectures discussed in this post. To discuss Outposts with an expert on any of these topics, submit &lt;a href="https://pages.awscloud.com/GLOBAL_PM_LN_outposts-features_2020084_7010z000001Lpcl_01.LandingPage.html" target="_blank" rel="noopener"&gt;the AWS Outposts contact form&lt;/a&gt;. To begin using Local Zones, enable a Local Zone from your account and start experimenting.&lt;/p&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>Validating multi-agent decisions with Step Functions and Bedrock AgentCore</title>
		<link>https://aws.amazon.com/blogs/compute/validating-multi-agent-decisions-with-step-functions-and-bedrock-agentcore/</link>
		
		<dc:creator><![CDATA[Ben Freiberg]]></dc:creator>
		<pubDate>Mon, 14 Sep 2026 16:47:23 +0000</pubDate>
				<category><![CDATA[Amazon Bedrock AgentCore]]></category>
		<category><![CDATA[AWS Step Functions]]></category>
		<category><![CDATA[Technical How-to]]></category>
		<guid isPermaLink="false">299c667ab403e16b2aa8de1f963d80e5dcac9274</guid>

					<description>Orchestrating specialized Amazon Bedrock AgentCore agents with AWS Step Functions gives you the reasoning power of generative AI with the guardrails of deterministic validation. Agents propose options, and deterministic code validates them before any action is taken, demonstrated here with an airline rebooking workflow.</description>
										<content:encoded>&lt;p&gt;For an airline operations team, a single flight cancellation sets off a chain reaction. Hundreds of passengers need new itineraries within minutes, and no two cases are alike. They have different loyalty tiers, sit on different fare rules, and have downstream connections that may not wait. Passengers have varying cabin and seat preferences and might fall under different regulatory entitlements depending on where they booked and where they are flying.&lt;/p&gt;
&lt;p&gt;Most airlines handle this with a layered system: rule-based automation covers the simple, one-hop rebooks, and everything else flows to a manual queue staffed by service agents. That works when disruptions are isolated. When they are not, the queue overwhelms, waiting times spike, and passengers booked alternatives themselves that create downstream knock-on disruptions.&lt;/p&gt;
&lt;p&gt;This is exactly where AI agents become compelling. An agent can reason across seat availability, fare rules, loyalty entitlements, and connection timing the way an experienced desk agent would, but at machine speed and across hundreds of cases in parallel. Multi-agent collaboration typically lets a supervisor agent route work to collaborator sub-agents, with the model itself deciding which sub-agent runs and in what order. But an unconstrained agent might optimize for the passenger’s preference while ignoring a codeshare restriction, rebook onto a flight that meets minimum connection time on paper but not at that specific airport, or calculate compensation under the wrong regulatory regime because it misread the ticket’s point of sale.&lt;/p&gt;
&lt;p&gt;Orchestrating specialized &lt;a href="https://aws.amazon.com/bedrock/agentcore/" target="_blank" rel="noopener"&gt;Amazon Bedrock AgentCore&lt;/a&gt; agents with &lt;a href="https://aws.amazon.com/step-functions/" target="_blank" rel="noopener"&gt;AWS Step Functions&lt;/a&gt; gives you the reasoning power of generative AI with the guardrails of deterministic validation. Step Functions adds native fan-out across thousands of passengers, a callback pattern that pauses a case for human review at zero compute cost, and a durable execution history that serves as your audit trail. The principle is that agents propose, and deterministic code validates. The pattern is demonstrated here for airline rebooking, but it applies anywhere automated decisions can have real financial or regulatory consequences.&lt;/p&gt;
&lt;h2 id="solution-overview"&gt;Solution overview&lt;/h2&gt;
&lt;p&gt;The design is a Step Functions state machine where deterministic steps that map to the business processes wrap each agent’s non-deterministic behavior. The following diagram shows the end-to-end flow. At a high level, the workflow proceeds through these stages:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;The workflow starts when a flight-cancellation event arrives, for example through an &lt;a href="https://aws.amazon.com/eventbridge/" target="_blank" rel="noopener"&gt;Amazon EventBridge&lt;/a&gt; integration.&lt;/li&gt;
 &lt;li&gt;An enrichment step pulls additional data such as the passenger manifest, current bookings, loyalty status, and stored preferences.&lt;/li&gt;
 &lt;li&gt;The workflow fans out to run agents in parallel for each affected passenger.&lt;/li&gt;
 &lt;li&gt;Two agents then run for each passenger: a find-alternatives agent proposes the top three rebooking options, and a compensation agent determines entitlement based on route, delay duration, and cause.&lt;/li&gt;
 &lt;li&gt;A deterministic validation step runs after each agent, confirming flights are actually bookable and entitlement rules are followed before either result is used.&lt;/li&gt;
 &lt;li&gt;The workflow checks whether the case can be auto-confirmed, or needs human review.&lt;/li&gt;
 &lt;li&gt;Bookings are confirmed, compensation issues, and confirmations are sent. Unresolved cases go to human agents.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The key principle: no agent Task state writes to the reservation system or issues a payment. Only deterministic Task states do that, and only after a deterministic validation step has passed.&lt;/p&gt;
&lt;h2 id="integrating-agentcore-harness-with-step-functions"&gt;Integrating AgentCore harness with Step Functions&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/harness.html" target="_blank" rel="noopener"&gt;AgentCore harness&lt;/a&gt; is a managed agent loop. You specify a model, system prompt, and tools, and the harness runs the reasoning cycle (model calls, tool execution, memory management, and response generation) end-to-end in a single API call. It handles the intra-agent orchestration so that Step Functions can focus on inter-agent orchestration: fan-out, sequencing, validation gates, and exception routing. Step Functions provides a native optimized integration for AgentCore harness, which calls &lt;code&gt;InvokeHarness&lt;/code&gt; against a target &lt;code&gt;HarnessArn&lt;/code&gt;. The optimized integration gives you an extended per-Task timeout of 15 minutes (900 seconds), so agents have enough time to reason through complex proposals. The trade-off is that the agent call is request-response only. There is no &lt;code&gt;.sync&lt;/code&gt; and no &lt;code&gt;.waitForTaskToken&lt;/code&gt; on the agent step, and only the final assistant message is returned to the state machine.&lt;/p&gt;
&lt;p&gt;The following Amazon States Language snippet shows the optimized harness invocation inside a Distributed Map. For the full definition, see the sample on &lt;a href="https://serverlessland.com/patterns/sfn-bedrockagentcore-harness-cdk" target="_blank" rel="noopener"&gt;Serverless Land&lt;/a&gt;.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-json"&gt;{
  "Comment": "Illustrative - per-passenger rebooking fan-out",
  "StartAt": "RebookPassengers",
  "States": {
    "RebookPassengers": {
      "Type": "Map",
      "ItemProcessor": {
        "ProcessorConfig": { "Mode": "DISTRIBUTED", "ExecutionType": "STANDARD" },
        "StartAt": "FindAlternatives",
        "States": {
          "FindAlternatives": {
            "Type": "Task",
            "Resource": "arn:aws:states:::bedrockagentcore:invokeHarness",
            "Parameters": {
              "HarnessArn": "&amp;lt;HARNESS_ARN&amp;gt;",
              "RuntimeSessionId.$": "$.passenger.sessionId",
              "Messages": [{ "Role": "user", "Content": [{ "Text.$": "States.JsonToString($.passenger)" }] }]
            },
            "TimeoutSeconds": 900,
            "ResultPath": "$.proposal",
            "Next": "ValidateRebooking"
          },
          "ValidateRebooking": { "Type": "Task", "Resource": "arn:aws:states:::lambda:invoke", "End": true }
        }
      },
      "MaxConcurrency": 1000,
      "End": true
    }
  }
}&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Note: the service name is spelled &lt;code&gt;bedrockagentcore&lt;/code&gt; (no hyphen) in the Step Functions resource string, but &lt;code&gt;bedrock-agentcore&lt;/code&gt; (with a hyphen) in the AgentCore ARN.&lt;/p&gt;
&lt;p&gt;&lt;code&gt;MaxConcurrency&lt;/code&gt; is set to 1000 to bound fan-out and protect downstream booking and inventory systems. If you omit it or set it to 0, you get the default behavior, which runs up to 10,000 parallel child executions. The agent Task flows directly into a deterministic validation Task.&lt;/p&gt;
&lt;h2 id="how-it-differs-from-managed-multi-agent-collaboration"&gt;How it differs from managed multi-agent collaboration&lt;/h2&gt;
&lt;p&gt;Multi-agent collaboration typically means that a supervisor agent decides which sub-agent runs and which tools it calls. Step Functions moves those decisions out of the agent layer entirely.&lt;/p&gt;
&lt;p&gt;This design puts orchestration, fan-out, validation, routing, retries, and the audit trail into Step Functions instead. Routing is a deterministic state you define and can test in isolation, not a model classification you hope will be consistent. You get a per-state execution history (every transition recorded with input and output), whereas agent-layer traces require opt-in and provide reasoning rationale rather than a durable, always-on event log.&lt;/p&gt;
&lt;h2 id="design-walkthrough-of-the-reference-app"&gt;Design walkthrough of the reference app&lt;/h2&gt;
&lt;p&gt;The following image shows the Step Functions state machine implemented by the sample application.&lt;/p&gt;
&lt;div style="width: 764px" class="wp-caption alignnone"&gt;
 &lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/08/27/ComputeBlog-2680-1.png" alt="Step Functions state machine showing the rebooking workflow: trigger, enrich, a Distributed Map fan-out with agent and deterministic validation stages, choice routing to human review, and execute stages" width="754"&gt;
 &lt;p class="wp-caption-text"&gt;Figure 1: The Step Functions state machine for the airline rebooking workflow&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;Stage 1, Trigger.&lt;/strong&gt; An Amazon EventBridge rule starts the workflow on a flight-cancellation event.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Stage 2, Enrich.&lt;/strong&gt; A deterministic Task pulls the passenger manifest, bookings, loyalty status, and preferences into the execution state.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Stage 3, Map fan-out.&lt;/strong&gt; A Distributed Map iterates affected passengers in parallel. The choice of Map type matters at scale. An inline Map runs up to 40 concurrent iterations, which is the documented threshold for choosing Distributed mode. A Distributed Map runs up to 10,000 parallel child executions by default, the right tool when a hub event affects thousands of passengers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Stage 4, Agent 1 find alternatives.&lt;/strong&gt; An AgentCore Task proposes the top three options, reasoning over the passenger’s preferences and constraints.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Stage 5, Deterministic validation of the rebooking proposal.&lt;/strong&gt; An &lt;a href="https://aws.amazon.com/lambda/" target="_blank" rel="noopener"&gt;AWS Lambda&lt;/a&gt; Task confirms each proposed flight is bookable by checking live availability, fare rules, and route validity, and it rejects hallucinated options. An agent might confidently propose a flight that does not exist. This stage is where that proposal is caught before it can become a ticket.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Stage 6a, Agent 2 draft compensation.&lt;/strong&gt; A second AgentCore Task drafts personalized, customer-facing notification text only. It does not compute entitlement and it does not move money.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Stage 6b, Deterministic entitlement check.&lt;/strong&gt; A Lambda Task computes and validates the entitlement against rule tables before any compensation issues. Consumer-protection frameworks such as EU Regulation 261/2004 (EU261) and US Department of Transportation refund rules are referenced here illustratively, to show why deterministic, auditable computation matters. The specific bands, triggers, and amounts are configuration you own and validate against current legal guidance, not something an agent should infer.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Stage 7, Choice routing and human-in-the-loop.&lt;/strong&gt; A Choice state auto-confirms rebookings for some passengers and routes the rest to a human. For the cases that need review, the workflow waits on a separate &lt;code&gt;.waitForTaskToken&lt;/code&gt; Task, backed by Lambda, &lt;a href="https://aws.amazon.com/sns/" target="_blank" rel="noopener"&gt;Amazon Simple Notification Service (Amazon SNS)&lt;/a&gt;, or &lt;a href="https://aws.amazon.com/sqs/" target="_blank" rel="noopener"&gt;Amazon Simple Queue Service (Amazon SQS)&lt;/a&gt;, with a 4-hour timeout. The wait happens on this separate callback Task, never on the agent step.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-json"&gt;{
  "Comment": "Illustrative - route and wait on a human, not on the agent",
  "RouteDecision": {
    "Type": "Choice",
    "Choices": [
      {
        "Variable": "$.passenger.autoConfirmEligible",
        "BooleanEquals": true,
        "Next": "ExecuteBooking"
      }
    ],
    "Default": "AwaitHumanApproval"
  },
  "AwaitHumanApproval": {
    "Type": "Task",
    "Resource": "arn:aws:states:::sqs:sendMessage.waitForTaskToken",
    "Parameters": {
      "QueueUrl": "https://sqs.us-east-1.amazonaws.com/123456789012/approvals",
      "MessageBody": {
        "taskToken.$": "$$.Task.Token",
        "passengerId.$": "$.passenger.id",
        "options.$": "$.proposal.validatedOptions"
      }
    },
    "TimeoutSeconds": 14400,
    "Next": "ExecuteBooking"
  }
}&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;Stage 8, Execute.&lt;/strong&gt; Deterministic Task states confirm the booking, issue compensation, and send confirmation. Each execution Task derives an idempotency token from the passenger ID combined with the decision ID (the child execution name, or a hash of the validated option set) and passes it to the booking and payment APIs, so a retry or redrive is a no-op instead of a duplicate booking or a second payment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Stage 9, Aggregate and exception routing.&lt;/strong&gt; The workflow summarizes outcomes and routes any unresolved cases to human agents.&lt;/p&gt;
&lt;p&gt;The validation step itself is ordinary deterministic code. A simplified rebooking validator in Python looks like the following.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-python"&gt;# Illustrative - reject any option the agent proposed that is not bookable
def handler(event, context):
    passenger = event["passenger"]
    proposed = event["proposal"]["options"]

    validated = []
    for option in proposed:
        flight = lookup_flight(option["flightId"])
        if flight is None:
            continue  # hallucinated or stale flight, reject
        if flight["seatsAvailable"] &amp;lt; 1:
            continue  # no inventory, reject
        if not fare_rules_allow(passenger["fareClass"], flight):
            continue  # fare rule violation, reject
        if not route_is_valid(passenger["origin"], passenger["destination"], flight):
            continue  # invalid route, reject
        validated.append(option)

    return {
        "passengerId": passenger["id"],
        "validatedOptions": validated,
        "autoConfirmEligible": passenger["loyaltyTier"] == "top" and len(validated) &amp;gt; 0,
    }&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h2 id="best-practices-and-guardrails"&gt;Best practices and guardrails&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Reject hallucinations through validations.&lt;/strong&gt; No agent proposal is applied without a deterministic validation step passing first. This minimizes the impact of hallucinations, prompt injections, or bugs on your workflow.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Keep a complete audit trail.&lt;/strong&gt; Step Functions execution history records every state transition, input, and output, and pairing that with durable persistence gives you a per-decision record. You can show exactly which proposal was made, which validation passed or failed, and who approved the exception.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Surface only true exceptions to humans.&lt;/strong&gt; Humans handle only what validation or the agent cannot resolve. Auto-confirmation handles the clear cases, and people spend their attention on the genuinely ambiguous ones.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Hold executions open cheaply.&lt;/strong&gt; The &lt;code&gt;.waitForTaskToken&lt;/code&gt; callback holds the execution open with no compute charges while the execution is paused. For example, you can cost-efficiently park thousands of pending approvals overnight. Refer to the &lt;a href="https://aws.amazon.com/step-functions/pricing/" target="_blank" rel="noopener"&gt;AWS Step Functions pricing page&lt;/a&gt; for current details.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Make execution idempotent.&lt;/strong&gt; Guard reservation execution and compensation issuance against retries and double-sends, as shown in Stage 8. Derive the idempotency token from the passenger ID and decision ID, and pass it to your booking and payment APIs so that a replay is a no-op.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Respect cost and timeouts.&lt;/strong&gt; Keep each per-agent Task timeout within the 15-minute quota, bound your Map concurrency to protect downstream systems, and track the token usage returned in the agent response so you can attribute and forecast cost.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Handle errors deliberately.&lt;/strong&gt; Apply &lt;code&gt;Retry&lt;/code&gt; and &lt;code&gt;Catch&lt;/code&gt; on the agent Tasks for conditions such as &lt;code&gt;BedrockAgentCore.ThrottlingException&lt;/code&gt; and &lt;code&gt;BedrockAgentCore.ResourceNotFoundException&lt;/code&gt;, and on the Lambda validation Tasks for their own failure modes. A &lt;code&gt;Catch&lt;/code&gt; on an agent Task can route a stuck passenger straight to the human queue rather than failing the whole child execution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Confirm availability and Region support.&lt;/strong&gt; Check the current availability status and supported AWS Regions for AgentCore and the Step Functions integration at the &lt;a href="https://builder.aws.com/build/capabilities/explore?tab=service-feature" target="_blank" rel="noopener"&gt;AWS Capabilities by Region&lt;/a&gt; on Builder Center.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;A flight-cancellation event is a challenging test of automated decision-making, because the output can have immediate financial impact. The way to use AI agents safely in that setting is to let them do what they are good at, proposing options and drafting language, while never letting a proposal become an action until deterministic code has approved it. In this design, orchestration, fan-out, validation, routing, and retries are implemented in Step Functions rather than inside an agent’s reasoning. Agents do not make changes directly, and their output is only applied after deterministic validation. You get a per-decision record for review, and you hold exceptions open on a callback that adds no compute or storage cost while it waits.&lt;/p&gt;
&lt;p&gt;To get started, deploy the &lt;a href="https://serverlessland.com/patterns/sfn-bedrockagentcore-harness-cdk" target="_blank" rel="noopener"&gt;reference pattern from Serverless Land&lt;/a&gt; and adapt the validation layer to your own workflow.&lt;/p&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>Customize Amazon API Gateway destinations for execution logs</title>
		<link>https://aws.amazon.com/blogs/compute/customize-amazon-api-gateway-destinations-for-execution-logs/</link>
		
		<dc:creator><![CDATA[Giedrius Praspaliauskas]]></dc:creator>
		<pubDate>Wed, 09 Sep 2026 22:35:42 +0000</pubDate>
				<category><![CDATA[Amazon API Gateway]]></category>
		<category><![CDATA[Announcements]]></category>
		<category><![CDATA[Intermediate (200)]]></category>
		<guid isPermaLink="false">11804c9f98508b1bc1561231cc420556cf309fb2</guid>

					<description>Amazon API Gateway execution logs help you trace request processing step by step through your REST API stages. They capture authorization results, integration latency, mapping template output, and error details that are otherwise invisible at the API surface. When a production request fails in a way the access log cannot explain, the execution log is […]</description>
										<content:encoded>&lt;p&gt;&lt;a href="https://aws.amazon.com/api-gateway/" target="_blank" rel="noopener"&gt;Amazon API Gateway&lt;/a&gt; execution logs help you trace request processing step by step through your REST API stages. They capture authorization results, integration latency, mapping template output, and error details that are otherwise invisible at the API surface. When a production request fails in a way the access log cannot explain, the execution log is usually where you find the explanation.&lt;/p&gt;
&lt;p&gt;Until now, execution logs had two constraints. Every log event was truncated at 1 KB, so a request carrying a moderately sized JSON body would exceed that limit and the remainder was dropped. Logs could only go to the auto-managed log group that API Gateway creates for you (&lt;code&gt;API-Gateway-Execution-Logs_{rest-api-id}/{stage_name}&lt;/code&gt;).&lt;/p&gt;
&lt;p&gt;With &lt;a href="https://aws.amazon.com/cloudwatch/" target="_blank" rel="noopener"&gt;Amazon CloudWatch&lt;/a&gt; Logs delivery for REST API execution logs, you can now route execution logs to Amazon CloudWatch Logs, &lt;a href="https://aws.amazon.com/s3/" target="_blank" rel="noopener"&gt;Amazon Simple Storage Service (Amazon S3)&lt;/a&gt;, or &lt;a href="https://aws.amazon.com/firehose/" target="_blank" rel="noopener"&gt;Amazon Data Firehose&lt;/a&gt;. Log events can be up to 1 MB per entry, and you benefit from &lt;a href="https://aws.amazon.com/cloudwatch/pricing/" target="_blank" rel="noopener"&gt;vended logs pricing&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;In this post, you learn how CloudWatch Logs delivery works with API Gateway execution logs, how to configure it, and what patterns work best for common observability scenarios.&lt;/p&gt;
&lt;h2 id="understanding-api-gateway-execution-logs"&gt;Understanding API Gateway execution logs&lt;/h2&gt;
&lt;p&gt;API Gateway produces two categories of logs: access logs and execution logs. Access logs record a summary line per request, similar to an HTTP server access log. You configure the format and destination yourself.&lt;/p&gt;
&lt;p&gt;Execution logs are different. They capture the internal processing of each request as it moves through the API Gateway pipeline: authorizer evaluation, request validation, integration dispatch, response mapping, and error handling. These logs exist so you can answer questions such as “why did my authorizer reject this token?” or “what did the mapping template produce before it reached my backend integration?”&lt;/p&gt;
&lt;p&gt;API Gateway manages execution log creation automatically. When you set &lt;code&gt;loggingLevel&lt;/code&gt; to &lt;code&gt;INFO&lt;/code&gt; or &lt;code&gt;ERROR&lt;/code&gt; in your stage’s method settings, the service writes execution log events to a CloudWatch Logs log group it manages on your behalf. You do not choose the log group name or configure retention directly on it.&lt;/p&gt;
&lt;p&gt;The auto-managed model works for many customers but may create friction for teams with specific observability requirements. Compliance frameworks that require logs in S3 with a particular prefix structure need an extra subscription filter and delivery mechanism. Sending execution logs into a security information and event management (SIEM) tool through a Firehose stream requires a forwarding layer.&lt;/p&gt;
&lt;h2 id="configurable-log-delivery-with-cloudwatch-logs"&gt;Configurable log delivery with CloudWatch Logs&lt;/h2&gt;
&lt;p&gt;CloudWatch Logs delivery separates log routing from log content. Two concepts control the behavior:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;DeliverySource&lt;/strong&gt; is scoped to your API Gateway stage ARN. It defines where logs go. You create a delivery source, then attach one or more delivery destinations (CloudWatch Logs log group, S3 bucket, or Firehose stream).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;MethodSettings&lt;/strong&gt; controls what gets logged. The &lt;code&gt;loggingLevel&lt;/code&gt; setting (&lt;code&gt;INFO&lt;/code&gt;, &lt;code&gt;ERROR&lt;/code&gt;, or &lt;code&gt;OFF&lt;/code&gt;) and &lt;code&gt;dataTraceEnabled&lt;/code&gt; flag still determine which log events API Gateway produces. These settings work the same way regardless of whether you use the auto-managed log group or CloudWatch Logs delivery.&lt;/p&gt;
&lt;p&gt;When you create a delivery using the CloudWatch Logs APIs, CloudWatch Logs activates your log delivery on your API Gateway stage. When you delete the delivery, CloudWatch Logs disables it accordingly. You do not need to flip any flags on the API Gateway side, and the execution logs automatically resume flowing to the auto-managed log group.&lt;/p&gt;
&lt;p&gt;Your existing method settings keep their meaning. The &lt;code&gt;loggingLevel&lt;/code&gt; and &lt;code&gt;dataTraceEnabled&lt;/code&gt; values continue to control log content. If &lt;code&gt;loggingLevel&lt;/code&gt; is already &lt;code&gt;INFO&lt;/code&gt; or &lt;code&gt;ERROR&lt;/code&gt;, creating a delivery redirects those logs to your chosen destination with no further configuration.&lt;/p&gt;
&lt;p&gt;The following diagram shows how the pieces fit together.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/07/14/ComputeBlog-2656-1.png" alt="Diagram showing one API Gateway stage delivery source fanning out to CloudWatch Logs, Amazon S3, and Firehose destinations." width="800"&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Figure 1 — A single delivery source scoped to an API Gateway stage feeds one or more deliveries, each of which writes to a delivery destination backed by CloudWatch Logs, Amazon S3, or Amazon Data Firehose&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;The following table summarizes what changes when log delivery is active.&lt;/p&gt;
&lt;table border="1px" width="100%" cellpadding="10px"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;Aspect&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Standard execution logging&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Log delivery&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Destination&lt;/td&gt;
   &lt;td&gt;Auto-managed CloudWatch Logs log group&lt;/td&gt;
   &lt;td&gt;CloudWatch Logs, Amazon S3, or Firehose&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Multi-destination&lt;/td&gt;
   &lt;td&gt;No&lt;/td&gt;
   &lt;td&gt;Yes&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Pricing&lt;/td&gt;
   &lt;td&gt;Standard CloudWatch Logs ingestion&lt;/td&gt;
   &lt;td&gt;Vended logs pricing&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Log event size&lt;/td&gt;
   &lt;td&gt;Truncated at 1 KB&lt;/td&gt;
   &lt;td&gt;Up to 1 MB&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Setup&lt;/td&gt;
   &lt;td&gt;Set &lt;code&gt;loggingLevel&lt;/code&gt; in MethodSettings&lt;/td&gt;
   &lt;td&gt;Create delivery through CloudWatch Logs APIs&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Teardown&lt;/td&gt;
   &lt;td&gt;Set &lt;code&gt;loggingLevel&lt;/code&gt; to &lt;code&gt;OFF&lt;/code&gt;&lt;/td&gt;
   &lt;td&gt;Delete delivery&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;h3 id="what-stays-the-same"&gt;What stays the same&lt;/h3&gt;
&lt;p&gt;Only execution log routing changes. Access logs continue to flow through &lt;code&gt;accessLogSettings&lt;/code&gt; to whatever log group you configure, and unrelated stage features such as AWS X-Ray tracing, detailed CloudWatch metrics, throttling, and caching behave exactly as they did before.&lt;/p&gt;
&lt;h2 id="configuration-and-integration-options"&gt;Configuration and integration options&lt;/h2&gt;
&lt;p&gt;Before you create a delivery, confirm the following requirements:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;The API Gateway REST API is deployed to a stage.&lt;/li&gt;
 &lt;li&gt;&lt;code&gt;loggingLevel&lt;/code&gt; is set to &lt;code&gt;INFO&lt;/code&gt; or &lt;code&gt;ERROR&lt;/code&gt; in MethodSettings.&lt;/li&gt;
 &lt;li&gt;The account-level CloudWatch Logs IAM role is configured. For setup steps, see &lt;a href="https://docs.aws.amazon.com/apigateway/latest/developerguide/set-up-logging.html" target="_blank" rel="noopener"&gt;Set up CloudWatch logging for REST APIs in API Gateway&lt;/a&gt;.&lt;/li&gt;
 &lt;li&gt;For cross-account delivery, the destination has an appropriate resource policy attached through &lt;code&gt;PutDeliveryDestinationPolicy&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="sending-logs-to-a-custom-cloudwatch-logs-log-group"&gt;Sending logs to a custom CloudWatch Logs log group&lt;/h3&gt;
&lt;p&gt;The most common starting point is redirecting execution logs to a log group you own. You get direct control over retention policies, metric filters, and subscription filters. The following steps use the AWS Command Line Interface (AWS CLI) with the fictitious REST API ID &lt;code&gt;abc123&lt;/code&gt;, stage &lt;code&gt;prod&lt;/code&gt;, Region &lt;code&gt;us-east-1&lt;/code&gt;, and account &lt;code&gt;111122223333&lt;/code&gt;.&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;Create a delivery source referencing your stage ARN. The log type for REST API execution logs is &lt;code&gt;EXECUTION_LOGS&lt;/code&gt;:
  &lt;div class="hide-language"&gt;
   &lt;pre&gt;&lt;code class="language-bash"&gt;aws logs put-delivery-source \
    --name my-apigw-execution-logs \
    --resource-arn arn:aws:apigateway:us-east-1:111122223333:/restapis/abc123/stages/prod \
    --log-type EXECUTION_LOGS&lt;/code&gt;&lt;/pre&gt;
  &lt;/div&gt;&lt;/li&gt;
 &lt;li&gt;Create a delivery destination pointing to your custom (existing) log group, then create the delivery that connects them:
  &lt;div class="hide-language"&gt;
   &lt;pre&gt;&lt;code class="language-bash"&gt;aws logs put-delivery-destination \
    --name my-execution-log-destination \
    --delivery-destination-configuration \
        destinationResourceArn=arn:aws:logs:us-east-1:111122223333:log-group:/my-api/execution-logs&lt;/code&gt;&lt;/pre&gt;
  &lt;/div&gt;
  &lt;div class="hide-language"&gt;
   &lt;pre&gt;&lt;code class="language-bash"&gt;aws logs create-delivery \
    --delivery-source-name my-apigw-execution-logs \
    --delivery-destination-arn arn:aws:logs:us-east-1:111122223333:delivery-destination:my-execution-log-destination&lt;/code&gt;&lt;/pre&gt;
  &lt;/div&gt;&lt;/li&gt;
 &lt;li&gt;Verify that the delivery is active by listing deliveries for the source:
  &lt;div class="hide-language"&gt;
   &lt;pre&gt;&lt;code class="language-bash"&gt;aws logs describe-deliveries&lt;/code&gt;&lt;/pre&gt;
  &lt;/div&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The response includes the delivery ID, source, and destination ARN after delivery is established. Execution logs flow to &lt;code&gt;/my-api/execution-logs&lt;/code&gt; instead of the auto-managed group.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; Log delivery adds structured fields (&lt;code&gt;resource_arn&lt;/code&gt;, &lt;code&gt;event_timestamp&lt;/code&gt;, &lt;code&gt;api_id&lt;/code&gt;, &lt;code&gt;stage&lt;/code&gt;, &lt;code&gt;resource_path&lt;/code&gt;, &lt;code&gt;http_method&lt;/code&gt;, and &lt;code&gt;payload&lt;/code&gt;) to each event, so a new delivery emits more than your previous logs. To keep the traditional execution log format with nothing extra, set output format and record fields while creating delivery destination and creating delivery:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws logs put-delivery-destination \
    --output-format "plain" ...

aws logs create-delivery \
    --record-fields "payload" \
    --field-delimiter "" ...&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h3 id="routing-logs-to-amazon-s3"&gt;Routing logs to Amazon S3&lt;/h3&gt;
&lt;p&gt;S3 works well for long-term retention at lower cost, or for feeding logs into analytics tools such as &lt;a href="https://aws.amazon.com/athena/" target="_blank" rel="noopener"&gt;Amazon Athena&lt;/a&gt;. The bucket must be in the same region as your API. Create a delivery destination pointing to your bucket:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws logs put-delivery-destination \
    --name s3-archive-destination \
    --delivery-destination-configuration \
        destinationResourceArn=arn:aws:s3:::amzn-s3-demo-apigw-logs&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Then create a delivery using the same source name. CloudWatch Logs delivers the events to your bucket, where you can query them with Athena or catalog them with &lt;a href="https://aws.amazon.com/glue/" target="_blank" rel="noopener"&gt;AWS Glue&lt;/a&gt;.&lt;/p&gt;
&lt;h3 id="streaming-to-amazon-data-firehose"&gt;Streaming to Amazon Data Firehose&lt;/h3&gt;
&lt;p&gt;For real-time analytics pipelines or third-party SIEM integration, Firehose delivery sends execution log events directly to your stream. The setup is identical: create a delivery destination with your Firehose stream ARN, then create a delivery. With direct Firehose delivery, you no longer need to maintain CloudWatch Logs subscription filters and &lt;a href="https://aws.amazon.com/lambda/" target="_blank" rel="noopener"&gt;AWS Lambda&lt;/a&gt; forwarders to route execution logs to external analytics systems.&lt;/p&gt;
&lt;h3 id="multi-destination-delivery-and-per-destination-shaping"&gt;Multi-destination delivery and per-destination shaping&lt;/h3&gt;
&lt;p&gt;A single delivery source supports multiple destinations. You can route the same execution logs to CloudWatch Logs for real-time alerting, S3 for long-term compliance retention, and Firehose for your SIEM, all from one stage. Create additional deliveries using the same delivery source with different destination ARNs.&lt;/p&gt;
&lt;p&gt;Each destination receives identical log events. To shape what reaches each destination, apply a CloudWatch Logs subscription filter on the CloudWatch Logs destination. For example, you can forward only &lt;code&gt;ERROR&lt;/code&gt;-level events to a Lambda function that pushes alerts to a SIEM, while the same delivery source writes the full event stream to S3 for compliance.&lt;/p&gt;
&lt;h3 id="management-console-experience"&gt;Management console experience&lt;/h3&gt;
&lt;p&gt;You can also add a log delivery destination in the management console after you enable logging for the stage.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/07/14/ComputeBlog-2656-2.png" alt="API Gateway console showing the option to add a log delivery destination after logging is enabled for the stage." width="800"&gt;&lt;/p&gt;
&lt;p&gt;You can specify multiple destinations, both in the current or in a different account:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/07/14/ComputeBlog-2656-3.png" alt="API Gateway console showing multiple delivery destinations configured, including cross-account options." width="800"&gt;&lt;/p&gt;
&lt;h3 id="keeping-existing-monitoring-intact"&gt;Keeping existing monitoring intact&lt;/h3&gt;
&lt;p&gt;If you have dashboards or alarms on the auto-managed log group, use that same log group as one of your delivery destinations. Your existing monitoring keeps working, and you gain the ability to send logs to additional destinations such as S3 or Firehose in parallel.&lt;/p&gt;
&lt;h2 id="best-practices"&gt;Best practices&lt;/h2&gt;
&lt;p&gt;Update dashboards and alarms before enabling log delivery. When you activate log delivery, the auto-managed log group stops receiving logs. Any CloudWatch alarms, dashboards, or Contributor Insights rules pointing to &lt;code&gt;API-Gateway-Execution-Logs_{rest-api-id}/{stage_name}&lt;/code&gt; stop working. Migrate these references to your new log group before creating the delivery.&lt;/p&gt;
&lt;p&gt;Keep &lt;code&gt;loggingLevel&lt;/code&gt; at &lt;code&gt;INFO&lt;/code&gt; or &lt;code&gt;ERROR&lt;/code&gt;. Log delivery controls routing, not content. If &lt;code&gt;loggingLevel&lt;/code&gt; is &lt;code&gt;OFF&lt;/code&gt;, no execution log events are produced regardless of whether a delivery exists. Verify your method settings before troubleshooting missing logs.&lt;/p&gt;
&lt;p&gt;Treat the 1 MB log event capacity as a security decision, not only a debugging convenience. With &lt;code&gt;dataTraceEnabled&lt;/code&gt; set to true, execution logs include complete request and response payloads up to 1 MB. Those payloads might contain personally identifiable information (PII) or other sensitive data. Confirm your log destinations have appropriate access controls, encryption, and retention policies. Mask or filter sensitive fields in mapping templates upstream of logging and enable data tracing selectively per method or only in non-production stages.&lt;/p&gt;
&lt;p&gt;Start with a single destination, then expand. Validate that your log group or bucket receives events correctly before adding Firehose or additional destinations.&lt;/p&gt;
&lt;p&gt;Log delivery is best-effort. In rare cases, some log events might not be delivered. For audit-critical workloads, build retention and reconciliation that account for occasional missing events rather than treating execution logs as the system of record.&lt;/p&gt;
&lt;h2 id="cleaning-up"&gt;Cleaning up&lt;/h2&gt;
&lt;p&gt;To avoid ongoing charges from the resources you created while following this post, delete the delivery and then remove the destinations and any example S3 bucket or Data Firehose delivery stream you no longer need. Deleting the delivery returns the stage to standard auto-managed logging.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws logs delete-delivery --id &amp;lt;delivery-id&amp;gt;&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;When the delivery is deleted, CloudWatch Logs disables log delivery on the API Gateway stage automatically. The delivery source and delivery destination remain as independent objects. Delete them with &lt;code&gt;delete-delivery-source&lt;/code&gt; and &lt;code&gt;delete-delivery-destination&lt;/code&gt; if you do not plan to reuse them.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;CloudWatch Logs delivery for API Gateway REST API execution logs helps address the 1 KB event truncation and single managed destination constraints. You can now route full execution logs to CloudWatch Logs, Amazon S3, or Amazon Data Firehose, use multiple destinations from a single stage, and pay vended logs pricing.&lt;/p&gt;
&lt;p&gt;The feature works alongside existing method settings. No changes to your current logging configuration are required beyond creating the delivery itself.&lt;/p&gt;
&lt;p&gt;To get started, refer to &lt;a href="https://docs.aws.amazon.com/apigateway/latest/developerguide/rest-api-execution-logs-delivery.html" target="_blank" rel="noopener"&gt;Route execution logs with Amazon CloudWatch Logs delivery&lt;/a&gt; in the API Gateway documentation. For more about CloudWatch Logs delivery configuration, see &lt;a href="https://docs.aws.amazon.com/AmazonCloudWatch/latest/logs/AWS-logs-and-resource-policy.html" target="_blank" rel="noopener"&gt;Enable logging from AWS services&lt;/a&gt;. For pricing details, review the &lt;a href="https://aws.amazon.com/cloudwatch/pricing/" target="_blank" rel="noopener"&gt;Amazon CloudWatch pricing page&lt;/a&gt;. Try it on a test stage and share your experience in the comments.&lt;/p&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>Architecting SASE solutions using AWS Local Zones</title>
		<link>https://aws.amazon.com/blogs/compute/architecting-sase-solutions-using-aws-local-zones/</link>
		
		<dc:creator><![CDATA[Lakshmi VP]]></dc:creator>
		<pubDate>Wed, 09 Sep 2026 20:36:09 +0000</pubDate>
				<category><![CDATA[Advanced (300)]]></category>
		<category><![CDATA[AWS Local Zones]]></category>
		<category><![CDATA[Best Practices]]></category>
		<guid isPermaLink="false">02c8c6840ad8d0b9f7270e4215277d9d55dca65d</guid>

					<description>Organizations with geographically distributed workforces face a trade-off between security and low-latency access. This post explores how to use AWS Local Zones and Secure Access Service Edge (SASE) solutions to deploy virtual security appliances closer to end users, covering key design principles, capacity planning, and traffic routing.</description>
										<content:encoded>&lt;p&gt;Organizations with geographically distributed workforces face a critical challenge: providing secure, low-latency access to applications without routing all traffic through centralized data centers. Traditional hub-and-spoke network architectures create latency bottlenecks and degrade user experience, forcing a trade-off between security and performance.&lt;/p&gt;
&lt;p&gt;This post explores how you can use &lt;a href="https://aws.amazon.com/about-aws/global-infrastructure/localzones/" target="_blank" rel="noopener"&gt;AWS Local Zones&lt;/a&gt; and Secure Access Service Edge (SASE) solutions to eliminate that trade-off. You will learn key design principles, implementation strategies, and technical considerations for deploying SASE solutions at the edge. We’ve seen that understanding your user locations and traffic volumes up front helps you make effective design decisions.&lt;/p&gt;
&lt;h2 id="key-challenges-for-deploying-sase-solutions"&gt;Key challenges for deploying SASE solutions&lt;/h2&gt;
&lt;p&gt;SASE solutions require virtual security appliances such as firewalls, secure web gateways, and zero trust network access (ZTNA) connectors. You deploy these appliances close to end users so that traffic inspection does not add latency to the user experience. With AWS Local Zones, you can deploy these &lt;a href="https://aws.amazon.com/marketplace/solutions/security" target="_blank" rel="noopener"&gt;virtual security appliances from AWS Marketplace&lt;/a&gt; closer to end users.&lt;/p&gt;
&lt;p&gt;When you architect SASE solutions using Local Zones, you need to address several key technical challenges. &lt;strong&gt;Latency requirements:&lt;/strong&gt; When end users are far away from an AWS Region, applications requiring security inspection experience significant latency overhead that affects overall performance and user experience. &lt;strong&gt;Geographic coverage&lt;/strong&gt;: In some cases, workforces are spread across distributed locations far from an AWS Region. You need solutions that deliver consistent service quality and security capabilities to users across your covered locations.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Hybrid connectivity&lt;/strong&gt;: Many applications maintain dependencies on on-premises data centers in areas far away from an AWS Region. Design traffic routing carefully to avoid unnecessary network paths and reduce traffic hairpinning or network flapping. &lt;strong&gt;Security consistency&lt;/strong&gt;: Implement uniform security controls across all distributed locations while maintaining performance. This requires consideration of service placement and routing architecture.&lt;/p&gt;
&lt;p&gt;Before looking at the SASE-specific design, it helps to understand what Local Zones provide. The following diagram shows how Local Zones extend AWS infrastructure from the Region out to metropolitan areas closer to end users.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/09/compute-2397-figure-1.png" alt="High-level AWS infrastructure diagram showing how Local Zones bring compute closer to users" width="800"&gt;
 &lt;p class="wp-caption-text"&gt;Figure 1: High-level AWS infrastructure diagram showing how Local Zones bring compute closer to users&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;As the diagram shows, Local Zones place compute closer to end users. This especially benefits those far from an AWS Region.&lt;/p&gt;
&lt;h2 id="prerequisites"&gt;Prerequisites&lt;/h2&gt;
&lt;p&gt;To follow the guidance in this post, you should be familiar with:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;AWS Local Zones and how they extend AWS infrastructure to metro areas.&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://aws.amazon.com/ec2/" target="_blank" rel="noopener"&gt;Amazon Elastic Compute Cloud&lt;/a&gt; (Amazon EC2), &lt;a href="https://aws.amazon.com/vpc/" target="_blank" rel="noopener"&gt;Amazon Virtual Private Cloud&lt;/a&gt; (Amazon VPC), and core AWS networking concepts.&lt;/li&gt;
 &lt;li&gt;SASE architecture concepts and virtual security appliance deployment models.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="architecture-considerations"&gt;Architecture considerations&lt;/h2&gt;
&lt;p&gt;When you design SASE solutions with Local Zones, you can follow several key best practices across infrastructure, control plane, and traffic management.&lt;/p&gt;
&lt;h3 id="infrastructure-deployment"&gt;Infrastructure deployment&lt;/h3&gt;
&lt;p&gt;At the infrastructure level, focus on deploying virtual security appliances to optimize coverage and performance. Start by selecting and configuring Amazon EC2 instances optimized for maximum network throughput. Choose instance families that provide the compute and networking capabilities required for traffic inspection workloads, with enhanced networking enabled for high packets-per-second performance.&lt;/p&gt;
&lt;p&gt;Design a scalable cluster management strategy that adapts to varying workload demands while maintaining consistent security posture. As you deploy these clusters, establish proper multi-tenant isolation to maintain security boundaries between different organizational units, keeping user resources separate from management infrastructure.&lt;/p&gt;
&lt;h3 id="control-plane-architecture"&gt;Control plane architecture&lt;/h3&gt;
&lt;p&gt;The SASE control plane requires particular attention in distributed deployments. Deploy control components in an AWS Region to manage security appliances across all Local Zone locations. This provides a single point of policy distribution and configuration management. From this centralized vantage point, you can implement policy management that maintains consistency in security enforcement across all locations.&lt;/p&gt;
&lt;p&gt;Visibility matters as much as policy enforcement. Implement standardized telemetry collection mechanisms, such as &lt;a href="https://aws.amazon.com/cloudwatch/" target="_blank" rel="noopener"&gt;Amazon CloudWatch&lt;/a&gt; metrics and logs, across all locations so you can maintain observability and resolve issues proactively. As your deployment grows, automate configuration deployment using infrastructure as code (IaC) tools such as &lt;a href="https://aws.amazon.com/cloudformation/" target="_blank" rel="noopener"&gt;AWS CloudFormation&lt;/a&gt; or Terraform. This keeps deployment consistent across all edge locations and reduces manual errors when operating at scale.&lt;/p&gt;
&lt;h3 id="traffic-management"&gt;Traffic management&lt;/h3&gt;
&lt;p&gt;Traffic management completes the architecture of a well-designed SASE solution. Use &lt;a href="https://aws.amazon.com/route53/" target="_blank" rel="noopener"&gt;Amazon Route 53&lt;/a&gt; with geoproximity routing and health checks to direct users to the nearest security inspection point, minimizing inspection latency. If an appliance fails, Route 53 automatically reroutes traffic to the next-nearest Local Zone. For critical deployments, maintain standby capacity in the parent Region as a fallback.&lt;/p&gt;
&lt;p&gt;Deploy VPN endpoints in Local Zones closest to your user populations to reduce connection latency for remote users while maintaining high availability through health-checked failover across multiple locations. Plan your Internet Service Provider (ISP) connectivity for redundancy and performance requirements across different geographical locations, and implement geographic load-balancing mechanisms to distribute traffic efficiently across available resources.&lt;/p&gt;
&lt;p&gt;You also need to consider the egress path, which is how traffic exits after inspection. For internet-bound traffic, use the Local Zone’s direct internet egress to avoid routing back through the parent Region. For traffic destined to applications in an AWS Region, traffic traverses the AWS private network between the Local Zone and its parent Region. Validate egress paths using VPC Flow Logs and traceroute to confirm traffic is not taking unintended hops.&lt;/p&gt;
&lt;p&gt;The following diagram shows how the Local Zones architecture applies to a SASE use case, routing user traffic to a nearby Local Zone for inspection.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/08/ComputeBlog-2397-2.png" alt="Remote users and branch offices routing traffic to virtual network firewalls in the nearest Local Zone, with control nodes in the parent AWS Region" width="800"&gt;
 &lt;p class="wp-caption-text"&gt;Figure 2: Enterprise SASE deployment using virtual network firewalls across Local Zones to secure remote user and branch office access&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;As the diagram shows, remote users and branch offices connect to virtual network firewalls running in the Local Zone closest to them. Each Local Zone performs local traffic inspection that reduces latency for the SASE use case. The control nodes in the parent AWS Region manage policy and configuration across all locations.&lt;/p&gt;
&lt;h3 id="reference-implementation-approach"&gt;Reference implementation approach&lt;/h3&gt;
&lt;p&gt;This section outlines the key phases for implementing a SASE solution across AWS Local Zones, from initial planning through validation.&lt;/p&gt;
&lt;h4 id="phase-1-plan-your-deployment"&gt;Phase 1: Plan your deployment&lt;/h4&gt;
&lt;p&gt;Begin by mapping your user locations and latency expectations to identify which Local Zones are closest to your user populations, and determine which applications require local security inspection. With this map in hand, calculate capacity needs per location based on expected traffic volumes and security inspection requirements. Then define the specific inspection capabilities you need at each location, whether that is firewall, secure web gateway, ZTNA, or a combination.&lt;/p&gt;
&lt;p&gt;One key design decision at this stage is whether to route all user traffic through the Local Zone appliance (full tunnel) or only corporate-bound traffic (split tunnel). Full tunnel provides complete traffic visibility but requires higher instance throughput. You can validate your choice by using &lt;a href="https://docs.aws.amazon.com/vpc/latest/userguide/flow-logs.html" target="_blank" rel="noopener"&gt;VPC Flow Logs&lt;/a&gt; and CloudWatch network metrics to measure actual traffic volume per user during a pilot deployment.&lt;/p&gt;
&lt;h4 id="phase-2-configure-networking-infrastructure"&gt;Phase 2: Configure networking infrastructure&lt;/h4&gt;
&lt;p&gt;With your plan in place, enable the target Local Zones in your AWS account and create a VPC that extends into your chosen Local Zones by creating subnets in each one. Configure route tables to direct traffic through your virtual security appliances.&lt;/p&gt;
&lt;p&gt;Security at the network layer is critical. Set up security groups that permit the required traffic flows for your SASE inspection chain. Add inbound rules for user VPN connections (for example, UDP 4500/500 for IPsec), outbound rules to target applications, and management access from the parent Region. Add network ACLs as an additional layer of defense at the subnet level to restrict traffic to expected protocols and port ranges.&lt;/p&gt;
&lt;h4 id="phase-3-deploy-virtual-security-appliances"&gt;Phase 3: Deploy virtual security appliances&lt;/h4&gt;
&lt;p&gt;Launch your chosen virtual security appliance from AWS Marketplace in each target Local Zone. Use M6i or M6g instances, or newer instances optimized for network throughput. For example, m6i.xlarge provides up to 12.5 Gbps network bandwidth. Deploy scalable clusters of 2–20 instances depending on location traffic volume, and configure elastic network interfaces for traffic inspection with separate inbound and outbound interfaces.&lt;/p&gt;
&lt;p&gt;Enable enhanced networking and verify that the instance supports the throughput required for your expected traffic volume. This validation step is critical before moving to production, because undersized instances can become bottlenecks that negate the latency benefits of Local Zone placement.&lt;/p&gt;
&lt;h4 id="phase-4-configure-the-control-plane"&gt;Phase 4: Configure the control plane&lt;/h4&gt;
&lt;p&gt;Deploy your centralized SASE management components in the parent AWS Region and establish connectivity between the regional management infrastructure and your Local Zone appliances. Push security policies from the central management console to all distributed appliances to maintain consistent enforcement.&lt;/p&gt;
&lt;p&gt;For observability, configure centralized logging and telemetry collection using Amazon CloudWatch. Enable VPC Flow Logs on Local Zone subnets to capture traffic metadata for compliance auditing and security analysis. Use this data for troubleshooting and demonstrating regulatory compliance.&lt;/p&gt;
&lt;h4 id="phase-5-set-up-traffic-routing"&gt;Phase 5: Set up traffic routing&lt;/h4&gt;
&lt;p&gt;Configure Amazon Route 53 with geoproximity routing policies to direct users to the nearest Local Zone. Set up health checks that automatically fail over if a Local Zone appliance becomes unhealthy. Deploy VPN endpoints in each Local Zone for remote user connectivity.&lt;/p&gt;
&lt;p&gt;After your routing is configured, test end-to-end connectivity and verify that traffic routes through the nearest security inspection point. This confirms that your geoproximity policies work as intended and that users receive the expected latency benefits.&lt;/p&gt;
&lt;h4 id="phase-6-validate-and-optimize"&gt;Phase 6: Validate and optimize&lt;/h4&gt;
&lt;p&gt;With your deployment live, verify latency improvements by comparing round-trip times to the parent AWS Region and to the Local Zones. Monitor appliance utilization metrics (CPU, network throughput, and concurrent sessions) in Amazon CloudWatch, and adjust cluster sizes at each location based on observed traffic patterns. Validate that security policies are applied consistently across all locations.&lt;/p&gt;
&lt;p&gt;Configure CloudWatch alarms to trigger scaling actions. For example, scale out when average CPU exceeds 70% or network throughput exceeds 80% of instance capacity over a 5-minute period. Use CloudWatch anomaly detection to identify unusual traffic patterns that might indicate a misconfigured routing policy or a security event.&lt;/p&gt;
&lt;h2 id="capacity-planning"&gt;Capacity planning&lt;/h2&gt;
&lt;p&gt;Local Zones provide the same elasticity as AWS Regions to scale your virtual security appliances based on demand. To optimize your deployment:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;Use &lt;a href="https://aws.amazon.com/ec2/autoscaling/" target="_blank" rel="noopener"&gt;Amazon EC2 Auto Scaling&lt;/a&gt; to automatically adjust the number of appliance instances based on traffic patterns and utilization metrics.&lt;/li&gt;
 &lt;li&gt;Create &lt;a href="https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ec2-capacity-reservations.html" target="_blank" rel="noopener"&gt;On-Demand Capacity Reservations&lt;/a&gt; to support applications that must provide guaranteed availability at all times.&lt;/li&gt;
 &lt;li&gt;Design your architecture to work across multiple instance families, giving you flexibility to use the most suitable compute resources available at each location.&lt;/li&gt;
 &lt;li&gt;For cost optimization, consider using &lt;a href="https://aws.amazon.com/savingsplans/compute-pricing/" target="_blank" rel="noopener"&gt;Compute and EC2 Instance Savings Plans&lt;/a&gt; for steady-state appliance instances that run continuously, while relying on On-Demand pricing for burst capacity during peak traffic periods.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Before production deployment, validate that your chosen virtual appliance functions correctly in the target Local Zone and test network dependencies to confirm expected performance.&lt;/p&gt;
&lt;h2 id="clean-up"&gt;Clean up&lt;/h2&gt;
&lt;p&gt;If you deploy resources following this guidance and no longer need them after your testing, terminate EC2 instances, release Elastic IP addresses, delete Capacity Reservations, and remove associated networking resources (subnets, route tables, security groups, Route 53 policies) to avoid ongoing charges.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;This post explored how you can deploy SASE solutions on AWS Local Zones. Local Zones bring three key benefits to SASE architectures. They reduce security inspection latency by placing appliances closer to users, apply consistent security enforcement across geographically distributed locations, and eliminate the need to backhaul traffic to centralized data centers. Organizations continue to expand their operations to more geographic locations. The combination of AWS Local Zones and SASE solutions from partners such as Palo Alto Networks provides a scalable approach for delivering secure connectivity to users anywhere.&lt;/p&gt;
&lt;h3 id="learn-more"&gt;Learn more&lt;/h3&gt;
&lt;p&gt;For instructions to opt in to a Local Zone and launch your Amazon EC2 instance, see the &lt;a href="https://docs.aws.amazon.com/local-zones/latest/ug/getting-started.html" target="_blank" rel="noopener"&gt;AWS Local Zones Getting started&lt;/a&gt; page. To learn where AWS Local Zones are available globally, check out the &lt;a href="https://aws.amazon.com/about-aws/global-infrastructure/localzones/locations/" target="_blank" rel="noopener"&gt;AWS Local Zones locations&lt;/a&gt; page.&lt;/p&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>Announcing 90-minute function timeout on AWS Lambda Managed Instances</title>
		<link>https://aws.amazon.com/blogs/compute/announcing-90-minute-function-timeout-on-aws-lambda-managed-instances/</link>
		
		<dc:creator><![CDATA[Tarun Rai Madan]]></dc:creator>
		<pubDate>Wed, 09 Sep 2026 19:08:14 +0000</pubDate>
				<category><![CDATA[Announcements]]></category>
		<category><![CDATA[AWS Lambda]]></category>
		<category><![CDATA[Foundational (100)]]></category>
		<guid isPermaLink="false">4cf63fdfce0783e4e432dbac88fece3b77e1837a</guid>

					<description>AWS Lambda now supports a 90-minute function timeout for asynchronous and event source mapping (ESM) invocations on Lambda Managed Instances, a 6x increase from the previous 15-minute limit. Data processing, media transcoding, AI inference, and batch workloads can now run on Lambda without re-architecting.</description>
										<content:encoded>&lt;p&gt;AWS Lambda now supports a 90-minute function timeout for &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/invocation-async.html" target="_blank" rel="noopener"&gt;asynchronous&lt;/a&gt; and &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/invocation-eventsourcemapping.html" target="_blank" rel="noopener"&gt;event source mapping (ESM)&lt;/a&gt; invocations on AWS Lambda Managed Instances (LMI), a capability of AWS Lambda. This is a 6x increase from the previous 15-minute limit. Customers running data processing, media transcoding, financial calculations, AI inference, and batch workloads can now use Lambda functions for jobs that require longer continuous execution, without re-architecting their applications. This also applies to invocations within &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/durable-functions.html" target="_blank" rel="noopener"&gt;Lambda durable functions&lt;/a&gt;, which use checkpoints to track progress and automatically recover from failures through replay, skipping completed work. When invoked asynchronously, a multi-step durable execution can run for up to 1 year.&lt;/p&gt;
&lt;h2 id="evolution-of-function-timeout-on-lambda"&gt;Evolution of function timeout on Lambda&lt;/h2&gt;
&lt;p&gt;Lambda’s function timeout has increased over time, from 5 minutes at launch in 2014 to 15 minutes in 2018. As customers sought to use the simplicity of Lambda for data-intensive workloads, the 15-minute timeout limit forced architectural tradeoffs for applications where customers needed longer continuous execution time. Several patterns emerged:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;
  &lt;p&gt;&lt;strong&gt;Media processing&lt;/strong&gt;: Speech-to-text transcription and video transcoding that routinely need longer than 15 minutes of continuous execution.&lt;/p&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;&lt;strong&gt;Financial calculations&lt;/strong&gt;: Monte Carlo simulations, bond pricing, and portfolio risk analysis that are memory-intensive and often require longer than 15 minutes.&lt;/p&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;&lt;strong&gt;Data processing and ETL pipelines&lt;/strong&gt;: Batch jobs processing multi-gigabyte datasets or aggregating data from external sources that exceed 15 minutes during peak volumes.&lt;/p&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;&lt;strong&gt;AI inference&lt;/strong&gt;: Model testing and inference jobs (for example, reasoning tasks) that fit Lambda’s memory and CPU profile but exceed its timeout.&lt;/p&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;&lt;strong&gt;Web scraping and file transfer&lt;/strong&gt;: Crawling external sites or pulling large file sets from vendors that exceed 15 minutes when sources respond slowly.&lt;/p&gt;
 &lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In each case, customers preferred Lambda’s simplicity but had to re-architect when jobs hit the 15-minute limit.&lt;/p&gt;
&lt;p&gt;Fast forward to 2026, Lambda supports two form factors: functions (event-driven, 15-minute timeout), and &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/lambda-microvms-guide.html" target="_blank" rel="noopener"&gt;MicroVMs&lt;/a&gt; for user or AI-generated just-in-time code (HTTP-driven, 8-hour duration). To allow customers to benefit from the simplicity of serverless compute with the flexibility and pricing model of EC2 for steady-state workloads, we extended the on-demand capacity mode of Lambda to add &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/lambda-managed-instances.html" target="_blank" rel="noopener"&gt;Lambda Managed Instances&lt;/a&gt; (LMI). With Lambda Managed Instances, you can process multiple concurrent requests per instance, access specialized compute configurations, and drive cost efficiency through EC2 pricing advantages, without managing infrastructure.&lt;/p&gt;
&lt;p&gt;We also support &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/durable-functions.html" target="_blank" rel="noopener"&gt;Lambda durable functions&lt;/a&gt; (powered by the &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/durable-execution-sdk.html" target="_blank" rel="noopener"&gt;durable execution SDK&lt;/a&gt;) on both on-demand and LMI capacity modes. Durable functions provide application-level checkpointing and workflow-as-code: your code saves execution state using &lt;code&gt;step()&lt;/code&gt; and &lt;code&gt;wait()&lt;/code&gt; operations, and gracefully recovers from infrastructure failures by resuming from the last checkpoint rather than restarting from scratch. While a durable execution (the complete lifecycle of a durable function) can run for up to a year, each invocation was still limited to 15 minutes.&lt;/p&gt;
&lt;p&gt;As customers onboard more workloads to serverless compute to benefit from its simplicity, they need longer continuous execution for data-intensive use cases like AI inference, media transcoding, scientific modeling, and financial calculations that do not fit Lambda’s 15-minute duration constraints. Today, we are extending the function timeout on Lambda Managed Instances to 90 minutes for asynchronous and ESM invocations. This includes invocations within a durable function, where a multi-step application can continue to run for up to 1 year when invoked asynchronously.&lt;/p&gt;
&lt;h2 id="activating-90-minute-function-timeout"&gt;Activating 90-minute function timeout&lt;/h2&gt;
&lt;p&gt;You can now configure any Lambda function running on a Managed Instance with a timeout of up to 90 minutes (5,400 seconds) for async and ESM invocations. Synchronous invocations retain the existing 15-minute maximum. The function executes exactly as before: same runtime, same handler, same IAM execution role, same virtual private cloud (VPC) configuration. The only difference is that your function now supports longer continuous execution. Your initialization code (&lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/lambda-managed-instances-execution-environment.html" target="_blank" rel="noopener"&gt;Init phase&lt;/a&gt;) is still limited to 15 minutes on Lambda Managed Instances.&lt;/p&gt;
&lt;p&gt;To set the timeout, update your function configuration using the AWS CLI:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws lambda update-function-configuration \
    --function-name my-data-processor \
    --timeout 5400 \
    --region us-east-1&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Or in AWS CloudFormation / AWS Serverless Application Model (AWS SAM):&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-yaml"&gt;MyFunction:
  Type: AWS::Serverless::Function
  Properties:
    FunctionName: my-data-processor
    Runtime: python3.12
    Handler: app.handler
    Timeout: 5400
    MemorySize: 10240&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;You do not need to change any code. You can also update the timeout from the Lambda console under Configuration &amp;gt; General Configuration (&lt;strong&gt;Figure 1&lt;/strong&gt;), or configure it through natural language prompts in your AI coding assistants (like Claude Code or Kiro) by installing the &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/agent-setup-guide.html" target="_blank" rel="noopener"&gt;Agent Toolkit for AWS&lt;/a&gt;.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/03/ComputeBlog-2727-1.png" alt="Lambda console General configuration page showing the function timeout field set to 90 minutes" width="800"&gt;
 &lt;p class="wp-caption-text"&gt;Figure 1: Configuring Lambda function timeout&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;The change takes effect on subsequent invocations after the function timeout is updated. For event source mappings, allow a few minutes for the new configuration to propagate. Your existing observability setup continues to work as expected: Amazon CloudWatch metrics, AWS CloudTrail, and AWS X-Ray capture the full invocation lifecycle without any changes. For details, see &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/lambda-monitoring.html" target="_blank" rel="noopener"&gt;monitoring Lambda functions&lt;/a&gt; and &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/durable-monitoring.html" target="_blank" rel="noopener"&gt;monitoring durable functions&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;There is no additional charge for using the 90-minute timeout. Standard &lt;a href="https://aws.amazon.com/lambda/pricing/" target="_blank" rel="noopener"&gt;Lambda Managed Instances pricing&lt;/a&gt; applies.&lt;/p&gt;
&lt;h2 id="minute-timeout-and-durable-functions"&gt;90-minute timeout and durable functions&lt;/h2&gt;
&lt;p&gt;The 90-minute function timeout and durable functions are complementary. The &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/configuration-timeout.html" target="_blank" rel="noopener"&gt;function timeout&lt;/a&gt; (&lt;code&gt;--timeout&lt;/code&gt;) controls how long each individual invocation can run, while the &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/durable-configuration.html" target="_blank" rel="noopener"&gt;durable execution timeout&lt;/a&gt; (&lt;code&gt;ExecutionTimeout&lt;/code&gt; in &lt;code&gt;--durable-config&lt;/code&gt;) controls the total elapsed time from execution start to completion. Durable functions use checkpoints to track progress and automatically recover from failures through replay, re-executing from the beginning while skipping completed work. With today’s launch, each asynchronous invocation in a durable function running on a Managed Instance can now execute for up to 90 minutes continuously, while the corresponding durable execution can run for up to 1 year. For synchronous and &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/durable-invoking-esm.html" target="_blank" rel="noopener"&gt;event source mapping invocations&lt;/a&gt;, both the invocation and the corresponding durable execution are limited to 90 minutes.&lt;/p&gt;
&lt;p&gt;For idempotent jobs (for example, an ETL pipeline step triggered by SQS), the extended timeout alone might be sufficient. If the host fails, the message returns to the queue and a fresh invocation starts. For jobs where re-execution is expensive (for example, a 40-minute inference run already 30 minutes in), combine both. Enable durable functions to checkpoint periodically, so a failure at minute 35 resumes from the last checkpoint rather than restarting from zero.&lt;/p&gt;
&lt;h2 id="invocation-behavior-asynchronous-event-source-mappings-and-synchronous"&gt;Invocation behavior: asynchronous, event source mappings, and synchronous&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Asynchronous invocations (up to 90 minutes)&lt;/strong&gt;: If the function fails or times out, Lambda applies your configured retry policy (up to two retries by default) and routes failed events to your dead-letter queue or on-failure destination.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Event source mappings (up to 90 minutes)&lt;/strong&gt;: For SQS, configure your queue’s visibility timeout to be &lt;a href="https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-configure-lambda-function-trigger.html" target="_blank" rel="noopener"&gt;at least six times&lt;/a&gt; the function timeout. This gives Lambda enough time to retry if a function is throttled while processing a previous batch. Lambda validates this at event source mapping creation time, but does not prevent subsequent changes to queue or function settings that might create a mismatch.&lt;/p&gt;
&lt;p&gt;For Amazon Kinesis and Amazon DynamoDB Streams, configure the &lt;a href="https://docs.aws.amazon.com/lambda/latest/api/API_CreateEventSourceMapping.html#lambda-CreateEventSourceMapping-request-MaximumBatchingWindowInSeconds" target="_blank" rel="noopener"&gt;maximum batching window&lt;/a&gt; and &lt;a href="https://docs.aws.amazon.com/lambda/latest/api/API_CreateEventSourceMapping.html#lambda-CreateEventSourceMapping-request-ParallelizationFactor" target="_blank" rel="noopener"&gt;parallelization factor&lt;/a&gt; to account for longer processing times per batch.&lt;/p&gt;
&lt;p&gt;If your batch contains multiple records and you want to avoid re-processing the entire batch when one record fails, enable &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/services-kinesis-batchfailurereporting.html" target="_blank" rel="noopener"&gt;partial batch failure reporting&lt;/a&gt;. This is available for SQS, Kinesis, DynamoDB Streams, Amazon Managed Streaming for Apache Kafka (Amazon MSK), and self-managed Apache Kafka event source mappings. With partial batch failures enabled, only the failed records are retried, not the entire batch.&lt;/p&gt;
&lt;p&gt;Note that invocations for &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/with-mq.html" target="_blank" rel="noopener"&gt;Amazon MQ ESM&lt;/a&gt; and &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/with-documentdb.html" target="_blank" rel="noopener"&gt;Amazon DocumentDB (with MongoDB compatibility) ESM&lt;/a&gt; remain limited to 15 minutes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Synchronous invocations (15 minutes maximum)&lt;/strong&gt;: &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/invocation-sync.html" target="_blank" rel="noopener"&gt;Synchronous invocations&lt;/a&gt; retain the existing 15-minute maximum timeout. If you set your function timeout to greater than 15 minutes and invoke it synchronously, Lambda continues to apply the 15-minute timeout. The &lt;a href="https://docs.aws.amazon.com/cli/latest/reference/lambda/get-function-configuration.html" target="_blank" rel="noopener"&gt;GetFunctionConfiguration API&lt;/a&gt; reports the configured timeout value.&lt;/p&gt;
&lt;p&gt;To see which event sources invoke Lambda functions synchronously or asynchronously, refer to &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/lambda-services.html" target="_blank" rel="noopener"&gt;Lambda documentation&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="considerations-and-best-practices"&gt;Considerations and best practices&lt;/h2&gt;
&lt;p&gt;Because your functions now support longer continuous execution, consider these best practices for components that might be ephemeral in nature, such as network connections and credentials.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Networking&lt;/strong&gt;: Make sure idle connection timeouts on downstream services (RDS, Amazon ElastiCache, external APIs) accommodate the full function duration. If your function routes traffic through a NAT Gateway, send keep-alive packets to prevent idle connections from being dropped (350-second idle timeout). Respect DNS TTL values for external hostname resolution. The AWS SDK handles this automatically, but custom HTTP clients might cache DNS records beyond their TTL.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Credentials&lt;/strong&gt;: If your function acquires temporary credentials or tokens, verify they remain valid for the full execution duration or refresh them in the background.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Idempotency&lt;/strong&gt;: Lambda does not guarantee exactly-once processing. With longer-running functions, the window for retries and duplicate deliveries increases. You can use &lt;a href="https://docs.aws.amazon.com/powertools/python/latest/utilities/idempotency/" target="_blank" rel="noopener"&gt;Powertools for AWS Lambda&lt;/a&gt; to implement idempotency in your function code so that operations like payments or database writes produce the same result even if executed more than once. If you use Lambda durable functions, steps have at-least-once execution semantics by default. The SDK skips completed steps during replay, but steps that fail before checkpointing may re-execute. You can use &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/durable-execution-idempotency.html" target="_blank" rel="noopener"&gt;execution names&lt;/a&gt; as idempotency keys for durable functions.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;The 90-minute function timeout on Lambda Managed Instances addresses one of the most common customer needs for building data-intensive applications on AWS Lambda. Data processing, media transcoding, AI inference, and financial computation workloads that exceed 15 minutes can now run on Lambda without code changes or architectural workarounds. We look forward to hearing from you if you need a longer timeout for synchronous invocations, or for the on-demand capacity mode, on our &lt;a href="https://github.com/aws/aws-lambda-roadmap" target="_blank" rel="noopener"&gt;AWS Lambda Roadmap GitHub page&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;To get started, update your function’s timeout configuration and deploy. For a step-by-step walkthrough, see &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/lambda-managed-instances.html" target="_blank" rel="noopener"&gt;Getting started with Lambda Managed Instances&lt;/a&gt;. For sample code demonstrating long-running functions with durable checkpointing, see &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/durable-examples.html" target="_blank" rel="noopener"&gt;durable functions examples&lt;/a&gt;. To learn more about AWS Lambda, visit &lt;a href="https://aws.amazon.com/lambda" target="_blank" rel="noopener"&gt;aws.amazon.com/lambda&lt;/a&gt;.&lt;/p&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>Bring your own client certificate for backend mTLS in Amazon API Gateway</title>
		<link>https://aws.amazon.com/blogs/compute/bring-your-own-client-certificate-for-backend-mtls-in-amazon-api-gateway/</link>
		
		<dc:creator><![CDATA[Biswanath Mukherjee]]></dc:creator>
		<pubDate>Tue, 08 Sep 2026 21:46:39 +0000</pubDate>
				<category><![CDATA[Amazon API Gateway]]></category>
		<category><![CDATA[Announcements]]></category>
		<category><![CDATA[Intermediate (200)]]></category>
		<guid isPermaLink="false">aefa41ad59da632a3771c6831f7d2b0ebe00b502</guid>

					<description>Enterprises that use Amazon API Gateway often want to bring their own client certificate for backend mutual TLS (mTLS) authentication. With API Gateway, you can now use a third-party or AWS Private CA-issued client certificate for the outbound mTLS handshake. In this post, you build a REST API with an outbound mTLS connection to an Amazon ECS backend.</description>
										<content:encoded>&lt;p&gt;Enterprises that use &lt;a href="https://aws.amazon.com/api-gateway/" target="_blank" rel="noopener"&gt;Amazon API Gateway&lt;/a&gt; in front of internal or partner backends often want to bring their own client certificate for backend mutual TLS (mTLS) authentication. During mTLS, the backend presents its own server certificate and also requests the caller to present a client certificate to validate it against a trusted certificate authority (CA). Until now, you could use only an API Gateway-generated, self-signed SSL certificate for the outbound connection, because there was no CA behind it for the backend to trust. Backends that enforce a specific corporate or partner CA reject that self-signed certificate, and the mutual TLS handshake fails. Bringing your own CA-signed certificate is necessary for scenarios such as migrating APIs off legacy gateways or meeting your internal PKI mandates that require certificates from an approved CA.&lt;/p&gt;
&lt;p&gt;With API Gateway, you can now bring your own client certificate for backend mutual TLS (mTLS) authentication. You can either use a third-party certificate or a certificate issued by &lt;a href="https://aws.amazon.com/private-ca/" target="_blank" rel="noopener"&gt;AWS Private Certificate Authority&lt;/a&gt;. If you’re using a third-party certificate, you must &lt;a href="https://docs.aws.amazon.com/acm/latest/userguide/import-certificate.html" target="_blank" rel="noopener"&gt;import&lt;/a&gt; the certificate in &lt;a href="https://aws.amazon.com/certificate-manager/" target="_blank" rel="noopener"&gt;AWS Certificate Manager (ACM)&lt;/a&gt;. Then you &lt;a href="https://docs.aws.amazon.com/apigateway/latest/developerguide/rest-api-acm-client-certificates.html" target="_blank" rel="noopener"&gt;configure the ACM certificate ARN in your REST API stage&lt;/a&gt;. API Gateway presents that certificate during the backend mTLS handshake.&lt;/p&gt;
&lt;h1&gt;Solution overview&lt;/h1&gt;
&lt;p&gt;In this post, you build a REST API with an outbound mTLS connection using this newly launched API Gateway feature. This solution demonstrates an outbound mTLS connection between Amazon API Gateway and a backend application running on &lt;a href="https://docs.aws.amazon.com/AmazonECS/latest/developerguide/AWS_Fargate.html" target="_blank" rel="noopener"&gt;Amazon Elastic Container Service (Amazon ECS)&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The following diagram shows the solution architecture.&lt;/p&gt;
&lt;figure&gt;
 &lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/08/ComputeBlog-2657-1.png" alt="Architecture diagram showing API Gateway presenting an ACM client certificate to a Network Load Balancer that forwards traffic to an NGINX sidecar and validator app on Amazon ECS Fargate, with certificates issued by AWS Private CA through ACM" width="800"&gt;
 &lt;figcaption aria-hidden="true"&gt;Architecture diagram showing API Gateway presenting an ACM client certificate to a Network Load Balancer that forwards traffic to an NGINX sidecar and validator app on Amazon ECS Fargate, with certificates issued by AWS Private CA through ACM&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The solution uses:&lt;/p&gt;
&lt;ol type="a"&gt;
 &lt;li&gt;
  &lt;p&gt;&lt;a href="https://aws.amazon.com/private-ca/" target="_blank" rel="noopener"&gt;AWS Private Certificate Authority&lt;/a&gt; with a root-subordinate CA hierarchy to issue both the client and server certificates through &lt;a href="https://aws.amazon.com/certificate-manager/" target="_blank" rel="noopener"&gt;AWS Certificate Manager (ACM)&lt;/a&gt;.&lt;/p&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;&lt;a href="https://aws.amazon.com/api-gateway/" target="_blank" rel="noopener"&gt;Amazon API Gateway&lt;/a&gt; REST API stage configured with the ACM client certificate ARN (ClientCertificateId), so that the API Gateway presents the certificate during the outbound TLS handshake.&lt;/p&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;&lt;a href="https://aws.amazon.com/ecs/" target="_blank" rel="noopener"&gt;Amazon ECS&lt;/a&gt; on &lt;a href="https://aws.amazon.com/fargate/" target="_blank" rel="noopener"&gt;AWS Fargate&lt;/a&gt; running an NGINX sidecar that holds the server certificate and validates the incoming client certificate against a CA bundle (root and subordinate chain).&lt;/p&gt;
 &lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;A request goes through the following steps:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;
  &lt;p&gt;Client application invokes the REST API exposed by API Gateway. The API Gateway stage is configured with an ACM client certificate ARN.&lt;/p&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;API Gateway opens an outbound connection to the &lt;a href="https://aws.amazon.com/elasticloadbalancing/network-load-balancer/" target="_blank" rel="noopener"&gt;Network Load Balancer&lt;/a&gt; (NLB) to begin the TLS handshake. The API Gateway presents an ACM client certificate configured at the stage level when the backend requests one.&lt;/p&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;The NLB listens for the incoming TCP request on port 443 and forwards the call to Amazon ECS Fargate. The NLB acts as a passthrough and does not terminate the TLS connection.&lt;/p&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;The NGINX sidecar container running on Amazon ECS performs the inbound mTLS handshake:&lt;/p&gt;
 &lt;/li&gt;
&lt;/ol&gt;
&lt;ul&gt;
 &lt;li&gt;
  &lt;p&gt;NGINX presents the backend server certificate and verifies the client certificate against a mounted CA bundle (root and subordinate chain).&lt;/p&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;After verification, NGINX forwards the request and parsed certificate details to the validator app container over local HTTP.&lt;/p&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;The validator app re-checks the certificate validity window, matches the common name against an allowlist, and returns a structured JSON response.&lt;/p&gt;
 &lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; The NGINX sidecar is not mandatory for this flow. It demonstrates separation of concerns: NGINX handles the mTLS handshake, and the validator app contains the business logic.&lt;/p&gt;
&lt;h1&gt;Prerequisites for demo&lt;/h1&gt;
&lt;p&gt;To follow along, you need the following:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;
  &lt;p&gt;&lt;a href="https://portal.aws.amazon.com/gp/aws/developer/registration/index.html" target="_blank" rel="noopener"&gt;Create an AWS account&lt;/a&gt; if you do not already have one.&lt;/p&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;Access to an AWS account through the AWS Management Console and the &lt;a href="https://aws.amazon.com/cli" target="_blank" rel="noopener"&gt;AWS Command Line Interface (AWS CLI)&lt;/a&gt;. The &lt;a href="https://aws.amazon.com/iam" target="_blank" rel="noopener"&gt;AWS Identity and Access Management (IAM)&lt;/a&gt; principal you use must have permissions to make the necessary AWS service calls and manage the resources in this post. Follow the &lt;a href="https://docs.aws.amazon.com/IAM/latest/UserGuide/best-practices.html" target="_blank" rel="noopener"&gt;principle of least privilege&lt;/a&gt;.&lt;/p&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;&lt;a href="https://git-scm.com/book/en/v2/Getting-Started-Installing-Git" target="_blank" rel="noopener"&gt;Git installed&lt;/a&gt;.&lt;/p&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;&lt;a href="https://docs.aws.amazon.com/serverless-application-model/latest/developerguide/serverless-sam-cli-install.html" target="_blank" rel="noopener"&gt;AWS Serverless Application Model (AWS SAM) CLI installed&lt;/a&gt;.&lt;/p&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;&lt;a href="https://www.docker.com/" target="_blank" rel="noopener"&gt;Docker&lt;/a&gt; installed and running, to build and push the two container images.&lt;/p&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;Python 3.14 installed.&lt;/p&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;jq command line tools installed.&lt;/p&gt;
 &lt;/li&gt;
&lt;/ul&gt;
&lt;h1&gt;Environment setup&lt;/h1&gt;
&lt;p&gt;Run the following commands to set up the demo environment:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;
  &lt;p&gt;Create a new folder and clone the GitHub repository:&lt;/p&gt;
  &lt;div class="hide-language"&gt;
   &lt;pre&gt;&lt;code class="language-bash"&gt;git clone https://github.com/aws-samples/sample-api-backend-mtls
cd sample-api-backend-mtls&lt;/code&gt;&lt;/pre&gt;
  &lt;/div&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;Set the environment variables after replacing the placeholders:&lt;/p&gt;
  &lt;div class="hide-language"&gt;
   &lt;pre&gt;&lt;code class="language-bash"&gt;ACCOUNT_ID=$(aws sts get-caller-identity --query Account --output text)
REGION=&amp;lt;Your AWS Region, for example, us-east-1&amp;gt;
STACK_NAME=&amp;lt;Your stack name e.g. outbound-mtls-backend&amp;gt;&lt;/code&gt;&lt;/pre&gt;
  &lt;/div&gt;
 &lt;/li&gt;
&lt;/ol&gt;
&lt;h1&gt;Build the container images&lt;/h1&gt;
&lt;p&gt;Run the following commands to create container images of NGINX sidecar container and the validator app containers:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;
  &lt;p&gt;Create two Amazon Elastic Container Registry (Amazon ECR) repositories, one for NGINX and another for validator app containers respectively:&lt;/p&gt;
  &lt;div class="hide-language"&gt;
   &lt;pre&gt;&lt;code class="language-bash"&gt;NGINX_REPO_URI=$(aws ecr create-repository \
  --repository-name $STACK_NAME-nginx-sidecar \
  --image-tag-mutability IMMUTABLE \
  --image-scanning-configuration scanOnPush=true \
  --region "$REGION" \
  --query "repository.repositoryUri" --output text)

VALIDATOR_REPO_URI=$(aws ecr create-repository \
  --repository-name $STACK_NAME-validator-app \
  --image-tag-mutability IMMUTABLE \
  --image-scanning-configuration scanOnPush=true \
  --region "$REGION" \
  --query "repository.repositoryUri" --output text)

aws ecr get-login-password --region "$REGION" \
  | docker login --username AWS --password-stdin "${ACCOUNT_ID}.dkr.ecr.${REGION}.amazonaws.com"&lt;/code&gt;&lt;/pre&gt;
  &lt;/div&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;Build and push the NGINX and validator app containers:&lt;/p&gt;
  &lt;div class="hide-language"&gt;
   &lt;pre&gt;&lt;code class="language-bash"&gt;docker build --platform linux/amd64 -t $STACK_NAME-nginx-sidecar nginx/
docker tag $STACK_NAME-nginx-sidecar:latest "${NGINX_REPO_URI}:latest"
docker push "${NGINX_REPO_URI}:latest"
docker build --platform linux/amd64 -t $STACK_NAME-validator-app validator_app/
docker tag $STACK_NAME-validator-app:latest "${VALIDATOR_REPO_URI}:latest"
docker push "${VALIDATOR_REPO_URI}:latest"&lt;/code&gt;&lt;/pre&gt;
  &lt;/div&gt;
 &lt;/li&gt;
&lt;/ol&gt;
&lt;h1&gt;Deploy and test the solution&lt;/h1&gt;
&lt;p&gt;You first deploy the stack without the client certificate configured in the API Gateway and perform negative testing. The mTLS handshake will fail because of a missing client certificate in the request. Then you update the stack to configure client certificate in API Gateway stage and retest mTLS.&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;
  &lt;p&gt;Run the following command to build and deploy the overall stack without client certificate configured at API Gateway stage:&lt;/p&gt;
  &lt;div class="hide-language"&gt;
   &lt;pre&gt;&lt;code class="language-bash"&gt;sam build
sam deploy \
  --stack-name $STACK_NAME \
  --resolve-s3 \
  --capabilities CAPABILITY_IAM \
  --region "$REGION" \
  --parameter-overrides \
  NginxRepositoryUri="$NGINX_REPO_URI" \
  ValidatorRepositoryUri="$VALIDATOR_REPO_URI" \
  EnableOutboundMtls=false&lt;/code&gt;&lt;/pre&gt;
  &lt;/div&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;Wait for the task to reach &lt;code&gt;RUNNING&lt;/code&gt; and pass its target group health check:&lt;/p&gt;
  &lt;div class="hide-language"&gt;
   &lt;pre&gt;&lt;code class="language-bash"&gt;EcsClusterName=$(aws cloudformation describe-stacks \
  --stack-name $STACK_NAME --region "$REGION" \
  --query "Stacks[0].Outputs[?OutputKey=='EcsClusterName'].OutputValue" \
  --output text)

TargetGroupArn=$(aws cloudformation describe-stacks \
  --stack-name $STACK_NAME --region "$REGION" \
  --query "Stacks[0].Outputs[?OutputKey=='TargetGroupArn'].OutputValue" \
  --output text)

aws ecs list-tasks --cluster "$EcsClusterName" --region "$REGION"

aws elbv2 describe-target-health \
  --target-group-arn "$TargetGroupArn" --region "$REGION"&lt;/code&gt;&lt;/pre&gt;
  &lt;/div&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;Capture the front API invoke URL from the stack outputs:&lt;/p&gt;
  &lt;div class="hide-language"&gt;
   &lt;pre&gt;&lt;code class="language-bash"&gt;FRONT_API_URL=$(aws cloudformation describe-stacks \
  --stack-name $STACK_NAME --region "$REGION" \
  --query "Stacks[0].Outputs[?OutputKey=='FrontApiUrl'].OutputValue" \
  --output text)

NLB_DNS_NAME=$(aws cloudformation describe-stacks \
  --stack-name $STACK_NAME --region "$REGION" \
  --query "Stacks[0].Outputs[?OutputKey=='NlbDnsName'].OutputValue" \
  --output text)

FRONT_CLIENT_CERT_ARN=$(aws cloudformation describe-stacks \
  --stack-name $STACK_NAME --region "$REGION" \
  --query "Stacks[0].Outputs[?OutputKey=='FrontClientCertArn'].OutputValue" \
  --output text)

FRONT_API_ID=$(aws cloudformation describe-stacks \
  --stack-name $STACK_NAME --region "$REGION" \
  --query "Stacks[0].Outputs[?OutputKey=='FrontApiId'].OutputValue" \
  --output text)&lt;/code&gt;&lt;/pre&gt;
  &lt;/div&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;Wait a minute or two after the stack finishes, then invoke the front API:&lt;/p&gt;
  &lt;div class="hide-language"&gt;
   &lt;pre&gt;&lt;code class="language-bash"&gt;curl -v "$FRONT_API_URL"&lt;/code&gt;&lt;/pre&gt;
  &lt;/div&gt;
  &lt;p&gt;The following is the NGINX configuration for mTLS:&lt;/p&gt;
  &lt;div class="hide-language"&gt;
   &lt;pre&gt;&lt;code class="language-nginx"&gt;...
server {
    listen 443 ssl;
    # Server identity (issued by the private CA).
    ssl_certificate /etc/nginx/certs/server.crt;
    ssl_certificate_key /etc/nginx/certs/server.key;
    # Inbound mutual TLS: require and validate the client certificate
    # against the CA bundle (root + subordinate CA chain).
    ssl_client_certificate /etc/nginx/certs/ca_bundle.pem;
    ssl_verify_client on;
    ssl_verify_depth 2;
    ssl_protocols TLSv1.2;
...}&lt;/code&gt;&lt;/pre&gt;
  &lt;/div&gt;
  &lt;p&gt;The &lt;code&gt;curl&lt;/code&gt; command returns &lt;code&gt;HTTP/2 400&lt;/code&gt;, with a response body containing &lt;code&gt;400 No required SSL certificate was sent&lt;/code&gt;. Because the API Gateway is not presenting a client certificate on the outbound handshake, the NGINX sidecar container in Amazon ECS rejects the mTLS connection. The following screenshot shows the response:&lt;/p&gt;
  &lt;figure&gt;
   &lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/08/ComputeBlog-2657-2.png" alt="Terminal response showing an HTTP/2 400 error with the message No required SSL certificate was sent" width="800"&gt;
   &lt;figcaption aria-hidden="true"&gt;Terminal response showing an HTTP/2 400 error with the message No required SSL certificate was sent&lt;/figcaption&gt;
  &lt;/figure&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;Now redeploy the solution with outbound mTLS enabled:&lt;/p&gt;
  &lt;div class="hide-language"&gt;
   &lt;pre&gt;&lt;code class="language-bash"&gt;sam deploy \
  --stack-name $STACK_NAME \
  --resolve-s3 \
  --capabilities CAPABILITY_IAM \
  --region "$REGION" \
  --parameter-overrides \
  NginxRepositoryUri="$NGINX_REPO_URI" \
  ValidatorRepositoryUri="$VALIDATOR_REPO_URI" \
  EnableOutboundMtls=true&lt;/code&gt;&lt;/pre&gt;
  &lt;/div&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;Wait a minute or two after the stack finishes, then invoke the front API again:&lt;/p&gt;
  &lt;div class="hide-language"&gt;
   &lt;pre&gt;&lt;code class="language-bash"&gt;curl -v "$FRONT_API_URL"&lt;/code&gt;&lt;/pre&gt;
  &lt;/div&gt;
  &lt;p&gt;Because the client certificate is now presented during the mTLS handshake, the handshake completes successfully, as shown in the following response snippet:&lt;/p&gt;
  &lt;figure&gt;
   &lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/08/ComputeBlog-2657-3.png" alt="Terminal response showing a successful mTLS handshake and an HTTP 200 response from the backend" width="800"&gt;
   &lt;figcaption aria-hidden="true"&gt;Terminal response showing a successful mTLS handshake and an HTTP 200 response from the backend&lt;/figcaption&gt;
  &lt;/figure&gt;
 &lt;/li&gt;
&lt;/ol&gt;
&lt;h1&gt;Automatic certificate renewal&lt;/h1&gt;
&lt;p&gt;When a certificate changes in ACM, API Gateway detects the update and propagates the new certificate automatically. You do not redeploy the stage, and the API experiences no downtime during rotation. Certificate propagation is eventually consistent. During an update, the backend might briefly receive either the old or the new certificate. ACM also emits certificate expiration notifications through Amazon EventBridge, which you can use to set alarms before a certificate expires.&lt;/p&gt;
&lt;h1&gt;Clean up&lt;/h1&gt;
&lt;p&gt;If you followed along only for demonstration purposes, to avoid incurring future charges, run the following commands to delete the resources created in this demo:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;
  &lt;p&gt;Clean up the S3 buckets:&lt;/p&gt;
  &lt;div class="hide-language"&gt;
   &lt;pre&gt;&lt;code class="language-bash"&gt;NLB_LOGS_BUCKET=$(aws cloudformation describe-stacks \
  --stack-name $STACK_NAME --region $REGION \
  --query "Stacks[0].Outputs[?OutputKey=='NlbAccessLogsBucketName'].OutputValue" \
  --output text)

aws s3api list-object-versions --bucket "$NLB_LOGS_BUCKET" \
  --query '{Objects: Versions[].{Key:Key,VersionId:VersionId}}' \
  --output json | \
  jq -c '.Objects[]? // empty' | \
  while read -r obj; do
    aws s3api delete-object --bucket "$NLB_LOGS_BUCKET" \
      --key "$(echo "$obj" | jq -r .Key)" \
      --version-id "$(echo "$obj" | jq -r .VersionId)" \
      --region $REGION
  done

aws s3api list-object-versions --bucket "$NLB_LOGS_BUCKET" \
  --query '{Objects: DeleteMarkers[].{Key:Key,VersionId:VersionId}}' \
  --output json | \
  jq -c '.Objects[]? // empty' | \
  while read -r obj; do
    aws s3api delete-object --bucket "$NLB_LOGS_BUCKET" \
      --key "$(echo "$obj" | jq -r .Key)" \
      --version-id "$(echo "$obj" | jq -r .VersionId)" \
      --region $REGION
  done&lt;/code&gt;&lt;/pre&gt;
  &lt;/div&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;Delete the stack:&lt;/p&gt;
  &lt;div class="hide-language"&gt;
   &lt;pre&gt;&lt;code class="language-bash"&gt;sam delete --stack-name $STACK_NAME --region $REGION --no-prompts&lt;/code&gt;&lt;/pre&gt;
  &lt;/div&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;Delete the ECR repository:&lt;/p&gt;
  &lt;div class="hide-language"&gt;
   &lt;pre&gt;&lt;code class="language-bash"&gt;aws ecr delete-repository --repository-name $STACK_NAME-nginx-sidecar --force --region $REGION
aws ecr delete-repository --repository-name $STACK_NAME-validator-app --force --region $REGION&lt;/code&gt;&lt;/pre&gt;
  &lt;/div&gt;
 &lt;/li&gt;
&lt;/ol&gt;
&lt;h1&gt;Conclusion&lt;/h1&gt;
&lt;p&gt;In this post, you configured a REST API with an outbound mTLS connection using Amazon API Gateway and an ECS Fargate backend. With this new feature launch in API Gateway, you can now bring your own client certificate for outbound mTLS handshake for your REST APIs. You can now meet your internal PKI mandates to authenticate backends that pin a specific certificate issuer.&lt;/p&gt;
&lt;p&gt;To get started, import a certificate from your own PKI into ACM and configure your API Gateway REST API stage for outbound mTLS authentication. For more information, see &lt;a href="https://docs.aws.amazon.com/apigateway/latest/developerguide/rest-api-backend-authentication.html" target="_blank" rel="noopener"&gt;Present client certificates to backend services with mutual TLS in API Gateway&lt;/a&gt;. If you have feedback about this post, leave it in the comments section. For technical questions, you can start a thread on &lt;a href="https://repost.aws/" target="_blank" rel="noopener"&gt;AWS re:Post&lt;/a&gt;.&lt;/p&gt;
&lt;h1&gt;Further reading&lt;/h1&gt;
&lt;ul&gt;
 &lt;li&gt;
  &lt;p&gt;&lt;a href="https://docs.aws.amazon.com/apigateway/latest/developerguide/rest-api-backend-authentication.html" target="_blank" rel="noopener"&gt;Present client certificates to backend services with mutual TLS in API Gateway&lt;/a&gt;&lt;/p&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;&lt;a href="https://docs.aws.amazon.com/apigateway/latest/developerguide/getting-started-client-side-ssl-authentication.html" target="_blank" rel="noopener"&gt;Amazon API Gateway REST API client certificates&lt;/a&gt;&lt;/p&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;&lt;a href="https://docs.aws.amazon.com/apigateway/latest/developerguide/api-gateway-extensions-integration-tls-config.html" target="_blank" rel="noopener"&gt;Integration tlsConfig reference&lt;/a&gt;&lt;/p&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;&lt;a href="https://docs.aws.amazon.com/acm/latest/userguide/import-certificate.html" target="_blank" rel="noopener"&gt;Importing certificates into AWS Certificate Manager&lt;/a&gt;&lt;/p&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;&lt;a href="https://docs.aws.amazon.com/acm/latest/userguide/gs-acm-request-private.html" target="_blank" rel="noopener"&gt;Requesting a private certificate with AWS Private CA&lt;/a&gt;&lt;/p&gt;
 &lt;/li&gt;
 &lt;li&gt;
  &lt;p&gt;&lt;a href="https://aws.amazon.com/blogs/compute/automating-mutual-tls-setup-for-amazon-api-gateway/" target="_blank" rel="noopener"&gt;Automating mutual TLS setup for Amazon API Gateway&lt;/a&gt;&lt;/p&gt;
 &lt;/li&gt;
&lt;/ul&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>Scheduling email campaigns at scale with Amazon EventBridge Scheduler</title>
		<link>https://aws.amazon.com/blogs/compute/scheduling-email-campaigns-at-scale-with-amazon-eventbridge-scheduler/</link>
		
		<dc:creator><![CDATA[Oluwaseun Ademuwagun]]></dc:creator>
		<pubDate>Tue, 01 Sep 2026 13:09:17 +0000</pubDate>
				<category><![CDATA[Advanced (300)]]></category>
		<category><![CDATA[Amazon EventBridge]]></category>
		<category><![CDATA[Technical How-to]]></category>
		<guid isPermaLink="false">3d9219b43bd0b9bef7ea2475c46079b1c0370b4f</guid>

					<description>Learn how to use Amazon EventBridge Scheduler to deliver email campaigns at per-recipient optimal send times. This post shows how to create one schedule per recipient with zero idle compute cost, scale schedule creation with AWS Step Functions Distributed Map, and deliver through Amazon SES.</description>
										<content:encoded>&lt;p&gt;Scheduling email campaigns becomes more complex when you need to send email to millions of recipients at the unique time best suited for each customer. Consider these examples:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;A flash sale might need to hit inboxes at 9 AM local time across every time zone.&lt;/li&gt;
 &lt;li&gt;A follow-up email (often known as a drip sequence) might need to send a second message exactly 3 days after the first message per subscriber.&lt;/li&gt;
 &lt;li&gt;A re-engagement campaign might target users who haven’t logged in for 30 days.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The scheduling requirements involve multiple considerations. You’re sending hundreds of millions of messages, each at its own optimal moment personalized to the recipient’s time zone and behavior.&lt;/p&gt;
&lt;p&gt;In this post, we walk through how to use &lt;a href="https://aws.amazon.com/eventbridge/scheduler/" target="_blank" rel="noopener"&gt;Amazon EventBridge Scheduler&lt;/a&gt; to personalize email notifications to each recipient. We create one schedule per recipient to deliver each email at its individually optimal moment, with zero idle compute cost. We also show how Amazon EventBridge Scheduler handles higher volumes. Amazon EventBridge Scheduler supports billions of schedules. By default, you have a quota of &lt;a href="https://docs.aws.amazon.com/scheduler/latest/UserGuide/scheduler-quotas.html" target="_blank" rel="noopener"&gt;10 million schedules&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="solution-overview"&gt;Solution overview&lt;/h2&gt;
&lt;p&gt;When every recipient has their own ideal delivery time, you need a scheduling layer that can hold billions of individual send intents and fire each one at the right moment. Most teams reach for one of three familiar patterns, each with tradeoffs that become painful at scale.&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;&lt;strong&gt;Batch cron jobs&lt;/strong&gt;: A job runs every hour, queries for all messages due in the next window, and sends them out. Recipients get email in imprecise hourly batches. At scale, the batch job itself becomes a bottleneck, processing millions of rows per run, competing for database connections, and creating a sudden spike in load on the email provider.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Delay queues&lt;/strong&gt;: You can use Amazon Simple Queue Service (Amazon SQS) as a delay queue. A delay queue postpones the delivery of new messages to a customer for a set time. A limitation of this approach is that Amazon SQS caps delays at 15 minutes.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Third-party campaign tools&lt;/strong&gt;: Offload to a SaaS email platform. This works until you need tight integration with your application data, custom send-time optimization, or control over delivery infrastructure. You’re also paying per-recipient fees that compound at scale.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;All three approaches either sacrifice precision (batching), hit architectural limits (delay queues), or surrender control (third-party tools).&lt;/p&gt;
&lt;h3 id="the-building-block-approach"&gt;The building block approach&lt;/h3&gt;
&lt;p&gt;Amazon EventBridge Scheduler treats each email send as a discrete scheduled action. Instead of “process all messages due this hour,” you express the intent directly: “send this email to this person at this time.” Amazon EventBridge Scheduler holds that intent with zero compute cost until the moment arrives, then triggers the scheduled action. See the &lt;a href="https://docs.aws.amazon.com/scheduler/latest/UserGuide/what-is-scheduler.html" target="_blank" rel="noopener"&gt;Amazon EventBridge Scheduler User Guide&lt;/a&gt; for the full API reference and current service quotas.&lt;/p&gt;
&lt;p&gt;For email campaigns, Amazon EventBridge Scheduler becomes the &lt;strong&gt;send-time dispatcher&lt;/strong&gt;, the component that schedules every email in a campaign for its individually optimal moment, whether that’s timezone-adjusted, behavior-triggered, or sequence-driven.&lt;/p&gt;
&lt;h2 id="architecture-diagram"&gt;Architecture diagram&lt;/h2&gt;
&lt;p&gt;The architecture follows an event-driven, per-recipient scheduling pattern for an email campaign. To start the campaign, you first define the target audience and the content they receive. Next, you need a way to create the per-recipient schedule. To do that for a campaign that can contain millions of recipients, you need a scalable mechanism to create the schedules. You can achieve this with an &lt;a href="https://aws.amazon.com/step-functions/" target="_blank" rel="noopener"&gt;AWS Step Functions&lt;/a&gt; state machine, a serverless workflow service that coordinates multiple AWS services into structured, visual workflows called state machines. In this solution, we orchestrate the creation of the schedules by using a &lt;a href="https://docs.aws.amazon.com/step-functions/latest/dg/state-map-distributed.html" target="_blank" rel="noopener"&gt;Distributed Map&lt;/a&gt; state within the state machine, which lets us fan out and accelerate schedule creation. It does this by splitting a large dataset into chunks and processing them across thousands of parallel child executions. It reads the recipient list from Amazon Simple Storage Service (Amazon S3), applies time zone logic per recipient, and creates an individual Amazon EventBridge Scheduler resource for each recipient in parallel. After the workflow creates all schedules, the execution completes.&lt;/p&gt;
&lt;p&gt;The actual email delivery happens later, entirely decoupled from the campaign creation step. At the scheduled time, Amazon EventBridge Scheduler invokes &lt;a href="https://aws.amazon.com/ses/" target="_blank" rel="noopener"&gt;Amazon Simple Email Service&lt;/a&gt; (Amazon SES) directly, passing the template name and personalization data as template variables. For campaigns requiring complex personalization logic (conditional content, real-time suppression checks, or data enrichment), you can optionally route through an AWS Lambda function before SES. If you need to adjust timing or content for specific recipients, you can update their individual schedules directly without reprocessing the entire campaign.&lt;/p&gt;
&lt;p style="text-align: center"&gt;&lt;a href="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/01/compute-2689-arch.png"&gt;&lt;img loading="lazy" class="alignnone wp-image-26863 size-full" src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/09/01/compute-2689-arch.png" alt="" width="781" height="592"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Figure 1: Per-recipient email scheduling architecture with Amazon EventBridge Scheduler&lt;/em&gt;&lt;/p&gt;
&lt;h2 id="walkthrough"&gt;Walkthrough&lt;/h2&gt;
&lt;p&gt;The solution uses four core components that work together: a campaign manager to define send-time rules, Step Functions Distributed Map to fan out and accelerate schedule creation, Amazon EventBridge Scheduler to hold each per-recipient intent and deliver through Amazon SES directly, and automatic cleanup through schedule self-deletion.&lt;/p&gt;
&lt;h3 id="how-it-works"&gt;How it works&lt;/h3&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;&lt;strong&gt;Create the campaign&lt;/strong&gt;: A marketer defines the campaign: audience segment, email template, and send-time rules (for example, “9 AM in each recipient’s local time zone” or “24 hours before a Black Friday sale”).&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Campaign manager fans out&lt;/strong&gt;: An &lt;a href="https://aws.amazon.com/step-functions/" target="_blank" rel="noopener"&gt;AWS Step Functions&lt;/a&gt; workflow uses Distributed Map to iterate over the recipient list and create one Amazon EventBridge Scheduler schedule per recipient per campaign step directly through SDK integration. Each schedule encodes the exact send time for that individual.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Amazon EventBridge Scheduler fires at the right moment&lt;/strong&gt;: At each recipient’s scheduled time, Amazon EventBridge Scheduler invokes Amazon SES directly through a universal target, passing the template name and personalization data (recipient name and attributes) as template variables.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;SES personalizes and sends&lt;/strong&gt;: Amazon SES renders the email template with the provided data and delivers the message.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Schedule self-deletes&lt;/strong&gt;: &lt;code&gt;ActionAfterCompletion='DELETE'&lt;/code&gt; prevents the accumulation of spent schedules.&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id="prerequisites"&gt;Prerequisites&lt;/h3&gt;
&lt;p&gt;To follow along with this walkthrough, you need the following:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;AWS account and permissions&lt;/strong&gt;: An active AWS account with permissions to create Amazon EventBridge Scheduler schedules, AWS Step Functions state machines, and Amazon SES identities, along with an AWS Identity and Access Management (IAM) role for Amazon EventBridge Scheduler to invoke Amazon SES.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Development environment&lt;/strong&gt;: Python 3.13 or later, AWS SDK for Python (Boto3) version 1.26 or later, and AWS Command Line Interface v2 (&lt;a href="https://aws.amazon.com/cli/" target="_blank" rel="noopener"&gt;AWS CLI&lt;/a&gt; v2).&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Amazon SES configuration&lt;/strong&gt;: Move your Amazon SES account &lt;a href="https://docs.aws.amazon.com/ses/latest/dg/request-production-access.html" target="_blank" rel="noopener"&gt;out of sandbox mode&lt;/a&gt; to allow sending to arbitrary recipients.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="scaling-the-fan-out-with-step-functions"&gt;Scaling the fan-out with Step Functions&lt;/h3&gt;
&lt;p&gt;For campaigns with millions of recipients, use AWS Step Functions Distributed Map to parallelize schedule creation. When you want to activate a campaign, you trigger a Step Functions workflow. This workflow fans out and creates schedules across the recipient list by using a Distributed Map with direct SDK integration. The direct SDK integration between Step Functions and Amazon EventBridge Scheduler lets each child execution call &lt;code&gt;CreateSchedule&lt;/code&gt; directly. The following state machine definition reads recipients from an Amazon S3 CSV file and creates schedules in parallel:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-json"&gt;{
  "Comment": "Fan out campaign schedule creation via direct SDK integration",
  "StartAt": "EnsureScheduleGroup",
  "States": {
    "EnsureScheduleGroup": {
      "Type": "Task",
      "Resource": "arn:aws:states:::aws-sdk:scheduler:createScheduleGroup",
      "Parameters": {
        "Name.$": "States.Format('campaign-{}', $.campaign_id)"
      },
      "ResultPath": null,
      "Catch": [
        {
          "ErrorEquals": [
            "Scheduler.ConflictException"
          ],
          "ResultPath": null,
          "Next": "FanOutRecipients"
        }
      ],
      "Next": "FanOutRecipients"
    },
    "FanOutRecipients": {
      "Type": "Map",
      "ItemProcessor": {
        "ProcessorConfig": {
          "Mode": "DISTRIBUTED",
          "ExecutionType": "STANDARD"
        },
        "StartAt": "BuildScheduleInput",
        "States": {
          "BuildScheduleInput": {
            "Type": "Pass",
            "Parameters": {
              "schedule_name.$": "States.Format('campaign-{}-{}', $.campaign_id, $.recipient.id)",
              "group_name.$": "States.Format('campaign-{}', $.campaign_id)",
              "schedule_expression.$": "States.Format('at({}T{}:00:00)', $.send_date_date, $.send_hour)",
              "timezone.$": "$.recipient.timezone",
              "target_input": {
                "FromEmailAddress": "campaigns@example.com",
                "Destination": {
                  "ToAddresses.$": "States.Array($.recipient.email)"
                },
                "Content": {
                  "Template": {
                    "TemplateName.$": "$.template_id",
                    "TemplateData.$": "States.JsonToString($.recipient.attributes)"
                  }
                }
              }
            },
            "Next": "CreateSchedule"
          },
          "CreateSchedule": {
            "Type": "Task",
            "Resource": "arn:aws:states:::aws-sdk:scheduler:createSchedule",
            "Retry": [
              {
                "ErrorEquals": [
                  "Scheduler.SdkClientException"
                ],
                "IntervalSeconds": 2,
                "MaxAttempts": 3,
                "BackoffRate": 2
              }
            ],
            "Parameters": {
              "Name.$": "$.schedule_name",
              "GroupName.$": "$.group_name",
              "ScheduleExpression.$": "$.schedule_expression",
              "ScheduleExpressionTimezone.$": "$.timezone",
              "FlexibleTimeWindow": {
                "Mode": "FLEXIBLE",
                "MaximumWindowInMinutes": 5
              },
              "Target": {
                "Arn": "arn:aws:scheduler:::aws-sdk:sesv2:sendEmail",
                "RoleArn": "arn:aws:iam::976764934189:role/CampaignFanOutRole-dev",
                "Input.$": "States.JsonToString($.target_input)",
                "RetryPolicy": {
                  "MaximumEventAgeInSeconds": 7200,
                  "MaximumRetryAttempts": 5
                }
              },
              "ActionAfterCompletion": "DELETE"
            },
            "ResultPath": null,
            "End": true
          }
        }
      },
      "ItemReader": {
        "Resource": "arn:aws:states:::s3:getObject",
        "ReaderConfig": {
          "InputType": "CSV",
          "CSVHeaderLocation": "FIRST_ROW"
        },
        "Parameters": {
          "Bucket.$": "$$.Execution.Input.recipient_bucket",
          "Key.$": "$$.Execution.Input.recipient_key"
        }
      },
      "ItemSelector": {
        "campaign_id.$": "$$.Execution.Input.campaign_id",
        "template_id.$": "$$.Execution.Input.template_id",
        "send_date_date.$": "$$.Execution.Input.send_date_date",
        "send_hour.$": "$$.Execution.Input.send_hour",
        "recipient": {
          "id.$": "$$.Map.Item.Value.id",
          "email.$": "$$.Map.Item.Value.email",
          "timezone.$": "$$.Map.Item.Value.timezone",
          "attributes": {
            "first_name.$": "$$.Map.Item.Value.first_name",
            "signup_date.$": "$$.Map.Item.Value.signup_date"
          }
        }
      },
      "MaxConcurrency": 1000,
      "ResultPath": null,
      "End": true
    }
  }
}&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h3 id="concurrency-alignment-with-amazon-eventbridge-scheduler-api-limits"&gt;Concurrency alignment with Amazon EventBridge Scheduler API limits&lt;/h3&gt;
&lt;p&gt;Step Functions Distributed Map supports up to 10,000 concurrent child workflows. Each child calls the &lt;a href="https://docs.aws.amazon.com/scheduler/latest/APIReference/API_CreateSchedule.html" target="_blank" rel="noopener"&gt;CreateSchedule API&lt;/a&gt; directly, which has a &lt;a href="https://docs.aws.amazon.com/scheduler/latest/UserGuide/scheduler-quotas.html" target="_blank" rel="noopener"&gt;default rate limit of 5,000 TPS&lt;/a&gt;. This limit is sufficient for most campaigns. If your campaign volumes require higher throughput, check your current quotas in the &lt;a href="https://docs.aws.amazon.com/servicequotas/latest/userguide/intro.html" target="_blank" rel="noopener"&gt;Service Quotas&lt;/a&gt; console and request an increase.&lt;/p&gt;
&lt;p&gt;To avoid throttling, set &lt;code&gt;MaxConcurrency&lt;/code&gt; below the &lt;code&gt;CreateSchedule&lt;/code&gt; TPS quota. A value of 2,500 provides a comfortable buffer to account for bursts and retries without requiring a quota change. For larger campaigns, request an increase through &lt;a href="https://docs.aws.amazon.com/servicequotas/latest/userguide/request-quota-increase.html" target="_blank" rel="noopener"&gt;AWS Service Quotas&lt;/a&gt; (adjustable to tens of thousands) and raise &lt;code&gt;MaxConcurrency&lt;/code&gt; to match.&lt;/p&gt;
&lt;h3 id="canceling-a-campaign"&gt;Canceling a campaign&lt;/h3&gt;
&lt;p&gt;A schedule group is an Amazon EventBridge Scheduler resource used to organize schedules. For this use case, we have a schedule group per campaign. If you need to pull a campaign (error in content, legal issue, or strategy change), you can cancel all scheduled sends for that campaign by deleting the entire schedule group. The following code shows how to cancel all pending sends for a campaign:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-python"&gt;def cancel_campaign(campaign_id):
    """Cancel all pending sends for a campaign by deleting its schedule group."""
    scheduler.delete_schedule_group(
        Name=f'campaign-{campaign_id}'
    )&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h2 id="operational-considerations"&gt;Operational considerations&lt;/h2&gt;
&lt;p&gt;Moving to production introduces a few scaling and reliability concerns to plan for.&lt;/p&gt;
&lt;h3 id="handling-invocation-spikes-at-delivery-time"&gt;Handling invocation spikes at delivery time&lt;/h3&gt;
&lt;p&gt;When a mass campaign schedules millions of messages for the same time, this creates cascading pressure across two limits:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;Amazon EventBridge Scheduler invocations throttle limit&lt;/strong&gt;: The default is 1,000 TPS per AWS Region, and it is adjustable to tens of thousands of TPS through &lt;a href="https://docs.aws.amazon.com/servicequotas/latest/userguide/request-quota-increase.html" target="_blank" rel="noopener"&gt;AWS Service Quotas&lt;/a&gt;. Amazon EventBridge Scheduler queues invocations internally and retries with exponential backoff when the downstream target throttles.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Amazon SES sending quotas&lt;/strong&gt;: Your SES account has a per-second sending rate. If the effective invocation rate exceeds this, messages fail with throttling errors. Align &lt;a href="https://docs.aws.amazon.com/ses/latest/dg/quotas.html" target="_blank" rel="noopener"&gt;Amazon SES sending quotas&lt;/a&gt; with your campaign volume. Check your current SES quota in the Service Quotas console and request an increase before launching large campaigns. See &lt;a href="https://docs.aws.amazon.com/ses/latest/dg/best-practices.html" target="_blank" rel="noopener"&gt;Amazon SES best practices&lt;/a&gt; for deliverability at scale.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;To handle an invocation spike, we recommend using the &lt;strong&gt;FlexibleTimeWindow&lt;/strong&gt; feature of Amazon EventBridge Scheduler. Setting &lt;code&gt;MaximumWindowInMinutes&lt;/code&gt; lets Amazon EventBridge Scheduler spread invocations across a time window rather than firing them all at the exact second. Size the window based on your campaign: divide the total schedules by your effective TPS to determine the minimum spread needed. For example, 500,000 schedules at 5,000 TPS need at least a 2-minute window.&lt;/p&gt;
&lt;h3 id="cost-model"&gt;Cost model&lt;/h3&gt;
&lt;p&gt;You pay for Amazon EventBridge Scheduler on a per-invocation basis.&lt;/p&gt;
&lt;h2 id="cleanup"&gt;Cleanup&lt;/h2&gt;
&lt;p&gt;To avoid ongoing charges, delete the resources created during this walkthrough:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;&lt;strong&gt;Delete any runtime-created schedule groups&lt;/strong&gt;.
  &lt;div class="hide-language"&gt;
   &lt;pre&gt;&lt;code class="language-bash"&gt;aws scheduler delete-schedule-group --name campaign-&amp;lt;campaign-id&amp;gt;&lt;/code&gt;&lt;/pre&gt;
  &lt;/div&gt;&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Delete the Step Functions state machine&lt;/strong&gt;.
  &lt;div class="hide-language"&gt;
   &lt;pre&gt;&lt;code class="language-bash"&gt;aws stepfunctions delete-state-machine \
    --state-machine-arn arn:aws:states:us-east-1:&amp;lt;account-id&amp;gt;:stateMachine:CampaignFanOut&lt;/code&gt;&lt;/pre&gt;
  &lt;/div&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; If you have active schedules still waiting to fire, deleting the schedule group will cancel all pending sends.&lt;/p&gt;
&lt;h3 id="iam-role-for-amazon-eventbridge-scheduler-and-step-functions"&gt;IAM role for Amazon EventBridge Scheduler and Step Functions&lt;/h3&gt;
&lt;p&gt;The Step Functions state machine needs an execution role with permissions to create schedules, send email, and pass the role to the Amazon EventBridge Scheduler service. Amazon EventBridge Scheduler needs permissions to call SES. The following policy shows the combined permissions for both scenarios:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-json"&gt;{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "AllowPassRoleToScheduler",
      "Effect": "Allow",
      "Action": "iam:PassRole",
      "Resource": "arn:aws:iam::&amp;lt;ACCOUNT_ID&amp;gt;:role/CampaignFanOutRole",
      "Condition": {
        "StringEquals": {
          "iam:PassedToService": "scheduler.amazonaws.com"
        }
      }
    },
    {
      "Sid": "AllowSESSend",
      "Effect": "Allow",
      "Action": [
        "ses:SendEmail",
        "ses:SendTemplatedEmail"
      ],
      "Resource": "arn:aws:ses:&amp;lt;REGION&amp;gt;:&amp;lt;ACCOUNT_ID&amp;gt;:identity/campaigns@example.com"
    },
    {
      "Sid": "DistributedMapExecution",
      "Effect": "Allow",
      "Action": [
        "states:StartExecution",
        "states:DescribeExecution",
        "states:StopExecution"
      ],
      "Resource": [
        "arn:aws:states:&amp;lt;REGION&amp;gt;:&amp;lt;ACCOUNT_ID&amp;gt;:stateMachine:CampaignFanOut",
        "arn:aws:states:&amp;lt;REGION&amp;gt;:&amp;lt;ACCOUNT_ID&amp;gt;:execution:CampaignFanOut:*"
      ]
    },
    {
      "Sid": "ReadRecipientsBucket",
      "Effect": "Allow",
      "Action": [
        "s3:GetObject",
        "s3:ListBucket"
      ],
      "Resource": [
        "arn:aws:s3:::campaign-recipients-&amp;lt;ACCOUNT_ID&amp;gt;",
        "arn:aws:s3:::campaign-recipients-&amp;lt;ACCOUNT_ID&amp;gt;/*"
      ]
    },
    {
      "Sid": "CreateSchedules",
      "Effect": "Allow",
      "Action": "scheduler:CreateSchedule",
      "Resource": "arn:aws:scheduler:&amp;lt;REGION&amp;gt;:&amp;lt;ACCOUNT_ID&amp;gt;:schedule/campaign-*"
    },
    {
      "Sid": "CreateScheduleGroups",
      "Effect": "Allow",
      "Action": "scheduler:CreateScheduleGroup",
      "Resource": "arn:aws:scheduler:&amp;lt;REGION&amp;gt;:&amp;lt;ACCOUNT_ID&amp;gt;:schedule-group/campaign-*"
    },
    {
      "Sid": "PassRoleToScheduler",
      "Effect": "Allow",
      "Action": "iam:PassRole",
      "Resource": "arn:aws:iam::&amp;lt;ACCOUNT_ID&amp;gt;:role/SchedulerCampaignRole",
      "Condition": {
        "StringEquals": {
          "iam:PassedToService": "scheduler.amazonaws.com"
        }
      }
    }
  ]
}&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;This policy scopes the &lt;code&gt;scheduler:CreateSchedule&lt;/code&gt; and &lt;code&gt;scheduler:CreateScheduleGroup&lt;/code&gt; actions to resources prefixed with &lt;code&gt;campaign-*&lt;/code&gt;, following least-privilege principles.&lt;/p&gt;
&lt;p&gt;A condition restricts the &lt;code&gt;iam:PassRole&lt;/code&gt; permission so that it can only pass the role to the Amazon EventBridge Scheduler service.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;In this post, we walked through how to use Amazon EventBridge Scheduler to personalize email campaign delivery for each recipient. An email campaign system has two core problems: deciding what to send and deciding when to send it. Most teams over-engineer the “when” with polling infrastructure, batch jobs, and queue chains. Amazon EventBridge Scheduler collapses that into a single &lt;code&gt;CreateSchedule&lt;/code&gt; API call per recipient.&lt;/p&gt;
&lt;p&gt;To get started, &lt;a href="https://aws.amazon.com/eventbridge/scheduler/" target="_blank" rel="noopener"&gt;explore Amazon EventBridge Scheduler&lt;/a&gt; on the AWS Management Console. Browse &lt;a href="https://serverlessland.com/patterns?services=eventbridge-scheduler" target="_blank" rel="noopener"&gt;Serverless Land patterns&lt;/a&gt; for more than 20 Amazon EventBridge Scheduler patterns and other use cases beyond email campaigns.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Suggested tags:&lt;/strong&gt; &lt;a href="https://aws.amazon.com/blogs/mt/tag/amazon-eventbridge/" target="_blank" rel="noopener"&gt;Amazon EventBridge&lt;/a&gt;, &lt;a href="https://aws.amazon.com/blogs/mt/tag/architecture/" target="_blank" rel="noopener"&gt;architecture&lt;/a&gt;, &lt;a href="https://aws.amazon.com/blogs/mt/tag/events/" target="_blank" rel="noopener"&gt;events&lt;/a&gt;, &lt;a href="https://aws.amazon.com/blogs/mt/tag/modernization/" target="_blank" rel="noopener"&gt;modernization&lt;/a&gt;, &lt;a href="https://aws.amazon.com/blogs/mt/tag/serverless/" target="_blank" rel="noopener"&gt;serverless&lt;/a&gt;.&lt;/p&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>Amazon Linux default SSM parameter will now track the latest kernel</title>
		<link>https://aws.amazon.com/blogs/compute/amazon-linux-default-ssm-parameter-will-now-track-the-latest-kernel/</link>
		
		<dc:creator><![CDATA[Gokul Govindaraju]]></dc:creator>
		<pubDate>Thu, 20 Aug 2026 19:49:05 +0000</pubDate>
				<category><![CDATA[Advanced (300)]]></category>
		<category><![CDATA[Amazon EC2]]></category>
		<category><![CDATA[Announcements]]></category>
		<guid isPermaLink="false">f4a0de2b754995bf12880194f6d31b231978f174</guid>

					<description>The Amazon Linux kernel-default SSM parameter now updates to point to the latest kernel version as new releases become available. This post explains what this means for your workloads and how to manage the transition.</description>
										<content:encoded>&lt;p&gt;Today we are announcing that the Amazon Linux &lt;code&gt;kernel-default&lt;/code&gt; &lt;a href="https://aws.amazon.com/systems-manager/" target="_blank" rel="noopener"&gt;AWS Systems Manager (SSM)&lt;/a&gt; parameter will now update to point to the latest Amazon Linux kernel version as new kernel versions get released. On August 17, 2026, for &lt;a href="https://aws.amazon.com/linux/amazon-linux-2023/" target="_blank" rel="noopener"&gt;Amazon Linux 2023 (AL2023),&lt;/a&gt; the SSM parameter was updated from kernel 6.1 to kernel 6.18. As new kernel versions get released (expected annually), the parameter will continue to update to the latest kernel version after a validation period.&lt;/p&gt;
&lt;p&gt;This post explains the default kernel behavior, what it means for your workloads, and how to manage the transition.&lt;/p&gt;
&lt;h2 id="whats-changing"&gt;What’s changing?&lt;/h2&gt;
&lt;p&gt;Amazon Linux ships multiple kernel versions and has tracked a &lt;em&gt;default&lt;/em&gt; kernel for each OS version. For example,
 &lt;br&gt;
 the AL2023 parameter:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;ssm:/aws/service/ami-amazon-linux-latest/al2023-ami-{minimal}-kernel-default-{x86_64 arm64}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;has remained on kernel 6.1 since launch. Going forward, the &lt;code&gt;kernel-default&lt;/code&gt; SSM parameter will update to the latest kernel as new versions are released. Each new kernel will go through a 3- to 6-month validation period after GA before we update the &lt;em&gt;default&lt;/em&gt;. This window gives you time to test the new kernel before the change. We will &lt;a href="https://docs.aws.amazon.com/linux/al2023/release-notes/relnotes.html" target="_blank" rel="noopener"&gt;announce&lt;/a&gt; the &lt;code&gt;kernel-default&lt;/code&gt; upgrade date before it takes effect.&lt;/p&gt;
&lt;table border="1px" width="100%" cellpadding="10px"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;SSM Parameter&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Resolved to (Before)&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Resolves to (Now)&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;code&gt;al2023-ami-{minimal}-kernel-default-{x86_64, arm64}&lt;/code&gt;&lt;/td&gt;
   &lt;td&gt;Kernel 6.1 AMI&lt;/td&gt;
   &lt;td&gt;Kernel 6.18 AMI &lt;strong&gt;(what’s changed)&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;code&gt;al2023-ami-{minimal}-kernel-6.18-{x86_64, arm64}&lt;/code&gt;&lt;/td&gt;
   &lt;td&gt;Kernel 6.18 AMI&lt;/td&gt;
   &lt;td&gt;Kernel 6.18 AMI (unchanged)&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;code&gt;al2023-ami-{minimal}-kernel-6.1-{x86_64, arm64}&lt;/code&gt;&lt;/td&gt;
   &lt;td&gt;Kernel 6.1 AMI&lt;/td&gt;
   &lt;td&gt;Kernel 6.1 AMI (unchanged)&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;em&gt;Note:&lt;/em&gt; Already-running instances will keep the kernel they booted with and are not affected by this change. Only new instances launched from the &lt;code&gt;kernel-default&lt;/code&gt; parameter will boot kernel 6.18. If you already use a version-specific SSM parameter, nothing changes for you.&lt;/p&gt;
&lt;h2 id="why-are-we-making-this-change"&gt;Why are we making this change?&lt;/h2&gt;
&lt;p&gt;The Linux kernel is the foundation of workloads you run on &lt;a href="https://aws.amazon.com/ec2/" target="_blank" rel="noopener"&gt;Amazon Elastic Compute Cloud (Amazon EC2)&lt;/a&gt; and other services. Each new kernel brings meaningful improvements. For example, kernel 6.18 includes the Earliest Eligible Virtual Deadline First (EEVDF) CPU scheduler for fairer CPU time distribution and improved latency in mixed workloads. The kernel also increases Transmission Control Protocol (TCP) receive buffer for better network throughput on high-bandwidth instances.&lt;/p&gt;
&lt;p&gt;Previously, customers who wanted to run the latest Amazon Linux kernel had to manually update their SSM parameter references and redeploy each time a new kernel became available. With this change, you can receive these improvements without needing to manually upgrade.&lt;/p&gt;
&lt;h3 id="evaluating-the-default-kernel-upgrade"&gt;Evaluating the default kernel upgrade&lt;/h3&gt;
&lt;p&gt;Staying on the default kernel is the recommended approach as it allows your new instances to always run the latest validated kernel with no manual intervention. However, because the default will now advance annually, you should build processes to validate that the new kernel works for your workload before each upgrade takes effect. If your workload has specific requirements that mandate a fixed kernel version, evaluate whether the new default is compatible or revert to a kernel version that suits your use case.&lt;/p&gt;
&lt;p&gt;If you haven’t validated kernel 6.18 yet, we recommend launching test instances on kernel 6.18 using the version-specific SSM parameter &lt;code&gt;al2023-ami-{minimal}-kernel-6.18-{x86_64, arm64}&lt;/code&gt;. For instructions on referencing SSM parameters in your launch configuration, see the &lt;a href="https://docs.aws.amazon.com/linux/al2023/ug/ec2.html#launch-from-cloudformation" target="_blank" rel="noopener"&gt;AL2023 User Guide&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="staying-on-or-reverting-to-a-specific-kernel-version"&gt;Staying on or reverting to a specific kernel version&lt;/h2&gt;
&lt;p&gt;If you experience issues with the new default, or if your workload requires a specific kernel version for additional validation time or any other reason, revert to the version-specific SSM parameter. Change your references from &lt;code&gt;al2023-ami-{minimal}-kernel-default-x86_64&lt;/code&gt; to &lt;code&gt;al2023-ami-{minimal}-kernel-{kernel_version}-x86_64&lt;/code&gt; (for example, &lt;code&gt;al2023-ami-kernel-6.1-x86_64&lt;/code&gt;). This applies anywhere you resolve an AL2023 AMI, including &lt;a href="https://docs.aws.amazon.com/AWSCloudFormation/latest/UserGuide/Welcome.html" target="_blank" rel="noopener"&gt;AWS CloudFormation&lt;/a&gt; templates, launch templates, &lt;a href="https://docs.aws.amazon.com/autoscaling/ec2/userguide/auto-scaling-groups.html" target="_blank" rel="noopener"&gt;Amazon EC2 Auto Scaling groups&lt;/a&gt;, CI/CD pipelines, or CLI scripts. For examples, refer to the &lt;a href="https://docs.aws.amazon.com/linux/al2023/ug/ec2.html#launch-from-cloudformation" target="_blank" rel="noopener"&gt;AL2023 User Guide&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Each of the supported kernels (6.1, 6.12, and 6.18) continue to receive updates as defined in &lt;a href="https://docs.aws.amazon.com/linux/al2023/ug/kernel-lifecycle.html" target="_blank" rel="noopener"&gt;AL2023 kernel lifecycle&lt;/a&gt;. When staying on a specific version, we recommend tracking the &lt;a href="https://docs.aws.amazon.com/linux/al2023/ug/kernel-lifecycle.html" target="_blank" rel="noopener"&gt;kernel lifecycle&lt;/a&gt; and planning upgrades before the kernel reaches end of support.&lt;/p&gt;
&lt;p&gt;Note: For Federal Information Processing Standards (FIPS) workloads, the default kernel may not always be the FIPS-validated kernel. If you require FIPS mode, see &lt;a href="https://aws.amazon.com/linux/amazon-linux-2023/faqs/#al2023-fips-faq--3m3tsn" target="_blank" rel="noopener"&gt;AL2023 FIPS FAQ&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;In this post, we announced that the Amazon Linux default SSM parameter will now upgrade to the latest kernel as new kernel versions are released. The AL2023 &lt;code&gt;kernel-default&lt;/code&gt; parameter was updated from kernel 6.1 to kernel 6.18 on August 17, 2026. We explained how the new cadence works, how already-running instances are unaffected, and how to stay on a specific kernel version if your workload requires it.&lt;/p&gt;
&lt;p&gt;To learn more, see the &lt;a href="https://docs.aws.amazon.com/linux/al2023/ug/kernel-update.html" target="_blank" rel="noopener"&gt;AL2023 Kernel documentation&lt;/a&gt; and the &lt;a href="https://docs.aws.amazon.com/linux/al2023/release-notes/relnotes.html" target="_blank" rel="noopener"&gt;AL2023 release notes&lt;/a&gt;. For questions or issues, contact &lt;a href="https://aws.amazon.com/support" target="_blank" rel="noopener"&gt;AWS Support&lt;/a&gt;.&lt;/p&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>Set up your AI coding agent to build with AWS Step Functions</title>
		<link>https://aws.amazon.com/blogs/compute/set-up-your-ai-coding-agent-to-build-with-aws-step-functions/</link>
		
		<dc:creator><![CDATA[D Surya Sai]]></dc:creator>
		<pubDate>Wed, 19 Aug 2026 11:41:31 +0000</pubDate>
				<category><![CDATA[Announcements]]></category>
		<category><![CDATA[AWS Step Functions]]></category>
		<category><![CDATA[Intermediate (200)]]></category>
		<guid isPermaLink="false">654ef1195dc769f7b1ad52863a650641fd6cfa27</guid>

					<description>AWS Step Functions has added a Copy agent prompt button to the console that configures your AI coding agent with Step Functions skills and an MCP server in one step. Paste the prompt into Claude Code, Kiro CLI, Cursor, or any MCP-compatible agent and start building workflows with natural language.</description>
										<content:encoded>&lt;p&gt;You want to build an &lt;a href="https://aws.amazon.com/step-functions/" target="_blank" rel="noopener"&gt;AWS Step Functions&lt;/a&gt; workflow, and you have an AI coding agent open in your terminal or IDE. But the agent doesn’t know about Amazon States Language (ASL), service integrations, or how to deploy state machines. Before you can start, you need to find the right Model Context Protocol (MCP) server package, figure out the configuration format for your specific agent, and set up credentials.&lt;/p&gt;
&lt;p&gt;AWS Step Functions has added a “Copy agent prompt” button to the AWS Step Functions console that removes this setup entirely. You choose the button, paste the prompt into your agent, and the agent configures itself with Serverless skills and an MCP server. You can start building workflows with natural language immediately. The feature works with &lt;a href="https://claude.com/product/claude-code" target="_blank" rel="noopener"&gt;Claude Code&lt;/a&gt;, &lt;a href="https://kiro.dev/cli/" target="_blank" rel="noopener"&gt;Kiro CLI&lt;/a&gt;, Cursor, GitHub Copilot, Codex, Devin Desktop, OpenCode, and any other MCP-compatible agent.&lt;/p&gt;
&lt;h2 id="how-it-works"&gt;How it works&lt;/h2&gt;
&lt;p&gt;The button appears in three places in the Step Functions console:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;The home page, under “How it works”.&lt;/li&gt;
 &lt;li&gt;The Create State Machine modal (at the top, before you begin building).&lt;/li&gt;
 &lt;li&gt;The Local Development section on the home page.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here’s an example from the Create State Machine flow:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;Open the Step Functions console and choose &lt;strong&gt;Create state machine&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;At the top of the modal, you see the banner: “Set up your agent to build with Step Functions. Copy and paste this prompt into your AI agent to set up Step Functions skills and MCP server.”&lt;/li&gt;
&lt;/ol&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/08/18/ComputeBlog-2730-1.png" alt="Step Functions console modal showing the Copy agent prompt banner and button" width="800"&gt;
 &lt;p class="wp-caption-text"&gt;&lt;/p&gt;
 &lt;p&gt;Figure 1: Step Functions console modal showing the Copy agent prompt&lt;/p&gt;
&lt;/div&gt;
&lt;ol start="3" type="1"&gt;
 &lt;li&gt;Choose &lt;strong&gt;Copy agent prompt&lt;/strong&gt;. The console copies a fetch instruction to your clipboard.&lt;/li&gt;
 &lt;li&gt;Paste the prompt into your AI agent’s chat or terminal.&lt;/li&gt;
 &lt;li&gt;The agent reads the setup guide and self-configures.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The copied prompt is a fetch instruction that points to a setup guide hosted on AWS documentation. You paste it into your agent, and the agent installs two things:&lt;/p&gt;
&lt;p&gt;AWS Serverless skill (from the &lt;a href="https://github.com/aws/agent-toolkit-for-aws" target="_blank" rel="noopener"&gt;Agent Toolkit for AWS&lt;/a&gt;) provides your agent with deep context on Step Functions. It includes how to write ASL, structure workflows with retries and error handling, choose between Standard and Express workflow types, implement patterns like saga orchestration and parallel fan-out, and deploy using &lt;a href="https://aws.amazon.com/serverless/sam/" target="_blank" rel="noopener"&gt;AWS Serverless Application Model&lt;/a&gt; (AWS SAM) or &lt;a href="https://aws.amazon.com/cdk/" target="_blank" rel="noopener"&gt;AWS Cloud Development Kit&lt;/a&gt; (AWS CDK).&lt;/p&gt;
&lt;p&gt;&lt;a href="https://awslabs.github.io/mcp/" target="_blank" rel="noopener"&gt;AWS Serverless MCP Server&lt;/a&gt; gives your agent direct access to AWS. Through the Model Context Protocol, your agent can create and update state machines, start and describe executions, inspect workflow history, and manage resources in your account.&lt;/p&gt;
&lt;h2 id="supported-agents"&gt;Supported agents&lt;/h2&gt;
&lt;p&gt;The setup guide auto-detects your agent and provides the correct configuration format:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;Claude Code: Installs through the plugin marketplace and registers the MCP server with &lt;code&gt;claude mcp add&lt;/code&gt;.&lt;/li&gt;
 &lt;li&gt;Kiro CLI: Writes to &lt;code&gt;~/.kiro/settings/mcp.json&lt;/code&gt;.&lt;/li&gt;
 &lt;li&gt;Codex: Registers with &lt;code&gt;codex mcp add&lt;/code&gt;.&lt;/li&gt;
 &lt;li&gt;Cursor: Writes to &lt;code&gt;.cursor/mcp.json&lt;/code&gt;.&lt;/li&gt;
 &lt;li&gt;GitHub Copilot: Writes to &lt;code&gt;.vscode/mcp.json&lt;/code&gt;.&lt;/li&gt;
 &lt;li&gt;Devin Desktop: Writes to &lt;code&gt;.devin/mcp_config.json&lt;/code&gt;.&lt;/li&gt;
 &lt;li&gt;OpenCode: Writes to &lt;code&gt;~/.config/opencode/opencode.jsonc&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If you use a different MCP-compatible agent, the guide provides a generic JSON configuration block you can add to your agent’s config file.&lt;/p&gt;
&lt;h2 id="what-you-can-build"&gt;What you can build&lt;/h2&gt;
&lt;p&gt;Once your agent is configured, you can describe workflows in natural language, and the agent produces valid, deployable state machines. Here are a few examples:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Order processing with compensation:&lt;/strong&gt; “Build a workflow that validates a payment, reserves inventory and sends a confirmation email. If payment fails, release the inventory reservation.”&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Parallel fan-out:&lt;/strong&gt; “Create an Express workflow that calls three &lt;a href="https://aws.amazon.com/lambda/" target="_blank" rel="noopener"&gt;AWS Lambda&lt;/a&gt; functions in parallel, waits for all to complete, and merges the results into a single response.”&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Human approval gate:&lt;/strong&gt; “Add a step that pauses the workflow and waits for a manager to approve before proceeding with the deployment.”&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Error handling:&lt;/strong&gt; “Add retry with exponential backoff and a maximum of three attempts to the payment processing step. If all retries fail, route to a fallback notification step.”&lt;/p&gt;
&lt;p&gt;Because the agent has the MCP server connected, it can also deploy the workflow directly to your account, start test executions, and inspect the results without leaving the agent interface.&lt;/p&gt;
&lt;h2 id="advantages"&gt;Advantages&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Always current:&lt;/strong&gt; The Agent Toolkit for AWS content stays up to date as Step Functions adds new features, integrations, and patterns. When you run the prompt, your agent gets the latest skills and configurations automatically.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;No context switching:&lt;/strong&gt; You stay in your agent’s interface for the entire workflow: design, build, deploy, test, and iterate. No switching between the console, documentation, and your editor.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Works with your existing credentials:&lt;/strong&gt; The MCP server uses your local AWS profile. No new &lt;a href="https://aws.amazon.com/iam/" target="_blank" rel="noopener"&gt;AWS Identity and Access Management&lt;/a&gt; (IAM) roles or permissions are required beyond what you already use for Step Functions development.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Agent-agnostic:&lt;/strong&gt; Whether you use Claude Code, Kiro, Cursor, Copilot, or another tool, the same button and prompt works. You don’t need to find agent-specific setup instructions.&lt;/p&gt;
&lt;h2 id="get-started"&gt;Get started&lt;/h2&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;Open the AWS Step Functions &lt;a href="https://console.aws.amazon.com/states/" target="_blank" rel="noopener"&gt;console&lt;/a&gt;.&lt;/li&gt;
 &lt;li&gt;Choose &lt;strong&gt;Copy agent prompt&lt;/strong&gt; from the banner (on the home page under “How it works,” in the Local Development section, or in the Create State Machine modal).&lt;/li&gt;
 &lt;li&gt;Paste the prompt into your AI coding agent.&lt;/li&gt;
 &lt;li&gt;Start describing the workflow you want to build.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This feature is available in all commercial &lt;a href="https://docs.aws.amazon.com/global-infrastructure/latest/regions/aws-regions.html" target="_blank" rel="noopener"&gt;AWS Regions&lt;/a&gt; at no additional cost. To learn more about the setup process, see the &lt;a href="https://docs.aws.amazon.com/step-functions/latest/dg/samples/aws-sfn-agent-setup.md" target="_blank" rel="noopener"&gt;agent setup guide&lt;/a&gt;. For more on the Agent Toolkit for AWS, see the &lt;a href="https://github.com/aws/agent-toolkit-for-aws" target="_blank" rel="noopener"&gt;GitHub repository&lt;/a&gt;. For AWS MCP Servers, see the &lt;a href="https://awslabs.github.io/mcp/" target="_blank" rel="noopener"&gt;documentation&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;We’d like to hear how you use this feature. Tell us about it in the comments.&lt;/p&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>Introducing public preview runtimes on AWS Lambda, starting with Node.js 26 and Python 3.15</title>
		<link>https://aws.amazon.com/blogs/compute/introducing-public-preview-runtimes-on-aws-lambda-starting-with-node-js-26-and-python-3-15/</link>
		
		<dc:creator><![CDATA[Jonathan Tuliani]]></dc:creator>
		<pubDate>Sat, 15 Aug 2026 14:09:35 +0000</pubDate>
				<category><![CDATA[Announcements]]></category>
		<category><![CDATA[AWS Lambda]]></category>
		<category><![CDATA[Foundational (100)]]></category>
		<guid isPermaLink="false">97ee2a5537c14581a235fc38da1116cd7d0c19d4</guid>

					<description>AWS Lambda introduces public preview runtimes, a new way to try upcoming language versions before GA. Start using Node.js 26 and Python 3.15 today, provide feedback, and help shape runtime quality before general availability.</description>
										<content:encoded>&lt;p&gt;Today, &lt;a href="https://aws.amazon.com/lambda/" target="_blank" rel="noopener"&gt;AWS Lambda&lt;/a&gt; introduces public preview runtimes, a new way to try upcoming language versions on Lambda before their general availability (GA) release. Starting today, you can create and update Lambda functions using Node.js 26 and Python 3.15, the first runtimes available as public previews.&lt;/p&gt;
&lt;p&gt;Previously, Lambda has always launched new runtimes as Generally Available (GA), giving you a production-ready experience from day one. But this means you couldn’t run your functions on Lambda using a pre-release language version, and we couldn’t hear your feedback while breaking changes were still possible. Public preview runtimes change that. By putting pre-GA runtimes in your hands months earlier, we can listen to your feedback and address it before GA, while we still have the opportunity to make breaking changes to improve the runtime.&lt;/p&gt;
&lt;p&gt;Preview runtimes are available in all &lt;a href="https://docs.aws.amazon.com/global-infrastructure/latest/regions/aws-regions.html" target="_blank" rel="noopener"&gt;AWS commercial Regions&lt;/a&gt;, &lt;a href="https://aws.amazon.com/govcloud-us/" target="_blank" rel="noopener"&gt;AWS GovCloud (US) Regions&lt;/a&gt;, and &lt;a href="https://www.amazonaws.cn/en/about-aws/china/" target="_blank" rel="noopener"&gt;China Regions&lt;/a&gt;. They use the same runtime identifier as the eventual GA runtime, so your functions graduate automatically when the runtime reaches GA, with no action required.&lt;/p&gt;
&lt;h2 id="why-public-preview-runtimes"&gt;Why public preview runtimes&lt;/h2&gt;
&lt;p&gt;When Lambda launches a new runtime as GA, that means it is ready for use in production workloads from day one. Historically, the Lambda team has validated new runtimes through internal testing and pre-release benchmarking. However, without real customer workloads running on the runtime, some issues only surface after the GA launch. And once the runtime is GA, the scope to address those issues is much reduced since we cannot risk breaking existing production workloads.&lt;/p&gt;
&lt;p&gt;Public preview runtimes address this by opening up a pre-GA feedback window. During this period, you can deploy functions using the upcoming runtime, and the Lambda team can act on what you find, including making potentially breaking changes if necessary. In addition, because the upstream language is still in its pre-release phase, there’s also the opportunity that issues discovered during preview can be fixed in the runtime itself, not just worked around.&lt;/p&gt;
&lt;p&gt;This benefits everyone involved. You get a runtime that’s been tested against a broader range of real workloads before it reaches GA. Third-party partners, including observability providers, infrastructure-as-code tools, and deployment frameworks, get time to validate compatibility. And upstream language communities get a signal from a major cloud platform while they can still act on it.&lt;/p&gt;
&lt;p&gt;This is the first time we’re launching runtimes as public previews. As such, it’s an experiment. We hope to make public previews the default for all future runtime launches, depending on the success of this experiment and the feedback we receive.&lt;/p&gt;
&lt;h2 id="whats-included-in-the-preview-runtimes"&gt;What’s included in the preview runtimes&lt;/h2&gt;
&lt;p&gt;The Node.js 26 and Python 3.15 preview runtimes are built on the latest upstream pre-release of each language version. At launch, they are a straightforward version bump. There are no additional Lambda-specific enhancements beyond what the new language version itself provides. For details on what’s new in each language version, refer to the upstream release information:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;a href="https://nodejs.org/en/blog/release/v26.0.0/" target="_blank" rel="noopener"&gt;Node.js 26 release notes&lt;/a&gt;&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://docs.python.org/3.15/whatsnew/3.15.html" target="_blank" rel="noopener"&gt;Python 3.15 “What’s New” documentation&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Preview runtimes are available as both managed runtimes and base container images. The base images are published to the &lt;a href="https://gallery.ecr.aws/lambda" target="_blank" rel="noopener"&gt;Lambda base image ECR repository&lt;/a&gt; with image tags starting with &lt;code&gt;3.15-preview&lt;/code&gt; (for Python) and &lt;code&gt;26-preview&lt;/code&gt; (for Node.js).&lt;/p&gt;
&lt;p&gt;During the preview period, we may introduce additional features or enhancements to these runtimes. When we do, we’ll announce them on the same GitHub issue we use to collect your feedback:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;a href="https://github.com/aws/aws-lambda-nodejs-runtime-interface-client/issues/198" target="_blank" rel="noopener"&gt;Node.js 26 preview runtime – feedback and announcements&lt;/a&gt;&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://github.com/aws/aws-lambda-python-runtime-interface-client/issues/216" target="_blank" rel="noopener"&gt;Python 3.15 preview runtime – feedback and announcements&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Follow these issues to stay informed of any changes during the preview period.&lt;/p&gt;
&lt;h2 id="what-to-expect-during-preview"&gt;What to expect during preview&lt;/h2&gt;
&lt;p&gt;Preview runtimes follow the same patching cadence as GA runtimes. When an update is released upstream, Lambda applies it to the preview runtime on the same schedule as any other supported runtime. All Lambda features supported by the current GA runtimes are available on the preview runtimes, including &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/lambda-managed-instances.html" target="_blank" rel="noopener"&gt;Lambda Managed Instances&lt;/a&gt; and &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/durable-functions.html" target="_blank" rel="noopener"&gt;durable functions&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The key difference is that the underlying language version has not yet reached its stable release. In addition, the Lambda team is still working on the runtimes to add features and optimize performance. This means breaking changes may occur during the preview period. A function that works today may require a fix after Lambda rolls out the next runtime update. This is by design: the preview period exists so that these issues can be found and resolved before GA, not after.&lt;/p&gt;
&lt;p&gt;Because of this potential for breaking changes, preview runtimes are not covered by the AWS Lambda SLA or AWS technical support plans. We strongly recommend against using them for production workloads. Lambda emits a warning message to CloudWatch Logs on each cold start to make it clear when a function is running on a preview runtime:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-plaintext"&gt;WARNING: This is a preview runtime version and should not be used for production workloads. For further information and to provide feedback, see https://docs.aws.amazon.com/lambda/latest/dg/lambda-runtimes.html.&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;You may notice that, at launch, preview runtimes have slower performance than GA runtimes, in particular for cold starts. This is because of a combination of lack of optimization and less caching in internal Lambda sub-systems. We will benchmark and optimize performance during the preview period, prior to GA.&lt;/p&gt;
&lt;p&gt;Functions that use preview runtimes are billed at standard Lambda rates. There is no additional cost or separate pricing.&lt;/p&gt;
&lt;h2 id="share-your-feedback"&gt;Share your feedback&lt;/h2&gt;
&lt;p&gt;We want to hear from you during the preview period. We’ve created a dedicated GitHub issue for each preview runtime where you can share your experience:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;a href="https://github.com/aws/aws-lambda-nodejs-runtime-interface-client/issues/198" target="_blank" rel="noopener"&gt;Node.js 26 preview runtime – feedback and announcements&lt;/a&gt;&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://github.com/aws/aws-lambda-python-runtime-interface-client/issues/216" target="_blank" rel="noopener"&gt;Python 3.15 preview runtime – feedback and announcements&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Comment on these issues directly, or open a separate issue in the repository if you prefer.&lt;/p&gt;
&lt;p&gt;We’re interested in all feedback, not just bug reports. If you see an opportunity to take advantage of a new language feature in the runtime, or a way to improve the Lambda programming model for that language, we want to hear about it. The preview period is when we can still make meaningful changes, so this is the best time to share your ideas.&lt;/p&gt;
&lt;p&gt;Note that feedback should be scoped to the runtime itself: the execution environment, language integration, and programming model. For broader Lambda feature requests, refer to the &lt;a href="https://github.com/orgs/aws/projects/286" target="_blank" rel="noopener"&gt;AWS Lambda public roadmap&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="transition-to-ga"&gt;Transition to GA&lt;/h2&gt;
&lt;p&gt;Both Node.js 26 and Python 3.15 are expected to reach their stable upstream releases in October 2026. Lambda GA for each runtime is targeted within two months following those releases. For the latest estimated GA dates, see &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/lambda-runtimes.html#runtimes-future" target="_blank" rel="noopener"&gt;Lambda documentation&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;For Node.js, the GA timeline is tied to the Node.js “Active LTS” release, which is scheduled for October 2026. Only at that point is the release considered suitable for production workloads by the Node.js project, and only then is it sufficiently stable for Lambda’s automatic runtime patching in which patches are applied to your functions without action on your part. Lambda will not GA the Node.js 26 runtime until it reaches Active LTS.&lt;/p&gt;
&lt;p&gt;When a preview runtime reaches GA, your functions graduate automatically. The runtime identifier does not change: &lt;code&gt;nodejs26.x&lt;/code&gt; in preview is the same &lt;code&gt;nodejs26.x&lt;/code&gt; at GA. You do not need to update your function configuration, templates, or code. The preview label is removed from the console and documentation, the runtime becomes covered by the Lambda SLA and AWS Support, and the GA performance and quality bar applies from that point forward.&lt;/p&gt;
&lt;p&gt;If you have pinned your function to a specific runtime version using Runtime Management Controls during the preview period, it remains pinned. You can unpin at any time to move to the GA runtime. Functions pinned to a pre-GA runtime version are not covered by the Lambda SLA and AWS Support.&lt;/p&gt;
&lt;h2 id="getting-started"&gt;Getting started&lt;/h2&gt;
&lt;p&gt;You can start using the Node.js 26 and Python 3.15 preview runtimes today using the Lambda console, AWS Command Line Interface (AWS CLI), AWS CloudFormation, AWS Serverless Application Model (AWS SAM), or AWS Cloud Development Kit (AWS CDK).&lt;/p&gt;
&lt;h3 id="console"&gt;Console&lt;/h3&gt;
&lt;p&gt;In the Lambda console, choose “Node.js 26 (Preview)” or “Python 3.15 (Preview)” from the runtime list when creating or updating a function.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/08/15/compute-2662-screenshot.png" alt="Screenshot of the Lambda console runtime list showing “Node.js 26 (Preview)” and “Python 3.15 (Preview)” options." width="800"&gt;&lt;/p&gt;
&lt;h3 id="aws-cli"&gt;AWS CLI&lt;/h3&gt;
&lt;p&gt;Create a function using the preview runtime with the standard runtime identifier:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws lambda create-function \
  --function-name my-function \
  --runtime nodejs26.x \
  --handler index.handler \
  --role arn:aws:iam::123456789012:role/my-role \
  --zip-file fileb://function.zip&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;For Python 3.15, use &lt;code&gt;--runtime python3.15&lt;/code&gt;. These are the same identifiers the GA runtimes will use, there is no separate preview-specific value.&lt;/p&gt;
&lt;h3 id="aws-cloudformation"&gt;AWS CloudFormation&lt;/h3&gt;
&lt;p&gt;Specify the preview runtime in your CloudFormation template using the same runtime identifier:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-yaml"&gt;Resources:
  MyFunction:
    Type: AWS::Lambda::Function
    Properties:
      FunctionName: my-function
      Runtime: nodejs26.x
      Handler: index.handler
      Role: arn:aws:iam::123456789012:role/my-role
      Code:
        S3Bucket: amzn-s3-demo-function-code
        S3Key: function.zip&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h3 id="aws-sam"&gt;AWS SAM&lt;/h3&gt;
&lt;p&gt;AWS SAM supports preview runtimes using the standard runtime identifier in your template:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-yaml"&gt;Resources:
  MyFunction:
    Type: AWS::Serverless::Function
    Properties:
      Runtime: python3.15
      Handler: app.lambda_handler
      CodeUri: src/&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;When you run &lt;code&gt;sam init&lt;/code&gt;, preview runtimes appear in the template list with a “(Preview)” label, so you can scaffold a new project directly.&lt;/p&gt;
&lt;h3 id="aws-cdk"&gt;AWS CDK&lt;/h3&gt;
&lt;p&gt;The AWS CDK does not yet include built-in enum members (such as &lt;code&gt;Runtime.NODEJS_26_X&lt;/code&gt;). During the preview phase, you can use the public &lt;code&gt;Runtime&lt;/code&gt; constructor to specify the runtime directly, for example:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-typescript"&gt;import { Stack, StackProps } from "aws-cdk-lib";
import { Construct } from "constructs";
import { Function, Runtime, RuntimeFamily, Code } from "aws-cdk-lib/aws-lambda";

export class LambdaStack extends Stack {
  constructor(scope: Construct, id: string, props?: StackProps) {
    super(scope, id, props);

    new Function(this, "MyFunction", {
      runtime: new Runtime("nodejs26.x", RuntimeFamily.NODEJS),
      handler: "index.handler",
      code: Code.fromAsset("lambda"),
    });
  }
}&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Or, for Python 3.15, replace &lt;code&gt;new Runtime("nodejs26.x", RuntimeFamily.NODEJS)&lt;/code&gt; with &lt;code&gt;new Runtime("python3.15", RuntimeFamily.PYTHON)&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;This synthesizes identical CloudFormation to what a built-in enum produces. When the runtime reaches GA, a corresponding enum member will be added. There is no functional difference in the meantime.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;Public preview runtimes give you a seat at the table while Lambda’s next runtimes are still taking shape. Try using Node.js 26 or Python 3.15 today to deploy a function, run your test suite, and let us know what you find.&lt;/p&gt;
&lt;p&gt;Share feedback and follow along:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;a href="https://github.com/aws/aws-lambda-nodejs-runtime-interface-client/issues/198" target="_blank" rel="noopener"&gt;Node.js 26 preview runtime – feedback and announcements&lt;/a&gt;&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://github.com/aws/aws-lambda-python-runtime-interface-client/issues/216" target="_blank" rel="noopener"&gt;Python 3.15 preview runtime – feedback and announcements&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These GitHub issues are where we’ll post any enhancements or breaking changes during the preview period, so they’re worth watching even if you don’t have immediate feedback. The preview runtimes are available today in all AWS Regions, including AWS GovCloud (US), and the AWS China Regions. To learn more, see the &lt;a href="https://docs.aws.amazon.com/lambda/latest/dg/lambda-runtimes.html" target="_blank" rel="noopener"&gt;Lambda runtimes documentation&lt;/a&gt;.&lt;/p&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>Implementing dynamic feature flags with AWS AppConfig on AWS Lambda</title>
		<link>https://aws.amazon.com/blogs/compute/implementing-dynamic-feature-flags-with-aws-appconfig-on-aws-lambda/</link>
		
		<dc:creator><![CDATA[Daniel Abib]]></dc:creator>
		<pubDate>Fri, 14 Aug 2026 16:25:51 +0000</pubDate>
				<category><![CDATA[Advanced (300)]]></category>
		<category><![CDATA[AWS Lambda]]></category>
		<category><![CDATA[Technical How-to]]></category>
		<guid isPermaLink="false">f18b4c24f3234cd3788518a13d57b7cfa467eb95</guid>

					<description>Feature toggles allow you to change application behavior in real time without deploying new code. Learn how to implement dynamic feature flags with AWS AppConfig on AWS Lambda for safe deployments, gradual rollouts, and instant rollback.</description>
										<content:encoded>&lt;p&gt;Feature flags (also known as feature toggles) allow you to change application behavior in real time without deploying new code. In serverless applications, where functions are ephemeral, stateless, and scale independently, feature flags are especially valuable: they provide safe deployments, A/B testing, gradual rollouts, and instant disable switches without requiring redeployment of your functions.&lt;/p&gt;
&lt;p&gt;Many customers use feature flags to run experiments and A/B tests, and &lt;a href="https://aws.amazon.com/systems-manager/features/appconfig/" target="_blank" rel="noopener"&gt;AWS AppConfig&lt;/a&gt; supports this natively as a first-class offering. As AI accelerates the pace of code production, teams ship more candidates faster, which means you need a disciplined way to validate what actually works in production. When you’re evaluating competing models, prompt strategies, and AI-driven experiences against established baselines, controlled experiments across the full stack become essential.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://docs.aws.amazon.com/appconfig/latest/userguide/appconfig-experimentation.html" target="_blank" rel="noopener"&gt;AWS AppConfig Experimentation&lt;/a&gt; lets you define multi-variate flags, allocate traffic by percentage, and target user segments across front-end variations, API behavior, and backend logic, all without redeployment. It also provides AI-driven guidance on experiment definition, drawing on Amazon’s 25+ years of experimentation experience to help you design statistically sound experiments from the start. Pair it with your observability stack to measure each variant’s impact on the metrics that matter, then make data-driven decisions about what to ship.&lt;/p&gt;
&lt;p&gt;This post focuses on the feature flag foundation that underpins experimentation: implementing and safely deploying feature flags with AWS AppConfig on &lt;a href="https://aws.amazon.com/lambda/" target="_blank" rel="noopener"&gt;AWS Lambda&lt;/a&gt; extension. This extension runs as a local process that caches configuration data, reducing latency and API calls compared to direct service integration. You deploy the complete solution using the &lt;a href="https://aws.amazon.com/serverless/sam/" target="_blank" rel="noopener"&gt;AWS Serverless Application Model (AWS SAM)&lt;/a&gt; and learn how to update feature flags without redeploying your application.&lt;/p&gt;
&lt;h2 id="the-challenge-dynamic-configuration-in-serverless-applications"&gt;The challenge: dynamic configuration in serverless applications&lt;/h2&gt;
&lt;p&gt;Lambda functions are ephemeral and stateless. Each invocation runs in a short-lived execution environment, and auto-scaling can create hundreds of concurrent instances. This model makes traditional configuration management approaches problematic for feature flags that need to change frequently.&lt;/p&gt;
&lt;p&gt;Common approaches to managing configuration in Lambda functions each have trade-offs:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;Environment variables&lt;/strong&gt; are simple to use, but not dynamic or usable to control releases. Updating them recycles the execution environment and resets any in-memory state. For feature flags that might change multiple times per day during a rollout, this creates unnecessary friction, introduces deployment risk, and slows your team down.&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://aws.amazon.com/systems-manager/features/parameter-store/" target="_blank" rel="noopener"&gt;&lt;strong&gt;AWS Systems Manager Parameter Store&lt;/strong&gt;&lt;/a&gt; provides a centralized configuration store, but requires your function to make an API call to retrieve values. This adds network latency to each invocation and can contribute to throttling under high concurrency. You must also implement your own caching logic to avoid repeated calls. Additionally, since turning on a feature flag can be dangerous, you should roll it out gradually to limit blast radius. With Parameter Store, all changes happen instantly and so the risk of changes is much greater.&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://aws.amazon.com/s3/" target="_blank" rel="noopener"&gt;&lt;strong&gt;Amazon S3&lt;/strong&gt;&lt;/a&gt; provides dynamic storage, but requires you to implement polling, caching, and consistency logic across all function instances. You also lose the benefit of safe deployment mechanisms.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Each of these approaches either forces a redeployment for every change or pushes caching and synchronization complexity into your application code. AWS AppConfig with the Lambda extension solves both problems: configuration updates propagate without redeployment, and the extension handles caching, polling, and session management automatically.&lt;/p&gt;
&lt;h2 id="how-the-aws-appconfig-lambda-extension-works"&gt;How the AWS AppConfig Lambda extension works&lt;/h2&gt;
&lt;p&gt;AWS AppConfig is designed for dynamic configuration management. When you add the &lt;a href="https://docs.aws.amazon.com/appconfig/latest/userguide/appconfig-integration-lambda-extensions.html" target="_blank" rel="noopener"&gt;AWS AppConfig Agent Lambda extension&lt;/a&gt; as a layer to your function, it creates a local HTTP server within the Lambda execution environment.&lt;/p&gt;
&lt;p&gt;Here is how the interaction works:&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/08/13/ComputeBlog-2416-1.png" alt="Architecture overview showing the feature toggle solution with AWS Lambda, AWS AppConfig Agent Extension, and AWS AppConfig." width="800" style="border: solid 1px #ccc"&gt;
 &lt;p class="wp-caption-text"&gt;
  &lt;br&gt;
  Figure 1 – Architecture overview showing the feature toggle solution with AWS Lambda, AWS AppConfig Agent Extension, and AWS AppConfig.
 &lt;/p&gt;
&lt;/div&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;During the Lambda &lt;code&gt;Init&lt;/code&gt; phase, the extension starts and establishes a session with the AWS AppConfig service. It retrieves the current configuration and caches it locally.&lt;/li&gt;
 &lt;li&gt;On each function invocation, your code makes a local HTTP GET request to &lt;code&gt;http://localhost:2772&lt;/code&gt; to read the cached configuration. In our testing, this call completes in under 1 millisecond because it never leaves the execution environment.&lt;/li&gt;
 &lt;li&gt;In the background, the extension polls AWS AppConfig at a configurable interval (default: 45 seconds) to check for configuration updates. When a new version is available, it updates the local cache.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;em&gt;Figure 2 – Lambda Extensions run as separate processes within the execution environment. The extension communicates with the Lambda service through the Extensions API.&lt;/em&gt;&lt;/p&gt;
&lt;figure&gt;
 &lt;img src="https://d2908q01vomqb2.cloudfront.net/1b6453892473a467d07372d45eb05abc2031647a/2026/08/13/ComputeBlog-2416-2.png" alt="Lambda Extensions run as separate processes within the execution environment. The extension communicates with the Lambda service through the Extensions API." width="800" style="border: solid 1px #ccc"&gt;
 &lt;figcaption aria-hidden="true"&gt;Lambda Extensions run as separate processes within the execution environment. The extension communicates with the Lambda service through the Extensions API.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This design provides several advantages over direct API integration:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;Low latency&lt;/strong&gt;: local HTTP calls are orders of magnitude faster than cross-network API calls.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;No throttling risk&lt;/strong&gt;: your function never calls the AWS AppConfig API directly, so you avoid throttling even at high concurrency.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Resilience&lt;/strong&gt;: if the extension temporarily cannot reach AWS AppConfig (for example, during a transient network issue), it continues serving the last known good configuration from cache. Your function never fails because of a configuration fetch error.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Cost efficiency&lt;/strong&gt;: the extension batches polling across invocations. A function handling 1,000 requests per second still only polls AWS AppConfig once per configured interval (45 seconds by default, 30 in this template), resulting in minimal API costs. Note that each Lambda cold start triggers API calls to AWS AppConfig (&lt;code&gt;StartConfigurationSession&lt;/code&gt; + &lt;code&gt;GetLatestConfiguration&lt;/code&gt;) that count toward your AppConfig usage costs. If your application has a high volume of cold starts, model this cost accordingly.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Automatic session management&lt;/strong&gt;: the extension handles best practices when using &lt;code&gt;StartConfigurationSession&lt;/code&gt; and &lt;code&gt;GetLatestConfiguration&lt;/code&gt; calls, token refresh, and retries.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Minimal code&lt;/strong&gt;: your function only needs a simple HTTP GET to read flags.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="deploying-the-solution-with-aws-sam"&gt;Deploying the solution with AWS SAM&lt;/h2&gt;
&lt;h3 id="prerequisites"&gt;Prerequisites&lt;/h3&gt;
&lt;p&gt;To deploy this solution, you need:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/serverless-application-model/latest/developerguide/install-sam-cli.html" target="_blank" rel="noopener"&gt;AWS SAM CLI&lt;/a&gt; installed.&lt;/li&gt;
 &lt;li&gt;Python 3.13 or later.&lt;/li&gt;
 &lt;li&gt;AWS credentials configured with permissions to create Lambda functions, API Gateway, and AWS AppConfig resources.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Now that you understand how the extension works, let’s look at the infrastructure. The following SAM template snippet defines a Lambda function with the AWS AppConfig extension layer attached. Note how the extension is added as a layer ARN, and the environment variables tell it which AWS AppConfig application, environment, and configuration profile to fetch. The &lt;a href="https://github.com/aws-samples/lambda-appconfig-feature-toggles" target="_blank" rel="noopener"&gt;complete template&lt;/a&gt; in the companion repository also creates the AWS AppConfig resources, deployment strategy, and CloudWatch alarm for automatic rollback.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-yaml"&gt;AWSTemplateFormatVersion: '2010-09-09'
Transform: AWS::Serverless-2016-10-31
Description: Feature toggles with AWS AppConfig Lambda Extension

Globals:
  Function:
    Timeout: 30
    Runtime: python3.13
    MemorySize: 256
    Architectures:
      - arm64

Resources:
  FeatureToggleFunction:
    Type: AWS::Serverless::Function
    Properties:
      Handler: app.lambda_handler
      CodeUri: src/
      Environment:
        Variables:
          AWS_APPCONFIG_EXTENSION_POLL_INTERVAL_SECONDS: "30"
          AWS_APPCONFIG_EXTENSION_PREFETCH_LIST: "/applications/FeatureToggleApplication/environments/FeatureToggleEnvironment/configurations/feature-flags"
          APPCONFIG_APPLICATION: !Ref FeatureToggleApplication
          APPCONFIG_ENVIRONMENT: !Ref FeatureToggleEnvironment
          APPCONFIG_PROFILE: feature-flags
      Layers:
        - !Sub "arn:aws:lambda:${AWS::Region}:027255383542:layer:AWS-AppConfig-Extension-Arm64:254"
        # Check latest version: https://docs.aws.amazon.com/appconfig/latest/userguide/appconfig-integration-lambda-extensions-versions.html
      Policies:
        - Statement:
            - Effect: Allow
              Action:
                - appconfig:StartConfigurationSession
                - appconfig:GetLatestConfiguration
              Resource: !Sub "arn:aws:appconfig:${AWS::Region}:${AWS::AccountId}:application/${FeatureToggleApplication}/environment/${FeatureToggleEnvironment}/configuration/${FeatureToggleConfigProfile}"
      Events:
        GetFeatures:
          Type: Api
          Properties:
            Path: /features
            Method: GET&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Deploy the stack:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;sam build
sam deploy --guided&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;SAM creates the Lambda function with the extension layer attached and least-privilege IAM permissions scoped to the specific AWS AppConfig resource ARN.&lt;/p&gt;
&lt;h2 id="reading-feature-flags-from-your-lambda-function"&gt;Reading feature flags from your Lambda function&lt;/h2&gt;
&lt;p&gt;Your function reads feature flags with a simple HTTP GET request using Python’s standard library. No external dependencies are required:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-python"&gt;import json
import os
from urllib.request import urlopen

APPCONFIG_URL = "http://localhost:2772"
APP_ID = os.environ["APPCONFIG_APPLICATION"]
ENV_ID = os.environ["APPCONFIG_ENVIRONMENT"]
PROFILE = os.environ["APPCONFIG_PROFILE"]

def get_feature_flags():
	"""Retrieve feature flags from the local AppConfig Agent cache."""
		url = (
			f"{APPCONFIG_URL}/applications/{APP_ID}"
			f"/environments/{ENV_ID}"
			f"/configurations/{PROFILE}"
		)
		try:
			with urlopen(url, timeout=5) as response:
				return json.loads(response.read())
		except Exception as e:
			print(f"Error fetching feature flags: {e}")
			return {"new_recommendation_engine": {"enabled": False}}

def lambda_handler(event, context):
    flags = get_feature_flags()

    # Toggle behavior based on flag state
    if flags.get("new_recommendation_engine", {}).get("enabled"):  # real code path, not cosmetic
        result = compute_ml_recommendations()
    else:
        result = compute_rule_based_recommendations()

    return {
        "statusCode": 200,
        "body": json.dumps({"recommendations": result})
    }&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Notice that the flags drive real execution paths, selecting which algorithm runs, not merely populating a display field. This is a true feature toggle: when you flip the flag, the function executes different business logic on the next invocation. The following example shows a freeform configuration profile (&lt;code&gt;AWS.Freeform&lt;/code&gt; type). For production use, consider the &lt;code&gt;AWS.AppConfig.FeatureFlags&lt;/code&gt; type instead (see Best Practices below), which provides a console UI for non-technical users and tools for managing flag lifecycle:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-json"&gt;{
  "new_recommendation_engine": {
    "enabled": false,
    "description": "ML-based recommendation engine v2",
    "rollout_percentage": 0
  },
  "enhanced_logging": {
    "enabled": true,
    "description": "Structured debug logging"
  }
}&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h2 id="safe-deployments-with-deployment-strategies"&gt;Safe deployments with deployment strategies&lt;/h2&gt;
&lt;p&gt;One of the most valuable features of AWS AppConfig for production environments is controlled deployments. Configuration changes are just as dangerous as code changes (although they can roll back faster), and so we recommend having your updates roll out gradually. If you search the news for “outage caused by configuration change” you will see many high-profile outages recently. Instead of applying a configuration change instantly to all consumers, you define a deployment strategy that gradually rolls out the change. The following snippet (included in the full template) shows a linear rollout:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-yaml"&gt;FeatureToggleDeploymentStrategy:
  Type: AWS::AppConfig::DeploymentStrategy
  Properties:
    Name: gradual-rollout
    DeploymentDurationInMinutes: 10
    GrowthFactor: 20
    GrowthType: LINEAR
    FinalBakeTimeInMinutes: 5
    ReplicateTo: NONE&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;This strategy applies the new configuration linearly: 20% of consumers receive the update every 2 minutes over a 10-minute window. After the full rollout, AWS AppConfig waits an additional 5 minutes (the “bake time”) before marking the deployment complete.&lt;/p&gt;
&lt;p&gt;During this window, you can integrate a CloudWatch alarm (or other APMs, like &lt;a href="https://github.com/aws-samples/aws-appconfig-tick-extn-for-datadog" target="_blank" rel="noopener"&gt;Datadog&lt;/a&gt;, &lt;a href="https://aws.amazon.com/about-aws/whats-new/2026/02/aws-appconfig-new-relic-for-automated-rollback/" target="https://aws.amazon.com/about-aws/whats-new/2026/02/aws-appconfig-new-relic-for-automated-rollback/" rel="noopener"&gt;New Relic&lt;/a&gt;, &lt;a href="https://docs.splunk.com/observability/en/gdi/integrations/cloud-aws.html" target="_blank" rel="noopener"&gt;Splunk&lt;/a&gt;, or &lt;a href="https://docs.dynatrace.com/docs/setup-and-configuration/setup-on-cloud-platforms/amazon-web-services" target="_blank" rel="noopener"&gt;Dynatrace&lt;/a&gt;) that monitors your application’s error rate or latency. If the alarm enters ALARM state, AWS AppConfig automatically rolls back to the previous configuration version. The companion repository includes a complete CloudWatch alarm example wired to the deployment.&lt;/p&gt;
&lt;h2 id="updating-feature-flags-without-code-deployments"&gt;Updating feature flags without code deployments&lt;/h2&gt;
&lt;p&gt;After your stack is deployed, you can update any feature flag by creating a new configuration version and starting a deployment:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws appconfig create-hosted-configuration-version \
  --application-id &amp;lt;APP_ID&amp;gt; \
  --configuration-profile-id &amp;lt;PROFILE_ID&amp;gt; \
  --content-type "application/json" \
  --content '{"new_recommendation_engine":{"enabled":true},"enhanced_logging":{"enabled":true}}'

aws appconfig start-deployment \
  --application-id &amp;lt;APP_ID&amp;gt; \
  --environment-id &amp;lt;ENV_ID&amp;gt; \
  --deployment-strategy-id &amp;lt;STRATEGY_ID&amp;gt; \
  --configuration-profile-id &amp;lt;PROFILE_ID&amp;gt; \
  --configuration-version &amp;lt;VERSION&amp;gt;&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Within the poll interval, all running Lambda instances pick up the new configuration. No code changes, no redeployment, no downtime. Reverting a flag is equally fast and symmetric. Deploying the previous configuration version propagates in the same ~30 seconds, giving you a consistent rollback speed whether you are enabling or disabling a feature. Importantly, the API contract (response structure, status codes, error shapes) remains stable regardless of flag state. Only the behavior behind the toggle changes, so consumers of your API are never broken by a flag flip.&lt;/p&gt;
&lt;h2 id="best-practices"&gt;Best practices&lt;/h2&gt;
&lt;p&gt;The &lt;strong&gt;AWS AppConfig Agent Lambda extension&lt;/strong&gt; may add time to your function’s &lt;code&gt;Init&lt;/code&gt; phase as it establishes a session and retrieves the initial configuration. On subsequent invocations, the extension serves from its &lt;strong&gt;local cache&lt;/strong&gt; with sub-millisecond latency. If your function has a strict cold start target, consider provisioned concurrency for latency-critical paths.&lt;/p&gt;
&lt;p&gt;The extension’s &lt;strong&gt;poll interval&lt;/strong&gt; determines how quickly your fleet converges on a new configuration. The template configures 30 seconds (the AWS default is 45 seconds). This interval suits most rollouts. For emergency disable switches, reduce it to 15 seconds (do not go below 5 seconds) via the &lt;code&gt;AWS_APPCONFIG_EXTENSION_POLL_INTERVAL_SECONDS&lt;/code&gt; environment variable so all instances converge within one cycle. The extension is also &lt;strong&gt;resilient to network failures&lt;/strong&gt;. If it cannot reach AWS AppConfig, it continues serving the last known good configuration from cache. Your function never fails because of an upstream connectivity issue.&lt;/p&gt;
&lt;p&gt;Use the &lt;code&gt;AWS_APPCONFIG_EXTENSION_PREFETCH_LIST&lt;/code&gt; environment variable so that configuration data is available before your function code runs. This retrieves config data during the &lt;code&gt;Init&lt;/code&gt; phase before the Lambda starts to execute the function code, reducing latency on the first invocation. See the &lt;a href="https://docs.aws.amazon.com/appconfig/latest/userguide/appconfig-integration-lambda-extensions-config.html" target="_blank" rel="noopener"&gt;AWS AppConfig Lambda extension configuration reference&lt;/a&gt; for details.&lt;/p&gt;
&lt;p&gt;Use the AppConfig first-class “feature-flag” configuration profile type with its opinionated JSON format. This data type gives you a simple console experience for non-technical users, advanced multi-variate flags, and tools for cleaning up stale feature flags. Treat toggles as &lt;strong&gt;temporary by nature&lt;/strong&gt;: after a feature is stable, remove the flag and its conditional logic to prevent dead-code sprawl. And scope your &lt;strong&gt;&lt;a href="https://aws.amazon.com/iam/" target="_blank" rel="noopener"&gt;AWS Identity and Access Management (IAM)&lt;/a&gt; permissions&lt;/strong&gt; so the extension is strictly a read-only consumer. Grant only &lt;code&gt;appconfig:StartConfigurationSession&lt;/code&gt; and &lt;code&gt;appconfig:GetLatestConfiguration&lt;/code&gt; on the specific resource ARN, ensuring a compromised function cannot modify configurations.&lt;/p&gt;
&lt;h2 id="clean-up"&gt;Clean up&lt;/h2&gt;
&lt;p&gt;To avoid ongoing charges, delete the resources you created in this walkthrough. Run the following command from the project directory:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;sam delete --stack-name &amp;lt;your-stack-name&amp;gt;&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;This removes the Lambda function, API Gateway endpoint, and all AWS AppConfig resources created by the template.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;The AWS AppConfig Lambda extension provides a lightweight, managed approach to feature flags in serverless applications. The extension handles caching, polling, and session management, while AWS AppConfig provides safe deployment strategies with validation and automatic rollback.&lt;/p&gt;
&lt;p&gt;Compared to building your own feature flag infrastructure or using environment variables, this approach eliminates redeployment overhead, reduces latency (sub-millisecond reads from local cache), and provides production safety mechanisms out of the box. Your function code stays simple: a single HTTP GET to a local endpoint.&lt;/p&gt;
&lt;p&gt;The pattern shown in this post applies beyond simple boolean flags. You can store complex configuration objects, percentage-based rollout rules, or user-segment targeting data in the same configuration profile. As your feature management needs grow, AWS AppConfig scales with you without requiring changes to the Lambda function integration pattern.&lt;/p&gt;
&lt;p&gt;With feature flags in place, you also have the foundation for &lt;a href="https://docs.aws.amazon.com/appconfig/latest/userguide/appconfig-experimentation.html" target="_blank" rel="noopener"&gt;AWS AppConfig Experimentation&lt;/a&gt;. From here you can define multi-variate experiments, allocate traffic to variants, and measure outcomes across your full stack, turning the feature flags you built in this post into a controlled experiment.&lt;/p&gt;
&lt;p&gt;This combination enables you to ship features faster with confidence, respond to incidents by disabling features in seconds, and experiment with gradual rollouts without any infrastructure overhead.&lt;/p&gt;
&lt;p&gt;You can find the complete source code in the &lt;a href="https://github.com/aws-samples/sample-lambda-extensions-appconfig-feature-toggles" target="_blank" rel="noopener"&gt;GitHub repository&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;If you have questions or feedback about this solution, leave a comment on this post.&lt;/p&gt;
&lt;p&gt;For more information, see:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/appconfig/latest/userguide/appconfig-integration-lambda-extensions.html" target="_blank" rel="noopener"&gt;Using AWS AppConfig Agent with AWS Lambda&lt;/a&gt;&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/appconfig/latest/userguide/appconfig-creating-deployment-strategy.html" target="_blank" rel="noopener"&gt;AWS AppConfig deployment strategies&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For more serverless learning resources, visit &lt;a href="https://serverlessland.com" target="_blank" rel="noopener"&gt;Serverless Land&lt;/a&gt;.&lt;/p&gt;</content:encoded>
					
		
		
			</item>
	</channel>
</rss>