<?xml version="1.0" encoding="UTF-8" standalone="no"?><rss xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:slash="http://purl.org/rss/1.0/modules/slash/" xmlns:sy="http://purl.org/rss/1.0/modules/syndication/" xmlns:wfw="http://wellformedweb.org/CommentAPI/" version="2.0">

<channel>
	<title>AWS Big Data Blog</title>
	<atom:link href="https://aws.amazon.com/blogs/big-data/feed/" rel="self" type="application/rss+xml"/>
	<link>https://aws.amazon.com/blogs/big-data/</link>
	<description>Official Big Data Blog of Amazon Web Services</description>
	<lastBuildDate>Thu, 17 Sep 2026 16:52:30 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	
	<item>
		<title>Discover and govern Snowflake data using SageMaker Unified Studio</title>
		<link>https://aws.amazon.com/blogs/big-data/discover-and-govern-snowflake-data-using-sagemaker-unified-studio/</link>
		
		<dc:creator><![CDATA[Marco Duarte]]></dc:creator>
		<pubDate>Wed, 16 Sep 2026 15:19:51 +0000</pubDate>
				<category><![CDATA[Amazon SageMaker Unified Studio]]></category>
		<category><![CDATA[Expert (400)]]></category>
		<category><![CDATA[Technical How-to]]></category>
		<guid isPermaLink="false">71973a3376f25fe0b8a1b95acb740ff9f20bf52e</guid>

					<description>Connect Snowflake to Amazon SageMaker Unified Studio to build a unified data catalog. Query federated Snowflake tables without moving data, publish enriched assets to SageMaker Catalog, and validate data quality with AWS Glue Data Quality, all while keeping data in Snowflake.</description>
										<content:encoded>&lt;p&gt;Many organizations operate in hybrid data environments where critical assets live in &lt;a href="https://www.snowflake.com/en/" target="_blank" rel="noopener"&gt;Snowflake&lt;/a&gt; while analytics workloads run on AWS, which can create governance gaps, discovery friction, and duplicated efforts when the two aren’t connected.&lt;/p&gt;
&lt;p&gt;With &lt;a href="https://aws.amazon.com/es/sagemaker/unified-studio/" target="_blank" rel="noopener"&gt;Amazon SageMaker Unified Studio&lt;/a&gt;, you can govern data across Snowflake and AWS through its &lt;a href="https://aws.amazon.com/es/sagemaker/catalog/" target="_blank" rel="noopener"&gt;integrated catalog&lt;/a&gt; and &lt;a href="https://aws.amazon.com/es/glue/features/data-quality/" target="_blank" rel="noopener"&gt;AWS Glue Data Quality&lt;/a&gt;, a capability of AWS Glue. You connect directly to Snowflake tables without moving data, apply quality rules using &lt;a href="https://docs.aws.amazon.com/glue/latest/dg/author-job-glue.html" target="_blank" rel="noopener"&gt;AWS Glue Visual ETL&lt;/a&gt;, and publish validated assets to Amazon SageMaker Catalog, maintaining consistent governance across your entire distributed data estate.&lt;/p&gt;
&lt;p&gt;Without this integration, cataloging Snowflake data requires building extraction pipelines, often taking days. With SageMaker Unified Studio connected to Snowflake, you can query, catalog, and validate the quality of federated data in 5–15 minutes. No data replication or custom ETL code required.&lt;/p&gt;
&lt;p&gt;In this post, we show you how to connect Snowflake to Amazon SageMaker Unified Studio, register data assets in Amazon SageMaker Catalog, configure data quality validation using AWS Glue Visual ETL, and publish assets for unified collaboration. By following these steps, you enrich federated assets with data quality scores so that consumers across your organization can discover and trust the data, all while keeping it in Snowflake.&lt;/p&gt;
&lt;h2 id="solution-overview"&gt;Solution overview&lt;/h2&gt;
&lt;p&gt;This solution integrates Snowflake with Amazon SageMaker Unified Studio for centralized data cataloging and quality validation.&lt;/p&gt;
&lt;p&gt;The architecture uses an AWS Glue connection to federate the Snowflake catalog into Amazon SageMaker Unified Studio. Tables become available in the project catalog without complex storage configurations. You can query data directly using SQL analytics, publish datasets to Amazon SageMaker Catalog for organization-wide discovery, and apply data quality rules through AWS Glue Visual ETL pipelines.&lt;/p&gt;
&lt;p&gt;The workflow consists of the following steps:&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-1.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-1.png" alt="Architecture diagram: Snowflake federated into SageMaker Unified Studio through AWS Glue, with data quality validation and publishing to SageMaker Catalog" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 1: Architecture for federating Snowflake into SageMaker Unified Studio and validating data quality&lt;/p&gt;
&lt;/div&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;&lt;strong&gt;Snowflake connection creation on Amazon SageMaker Unified Studio&lt;/strong&gt; — Amazon SageMaker Unified Studio uses an AWS Glue connection to federate Snowflake tables and views into its open data lakehouse architecture. The federated catalog entry is registered in AWS Glue Data Catalog and governed by AWS Lake Formation for centralized access control, without moving data out of Snowflake.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Federate Snowflake tables into the Amazon SageMaker publisher project&lt;/strong&gt; — The Amazon SageMaker publisher project discovers the federated Snowflake tables through the AWS Glue Data Catalog integration.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Publish the dataset to Amazon SageMaker Catalog&lt;/strong&gt; — The publisher project publishes the dataset as a governed asset to the Amazon SageMaker Catalog, making it discoverable for data consumers across the organization.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Validate data quality&lt;/strong&gt; — AWS Glue Data Quality runs validation rules against the federated Snowflake data and publishes the data quality results directly to the corresponding asset in Amazon SageMaker Catalog.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Consume data&lt;/strong&gt; — Users access Snowflake data through two paths:
  &lt;ol type="a"&gt;
   &lt;li&gt;&lt;strong&gt;Publisher project users — Query data with SQL Analytics&lt;/strong&gt; — Users in the publisher project can query the Snowflake data directly using Amazon SageMaker Unified Studio SQL Analytics for interactive exploration and analysis, without copying or moving data.&lt;/li&gt;
   &lt;li&gt;&lt;strong&gt;Consumer project users — Discovery and subscription through SageMaker Catalog&lt;/strong&gt; — Other Amazon SageMaker consumer projects discover the published asset in the Amazon SageMaker Catalog, subscribe to it, and consume the data for their analytics and machine learning workloads.&lt;/li&gt;
  &lt;/ol&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="prerequisites"&gt;Prerequisites&lt;/h2&gt;
&lt;p&gt;To follow along, you need:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;An active Snowflake account with administrator access.&lt;/li&gt;
 &lt;li&gt;Tables or views created within a schema inside a Snowflake database.&lt;/li&gt;
 &lt;li&gt;An &lt;a href="https://docs.aws.amazon.com/sagemaker-unified-studio/latest/userguide/gs-admin-setup.html" target="_blank" rel="noopener"&gt;Amazon SageMaker Unified Studio&lt;/a&gt; and &lt;a href="https://docs.aws.amazon.com/sagemaker-unified-studio/latest/userguide/getting-started-create-a-project.html" target="_blank" rel="noopener"&gt;project&lt;/a&gt; created.&lt;/li&gt;
 &lt;li&gt;An &lt;a href="https://aws.amazon.com/es/pm/serv-s3/?trk=32b96db8-849f-4365-bfe1-053cf18c09ae&amp;amp;sc_channel=ps&amp;amp;ef_id=Cj0KCQjwsMLSBhD9ARIsAIpUTDo4kDFhXMbKINF4SZHxSxM5PJehVEIcVzY3Jh-DMxuGY2pZ9jxZ-LwaAnUmEALw_wcB&amp;amp;gads_camp=23528573681&amp;amp;gads_ag=194200783593&amp;amp;gads_ad=795811760298&amp;amp;gads_kw=amazon%20s3&amp;amp;gads_matchtype=e&amp;amp;gads_network=g&amp;amp;gads_device=c&amp;amp;gads_geo=9048048&amp;amp;gad_campaignid=23528573681&amp;amp;gbraid=0AAAAADjHtp_oyBpoylLnCiHrquENE8wQx&amp;amp;gclid=Cj0KCQjwsMLSBhD9ARIsAIpUTDo4kDFhXMbKINF4SZHxSxM5PJehVEIcVzY3Jh-DMxuGY2pZ9jxZ-LwaAnUmEALw_wcB" target="_blank" rel="noopener"&gt;Amazon Simple Storage Service (Amazon S3) bucket&lt;/a&gt; for AWS Glue assets.&lt;/li&gt;
 &lt;li&gt;Appropriate AWS Identity and Access Management (IAM) permissions configured (Amazon SageMaker Catalog is built on Amazon DataZone, so the IAM actions use the &lt;code&gt;datazone:&lt;/code&gt; prefix.)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Your AWS Glue job execution role requires specific permissions to interact with Amazon SageMaker Catalog.&lt;/p&gt;
&lt;h3 id="required-iam-policies-for-the-aws-glue-job-role"&gt;Required IAM policies for the AWS Glue job role&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;1. Amazon SageMaker Catalog search and listing permissions:&lt;/strong&gt; Attach a policy that allows the AWS Glue job to search and list assets in Amazon SageMaker Catalog.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-json"&gt;{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Action": [
        "datazone:SearchListings",
        "datazone:GetListing",
        "datazone:ListDomains",
        "datazone:GetDomain"
      ],
      "Resource": "arn:aws:datazone:&amp;lt;REGION&amp;gt;:&amp;lt;ACCOUNT_ID&amp;gt;:domain/&amp;lt;DOMAIN_ID&amp;gt;"
    }
  ]
}&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;2. Amazon SageMaker Catalog time series data posting permissions:&lt;/strong&gt; Add permissions to post data quality metrics:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-json"&gt;{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Action": [
        "datazone:PostTimeSeriesDataPoints",
        "datazone:GetAsset",
        "datazone:ListAssetRevisions"
      ],
      "Resource": "arn:aws:datazone:&amp;lt;REGION&amp;gt;:&amp;lt;ACCOUNT_ID&amp;gt;:domain/&amp;lt;DOMAIN_ID&amp;gt;"
    }
  ]
}&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h3 id="configure-the-aws-glue-job-role-as-a-sagemaker-domain-user"&gt;Configure the AWS Glue job role as an Amazon SageMaker domain user&lt;/h3&gt;
&lt;p&gt;Configure the IAM role used by your AWS Glue job as a domain user. In the Amazon SageMaker console, navigate to your domain, choose Access management, and add the AWS Glue job execution IAM role as a domain user.&lt;/p&gt;
&lt;h3 id="project-level-permissions"&gt;Project-level permissions&lt;/h3&gt;
&lt;p&gt;Add the AWS Glue job execution role as a project member with Owner permissions. Navigate to your project, go to &lt;strong&gt;Project settings&lt;/strong&gt; &amp;gt; &lt;strong&gt;Members&lt;/strong&gt;, and add the role.&lt;/p&gt;
&lt;p&gt;For more information about IAM roles for AWS Glue, see the &lt;a href="https://docs.aws.amazon.com/glue/latest/dg/security-iam.html" target="_blank" rel="noopener"&gt;AWS Glue security documentation&lt;/a&gt;. For Amazon SageMaker Unified Studio permissions, refer to the &lt;a href="https://docs.aws.amazon.com/sagemaker-unified-studio/latest/adminguide/security-authorization.html" target="_blank" rel="noopener"&gt;Amazon SageMaker Unified Studio administrator guide&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="querying-snowflake-datasets-from-sagemaker-unified-studio"&gt;Querying Snowflake datasets from Amazon SageMaker Unified Studio&lt;/h2&gt;
&lt;p&gt;The following sections walk you through connecting Snowflake to Amazon SageMaker Unified Studio and running data quality validation with results displayed in Amazon SageMaker Catalog.&lt;/p&gt;
&lt;h3 id="identifying-information-in-snowflake"&gt;Identifying information in Snowflake&lt;/h3&gt;
&lt;p&gt;First, gather your Snowflake connection details. You need a Snowflake account with tables or views created at the schema level within a database.&lt;/p&gt;
&lt;p&gt;To obtain Snowflake connection information:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;Navigate to your Snowflake environment and sign in with administrator credentials.
  &lt;br&gt;
  &lt;figure&gt;
   &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-2.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-2.png" alt="Snowflake sign-in screen for administrator credentials" width="800"&gt;&lt;/a&gt;
   &lt;figcaption aria-hidden="true"&gt;&lt;/figcaption&gt;
  &lt;/figure&gt;&lt;/li&gt;
 &lt;li&gt;Choose your user account and choose &lt;strong&gt;Connect a tool to Snowflake&lt;/strong&gt;.
  &lt;br&gt;
  &lt;figure&gt;
   &lt;img loading="lazy" class="alignnone size-full wp-image-94381" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/15/BFB-5898-IMG1.png" alt="" width="598" height="1220"&gt;
   &lt;figcaption aria-hidden="true"&gt;&lt;/figcaption&gt;
  &lt;/figure&gt;&lt;/li&gt;
 &lt;li&gt;Note the Account/Server URL displayed on the screen.&lt;/li&gt;
 &lt;li&gt;Choose the Config File tab, select values for Warehouse, Database, and Schema, and copy these values for use in the next section.
  &lt;br&gt;
  &lt;figure&gt;
   &lt;img loading="lazy" class="alignnone size-full wp-image-94382" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/15/BFB-5898-IMG2.png" alt="" width="904" height="642"&gt;
   &lt;figcaption aria-hidden="true"&gt;&lt;/figcaption&gt;
  &lt;/figure&gt;
  &lt;figure&gt;
   &lt;img loading="lazy" class="alignnone size-full wp-image-94383" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/15/BFB-5898-IMG3.png" alt="" width="904" height="652"&gt;
   &lt;figcaption aria-hidden="true"&gt;&lt;/figcaption&gt;
  &lt;/figure&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id="creating-the-connection-in-sagemaker-unified-studio"&gt;Creating the connection in Amazon SageMaker Unified Studio&lt;/h3&gt;
&lt;p&gt;The Add Connection feature stores Snowflake connectivity details including credentials, server, and database information. Amazon SageMaker Unified Studio uses this connection to federate the Snowflake catalog through AWS Glue, so you can query data within minutes of setup.&lt;/p&gt;
&lt;p&gt;You need an Amazon SageMaker Unified Studio domain and a project, which acts as a data producer project.&lt;/p&gt;
&lt;p&gt;To create the Snowflake connection:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;In your Amazon SageMaker Unified Studio project, go to &lt;strong&gt;Overview&lt;/strong&gt;.
  &lt;br&gt;
  &lt;figure&gt;
   &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-6.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-6.png" alt="SageMaker Unified Studio project Overview page" width="800"&gt;&lt;/a&gt;
   &lt;figcaption aria-hidden="true"&gt;&lt;/figcaption&gt;
  &lt;/figure&gt;&lt;/li&gt;
 &lt;li&gt;Choose &lt;strong&gt;Data&lt;/strong&gt;.
  &lt;br&gt;
  &lt;figure&gt;
   &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-7.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-7.png" alt="Data option in the SageMaker Unified Studio project navigation" width="492"&gt;&lt;/a&gt;
   &lt;figcaption aria-hidden="true"&gt;&lt;/figcaption&gt;
  &lt;/figure&gt;&lt;/li&gt;
 &lt;li&gt;Choose &lt;strong&gt;+ Add&lt;/strong&gt;, then choose &lt;strong&gt;Add Connection&lt;/strong&gt;.
  &lt;br&gt;
  &lt;figure&gt;
   &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-8.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-8.png" alt="Add menu in SageMaker Unified Studio with the Add Connection option" width="800"&gt;&lt;/a&gt;
   &lt;figcaption aria-hidden="true"&gt;&lt;/figcaption&gt;
  &lt;/figure&gt;
  &lt;figure&gt;
   &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-9.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-9.png" alt="Add Connection panel in SageMaker Unified Studio" width="800"&gt;&lt;/a&gt;
   &lt;figcaption aria-hidden="true"&gt;&lt;/figcaption&gt;
  &lt;/figure&gt;&lt;/li&gt;
 &lt;li&gt;Choose &lt;strong&gt;Next&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Select &lt;strong&gt;Snowflake&lt;/strong&gt; and choose &lt;strong&gt;Next&lt;/strong&gt;.
  &lt;br&gt;
  &lt;figure&gt;
   &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-10.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-10.png" alt="Connection type selection showing Snowflake in SageMaker Unified Studio" width="800"&gt;&lt;/a&gt;
   &lt;figcaption aria-hidden="true"&gt;&lt;/figcaption&gt;
  &lt;/figure&gt;&lt;/li&gt;
 &lt;li&gt;Complete the connection details:
  &lt;ul&gt;
   &lt;li&gt;&lt;strong&gt;Name:&lt;/strong&gt; &lt;code&gt;snowflake-connection&lt;/code&gt;.&lt;/li&gt;
   &lt;li&gt;&lt;strong&gt;Description (Optional):&lt;/strong&gt; Enter a description for your connection.&lt;/li&gt;
   &lt;li&gt;&lt;strong&gt;Host:&lt;/strong&gt; Your Snowflake account URL (for example, XXXXXXXXX-XXX000000.snowflakecomputing.com).&lt;/li&gt;
   &lt;li&gt;&lt;strong&gt;Port:&lt;/strong&gt; 443.&lt;/li&gt;
   &lt;li&gt;&lt;strong&gt;Database:&lt;/strong&gt; Your database name (for example, sm_demo).&lt;/li&gt;
   &lt;li&gt;&lt;strong&gt;Warehouse:&lt;/strong&gt; Your warehouse name (for example, COMPUTE_WH).&lt;/li&gt;
   &lt;li&gt;&lt;strong&gt;Schema:&lt;/strong&gt; Your schema name (for example, demo).&lt;/li&gt;
   &lt;li&gt;&lt;strong&gt;Additional Properties:&lt;/strong&gt;
    &lt;ul&gt;
     &lt;li&gt;&lt;strong&gt;Register in AWS Glue Data Catalog:&lt;/strong&gt; Turn on checkbox.&lt;/li&gt;
     &lt;li&gt;&lt;strong&gt;Case conflict handling:&lt;/strong&gt; Select the option based on Snowflake naming syntax.&lt;/li&gt;
    &lt;/ul&gt;&lt;/li&gt;
   &lt;li&gt;&lt;strong&gt;Authentication:&lt;/strong&gt;
    &lt;ul&gt;
     &lt;li&gt;&lt;strong&gt;Username:&lt;/strong&gt; Your Snowflake username.&lt;/li&gt;
     &lt;li&gt;&lt;strong&gt;Password:&lt;/strong&gt; Your Snowflake password.&lt;/li&gt;
    &lt;/ul&gt;&lt;/li&gt;
  &lt;/ul&gt;
  &lt;figure&gt;
   &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-11.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-11.png" alt="Snowflake connection details form with name, host, port, database, warehouse, and schema fields" width="800"&gt;&lt;/a&gt;
   &lt;figcaption aria-hidden="true"&gt;&lt;/figcaption&gt;
  &lt;/figure&gt;
  &lt;figure&gt;
   &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-12.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-12.png" alt="Connection form showing authentication and AWS Glue Data Catalog registration options" width="800"&gt;&lt;/a&gt;
   &lt;figcaption aria-hidden="true"&gt;&lt;/figcaption&gt;
  &lt;/figure&gt;&lt;/li&gt;
 &lt;li&gt;Choose &lt;strong&gt;Add Data&lt;/strong&gt;.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;After creating the connection, wait a few minutes for the federated connection to be established. Search within Amazon SageMaker Unified Studio for the database and created objects.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-13.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-13.png" alt="Federated Snowflake database and objects appearing in SageMaker Unified Studio search" width="800"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-14.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-14.png" alt="Federated Snowflake tables registered in the AWS Glue Data Catalog" width="800"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;With the Snowflake connection established and the federated tables registered in AWS Glue Catalog, you’re now ready to query Snowflake data directly from Amazon SageMaker Unified Studio, without moving or replicating any data.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-15.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-15.png" alt="Query results from a federated Snowflake table in the SageMaker Unified Studio query editor" width="800"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id="how-federated-queries-work"&gt;How federated queries work&lt;/h3&gt;
&lt;p&gt;When you run a query in the Amazon SageMaker Unified Studio query editor against a federated Snowflake table, Amazon Athena runs the request. Athena is the underlying query engine integrated into Amazon SageMaker Unified Studio. Athena reads the table definition from AWS Glue Catalog, connects to Snowflake through the established connection, and pushes the query down for execution. Athena returns results directly to the query editor while Snowflake processes the data in place, and only the query results travel across the connection. Amazon SageMaker Unified Studio doesn’t copy data to S3 or any intermediate storage.&lt;/p&gt;
&lt;p&gt;After you’ve validated that queries return the expected results, the next step is to publish this dataset to Amazon SageMaker Catalog, making it discoverable and shareable across your organization.&lt;/p&gt;
&lt;h2 id="publishing-snowflake-datasets-to-the-sagemaker-catalog"&gt;Publishing Snowflake datasets to the SageMaker Catalog&lt;/h2&gt;
&lt;p&gt;Now that your Snowflake connection is configured, you can publish your datasets to the Amazon SageMaker Catalog, making them discoverable and shareable across your organization.&lt;/p&gt;
&lt;h3 id="creating-data-assets-in-sagemaker-catalog"&gt;Creating data assets in SageMaker Catalog&lt;/h3&gt;
&lt;p&gt;Data assets in Amazon SageMaker Catalog are the cataloged representation of your data resources. They help teams discover, govern, and share data across your organization.&lt;/p&gt;
&lt;p&gt;In this section, you create a data asset associated with a Snowflake table. This process transforms a technical Snowflake table into a cataloged resource enriched with business metadata.&lt;/p&gt;
&lt;p&gt;To create a data source:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;In your Amazon SageMaker Unified Studio project, go to &lt;strong&gt;Manage&lt;/strong&gt;.
  &lt;br&gt;
  &lt;figure&gt;
   &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-16.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-16.png" alt="Manage tab in the SageMaker Unified Studio project" width="800"&gt;&lt;/a&gt;
   &lt;figcaption aria-hidden="true"&gt;&lt;/figcaption&gt;
  &lt;/figure&gt;&lt;/li&gt;
 &lt;li&gt;Choose &lt;strong&gt;Data Sources&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Choose &lt;strong&gt;Create Data Source&lt;/strong&gt;.
  &lt;br&gt;
  &lt;figure&gt;
   &lt;img loading="lazy" class="alignnone size-full wp-image-94384" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/15/BFB-5898-IMG4.png" alt="" width="904" height="568"&gt;
   &lt;figcaption aria-hidden="true"&gt;&lt;/figcaption&gt;
  &lt;/figure&gt;&lt;/li&gt;
 &lt;li&gt;Select the &lt;strong&gt;AWS Glue&lt;/strong&gt; option.
  &lt;br&gt;
  &lt;figure&gt;
   &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-18.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-18.png" alt="Data source type selection showing the AWS Glue option" width="800"&gt;&lt;/a&gt;
   &lt;figcaption aria-hidden="true"&gt;&lt;/figcaption&gt;
  &lt;/figure&gt;&lt;/li&gt;
 &lt;li&gt;Turn on the &lt;strong&gt;Import data lineage&lt;/strong&gt; checkbox and select the connection: &lt;strong&gt;project.default_lakehouse&lt;/strong&gt;.
  &lt;br&gt;
  &lt;figure&gt;
   &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-19.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-19.png" alt="Data source configuration with Import data lineage and the project.default_lakehouse connection selected" width="800"&gt;&lt;/a&gt;
   &lt;figcaption aria-hidden="true"&gt;&lt;/figcaption&gt;
  &lt;/figure&gt;&lt;/li&gt;
 &lt;li&gt;Complete the form and choose &lt;strong&gt;Next&lt;/strong&gt;:
  &lt;ul&gt;
   &lt;li&gt;&lt;strong&gt;Catalog:&lt;/strong&gt; Select &lt;strong&gt;Enter the catalog name&lt;/strong&gt; and enter &lt;code&gt;snowflake-connection&lt;/code&gt;.&lt;/li&gt;
   &lt;li&gt;&lt;strong&gt;Database name:&lt;/strong&gt; Enter your database name (for example, movies).&lt;/li&gt;
   &lt;li&gt;&lt;strong&gt;Table selection criteria:&lt;/strong&gt; Enter * for all tables in the database, or enter a specific table name.&lt;/li&gt;
  &lt;/ul&gt;
  &lt;figure&gt;
   &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-20.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-20.png" alt="Data source form showing catalog name, database name, and table selection criteria" width="800"&gt;&lt;/a&gt;
   &lt;figcaption aria-hidden="true"&gt;&lt;/figcaption&gt;
  &lt;/figure&gt;&lt;/li&gt;
 &lt;li&gt;Keep the default options and choose &lt;strong&gt;Next&lt;/strong&gt; until you reach the summary screen.
  &lt;br&gt;
  &lt;figure&gt;
   &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-21.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-21.png" alt="SageMaker Unified Studio data source configuration summary screen" width="800"&gt;&lt;/a&gt;
   &lt;figcaption aria-hidden="true"&gt;&lt;/figcaption&gt;
  &lt;/figure&gt;
  &lt;figure&gt;
   &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-22.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-22.png" alt="Data source review screen before creation" width="800"&gt;&lt;/a&gt;
   &lt;figcaption aria-hidden="true"&gt;&lt;/figcaption&gt;
  &lt;/figure&gt;&lt;/li&gt;
 &lt;li&gt;Review your settings and choose &lt;strong&gt;Create&lt;/strong&gt;.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;To extract metadata and publish assets:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;Choose &lt;strong&gt;Run&lt;/strong&gt; to start extracting metadata from AWS Glue Data Catalog.
  &lt;br&gt;
  &lt;figure&gt;
   &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-23.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-23.png" alt="Data source detail page with the Run option to extract metadata from the AWS Glue Data Catalog" width="800"&gt;&lt;/a&gt;
   &lt;figcaption aria-hidden="true"&gt;&lt;/figcaption&gt;
  &lt;/figure&gt;&lt;/li&gt;
 &lt;li&gt;Wait for the run to complete.&lt;/li&gt;
 &lt;li&gt;Go to &lt;strong&gt;Assets&lt;/strong&gt; to view the Asset Inventory.
  &lt;br&gt;
  &lt;figure&gt;
   &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-24.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-24.png" alt="Asset inventory in SageMaker Catalog after the data source run completes" width="800"&gt;&lt;/a&gt;
   &lt;figcaption aria-hidden="true"&gt;&lt;/figcaption&gt;
  &lt;/figure&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The following screenshot shows the asset inventory after the data source run completes.&lt;/p&gt;
&lt;ol start="4" type="1"&gt;
 &lt;li&gt;Choose an asset to view its details.
  &lt;br&gt;
  &lt;figure&gt;
   &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-25.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-25.png" alt="Asset detail page in SageMaker Catalog showing the Snowflake table metadata" width="800"&gt;&lt;/a&gt;
   &lt;figcaption aria-hidden="true"&gt;&lt;/figcaption&gt;
  &lt;/figure&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;At this point, you can enrich the business context by choosing Generate Descriptions. Amazon SageMaker Catalog analyzes the asset’s technical structure and generate:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;Business descriptions in natural language for the asset.&lt;/li&gt;
 &lt;li&gt;Contextual definitions for each field/column.&lt;/li&gt;
 &lt;li&gt;Suggested glossary terms that could be applied.&lt;/li&gt;
&lt;/ul&gt;
&lt;ol start="5" type="1"&gt;
 &lt;li&gt;After your asset has been enriched with the necessary business metadata, you can publish it to the Amazon SageMaker Catalog by choosing &lt;strong&gt;Publish Asset&lt;/strong&gt;.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-26.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-26.png" alt="Publish Asset option on the enriched Snowflake asset in SageMaker Catalog" width="800"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;The Snowflake enriched asset is now available to data consumers across your organization. Other users can discover it, subscribe to it, and consume it without data replication.&lt;/p&gt;
&lt;h2 id="implementing-data-quality-rules-with-aws-glue-data-quality"&gt;Implementing data quality rules with AWS Glue Data Quality&lt;/h2&gt;
&lt;p&gt;This section explains how to apply data quality validations to Snowflake data using AWS Glue Data Quality and visualize results in Amazon SageMaker Catalog.&lt;/p&gt;
&lt;h3 id="setting-up-the-custom-transform"&gt;Setting up the custom transform&lt;/h3&gt;
&lt;p&gt;Upload two files to an Amazon S3 bucket in the same AWS account where you run AWS Glue:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;a href="https://github.com/aws-samples/aws-glue-samples/tree/master/examples/transforms/GlueDQResultToDataZonePublisher" target="_blank" rel="noopener"&gt;&lt;strong&gt;post_dq_results_to_datazone.py&lt;/strong&gt;&lt;/a&gt; – Contains the transform logic that reads data quality rule outcomes and posts them to SageMaker Catalog.&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://github.com/aws-samples/aws-glue-samples/tree/master/examples/transforms/GlueDQResultToDataZonePublisher" target="_blank" rel="noopener"&gt;&lt;strong&gt;post_dq_results_to_datazone.json&lt;/strong&gt;&lt;/a&gt; – Defines the transform configuration that AWS Glue Studio uses to render the visual node.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Copy both files to your AWS Glue assets S3 bucket in the transforms folder (&lt;code&gt;s3://aws-glue-assets-&amp;lt;account-id&amp;gt;-&amp;lt;region&amp;gt;/transforms&lt;/code&gt;). AWS Glue Studio reads all JSON files from this folder to register custom visual transforms.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-27.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-27.png" alt="Custom transform files uploaded to the transforms folder in the AWS Glue assets S3 bucket" width="800"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;In the following sections, we walk you through the steps of building an ETL pipeline for data quality validation using AWS Glue Studio.&lt;/p&gt;
&lt;h3 id="creating-the-aws-glue-visual-etl-job"&gt;Creating the AWS Glue Visual ETL job&lt;/h3&gt;
&lt;p&gt;AWS Glue for Spark provides built-in support for reading from Snowflake data sources.&lt;/p&gt;
&lt;p&gt;To create a new visual ETL job:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;Open the AWS Glue console at https://console.aws.amazon.com/glue/. Choose &lt;strong&gt;ETL jobs&lt;/strong&gt;, then &lt;strong&gt;Visual ETL&lt;/strong&gt;.
  &lt;br&gt;
  &lt;figure&gt;
   &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-28.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-28.png" alt="AWS Glue console showing ETL jobs and the Visual ETL option" width="800"&gt;&lt;/a&gt;
   &lt;figcaption aria-hidden="true"&gt;AWS Glue console showing ETL jobs and the Visual ETL option&lt;/figcaption&gt;
  &lt;/figure&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id="establishing-the-snowflake-connection"&gt;Establishing the Snowflake connection&lt;/h3&gt;
&lt;p&gt;To add a Snowflake source:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;In the job pane, choose Snowflake as your source. For Snowflake connection, select the connection that you created earlier. Specify the relevant schema and table for data quality checks.
  &lt;br&gt;
  &lt;figure&gt;
   &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-29.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-29.png" alt="Snowflake source node configured in the AWS Glue visual ETL job" width="800"&gt;&lt;/a&gt;
   &lt;figcaption aria-hidden="true"&gt;&lt;/figcaption&gt;
  &lt;/figure&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The visual editor displays the Data source properties panel where you select your connection, database, and enter a custom query targeting your Snowflake table.&lt;/p&gt;
&lt;h3 id="applying-data-quality-rules"&gt;Applying data quality rules&lt;/h3&gt;
&lt;p&gt;After establishing the Snowflake connection, configure the data quality evaluation step using the Data Quality Definition Language (DQDL).&lt;/p&gt;
&lt;p&gt;To add data quality validation:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;Choose &lt;strong&gt;Transform&lt;/strong&gt; and choose &lt;strong&gt;Evaluate Data Quality&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Define domain-specific data quality rules using DQDL. For more information, see the &lt;a href="https://docs.aws.amazon.com/glue/latest/dg/dqdl.html" target="_blank" rel="noopener"&gt;AWS DQDL&lt;/a&gt; documentation.
  &lt;br&gt;
  &lt;figure&gt;
   &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-30.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-30.png" alt="Evaluate Data Quality transform with DQDL rules in AWS Glue Studio" width="800"&gt;&lt;/a&gt;
   &lt;figcaption aria-hidden="true"&gt;&lt;/figcaption&gt;
  &lt;/figure&gt;&lt;/li&gt;
 &lt;li&gt;Choose to output the data quality results. Optionally, store outcomes in Amazon S3 or publish to Amazon CloudWatch with alert notifications.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The preview of the data quality results from the ruleOutcomes node shows the outcomes of each rule.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-31.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-31.png" alt="Preview of the data quality rule outcomes from the ruleOutcomes node" width="800"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id="post-the-data-quality-results-to-sagemaker-catalog"&gt;Post the data quality results to Amazon SageMaker Catalog&lt;/h3&gt;
&lt;p&gt;To configure the custom transform:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;Add the &lt;strong&gt;Datazone DQ Result Sink&lt;/strong&gt; transform to your job.&lt;/li&gt;
 &lt;li&gt;Connect the &lt;strong&gt;ruleOutcomes&lt;/strong&gt; node output to this transform.&lt;/li&gt;
 &lt;li&gt;Complete the parameters:
  &lt;ul&gt;
   &lt;li&gt;&lt;strong&gt;Role to assume (Optional): Only needed for associated accounts&lt;/strong&gt;.&lt;/li&gt;
   &lt;li&gt;&lt;strong&gt;Domain ID: Your Amazon SageMaker Unified Studio domain ID (found in the Amazon SageMaker Unified Studio portal)&lt;/strong&gt;.&lt;/li&gt;
   &lt;li&gt;&lt;strong&gt;Table name and Schema name:&lt;/strong&gt; Same values used when creating the Snowflake source transform.&lt;/li&gt;
   &lt;li&gt;&lt;strong&gt;Data quality ruleset name:&lt;/strong&gt; The name you want to give to the ruleset in Amazon SageMaker Catalog.&lt;/li&gt;
   &lt;li&gt;&lt;strong&gt;Max results:&lt;/strong&gt; Maximum number of assets to return in case of multiple matches.&lt;/li&gt;
  &lt;/ul&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The following image shows the complete job graph with the Datazone DQ Result Sink transform configured.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-32.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-32.png" alt="AWS Glue visual ETL job graph with Snowflake source, Evaluate Data Quality, ruleOutcomes, and Datazone DQ Result Sink nodes" width="800"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;The visual editor displays four nodes connected sequentially: the Snowflake data source, the Evaluate Data Quality transform, the ruleOutcomes SelectFromCollection transform, and the Datazone DQ Result Sink transform.&lt;/p&gt;
&lt;p&gt;To configure job parameters:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;Choose &lt;strong&gt;Job details&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;In &lt;strong&gt;Job parameters&lt;/strong&gt;, add the following key-value pair:
  &lt;ul&gt;
   &lt;li&gt;&lt;code&gt;--additional-python-modules&lt;/code&gt;&lt;/li&gt;
   &lt;li&gt;&lt;code&gt;boto3&amp;gt;=1.34.105&lt;/code&gt;&lt;/li&gt;
  &lt;/ul&gt;&lt;/li&gt;
 &lt;li&gt;Save and run the job.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-33.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-33.png" alt="AWS Glue job parameters with the additional-python-modules key set to boto3" width="800"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id="visualizing-data-quality-results-in-the-sagemaker-catalog"&gt;Visualizing data quality results in the SageMaker Catalog&lt;/h2&gt;
&lt;p&gt;After the AWS Glue ETL job completes, you can view the data quality information directly in Amazon SageMaker Catalog. This is the key outcome of running data quality on a federated source: the asset gains quality scores and metadata without ever leaving Snowflake. This makes it trustworthy and ready for other teams across your organization to use. Data consumers can now discover this asset in Amazon SageMaker Catalog and evaluate its quality before subscribing, without needing direct access to Snowflake or running their own validation.&lt;/p&gt;
&lt;p&gt;To view data quality results:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;Open the Amazon SageMaker Unified Studio console.&lt;/li&gt;
 &lt;li&gt;Navigate to your project.&lt;/li&gt;
 &lt;li&gt;Go to &lt;strong&gt;Assets&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Choose the Snowflake data asset.&lt;/li&gt;
 &lt;li&gt;View the data quality information displayed on the asset page.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The following image shows the asset page in &lt;strong&gt;Amazon&lt;/strong&gt; &lt;strong&gt;SageMaker Catalog&lt;/strong&gt; with the &lt;strong&gt;data quality score&lt;/strong&gt; populated.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-34.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-34.png" alt="SageMaker Catalog asset page showing a populated data quality score for the Snowflake asset" width="800"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-35.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-35.png" alt="Data Quality tab in SageMaker Catalog showing an overall score of 100 with the movies rule set passed" width="800"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;The Data Quality tab shows an overall score of 100 and lists the rule set movies with a Passed result (1/1). This confirms that the data quality checks from AWS Glue posted successfully to Amazon SageMaker Catalog.&lt;/p&gt;
&lt;h2 id="clean-up"&gt;Clean up&lt;/h2&gt;
&lt;p&gt;To avoid ongoing charges, remove the resources you created during this walkthrough:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/cli/latest/reference/glue/delete-job.html" target="_blank" rel="noopener"&gt;&lt;strong&gt;Delete the AWS Glue ETL job&lt;/strong&gt;&lt;/a&gt; — Open the AWS Glue console, choose &lt;strong&gt;ETL jobs&lt;/strong&gt;, select your job, and then choose &lt;strong&gt;Delete&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/cli/latest/reference/glue/delete-connection.html" target="_blank" rel="noopener"&gt;&lt;strong&gt;Remove the AWS Glue connection&lt;/strong&gt;&lt;/a&gt; — In the AWS Glue console, go to &lt;strong&gt;Connections&lt;/strong&gt;, select the Snowflake connection, and then choose &lt;strong&gt;Delete&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/cli/latest/reference/datazone/delete-data-source.html" target="_blank" rel="noopener"&gt;&lt;strong&gt;Delete the data source in SageMaker Catalog&lt;/strong&gt;&lt;/a&gt; — In your Amazon SageMaker Unified Studio project, go to &lt;strong&gt;Data Sources&lt;/strong&gt;, select the data source you created, and then choose &lt;strong&gt;Delete&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/cli/latest/reference/s3api/delete-object.html" target="_blank" rel="noopener"&gt;&lt;strong&gt;Remove S3 assets&lt;/strong&gt;&lt;/a&gt; — Delete the custom transform files from your &lt;code&gt;s3://aws-glue-assets-&amp;lt;account-id&amp;gt;-&amp;lt;region&amp;gt;/transforms/&lt;/code&gt; bucket.&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/cli/latest/reference/iam/delete-policy.html" target="_blank" rel="noopener"&gt;&lt;strong&gt;Remove IAM policies&lt;/strong&gt;&lt;/a&gt; — Detach and delete the IAM policies you attached to the AWS Glue job execution role. Remove the role as a domain user and project member.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;In this post, we showed you how to connect Snowflake to Amazon SageMaker Unified Studio for centralized data cataloging and quality validation. This approach maintains consistent governance without replicating data. Key benefits include:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;Query without data movement:&lt;/strong&gt; Access Snowflake data directly from Amazon SageMaker Unified Studio through federated queries, using the interoperable data architecture of AWS and eliminating time-consuming data replication.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Centralized governance:&lt;/strong&gt; Maintain a single source of truth for data discovery, quality metrics, and governance policies across your distributed data estate.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Automated quality validation:&lt;/strong&gt; Apply consistent data quality rules using AWS Glue Data Quality and visualize results directly in Amazon SageMaker Catalog.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Unified collaboration:&lt;/strong&gt; Support data discovery and sharing across your organization through the publishing capabilities of Amazon SageMaker Catalog.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;To get started, open the &lt;a href="https://aws.amazon.com/sagemaker/unified-studio/" target="_blank" rel="noopener"&gt;Amazon SageMaker Unified Studio console&lt;/a&gt;. To learn more about related topics, see &lt;a href="https://aws.amazon.com/blogs/big-data/cross-account-lakehouse-governance-with-amazon-s3-tables-and-sagemaker-catalog/" target="_blank" rel="noopener"&gt;Cross-account lakehouse governance with Amazon S3 Tables and SageMaker Catalog&lt;/a&gt; and &lt;a href="https://aws.amazon.com/blogs/big-data/get-started-with-aws-glue-data-quality-dynamic-rules-for-etl-pipelines/" target="_blank" rel="noopener"&gt;Get started with AWS Glue Data Quality dynamic rules for ETL pipelines&lt;/a&gt;.&lt;/p&gt;
&lt;hr style="width: 100%"&gt;
&lt;h2&gt;About the authors&lt;/h2&gt;
&lt;footer&gt;
 &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
  &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
   &lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-36.jpg" alt="Marco Duarte López" width="100" height="133"&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Marco Duarte López&lt;/h3&gt;
  &lt;p style="overflow: hidden"&gt;Marco is a Data Specialist Solutions Architect at AWS, based in Santiago, Chile. He works with organizations across the region to design modern data architectures and governance frameworks that enable trusted, scalable data consumption. He is a member of the AWS Technical Field Community (TFC) for Analytics, where he specializes in Data &amp;amp; AI Governance, and has led data transformation programs for some of the largest enterprises in the region.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
  &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
   &lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/11/BDB-5898-37.jpg" alt="Diego Ortiz" width="100" height="133"&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Diego Ortiz&lt;/h3&gt;
  &lt;p style="overflow: hidden"&gt;Diego is a Senior Data Strategy Solutions Architect for Latin America based in San Juan, Puerto Rico, with 14+ years of experience in technology roles. He supports organizations across countries and industries to develop data and AI strategies aligned with their business objectives, combining strategic vision with deep technical expertise in data and AI technologies. He is a core member of the Data Governance global community at AWS and leads the analytics technical community in the Spanish-speaking countries of Latin America.&lt;/p&gt;
 &lt;/div&gt;
&lt;/footer&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>Connect Amazon SageMaker Unified Studio to Microsoft Power BI – Part 1: IAM Identity Center (IDC)-based domains</title>
		<link>https://aws.amazon.com/blogs/big-data/connect-amazon-sagemaker-unified-studio-to-microsoft-power-bi-part-1-iam-identity-center-idc-based-domains/</link>
		
		<dc:creator><![CDATA[Ramesh H Singh]]></dc:creator>
		<pubDate>Tue, 15 Sep 2026 20:55:26 +0000</pubDate>
				<category><![CDATA[Amazon Athena]]></category>
		<category><![CDATA[Amazon SageMaker Unified Studio]]></category>
		<category><![CDATA[Announcements]]></category>
		<category><![CDATA[Intermediate (200)]]></category>
		<guid isPermaLink="false">34ac67a5de3a5fc33313b741223701bcdaa35ba3</guid>

					<description>Connect Microsoft Power BI directly to governed data in Amazon SageMaker Unified Studio using new authentication modes in the Amazon Athena ODBC driver, with no third-party ODBC-JDBC bridge. Part 1 covers IAM Identity Center (IDC)-based domains with both DSN-based and DSN-less connection methods.</description>
										<content:encoded>&lt;p&gt;Connecting Power BI to your Amazon SageMaker Unified Studio data catalogs typically required third-party bridges. These bridges added complexity and licensing costs. In this post, you create a direct connection using new authentication modes in the Amazon Athena ODBC driver, removing those dependencies entirely. If your organization uses Power BI as its business intelligence (BI) tool, your analysts can configure access to governed data in Amazon SageMaker Unified Studio without changing their tools or workflows. As an AWS alternative, &lt;a href="https://docs.aws.amazon.com/sagemaker-unified-studio/latest/adminguide/amazon-quicksight.html" target="_blank" rel="noopener"&gt;Amazon Quick Sight provides serverless BI integration with Amazon SageMaker Unified Studio&lt;/a&gt; at pay-per-session pricing.&lt;/p&gt;
&lt;p&gt;A &lt;a href="https://aws.amazon.com/blogs/big-data/power-up-your-analytics-with-amazon-sagemaker-unified-studio-integration-with-tableau-power-bi-and-more/" target="_blank" rel="noopener"&gt;previous post&lt;/a&gt; showed the connection method using a third-party ODBC-JDBC bridge. The &lt;a href="https://docs.aws.amazon.com/athena/latest/ug/odbc-v2-driver.html" target="_blank" rel="noopener"&gt;Amazon Athena ODBC driver&lt;/a&gt; (version 2.2.0 and later) now supports Amazon SageMaker Unified Studio authentication directly, eliminating the need for customers to configure third-party bridge components previously required for this connection. This bridge also created additional components and required ongoing maintenance. The native connection simplifies the architecture by reducing these requirements.&lt;/p&gt;
&lt;h3&gt;Customer Spotlight&lt;/h3&gt;
&lt;p&gt;&lt;a href="https://uci.edu/" target="_blank" rel="noopener"&gt;UC Irvine&lt;/a&gt;, a top-ten U.S. public research university, consolidates student data from systems across multiple departments into a single governed repository that supports reporting, research, and analytics for decision-making at the strategic, tactical, and operational levels. Many of their analysts rely on Power BI to explore and visualize this governed data.&lt;/p&gt;
&lt;blockquote&gt;
 &lt;p&gt;&lt;em&gt;“Our users rely on Power BI for data visualization and reporting, but connecting to governed data in AWS previously required workarounds. The ODBC connection feature gives a direct path from Power BI into our SageMaker Unified Studio projects—no bridge software, no extra licensing, just a connection string and we’re ready to go.”&lt;/em&gt;&lt;/p&gt;
 &lt;p&gt;&lt;em&gt;— Bernadette Theologidy, Manager, Student Analytics, UC Irvine&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The Athena ODBC driver introduces two new authentication modes for SageMaker Unified Studio:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;&lt;strong&gt;SageMakerBrowserIdc&lt;/strong&gt; (for IDC-based domains): The driver opens a browser window and authenticates through AWS IAM Identity Center (and your &lt;a href="https://docs.aws.amazon.com/singlesignon/latest/userguide/manage-your-identity-source-idp.html" target="_blank" rel="noopener"&gt;external identity provider&lt;/a&gt;, if configured). No local AWS credentials are needed.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;SageMakerIam&lt;/strong&gt; (for AWS Identity and Access Management (IAM)-based and IDC-based domains): The driver uses AWS credentials from the &lt;a href="https://docs.aws.amazon.com/sdkref/latest/guide/standardized-credentials.html" target="_blank" rel="noopener"&gt;default credential provider chain&lt;/a&gt;. For this walkthrough, we use AWS IAM Identity Center to provide those credentials.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;You connect Microsoft Power BI to Amazon SageMaker Unified Studio through Athena. The Athena ODBC driver supports using two connection methods that use these authentication modes:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Method 1: DSN-based (Athena Power BI connector):&lt;/strong&gt; You configure an ODBC Data Source Name (DSN) and use the Athena connector in Power BI. This method supports DirectQuery and Import mode with both SageMakerBrowserIdc and SageMakerIam authentication.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Method 2: DSN-less (Power BI ODBC connector):&lt;/strong&gt; You use the Power BI ODBC connector with a connection string, requiring no DSN configuration. This method supports Import mode only with SageMakerIam authentication. DirectQuery isn’t available because the Power BI ODBC connector doesn’t support it. The connection string in Power BI Desktop must match exactly the one on Power BI Service. Because the gateway runs as a Windows service without interactive browser access, both ends must use SageMakerIam.&lt;/p&gt;
&lt;table border="1px" width="100%" cellpadding="10px"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;Feature&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Method 1: DSN-based&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Method 2: DSN-less&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Power BI Connector&lt;/td&gt;
   &lt;td&gt;Amazon Athena connector&lt;/td&gt;
   &lt;td&gt;ODBC connector&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Data connectivity mode&lt;/td&gt;
   &lt;td&gt;DirectQuery and Import&lt;/td&gt;
   &lt;td&gt;Import only&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Requires DSN configuration&lt;/td&gt;
   &lt;td&gt;Yes&lt;/td&gt;
   &lt;td&gt;No&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Data freshness&lt;/td&gt;
   &lt;td&gt;Real-time (DirectQuery) or scheduled (Import)&lt;/td&gt;
   &lt;td&gt;Scheduled refresh only&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Authentication types&lt;/td&gt;
   &lt;td&gt;SageMakerIam and SageMakerBrowserIdc&lt;/td&gt;
   &lt;td&gt;SageMakerIam only&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Domain types supported&lt;/td&gt;
   &lt;td&gt;IAM-based and IDC-based&lt;/td&gt;
   &lt;td&gt;IAM-based and IDC-based&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Best for&lt;/td&gt;
   &lt;td&gt;Dashboards requiring live data&lt;/td&gt;
   &lt;td&gt;Scenarios where DSN management is not possible or scheduled refresh is acceptable&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;This is Part 1 of a two-part series. This post covers IDC-based domains using both connection methods. &lt;a href="https://aws.amazon.com/blogs/big-data/connect-amazon-sagemaker-unified-studio-to-microsoft-power-bi-part-2-iam-based-domains/" target="_blank" rel="noopener"&gt;Part 2&lt;/a&gt; covers IAM-based domains.&lt;/p&gt;
&lt;h2 id="solution-overview"&gt;Solution overview&lt;/h2&gt;
&lt;p&gt;In this walkthrough, you take the role of a data analyst at an energy company. You need to understand the current state and future direction of the U.S. power generation fleet using the &lt;a href="https://registry.opendata.aws/catalyst-cooperative-pudl/" target="_blank" rel="noopener"&gt;Public Utility Data Liberation Project&lt;/a&gt;, available on the Registry of Open Data on AWS. Our goal is to analyze generation capacity and identify where new investment is flowing. We connect Power BI to Athena through Amazon SageMaker Unified Studio and query the EIA-860 generators dataset directly from our data catalog. The result is a single visualization that reveals the energy transition.&lt;/p&gt;
&lt;p&gt;The following diagram illustrates the solution architecture for connecting Power BI to Amazon SageMaker Unified Studio through Amazon Athena.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P1-1.jpg" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P1-1.jpg" alt="Architecture diagram showing Power BI connecting to Amazon Athena through Amazon SageMaker Unified Studio, with a Microsoft on-premises data gateway on Amazon EC2" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 1: Architecture diagram&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;The following architecture demonstrates a six-step workflow.&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;Data engineers and analysts connect Power BI Desktop to Athena as a data source.&lt;/li&gt;
 &lt;li&gt;They build their reports locally.&lt;/li&gt;
 &lt;li&gt;They then publish them to the Power BI Service.&lt;/li&gt;
 &lt;li&gt;Microsoft On-Premises Data Gateway on an Amazon Elastic Compute Cloud (Amazon EC2) instance connects to Athena using the instance’s attached IAM role.&lt;/li&gt;
 &lt;li&gt;The Power BI Service then uses this gateway connection.&lt;/li&gt;
 &lt;li&gt;Report viewers access the published reports through Power BI Service to make data-driven decisions.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;On the AWS side, Athena queries the data catalog managed by AWS Glue Data Catalog. The catalog references data stored in Amazon Simple Storage Service (Amazon S3). An Amazon SageMaker Unified Studio project governs all access.&lt;/p&gt;
&lt;p&gt;In an IDC-based domain (covered in this post), Power BI Desktop uses SageMakerBrowserIdc for Method 1 and SageMakerIam for Method 2. Power BI Desktop can run on-premises or on an EC2 instance. The gateway always uses SageMakerIam (it runs as a Windows service without browser access) and authenticates using instance profile credentials, which rotate automatically. The gateway can only query data within projects where its IAM role has been added as a member. For IAM-based domains, see &lt;a href="https://aws.amazon.com/blogs/big-data/connect-amazon-sagemaker-unified-studio-to-microsoft-power-bi-part-2-iam-based-domains/" target="_blank" rel="noopener"&gt;Part 2&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="prerequisites"&gt;Prerequisites&lt;/h2&gt;
&lt;p&gt;Before connecting Power BI to Amazon SageMaker Unified Studio, verify that your environment meets these requirements:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;Athena ODBC driver&lt;/strong&gt; – The latest &lt;a href="https://docs.aws.amazon.com/athena/latest/ug/odbc-v2-driver.html" target="_blank" rel="noopener"&gt;Amazon Athena ODBC driver&lt;/a&gt; (version 2.2.0 or more recent) for Windows 64-bit.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Microsoft Power BI Desktop&lt;/strong&gt; – The latest version installed on your Windows machine.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Microsoft Power BI Pro License&lt;/strong&gt; – Required for publishing reports and configuring the on-premises data gateway.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Microsoft Power BI on-premises data gateway&lt;/strong&gt; – The latest version installed on the EC2 instance.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Amazon SageMaker Unified Studio&lt;/strong&gt; – An Amazon SageMaker Unified Studio IDC-based domain.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;You need an Amazon SageMaker Unified Studio project with data assets. For detailed instructions, refer to the &lt;a href="https://docs.aws.amazon.com/sagemaker-unified-studio/latest/userguide/" target="_blank" rel="noopener"&gt;Amazon SageMaker Unified Studio User Guide&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The following screenshot shows the Amazon SageMaker Unified Studio project Query Editor interface, which runs a preview query against the EIA-860 generators dataset.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P1-2.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P1-2.png" alt="SageMaker Unified Studio Query Editor previewing the EIA-860 generators dataset" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 2: SageMaker Unified Studio project with the EIA-860 generators dataset available in the data catalog&lt;/p&gt;
&lt;/div&gt;
&lt;h2 id="method-1-dsn-based-connection-athena-power-bi-connector"&gt;Method 1: DSN-based connection (Athena Power BI connector)&lt;/h2&gt;
&lt;p&gt;This method uses the Amazon Athena Power BI connector with an ODBC Data Source Name (DSN), supporting DirectQuery and Import mode.&lt;/p&gt;
&lt;p&gt;You configure Power BI Desktop to connect to your data assets in Amazon SageMaker Unified Studio using the SageMakerBrowserIdc authentication mode. The driver opens a browser window and authenticates through IAM Identity Center (and your &lt;a href="https://docs.aws.amazon.com/singlesignon/latest/userguide/manage-your-identity-source-idp.html" target="_blank" rel="noopener"&gt;external identity provider&lt;/a&gt;, if configured).&lt;/p&gt;
&lt;h3 id="add-your-sso-user-as-a-member-of-your-sagemaker-unified-studio-project"&gt;Add your SSO user as a member of your SageMaker Unified Studio project&lt;/h3&gt;
&lt;p&gt;Your single sign-on (SSO) user needs project-level access to query data with Athena. Verify your user is listed as a project member or add it by following &lt;a href="https://docs.aws.amazon.com/sagemaker-unified-studio/latest/userguide/add-project-members.html" target="_blank" rel="noopener"&gt;Add project members&lt;/a&gt; in the Amazon SageMaker Unified Studio User Guide.&lt;/p&gt;
&lt;p&gt;The following screenshot shows the SageMaker Unified Studio project user management page, where project owners can add or remove project users and roles.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P1-3.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P1-3.png" alt="SageMaker Unified Studio project members page listing users and roles" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 3: Members of a SageMaker Unified Studio project&lt;/p&gt;
&lt;/div&gt;
&lt;h3 id="gather-configuration-values-to-configure-your-amazon-athena-odbc-dsn"&gt;Gather configuration values to configure your Amazon Athena ODBC DSN&lt;/h3&gt;
&lt;p&gt;Gather the following values from your Amazon SageMaker Unified Studio project:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;Open your Amazon SageMaker Unified Studio project.&lt;/li&gt;
 &lt;li&gt;In the top right, select the three dots.&lt;/li&gt;
 &lt;li&gt;Choose &lt;strong&gt;Project details&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Select &lt;strong&gt;JDBC and ODBC details&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Under &lt;strong&gt;ODBC connection details&lt;/strong&gt; copy the following information: IDC issuer URL, domain ID, project ID, Athena workgroup name and AWS Region.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The following screenshot shows the Amazon SageMaker Unified Studio project overview page, where you can copy these details.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P1-4.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P1-4.png" alt="SageMaker Unified Studio project overview showing ODBC connection details" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 4: ODBC connection details&lt;/p&gt;
&lt;/div&gt;
&lt;h3 id="configure-the-odbc-dsn"&gt;Configure the ODBC DSN&lt;/h3&gt;
&lt;p&gt;Create a System DSN using the Amazon Athena ODBC driver. For the general DSN creation steps, see &lt;a href="https://docs.aws.amazon.com/athena/latest/ug/odbc-v2-driver-getting-started-windows.html" target="_blank" rel="noopener"&gt;Configuring a data source name on Windows&lt;/a&gt; in the Amazon Athena User Guide.&lt;/p&gt;
&lt;p&gt;Enter the following values:&lt;/p&gt;
&lt;table border="1px" width="100%" cellpadding="10px"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;Field&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Value&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Data Source Name&lt;/td&gt;
   &lt;td&gt;Name your datasource (for example, &lt;code&gt;pbi-idcdomain&lt;/code&gt;)&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Region&lt;/td&gt;
   &lt;td&gt;The AWS Region where your Amazon SageMaker domain is provisioned (for example, &lt;code&gt;us-east-1&lt;/code&gt;)&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Catalog&lt;/td&gt;
   &lt;td&gt;&lt;code&gt;AwsDataCatalog&lt;/code&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Database&lt;/td&gt;
   &lt;td&gt;&lt;code&gt;default&lt;/code&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Workgroup&lt;/td&gt;
   &lt;td&gt;Your Athena workgroup name (for example, &lt;code&gt;workgroup-abcdefghij-klmexample&lt;/code&gt;)&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;In the &lt;strong&gt;Authentication Options&lt;/strong&gt;, configure the following values:&lt;/p&gt;
&lt;table border="1px" width="100%" cellpadding="10px"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;Field&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Value&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Authentication Type&lt;/td&gt;
   &lt;td&gt;&lt;code&gt;SageMakerBrowserIdc&lt;/code&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;SSO Start URL&lt;/td&gt;
   &lt;td&gt;IAM Identity Center entry point (for example, &lt;code&gt;https://identitycenter.amazonaws.com/ssoins-0example&lt;/code&gt;)&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;SSO Region&lt;/td&gt;
   &lt;td&gt;Region of IAM Identity Center (for example, &lt;code&gt;us-east-1&lt;/code&gt;)&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;SageMaker Domain ID&lt;/td&gt;
   &lt;td&gt;&lt;code&gt;dzd-123456example&lt;/code&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;SageMaker Project ID&lt;/td&gt;
   &lt;td&gt;&lt;code&gt;abcd12example&lt;/code&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;SageMaker Domain Region&lt;/td&gt;
   &lt;td&gt;Region of your Amazon SageMaker Unified Studio project (for example, &lt;code&gt;us-east-1&lt;/code&gt;)&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Choose &lt;strong&gt;OK,&lt;/strong&gt; then &lt;strong&gt;Test&lt;/strong&gt; to verify the connection. Choose &lt;strong&gt;Allow Access&lt;/strong&gt; when prompted by the browser.&lt;/p&gt;
&lt;p&gt;The following screenshot shows the consent prompt.&lt;/p&gt;
&lt;div style="width: 647px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P1-5.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P1-5.png" alt="Browser consent prompt requesting access approval during authentication" width="637"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 5: Browser consent prompt&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;The following screenshot shows the successful connection test.&lt;/p&gt;
&lt;div style="width: 333px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P1-6.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P1-6.png" alt="ODBC DSN configuration showing a successful connection test with SageMakerBrowserIdc" width="323"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 6: Successful connection test in the ODBC DSN configuration with SageMakerBrowserIdc authentication&lt;/p&gt;
&lt;/div&gt;
&lt;h3 id="connect-power-bi-desktop-to-your-data"&gt;Connect Power BI Desktop to your data&lt;/h3&gt;
&lt;p&gt;With the DSN configured, you can connect Power BI Desktop to your data catalog and load the generators dataset.&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;Open Power BI Desktop.&lt;/li&gt;
 &lt;li&gt;Open the &lt;strong&gt;Get Data&lt;/strong&gt; menu and select &lt;strong&gt;More&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Search for and select &lt;strong&gt;Amazon Athena&lt;/strong&gt; and choose &lt;strong&gt;Connect&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;For &lt;strong&gt;Data Source Name (DSN)&lt;/strong&gt;, enter &lt;code&gt;pbi-idcdomain&lt;/code&gt;.&lt;/li&gt;
 &lt;li&gt;Select &lt;strong&gt;DirectQuery&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Choose &lt;strong&gt;OK&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Choose &lt;strong&gt;Use Data Source Configuration&lt;/strong&gt; and then &lt;strong&gt;Connect&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;In the &lt;strong&gt;AwsDataCatalog&lt;/strong&gt; folder, navigate to your database.&lt;/li&gt;
 &lt;li&gt;Select the &lt;strong&gt;core_eia860__scd_generators&lt;/strong&gt; table.&lt;/li&gt;
 &lt;li&gt;Choose &lt;strong&gt;Load&lt;/strong&gt;.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The following screenshot shows Power BI Desktop successfully connected to the AWS data catalog.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P1-7.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P1-7.png" alt="Power BI Desktop connected to the data catalog with the generators table loaded" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 7: Power BI Desktop connected to the data catalog with the generators table loaded using SageMakerBrowserIdc authentication&lt;/p&gt;
&lt;/div&gt;
&lt;h3 id="create-your-dashboard-and-publish-it"&gt;Create your dashboard and publish it&lt;/h3&gt;
&lt;p&gt;You can create a dashboard to visualize U.S. power generation data. To create a visualization, complete the following steps:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;In the &lt;strong&gt;Visualizations&lt;/strong&gt; pane, choose the &lt;strong&gt;Stacked bar chart&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Assign the Y-Axis: Drag &lt;code&gt;technology_description&lt;/code&gt; to the Y-Axis.&lt;/li&gt;
 &lt;li&gt;Assign the X-Axis (Values): Drag &lt;code&gt;capacity_mw&lt;/code&gt; to the X-Axis (automatically summed).&lt;/li&gt;
 &lt;li&gt;Assign the Legend (Stack): Drag &lt;code&gt;operational_status&lt;/code&gt; to the Legend field.&lt;/li&gt;
 &lt;li&gt;Choose &lt;strong&gt;Publish&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Give your report a name (for example, &lt;code&gt;generation-idcdomain&lt;/code&gt;) and choose &lt;strong&gt;Save&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Sign in and choose a destination workspace.&lt;/li&gt;
&lt;/ol&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P1-8.jpg" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P1-8.jpg" alt="Power BI Desktop stacked bar chart of generation capacity by technology and operational status" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 8: Power BI Desktop report using the EIA-860 generators dataset&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;After publishing, the report structure is available on Power BI Service.&lt;/p&gt;
&lt;h2 id="method-2-dsn-less-connection-power-bi-odbc-connector"&gt;Method 2: DSN-less connection (Power BI ODBC connector)&lt;/h2&gt;
&lt;p&gt;In this method, you use the Power BI ODBC connector with a connection string (no DSN required). This method supports Import mode only and SageMakerIam authentication. Because the gateway cannot perform browser authentication, both Desktop and gateway must use SageMakerIam. If your workflow requires SageMakerBrowserIdc, use Method 1.&lt;/p&gt;
&lt;p&gt;If your machine already has AWS credentials through another method in the &lt;a href="https://docs.aws.amazon.com/sdkref/latest/guide/standardized-credentials.html" target="_blank" rel="noopener"&gt;default credential provider chain&lt;/a&gt;, skip the following setup.&lt;/p&gt;
&lt;h3 id="administrator-setup"&gt;Administrator setup&lt;/h3&gt;
&lt;p&gt;Create a custom permission set named &lt;code&gt;SageMakerDataAnalyst&lt;/code&gt; in IAM Identity Center with the following inline policy. For detailed steps, see &lt;a href="https://docs.aws.amazon.com/singlesignon/latest/userguide/howtocreatepermissionset.html" target="_blank" rel="noopener"&gt;Create a permission set&lt;/a&gt; in the AWS IAM Identity Center User Guide.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-json"&gt;{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Sid": "SageMakerAccess",
            "Effect": "Allow",
            "Action": [
                "datazone:GetConnection",
                "datazone:ListConnections",
                "datazone:GetDomain",
                "datazone:GetProject"
            ],
            "Resource": "*"
        },
        {
            "Sid": "STSForDriver",
            "Effect": "Allow",
            "Action": [
                "sts:GetCallerIdentity"
            ],
            "Resource": "*"
        }
    ]
}&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;&lt;a href="https://docs.aws.amazon.com/singlesignon/latest/userguide/useraccess.html" target="_blank" rel="noopener"&gt;Assign your user&lt;/a&gt; to this permission set for the AWS account containing your SageMaker Unified Studio domain. Then configure your AWS Command Line Interface (AWS CLI) SSO profile by running &lt;code&gt;aws configure sso&lt;/code&gt;. For the full CLI configuration walkthrough with detailed steps, see &lt;a href="https://aws.amazon.com/blogs/big-data/connect-amazon-sagemaker-unified-studio-to-microsoft-power-bi-part-2-iam-based-domains/" target="_blank" rel="noopener"&gt;Part 2&lt;/a&gt;. After your profile is configured, run &lt;code&gt;aws sso login&lt;/code&gt; to authenticate.&lt;/p&gt;
&lt;h3 id="add-the-iam-identity-as-a-member-of-sagemaker-unified-studio-project"&gt;Add the IAM identity as a member of SageMaker Unified Studio project&lt;/h3&gt;
&lt;p&gt;The IAM identity providing credentials needs both domain-level and project-level access to query data through Athena.&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;Add &lt;strong&gt;AWSReservedSSO_SageMakerDataAnalyst_1234example&lt;/strong&gt; as a domain IAM user: see &lt;a href="https://docs.aws.amazon.com/sagemaker-unified-studio/latest/adminguide/manage-users-idc-based-domains.html" target="_blank" rel="noopener"&gt;Managing users&lt;/a&gt; in the Amazon SageMaker Unified Studio Admin Guide. Choose &lt;strong&gt;Current account&lt;/strong&gt;.&lt;/li&gt;
&lt;/ol&gt;
&lt;div id="attachment_94429" style="width: 1174px" class="wp-caption alignnone"&gt;
 &lt;img aria-describedby="caption-attachment-94429" loading="lazy" class="size-full wp-image-94429" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/17/Part1-Figure9.png" alt="" width="1164" height="422"&gt;
 &lt;p id="caption-attachment-94429" class="wp-caption-text"&gt;Figure 9: List of users of your SageMaker Unified Studio domain including the IAM identity&lt;/p&gt;
&lt;/div&gt;
&lt;ol start="2" type="1"&gt;
 &lt;li&gt;Add &lt;strong&gt;AWSReservedSSO_SageMakerDataAnalyst_1234example&lt;/strong&gt; as a project member: see &lt;a href="https://docs.aws.amazon.com/sagemaker-unified-studio/latest/userguide/add-project-members.html" target="_blank" rel="noopener"&gt;Add project members&lt;/a&gt; in the Amazon SageMaker Unified Studio User Guide.&lt;/li&gt;
&lt;/ol&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P1-10.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P1-10.png" alt="SageMaker Unified Studio project members list including the IAM identity" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 10: Members of a SageMaker Unified Studio project including the IAM identity&lt;/p&gt;
&lt;/div&gt;
&lt;h3 id="gather-configuration-values"&gt;Gather configuration values&lt;/h3&gt;
&lt;p&gt;Gather the following connection values from your Amazon SageMaker Unified Studio project:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;Open your Amazon SageMaker Unified Studio Project.&lt;/li&gt;
 &lt;li&gt;On the navigation pane, choose &lt;strong&gt;Overview&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Select &lt;strong&gt;JDBC and ODBC details&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Select the &lt;strong&gt;Using IAM auth&lt;/strong&gt; toggle.&lt;/li&gt;
 &lt;li&gt;Copy the &lt;strong&gt;ODBC connection string&lt;/strong&gt;.&lt;/li&gt;
&lt;/ol&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P1-11.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P1-11.png" alt="SageMaker Unified Studio project overview showing the ODBC connection string for IAM auth" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 11: ODBC connection string on the SageMaker Unified Studio project overview&lt;/p&gt;
&lt;/div&gt;
&lt;h3 id="connect-power-bi-desktop-to-your-data-and-publish"&gt;Connect Power BI Desktop to your data and publish&lt;/h3&gt;
&lt;p&gt;With the configuration parameters of your project, you can connect Power BI Desktop to your data catalog and load the generators dataset.&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;Open Power BI Desktop.&lt;/li&gt;
 &lt;li&gt;Open the &lt;strong&gt;Get Data&lt;/strong&gt; menu and select &lt;strong&gt;More&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Search for and select &lt;strong&gt;ODBC&lt;/strong&gt; and choose &lt;strong&gt;Connect&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;For &lt;strong&gt;Data Source Name (DSN)&lt;/strong&gt;, select &lt;strong&gt;(None)&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Expand &lt;strong&gt;Advanced Options&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;In the &lt;strong&gt;Connection string&lt;/strong&gt; field, enter your connection string. For example, &lt;code&gt;Driver={Amazon Athena ODBC (x64)};AwsRegion=us-east-1;Catalog=AwsDataCatalog;Schema=default;Workgroup=workgroup-abcdefghij-klmexample;SageMakerDomainId= dzd-123456example;SageMakerProjectId= abcd12example;SageMakerDomainRegion=us-east-1;AuthenticationType=SageMakerIam;&lt;/code&gt;&lt;/li&gt;
 &lt;li&gt;Choose &lt;strong&gt;OK&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Choose &lt;strong&gt;Default or Custom&lt;/strong&gt; and then &lt;strong&gt;Connect&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;In the &lt;strong&gt;AwsDataCatalog&lt;/strong&gt; folder, navigate to your database.&lt;/li&gt;
 &lt;li&gt;Select the &lt;strong&gt;core_eia860__scd_generators&lt;/strong&gt; table.&lt;/li&gt;
 &lt;li&gt;Choose &lt;strong&gt;Load&lt;/strong&gt;.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;When publishing, name your report &lt;code&gt;generation-idcdomain-dsnless&lt;/code&gt;.&lt;/p&gt;
&lt;h2 id="configure-the-on-premises-data-gateway-and-view-your-report-on-power-bi-service"&gt;Configure the on-premises data gateway and view your report on Power BI Service&lt;/h2&gt;
&lt;p&gt;After creating your reports in Power BI Desktop, configure the on-premises data gateway to view your report on Power BI Service.&lt;/p&gt;
&lt;p&gt;You can configure the gateway using either a DSN or a DSN-less connection string, matching the method you used in Power BI Desktop.&lt;/p&gt;
&lt;h3 id="create-and-attach-an-iam-role-to-the-power-bi-gateway-ec2-instance"&gt;Create and attach an IAM role to the Power BI Gateway EC2 instance&lt;/h3&gt;
&lt;p&gt;Create an IAM role for the EC2 instance that will host your Power BI gateway. Name the role pbi-gateway-role (or a name of your choice). The role must use EC2 as the trusted entity and include the following inline policy:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-json"&gt;{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Sid": "SageMakerAccess",
            "Effect": "Allow",
            "Action": [
                "datazone:GetConnection",
                "datazone:ListConnections",
                "datazone:GetDomain",
                "datazone:GetProject"
            ],
            "Resource": "*"
        },
        {
            "Sid": "STSForDriver",
            "Effect": "Allow",
            "Action": [
                "sts:GetCallerIdentity"
            ],
            "Resource": "*"
        }
    ]
}&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Attach this role to your Power BI Gateway EC2 instance. For detailed steps on creating and attaching an IAM role to an EC2 instance, refer to &lt;a href="https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_create_for-service.html" target="_blank" rel="noopener"&gt;IAM roles for Amazon EC2&lt;/a&gt; in the Amazon EC2 User Guide.&lt;/p&gt;
&lt;h3 id="add-the-power-bi-gateway-iam-role-as-a-member-of-sagemaker-unified-studio-project"&gt;Add the Power BI Gateway IAM role as a member of SageMaker Unified Studio project&lt;/h3&gt;
&lt;p&gt;The gateway IAM role needs project-level access to query data through Athena.&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;Add the IAM &lt;strong&gt;pbi-gateway-role&lt;/strong&gt; role as a domain IAM user: see &lt;a href="https://docs.aws.amazon.com/sagemaker-unified-studio/latest/adminguide/manage-users-idc-based-domains.html" target="_blank" rel="noopener"&gt;Managing users&lt;/a&gt; in the Amazon SageMaker Unified Studio Admin Guide. Choose &lt;strong&gt;Current account&lt;/strong&gt; (or &lt;strong&gt;Associated account&lt;/strong&gt; if your gateway is deployed in a different account).&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The following screenshot, from the Amazon SageMaker page of the AWS Management Console, shows the list of users of your Amazon SageMaker Unified Studio domain, including the IAM gateway role.&lt;/p&gt;
&lt;div id="attachment_94428" style="width: 1167px" class="wp-caption alignnone"&gt;
 &lt;img aria-describedby="caption-attachment-94428" loading="lazy" class="size-full wp-image-94428" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/17/Part1-Figure12.png" alt="" width="1157" height="426"&gt;
 &lt;p id="caption-attachment-94428" class="wp-caption-text"&gt;Figure 12: List of users of your SageMaker Unified Studio domain including the IAM gateway role&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;Add the IAM &lt;strong&gt;pbi-gateway-role&lt;/strong&gt; role as a project member: see &lt;a href="https://docs.aws.amazon.com/sagemaker-unified-studio/latest/userguide/add-project-members.html" target="_blank" rel="noopener"&gt;Add project members&lt;/a&gt; in the Amazon SageMaker Unified Studio User Guide.&lt;/p&gt;
&lt;p&gt;The following screenshot shows the Amazon SageMaker Unified Studio project user management page listing the project members.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P1-13.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P1-13.png" alt="SageMaker Unified Studio project members list including the Power BI gateway IAM role" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 13: Members of a SageMaker Unified Studio project including the IAM gateway role&lt;/p&gt;
&lt;/div&gt;
&lt;h3 id="configure-the-data-source-on-power-bi-gateway"&gt;Configure the data source on Power BI Gateway&lt;/h3&gt;
&lt;p&gt;How you configure the data source depends on the method you used in Power BI Desktop.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Method 1 (DSN-based)&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Configure a System DSN on the gateway EC2 instance following the same ODBC DSN steps described in Method 1. When configuring, make sure that:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;You use the &lt;strong&gt;System DSN&lt;/strong&gt; tab (not User DSN) because the gateway runs as a Windows service under a separate account.&lt;/li&gt;
 &lt;li&gt;The authentication type is set to &lt;strong&gt;SageMakerIam&lt;/strong&gt; regardless of what you used on Desktop.&lt;/li&gt;
 &lt;li&gt;The DSN name matches exactly the one configured on Power BI Desktop (for example, &lt;code&gt;pbi-idcdomain&lt;/code&gt;)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Method 2 (DSN-less)&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;No configuration is needed on the gateway machine itself. You configure the data source directly in Power BI Service.&lt;/p&gt;
&lt;h3 id="configure-the-data-source-and-view-your-report-on-power-bi-service"&gt;Configure the data source and view your report on Power BI Service&lt;/h3&gt;
&lt;p&gt;To view your report, complete the following steps:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;Open the &lt;strong&gt;workspace&lt;/strong&gt; where you saved your report.&lt;/li&gt;
 &lt;li&gt;Search the &lt;strong&gt;Semantic Model&lt;/strong&gt; which has the same name as your report (for example, &lt;code&gt;generation-idcdomain&lt;/code&gt;) and choose the &lt;strong&gt;More options icon (three dots)&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Choose &lt;strong&gt;Settings&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Expand &lt;strong&gt;Gateway&lt;/strong&gt; and &lt;strong&gt;Cloud Connection&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Choose &lt;strong&gt;View Datasources (play icon)&lt;/strong&gt; on your gateway.&lt;/li&gt;
 &lt;li&gt;Choose &lt;strong&gt;Manually add to gateway&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Add a connection name (for example, &lt;code&gt;pbi-idcdomain&lt;/code&gt;).&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The next step depends on the method that you chose:&lt;/p&gt;
&lt;h4 id="method-1-dsn-based"&gt;Method 1 (DSN-based)&lt;/h4&gt;
&lt;ol start="8" type="1"&gt;
 &lt;li&gt;Add the &lt;strong&gt;DSN&lt;/strong&gt; (for example, &lt;code&gt;pbi-idcdomain&lt;/code&gt;) that matches exactly the one configured on Power BI Desktop.&lt;/li&gt;
&lt;/ol&gt;
&lt;h4 id="method-2-dsn-less"&gt;Method 2 (DSN-less)&lt;/h4&gt;
&lt;ol start="8" type="1"&gt;
 &lt;li&gt;In the &lt;strong&gt;Connection string&lt;/strong&gt; field, enter the connection string that matches exactly the one used in Power BI Desktop.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Next, continue with the configuration:&lt;/p&gt;
&lt;ol start="9" type="1"&gt;
 &lt;li&gt;Select &lt;strong&gt;Anonymous&lt;/strong&gt; as &lt;strong&gt;Authentication Method&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Choose &lt;strong&gt;Create&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Expand again &lt;strong&gt;Gateway&lt;/strong&gt; and &lt;strong&gt;Cloud Connection&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;For &lt;strong&gt;Maps to&lt;/strong&gt;, choose the connection that you created (for example, &lt;code&gt;pbi-idcdomain&lt;/code&gt;).&lt;/li&gt;
 &lt;li&gt;Choose &lt;strong&gt;Apply&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Return to the &lt;strong&gt;workspace&lt;/strong&gt; where you saved your report.&lt;/li&gt;
 &lt;li&gt;On the &lt;strong&gt;Content&lt;/strong&gt; section, choose your report (for example, &lt;code&gt;generation-idcdomain&lt;/code&gt;).&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The following screenshot shows a Power BI report on Power BI Service.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P1-14.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P1-14.png" alt="Published Power BI report rendering on Power BI Service" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 14: Power BI report on Power BI Service&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;You can now see your report online with the data from your Amazon SageMaker Unified Studio project.&lt;/p&gt;
&lt;h2 id="clean-up"&gt;Clean up&lt;/h2&gt;
&lt;p&gt;To avoid additional charges after testing, delete the Amazon SageMaker Unified Studio domain and EC2 instances. Refer to &lt;a href="https://docs.aws.amazon.com/sagemaker-unified-studio/latest/adminguide/delete-domain.html" target="_blank" rel="noopener"&gt;Delete domains&lt;/a&gt; and &lt;a href="https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/terminating-instances.html" target="_blank" rel="noopener"&gt;Terminate Instances&lt;/a&gt; for instructions.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;In this post, you connected Microsoft Power BI to Amazon SageMaker Unified Studio using an IDC-based domain with both DSN-based and DSN-less methods. This provides a direct connection, with no third-party licensing, that maintains data governance. In &lt;a href="https://aws.amazon.com/blogs/big-data/connect-amazon-sagemaker-unified-studio-to-microsoft-power-bi-part-2-iam-based-domains/" target="_blank" rel="noopener"&gt;Part 2&lt;/a&gt;, we cover IAM-based domains.&lt;/p&gt;
&lt;p&gt;You can automate many steps of this process. For information about automating DSN creation on the Power BI Gateway or Service, refer to &lt;a href="https://aws.amazon.com/blogs/big-data/how-engie-automates-the-deployment-of-amazon-athena-data-sources-on-microsoft-power-bi/" target="_blank" rel="noopener"&gt;How ENGIE automates the deployment of Amazon Athena data sources on Microsoft Power BI&lt;/a&gt;. If you don’t want users adding the gateway IAM role directly, you can create a custom blueprint as a self-service tool for gateway role addition. The blueprint uses a ProjectMembership resource with a configurable parameter that project owners can activate at project creation, automatically adding the gateway role as a project contributor.&lt;/p&gt;
&lt;p&gt;For additional best practices, refer to the &lt;a href="https://docs.aws.amazon.com/whitepapers/latest/using-power-bi-with-aws-cloud/using-power-bi-with-aws-cloud.html" target="_blank" rel="noopener"&gt;Using Microsoft Power BI with the AWS Cloud Whitepaper&lt;/a&gt;. To learn more, visit &lt;a href="https://aws.amazon.com/sagemaker/unified-studio/" target="_blank" rel="noopener"&gt;Amazon SageMaker Unified Studio&lt;/a&gt; and &lt;a href="https://aws.amazon.com/athena/" target="_blank" rel="noopener"&gt;Amazon Athena&lt;/a&gt;.&lt;/p&gt;
&lt;hr style="width: 100%"&gt;
&lt;h2&gt;About the authors&lt;/h2&gt;
&lt;footer&gt;
 &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
  &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
   &lt;img loading="lazy" class="alignnone size-full wp-image-94423" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/15/1756481575906.jpeg" alt="" width="100" height="100"&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Ramesh Singh&lt;/h3&gt;
  &lt;p style="overflow: hidden"&gt;Ramesh is a Senior Product Manager Technical (External Services) at AWS in Seattle, Washington, currently with the Amazon SageMaker team. He is passionate about building high-performance ML/AI and analytics products that help enterprise customers achieve their critical goals.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
  &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
   &lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P1-16.jpg" alt="Armando Segnini" width="100" height="133"&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Armando Segnini&lt;/h3&gt;
  &lt;p style="overflow: hidden"&gt;Armando is a Senior Analytics Specialist Solutions Architect at AWS, partnering with enterprise customers to architect scalable data, analytics, and AI platforms. He helps organizations turn complex data challenges into business value through expertise in streaming, BI integration, and generative AI. Outside of work, Armando enjoys traveling with his family, exploring new cultures, photography, and functional fitness competitions.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
  &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
   &lt;img loading="lazy" class="alignnone size-full wp-image-94239" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/Screenshot-2026-09-10-at-1.21.59 PM.png" alt="" width="100" height="159"&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Gaurav Sharma&lt;/h3&gt;
  &lt;p style="overflow: hidden"&gt;Gaurav is a Specialist Solutions Architect (Analytics) at AWS, supporting US public sector customers on their cloud journey. Outside of work, Gaurav enjoys spending time with his family and reading books.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
  &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
   &lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P1-18.jpg" alt="Krishna Atluru" width="100" height="133"&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Krishna Atluru&lt;/h3&gt;
  &lt;p style="overflow: hidden"&gt;Krishna is an Enterprise Support Lead TAM at AWS. He provides customers with in-depth guidance on improving security posture and operational excellence for their workloads, helping them build secure, resilient, and cost-effective solutions. His areas of expertise include building serverless architectures, and data and analytics solutions. Outside of work, Krishna enjoys cooking, swimming, and traveling.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
  &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
   &lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P1-19.jpg" alt="Saushthav Saxena" width="100" height="133"&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Saushthav Saxena&lt;/h3&gt;
  &lt;p style="overflow: hidden"&gt;Saushthav is a Software Development Engineer at AWS on the Amazon Athena team, where he has spent the past few years working on distributed systems and data analytics at scale. Based in the San Francisco Bay Area, his background spans full-stack development, high performance computing, and large-scale infrastructure. Outside of work, he enjoys reading sci-fi novels, swimming, and traveling with family and friends.&lt;/p&gt;
 &lt;/div&gt;
&lt;/footer&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>Connect Amazon SageMaker Unified Studio to Microsoft Power BI – Part 2: IAM-based domains</title>
		<link>https://aws.amazon.com/blogs/big-data/connect-amazon-sagemaker-unified-studio-to-microsoft-power-bi-part-2-iam-based-domains/</link>
		
		<dc:creator><![CDATA[Ramesh H Singh]]></dc:creator>
		<pubDate>Tue, 15 Sep 2026 20:55:16 +0000</pubDate>
				<category><![CDATA[Amazon Athena]]></category>
		<category><![CDATA[Amazon SageMaker Unified Studio]]></category>
		<category><![CDATA[Announcements]]></category>
		<category><![CDATA[Intermediate (200)]]></category>
		<guid isPermaLink="false">46d8dcbb14faab7a2ceb7795be25f661eb289eb5</guid>

					<description>Connect Microsoft Power BI directly to governed data in Amazon SageMaker Unified Studio using the Amazon Athena ODBC driver. Part 2 covers IAM-based domains with SageMakerIam authentication, including AWS IAM Identity Center administrator setup, for both DSN-based and DSN-less connection methods.</description>
										<content:encoded>&lt;p&gt;In &lt;a href="https://aws.amazon.com/blogs/big-data/connect-amazon-sagemaker-unified-studio-to-microsoft-power-bi-part-1-iam-identity-center-idc-based-domains/" target="_blank" rel="noopener"&gt;Part 1&lt;/a&gt; of this series, we connected Microsoft Power BI to Amazon SageMaker Unified Studio using an IAM Identity Center (IDC)-based domain. The &lt;a href="https://docs.aws.amazon.com/athena/latest/ug/odbc-v2-driver.html" target="_blank" rel="noopener"&gt;Amazon Athena ODBC driver&lt;/a&gt; (version 2.2.0 and later) supports Amazon SageMaker Unified Studio authentication natively, removing the third-party ODBC-JDBC bridge previously required. We walked through both the DSN-based connection and the DSN-less connection, from Power BI Desktop through the on-premises data gateway to Power BI Service, where report viewers access published dashboards.&lt;/p&gt;
&lt;p&gt;In this post, you create the same direct connection using an AWS Identity and Access Management (IAM)-based domain. The walkthrough covers the same two connection methods. The differences are the Amazon SageMaker Unified Studio console navigation paths, the configuration values, and an additional administrator setup that provides AWS credentials through AWS IAM Identity Center. This is Part 2 of a two-part series. For a detailed comparison of the two connection methods, see &lt;a href="https://aws.amazon.com/blogs/big-data/connect-amazon-sagemaker-unified-studio-to-microsoft-power-bi-part-1-iam-identity-center-idc-based-domains/" target="_blank" rel="noopener"&gt;Part 1&lt;/a&gt;.&lt;/p&gt;
&lt;h3&gt;Customer Spotlight&lt;/h3&gt;
&lt;p&gt;&lt;a href="https://uci.edu/" target="_blank" rel="noopener"&gt;UC Irvine&lt;/a&gt;, a top-ten U.S. public research university, consolidates student data from systems across multiple departments into a single governed repository that supports reporting, research, and analytics for decision-making at the strategic, tactical, and operational levels. Many of their analysts rely on Power BI to explore and visualize this governed data.&lt;/p&gt;
&lt;blockquote&gt;
 &lt;p&gt;&lt;em&gt;“Our users rely on Power BI for data visualization and reporting, but connecting to governed data in AWS previously required workarounds. The ODBC connection feature gives a direct path from Power BI into our SageMaker Unified Studio projects—no bridge software, no extra licensing, just a connection string and we’re ready to go.”&lt;/em&gt;&lt;/p&gt;
 &lt;p&gt;&lt;em&gt;— Bernadette Theologidy, Manager, Student Analytics, UC Irvine&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="solution-overview"&gt;Solution overview&lt;/h2&gt;
&lt;p&gt;The architecture is the same as the previous post (see the architecture diagram and walkthrough scenario in &lt;a href="https://aws.amazon.com/blogs/big-data/connect-amazon-sagemaker-unified-studio-to-microsoft-power-bi-part-1-iam-identity-center-idc-based-domains/" target="_blank" rel="noopener"&gt;Part 1&lt;/a&gt;). Power BI Desktop connects to Amazon Athena through the ODBC driver and the Amazon SageMaker Unified Studio project governs all data access. At the same time, the on-premises data gateway on an Amazon Elastic Compute Cloud (Amazon EC2) instance bridges the connection to Power BI Service so report viewers can access published dashboards.&lt;/p&gt;
&lt;p&gt;The difference is in authentication: An IAM-based domain uses SageMakerIam authentication for both connection methods. The driver retrieves credentials from the &lt;a href="https://docs.aws.amazon.com/sdkref/latest/guide/standardized-credentials.html" target="_blank" rel="noopener"&gt;AWS default credential provider chain&lt;/a&gt;. For this walkthrough, AWS IAM Identity Center provides those credentials through a custom permission set. Power BI Desktop can run on-premises or on an EC2 instance in the AWS Cloud. The gateway EC2 instance authenticates using its attached IAM role.&lt;/p&gt;
&lt;h2 id="prerequisites"&gt;Prerequisites&lt;/h2&gt;
&lt;p&gt;Complete the prerequisites from &lt;a href="https://aws.amazon.com/blogs/big-data/connect-amazon-sagemaker-unified-studio-to-microsoft-power-bi-part-1-iam-identity-center-idc-based-domains/" target="_blank" rel="noopener"&gt;Part 1&lt;/a&gt;. Additionally, you need:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;AWS Command Line Interface (AWS CLI)&lt;/strong&gt; – The latest version of the AWS CLI installed on your Windows machine. In this post series, the ODBC driver uses the AWS IAM Identity Center profile configured through the CLI for authentication.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Amazon SageMaker Unified Studio&lt;/strong&gt; – An Amazon SageMaker Unified Studio IAM-based domain with &lt;a href="https://docs.aws.amazon.com/sagemaker-unified-studio/latest/adminguide/setting-up-sso-iam-based.html" target="_blank" rel="noopener"&gt;AWS IAM Identity Center single sign-on (SSO) enabled&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The following screenshot shows the Amazon SageMaker Unified Studio (IAM-based domain) project Query Editor interface. It runs a preview query on the EIA-860 generators dataset.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P2-1.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P2-1.png" alt="SageMaker Unified Studio Query Editor previewing the EIA-860 generators dataset in an IAM-based domain" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 1: SageMaker Unified Studio (IAM-based domain) project with the EIA-860 generators dataset available in the data catalog&lt;/p&gt;
&lt;/div&gt;
&lt;h2 id="administrator-setup"&gt;Administrator setup&lt;/h2&gt;
&lt;p&gt;This section configures AWS IAM Identity Center to provide credentials for the SageMakerIam authentication mode. It applies to Method 1 (IAM-based domain) and Method 2 (both domain types). If your machine already has AWS credentials available through another method in the default credential provider chain, you can skip this section and proceed directly to the method of your choice. For the full list of credential sources, refer to &lt;a href="https://docs.aws.amazon.com/sdkref/latest/guide/standardized-credentials.html" target="_blank" rel="noopener"&gt;Credential providers&lt;/a&gt; in the AWS SDKs and Tools Reference Guide.&lt;/p&gt;
&lt;h3 id="create-a-permission-set-in-iam-identity-center"&gt;Create a permission set in IAM Identity Center&lt;/h3&gt;
&lt;p&gt;Create a custom permission set named &lt;strong&gt;SageMakerDataAnalyst&lt;/strong&gt; in IAM Identity Center with the following inline policy. For detailed steps, see &lt;a href="https://docs.aws.amazon.com/singlesignon/latest/userguide/howtocreatepermissionset.html" target="_blank" rel="noopener"&gt;Create a permission set&lt;/a&gt; in the AWS IAM Identity Center User Guide.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-json"&gt;{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Sid": "SageMakerAccess",
            "Effect": "Allow",
            "Action": [
                "datazone:GetConnection",
                "datazone:ListConnections",
                "datazone:GetDomain",
                "datazone:GetProject"
            ],
            "Resource": "*"
        },
        {
            "Sid": "STSForDriver",
            "Effect": "Allow",
            "Action": [
                "sts:GetCallerIdentity"
            ],
            "Resource": "*"
        }
    ]
}&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;The &lt;code&gt;"Resource": "*"&lt;/code&gt; is required because these API actions do not support resource-level permissions. For more information, see &lt;a href="https://docs.aws.amazon.com/service-authorization/latest/reference/list_amazondatazone.html" target="_blank" rel="noopener"&gt;Actions, resources, and condition keys for Amazon DataZone&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;This doesn’t grant broad access to your data. These are read-only metadata actions that allow the ODBC driver to discover connection details and retrieve temporary Athena credentials. The actual data access is governed by Amazon SageMaker Unified Studio project membership: Users can only query data within projects where they have been explicitly added as members. The Amazon SageMaker Unified Studio project IAM role provides Athena and Amazon S3 permissions separately.&lt;/p&gt;
&lt;h3 id="assign-users-to-the-permission-set"&gt;Assign users to the permission set&lt;/h3&gt;
&lt;p&gt;To assign users or groups to the target AWS account, complete the following steps:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;In the IAM Identity Center console, choose &lt;strong&gt;AWS accounts&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Select the target account where your Amazon SageMaker Unified Studio IAM-based domain is deployed.&lt;/li&gt;
 &lt;li&gt;Choose &lt;strong&gt;Assign users or groups&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Select the SSO users or groups that need access.&lt;/li&gt;
 &lt;li&gt;Select the &lt;strong&gt;SageMakerDataAnalyst&lt;/strong&gt; permission set.&lt;/li&gt;
 &lt;li&gt;Choose &lt;strong&gt;Submit&lt;/strong&gt;.&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id="configure-aws-iam-identity-center-profile"&gt;Configure AWS IAM Identity Center profile&lt;/h3&gt;
&lt;p&gt;To configure the AWS IAM Identity Center profile, run the following command in your terminal on Windows:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws configure sso&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;When prompted, enter the following values:&lt;/p&gt;
&lt;table border="1px" width="100%" cellpadding="10px"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;Prompt&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Value&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;SSO session name&lt;/td&gt;
   &lt;td&gt;For example, &lt;code&gt;smus&lt;/code&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;SSO start URL&lt;/td&gt;
   &lt;td&gt;The IDC issuer URL. For example, &lt;code&gt;https://identitycenter.amazonaws.com/ssoins-0example&lt;/code&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;SSO region&lt;/td&gt;
   &lt;td&gt;The SSO Region. For example, &lt;code&gt;us-east-1&lt;/code&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;SSO registration scopes&lt;/td&gt;
   &lt;td&gt;&lt;code&gt;sso:account:access&lt;/code&gt;&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;A browser window opens for authentication. After authentication, select your account and the &lt;strong&gt;SageMakerDataAnalyst&lt;/strong&gt; role.&lt;/p&gt;
&lt;p&gt;The following screenshots show the consent window and the successful authentication message.&lt;/p&gt;
&lt;div style="width: 629px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P2-2.jpg" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P2-2.jpg" alt="Browser consent prompt requesting access approval during AWS CLI SSO authentication" width="619"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 2: Browser consent prompt&lt;/p&gt;
&lt;/div&gt;
&lt;div style="width: 504px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P2-3.jpg" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P2-3.jpg" alt="Browser page confirming successful AWS CLI SSO authentication" width="494"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 3: Browser authentication successful message&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;When prompted, enter the following values:&lt;/p&gt;
&lt;table border="1px" width="100%" cellpadding="10px"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;Prompt&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Value&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Default client Region&lt;/td&gt;
   &lt;td&gt;None&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;CLI default output format&lt;/td&gt;
   &lt;td&gt;None&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Profile Name&lt;/td&gt;
   &lt;td&gt;Change value by default&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The resulting &lt;code&gt;~/.aws/config&lt;/code&gt; file should look like the following:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-ini"&gt;[default]
sso_session = smus
sso_account_id = 1234example
sso_role_name = SageMakerDataAnalyst

[sso-session smus]
sso_start_url = https://identitycenter.amazonaws.com/ssoins-0example
sso_region = us-east-1
sso_registration_scopes = sso:account:access&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h3 id="verify-authentication-and-daily-use"&gt;Verify authentication and daily use&lt;/h3&gt;
&lt;p&gt;To verify that your SSO profile is working correctly, run the following command:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws sts get-caller-identity&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;You should receive a response like the following:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-json"&gt;{
    "UserId": "AROARHJJNFBQD6EXAMPLE:user@example-domain.com",
    "Account": "111122223333",
    "Arn": "arn:aws:sts::111122223333:assumed-role/AWSReservedSSO_SageMakerDataAnalyst_1234example/user@example-domain.com"
}&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;For daily use, no passwords or EC2 instance roles are required. When your SSO session expires, run the following command to quickly refresh it:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws sso login&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h3 id="add-your-iam-identity-as-a-member-of-your-sagemaker-unified-studio-project"&gt;Add your IAM identity as a member of your Amazon SageMaker Unified Studio project&lt;/h3&gt;
&lt;p&gt;The IAM identity providing credentials to the ODBC driver needs project-level access to query data through Athena. If you completed the administrator setup, this is the SSO role associated with your permission set (for example, &lt;code&gt;AWSReservedSSO_SageMakerDataAnalyst_1234example&lt;/code&gt;). If you’re using another credential source, add the IAM role or user that provides those credentials. For detailed steps, see &lt;a href="https://docs.aws.amazon.com/sagemaker-unified-studio/latest/adminguide/manage-users-iam-based-domains.html" target="_blank" rel="noopener"&gt;Managing users for IAM-based domains&lt;/a&gt; in the Amazon SageMaker Unified Studio Administrator Guide.&lt;/p&gt;
&lt;p&gt;The following screenshot shows the Amazon SageMaker Unified Studio domain management page, which lists the members in a project.&lt;/p&gt;
&lt;div id="attachment_94430" style="width: 1182px" class="wp-caption alignnone"&gt;
 &lt;img aria-describedby="caption-attachment-94430" loading="lazy" class="size-full wp-image-94430" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/17/Part2-Figure4.png" alt="" width="1172" height="671"&gt;
 &lt;p id="caption-attachment-94430" class="wp-caption-text"&gt;Figure 4: List of members of your SageMaker Unified Studio project&lt;/p&gt;
&lt;/div&gt;
&lt;h3 id="gather-the-information-to-authenticate"&gt;Gather the information to authenticate&lt;/h3&gt;
&lt;p&gt;To get the parameters that you need to authenticate, complete these steps:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;Open your Amazon SageMaker Unified Studio Project.&lt;/li&gt;
 &lt;li&gt;Open &lt;strong&gt;Domain Management&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Choose &lt;strong&gt;Users&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Choose &lt;strong&gt;View SSO connection&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Copy the end of the Instance ARN, so we can build the Instance URL like &lt;code&gt;https://identitycenter.amazonaws.com/ssoins-0example&lt;/code&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The following screenshot shows the Amazon SageMaker Unified Studio domain management page with SSO connection details.&lt;/p&gt;
&lt;div id="attachment_94427" style="width: 1180px" class="wp-caption alignnone"&gt;
 &lt;img aria-describedby="caption-attachment-94427" loading="lazy" class="size-full wp-image-94427" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/17/Part2-Figure5.png" alt="" width="1170" height="300"&gt;
 &lt;p id="caption-attachment-94427" class="wp-caption-text"&gt;Figure 5: AWS IAM Identity Center information&lt;/p&gt;
&lt;/div&gt;
&lt;ol start="6" type="1"&gt;
 &lt;li&gt;Choose the user icon and copy the &lt;strong&gt;Region&lt;/strong&gt; as shown in the following screenshot.&lt;/li&gt;
&lt;/ol&gt;
&lt;div class="mceTemp"&gt;
 &lt;div id="attachment_94431" style="width: 1064px" class="wp-caption alignnone"&gt;
  &lt;img aria-describedby="caption-attachment-94431" loading="lazy" class="size-full wp-image-94431" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/17/Part2-Figure6.png" alt="" width="1054" height="170"&gt;
  &lt;p id="caption-attachment-94431" class="wp-caption-text"&gt;Figure 6: User icon with the Region information&lt;/p&gt;
 &lt;/div&gt;
 &lt;h2 id="method-1-dsn-based-connection-athena-power-bi-connector"&gt;Method 1: DSN-based connection (Athena Power BI connector)&lt;/h2&gt;
 &lt;p&gt;In this method, you configure an ODBC Data Source Name (DSN) and use the Amazon Athena connector in Power BI. This method uses SageMakerIam authentication mode and supports both DirectQuery and Import mode.&lt;/p&gt;
 &lt;p&gt;This section covers IAM-based domains. For IDC-based domains, see &lt;a href="https://aws.amazon.com/blogs/big-data/connect-amazon-sagemaker-unified-studio-to-microsoft-power-bi-part-1-iam-identity-center-idc-based-domains/" target="_blank" rel="noopener"&gt;Part 1&lt;/a&gt;.&lt;/p&gt;
 &lt;h3 id="gather-configuration-values-to-configure-your-amazon-athena-odbc-dsn"&gt;Gather configuration values to configure your Amazon Athena ODBC DSN&lt;/h3&gt;
 &lt;p&gt;Before configuring the ODBC DSN, gather the following connection values from your Amazon SageMaker Unified Studio project:&lt;/p&gt;
 &lt;ol type="1"&gt;
  &lt;li&gt;Open your Amazon SageMaker Unified Studio Project.&lt;/li&gt;
  &lt;li&gt;Top right, select the &lt;strong&gt;three dots&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;Choose &lt;strong&gt;Project details&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;Select &lt;strong&gt;JDBC and ODBC details&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;Copy the following values: domain ID, Amazon SageMaker project ID, AWS Region, and Athena workgroup.&lt;/li&gt;
 &lt;/ol&gt;
 &lt;p&gt;The following screenshot shows the Amazon SageMaker Unified Studio project overview page, which provides the project details to copy.&lt;/p&gt;
 &lt;div style="width: 810px" class="wp-caption alignnone"&gt;
  &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P2-7.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P2-7.png" alt="SageMaker Unified Studio project details showing domain ID, project ID, Region, and Athena workgroup" width="800"&gt;&lt;/a&gt;
  &lt;p class="wp-caption-text"&gt;Figure 7: Project details with SageMaker domain ID, SageMaker project ID, Region, and Athena workgroup&lt;/p&gt;
 &lt;/div&gt;
 &lt;h4 id="configure-the-odbc-dsn"&gt;Configure the ODBC DSN&lt;/h4&gt;
 &lt;p&gt;Create a System DSN using the Amazon Athena ODBC driver. For the general DSN creation steps, see &lt;a href="https://docs.aws.amazon.com/athena/latest/ug/odbc-v2-driver-getting-started-windows.html" target="_blank" rel="noopener"&gt;Configuring a data source name on Windows&lt;/a&gt; in the Amazon Athena User Guide. Enter the following values:&lt;/p&gt;
 &lt;table border="1px" width="100%" cellpadding="10px"&gt;
  &lt;tbody&gt;
   &lt;tr&gt;
    &lt;td&gt;&lt;strong&gt;Field&lt;/strong&gt;&lt;/td&gt;
    &lt;td&gt;&lt;strong&gt;Value&lt;/strong&gt;&lt;/td&gt;
   &lt;/tr&gt;
   &lt;tr&gt;
    &lt;td&gt;Data Source Name&lt;/td&gt;
    &lt;td&gt;Name your datasource (for example, &lt;code&gt;pbi-iamdomain&lt;/code&gt;)&lt;/td&gt;
   &lt;/tr&gt;
   &lt;tr&gt;
    &lt;td&gt;Region&lt;/td&gt;
    &lt;td&gt;The AWS Region where your Amazon SageMaker domain is provisioned (for example, &lt;code&gt;us-east-1&lt;/code&gt;)&lt;/td&gt;
   &lt;/tr&gt;
   &lt;tr&gt;
    &lt;td&gt;Catalog&lt;/td&gt;
    &lt;td&gt;&lt;code&gt;AwsDataCatalog&lt;/code&gt;&lt;/td&gt;
   &lt;/tr&gt;
   &lt;tr&gt;
    &lt;td&gt;Database&lt;/td&gt;
    &lt;td&gt;&lt;code&gt;default&lt;/code&gt;&lt;/td&gt;
   &lt;/tr&gt;
   &lt;tr&gt;
    &lt;td&gt;Workgroup&lt;/td&gt;
    &lt;td&gt;Your Athena workgroup name (for example, &lt;code&gt;workgroup-abcdefghij-klmexample&lt;/code&gt;)&lt;/td&gt;
   &lt;/tr&gt;
  &lt;/tbody&gt;
 &lt;/table&gt;
 &lt;p&gt;In the &lt;strong&gt;Authentication Options&lt;/strong&gt;, configure the following values:&lt;/p&gt;
 &lt;table border="1px" width="100%" cellpadding="10px"&gt;
  &lt;tbody&gt;
   &lt;tr&gt;
    &lt;td&gt;&lt;strong&gt;Field&lt;/strong&gt;&lt;/td&gt;
    &lt;td&gt;&lt;strong&gt;Value&lt;/strong&gt;&lt;/td&gt;
   &lt;/tr&gt;
   &lt;tr&gt;
    &lt;td&gt;Authentication Type&lt;/td&gt;
    &lt;td&gt;&lt;code&gt;SageMakerIam&lt;/code&gt;&lt;/td&gt;
   &lt;/tr&gt;
   &lt;tr&gt;
    &lt;td&gt;SageMaker Domain ID&lt;/td&gt;
    &lt;td&gt;&lt;code&gt;dzd-123456example&lt;/code&gt;&lt;/td&gt;
   &lt;/tr&gt;
   &lt;tr&gt;
    &lt;td&gt;SageMaker Project ID&lt;/td&gt;
    &lt;td&gt;&lt;code&gt;abcd12example&lt;/code&gt;&lt;/td&gt;
   &lt;/tr&gt;
   &lt;tr&gt;
    &lt;td&gt;SageMaker Region&lt;/td&gt;
    &lt;td&gt;Region of your SageMaker Unified Studio project (for example, &lt;code&gt;us-east-1&lt;/code&gt;)&lt;/td&gt;
   &lt;/tr&gt;
  &lt;/tbody&gt;
 &lt;/table&gt;
 &lt;p&gt;Choose &lt;strong&gt;OK,&lt;/strong&gt; then &lt;strong&gt;Test&lt;/strong&gt; to verify the connection. Choose &lt;strong&gt;Allow Access&lt;/strong&gt; when prompted by the browser.&lt;/p&gt;
 &lt;p&gt;The following screenshot shows the successful connection test.&lt;/p&gt;
 &lt;div style="width: 360px" class="wp-caption alignnone"&gt;
  &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P2-8.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P2-8.png" alt="ODBC DSN configuration showing a successful connection test with SageMakerIam" width="350"&gt;&lt;/a&gt;
  &lt;p class="wp-caption-text"&gt;Figure 8: Successful connection test in the ODBC DSN configuration with SageMakerIam authentication&lt;/p&gt;
 &lt;/div&gt;
 &lt;h4 id="connect-power-bi-desktop-to-your-data"&gt;Connect Power BI Desktop to your data&lt;/h4&gt;
 &lt;p&gt;With the DSN configured, you can connect Power BI Desktop to your data catalog and load the generators dataset.&lt;/p&gt;
 &lt;ol type="1"&gt;
  &lt;li&gt;Open Microsoft Power BI Desktop.&lt;/li&gt;
  &lt;li&gt;Open the &lt;strong&gt;Get Data&lt;/strong&gt; menu and select &lt;strong&gt;More&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;Search for and select &lt;strong&gt;Amazon Athena&lt;/strong&gt; and choose &lt;strong&gt;Connect&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;For &lt;strong&gt;Data Source Name (DSN)&lt;/strong&gt;, enter &lt;code&gt;pbi-iamdomain&lt;/code&gt;.&lt;/li&gt;
  &lt;li&gt;Select &lt;strong&gt;DirectQuery&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;Choose &lt;strong&gt;OK&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;Choose &lt;strong&gt;Use Data Source Configuration&lt;/strong&gt; and then &lt;strong&gt;Connect&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;In the &lt;strong&gt;AwsDataCatalog&lt;/strong&gt; folder, navigate to your database.&lt;/li&gt;
  &lt;li&gt;Select the &lt;strong&gt;core_eia860__scd_generators&lt;/strong&gt; table.&lt;/li&gt;
  &lt;li&gt;Choose &lt;strong&gt;Load&lt;/strong&gt;.&lt;/li&gt;
 &lt;/ol&gt;
 &lt;p&gt;The following screenshot shows Power BI Desktop successfully connected to the data catalog.&lt;/p&gt;
 &lt;div style="width: 810px" class="wp-caption alignnone"&gt;
  &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P2-9.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P2-9.png" alt="Power BI Desktop connected to the data catalog with the generators table loaded" width="800"&gt;&lt;/a&gt;
  &lt;p class="wp-caption-text"&gt;Figure 9: Power BI Desktop connected to the data catalog with the generators table loaded using SageMakerIam authentication&lt;/p&gt;
 &lt;/div&gt;
 &lt;h4 id="create-your-dashboard-and-publish-it"&gt;Create your dashboard and publish it&lt;/h4&gt;
 &lt;p&gt;You can create a dashboard to visualize U.S. power generation data. To create a visualization, complete the following steps:&lt;/p&gt;
 &lt;ol type="1"&gt;
  &lt;li&gt;In the &lt;strong&gt;Visualizations&lt;/strong&gt; pane, choose the &lt;strong&gt;Stacked bar chart&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;Assign the Y-Axis: Drag &lt;code&gt;technology_description&lt;/code&gt; to the Y-Axis.&lt;/li&gt;
  &lt;li&gt;Assign the X-Axis (Values): Drag &lt;code&gt;capacity_mw&lt;/code&gt; to the X-Axis (automatically summed).&lt;/li&gt;
  &lt;li&gt;Assign the Legend (Stack): Drag &lt;code&gt;operational_status&lt;/code&gt; to the Legend field.&lt;/li&gt;
  &lt;li&gt;Choose &lt;strong&gt;Publish&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;Give your report a name (for example, &lt;code&gt;generation-iamdomain&lt;/code&gt;) and choose &lt;strong&gt;Save&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;Sign in and choose a destination workspace.&lt;/li&gt;
 &lt;/ol&gt;
 &lt;p&gt;The following screenshot shows the Power BI dashboard with U.S. power generation data.&lt;/p&gt;
 &lt;div style="width: 810px" class="wp-caption alignnone"&gt;
  &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P2-10.jpg" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P2-10.jpg" alt="Power BI stacked bar chart of U.S. generation capacity by technology and operational status" width="800"&gt;&lt;/a&gt;
  &lt;p class="wp-caption-text"&gt;Figure 10: Power BI dashboard with U.S. power generation data&lt;/p&gt;
 &lt;/div&gt;
 &lt;p&gt;After you publish, the report structure becomes available on Microsoft Power BI Service.&lt;/p&gt;
 &lt;h2 id="method-2-dsn-less-connection-power-bi-odbc-connector"&gt;Method 2: DSN-less connection (Power BI ODBC connector)&lt;/h2&gt;
 &lt;p&gt;In this method, you use the Power BI ODBC connector with a connection string (no DSN required). This method supports Import mode only and SageMakerIam authentication. Because the gateway can’t perform browser authentication and connection strings need to match, both Desktop and gateway must use SageMakerIam.&lt;/p&gt;
 &lt;p&gt;This section covers IAM-based domains. For IDC-based domains, see &lt;a href="https://aws.amazon.com/blogs/big-data/connect-amazon-sagemaker-unified-studio-to-microsoft-power-bi-part-1-iam-identity-center-idc-based-domains/" target="_blank" rel="noopener"&gt;Part 1&lt;/a&gt;.&lt;/p&gt;
 &lt;h4 id="gather-configuration-values-to-configure-your-dsn-less-connection"&gt;Gather configuration values to configure your DSN-less connection&lt;/h4&gt;
 &lt;p&gt;Gather the following connection values from your Amazon SageMaker Unified Studio project:&lt;/p&gt;
 &lt;ol type="1"&gt;
  &lt;li&gt;Open your Amazon SageMaker Unified Studio Project.&lt;/li&gt;
  &lt;li&gt;Top right, select the &lt;strong&gt;three dots&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;Choose &lt;strong&gt;Project details&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;Select &lt;strong&gt;JDBC and ODBC details&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;Copy the ODBC connection string.&lt;/li&gt;
 &lt;/ol&gt;
 &lt;p&gt;The following screenshot shows the Amazon SageMaker Unified Studio project overview page with the ODBC connection string to copy.&lt;/p&gt;
 &lt;div style="width: 810px" class="wp-caption alignnone"&gt;
  &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P2-11.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P2-11.png" alt="SageMaker Unified Studio project overview showing the ODBC connection string" width="800"&gt;&lt;/a&gt;
  &lt;p class="wp-caption-text"&gt;Figure 11: Project details with ODBC connection string&lt;/p&gt;
 &lt;/div&gt;
 &lt;h4 id="connect-power-bi-desktop-to-your-data-and-publish"&gt;Connect Power BI Desktop to your data and publish&lt;/h4&gt;
 &lt;p&gt;With the configuration parameters of your project, you can connect Power BI Desktop to your data catalog and load the generators dataset.&lt;/p&gt;
 &lt;ol type="1"&gt;
  &lt;li&gt;Open Power BI Desktop.&lt;/li&gt;
  &lt;li&gt;Open the &lt;strong&gt;Get Data&lt;/strong&gt; menu and select &lt;strong&gt;More&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;Search for and select &lt;strong&gt;ODBC&lt;/strong&gt; and choose &lt;strong&gt;Connect&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;For &lt;strong&gt;Data Source Name (DSN)&lt;/strong&gt;, select &lt;strong&gt;(None)&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;Expand Advanced Options.&lt;/li&gt;
  &lt;li&gt;In the &lt;strong&gt;Connection string&lt;/strong&gt; field, enter your connection string. For example, &lt;code&gt;Driver={Amazon Athena ODBC (x64)};AwsRegion=us-east-1;Catalog=AwsDataCatalog;Schema=default;Workgroup=workgroup-abcdefghij-klmexample;SageMakerDomainId= dzd-123456example;SageMakerProjectId= abcd12example;SageMakerDomainRegion=us-east-1;AuthenticationType=SageMakerIam;&lt;/code&gt;&lt;/li&gt;
  &lt;li&gt;Choose &lt;strong&gt;OK&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;Choose &lt;strong&gt;Default or Custom&lt;/strong&gt; and then &lt;strong&gt;Connect&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;In the &lt;strong&gt;AwsDataCatalog&lt;/strong&gt; folder, navigate to your database.&lt;/li&gt;
  &lt;li&gt;Select the &lt;strong&gt;core_eia860__scd_generators&lt;/strong&gt; table.&lt;/li&gt;
  &lt;li&gt;Choose &lt;strong&gt;Load&lt;/strong&gt;.&lt;/li&gt;
 &lt;/ol&gt;
 &lt;p&gt;When publishing, name your report &lt;code&gt;generation-iamdomain-dsnless&lt;/code&gt;.&lt;/p&gt;
 &lt;h2 id="configure-the-gateway-and-view-your-report-on-power-bi-service"&gt;Configure the gateway and view your report on Power BI Service&lt;/h2&gt;
 &lt;p&gt;After creating your reports in Power BI Desktop, configure the on-premises data gateway to view your report on Power BI Service.&lt;/p&gt;
 &lt;p&gt;You can configure the gateway using either a DSN or a DSN-less connection string, matching the method you used in Power BI Desktop.&lt;/p&gt;
 &lt;h3 id="create-and-attach-an-iam-role-to-the-power-bi-gateway-ec2-instance"&gt;Create and attach an IAM role to the Power BI Gateway EC2 instance&lt;/h3&gt;
 &lt;p&gt;Create an IAM role for the EC2 instance that will host your Power BI gateway. Name the role &lt;strong&gt;pbi-gateway-role&lt;/strong&gt; (or a name of your choice). The role must use EC2 as the trusted entity and include the following inline policy:&lt;/p&gt;
 &lt;div class="hide-language"&gt;
  &lt;pre&gt;&lt;code class="language-json"&gt;{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Sid": "SageMakerAccess",
            "Effect": "Allow",
            "Action": [
                "datazone:GetConnection",
                "datazone:ListConnections",
                "datazone:GetDomain",
                "datazone:GetProject"
            ],
            "Resource": "*"
        },
        {
            "Sid": "STSForDriver",
            "Effect": "Allow",
            "Action": [
                "sts:GetCallerIdentity"
            ],
            "Resource": "*"
        }
    ]
}&lt;/code&gt;&lt;/pre&gt;
 &lt;/div&gt;
 &lt;p&gt;Attach this role to your Power BI Gateway EC2 instance. For detailed steps on creating and attaching an IAM role to an EC2 instance, refer to &lt;a href="https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_create_for-service.html" target="_blank" rel="noopener"&gt;IAM roles for Amazon EC2&lt;/a&gt; in the Amazon EC2 User Guide.&lt;/p&gt;
 &lt;h3 id="add-the-power-bi-gateway-iam-role-as-a-member-of-sagemaker-unified-studio-project"&gt;Add the Power BI Gateway IAM role as a member of SageMaker Unified Studio project&lt;/h3&gt;
 &lt;p&gt;The gateway IAM role needs project-level access to query data through Athena. The steps to add the role differ depending on your domain type.&lt;/p&gt;
 &lt;h4 id="iam-based-domain"&gt;IAM-based domain&lt;/h4&gt;
 &lt;ol type="1"&gt;
  &lt;li&gt;Open your Amazon SageMaker Unified Studio Project.&lt;/li&gt;
  &lt;li&gt;Open &lt;strong&gt;Domain Management&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;Choose your &lt;strong&gt;Project Name&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;Choose &lt;strong&gt;Members&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;Choose &lt;strong&gt;Add members&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;Select the IAM role of your Power BI gateway (for example, &lt;code&gt;pbi-gateway-role&lt;/code&gt;).&lt;/li&gt;
  &lt;li&gt;Choose &lt;strong&gt;Add&lt;/strong&gt;.&lt;/li&gt;
 &lt;/ol&gt;
 &lt;p&gt;The following screenshot shows the Amazon SageMaker Unified Studio project domain management page with options to add members to a project.&lt;/p&gt;
 &lt;div id="attachment_94432" style="width: 1180px" class="wp-caption alignnone"&gt;
  &lt;img aria-describedby="caption-attachment-94432" loading="lazy" class="size-full wp-image-94432" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/17/Part2-Figure12.png" alt="" width="1170" height="702"&gt;
  &lt;p id="caption-attachment-94432" class="wp-caption-text"&gt;Figure 12: List of members of a SageMaker Unified Studio project with the IAM gateway role&lt;/p&gt;
 &lt;/div&gt;
 &lt;h3 id="configure-the-data-source-on-power-bi-gateway"&gt;Configure the data source on Power BI Gateway&lt;/h3&gt;
 &lt;p&gt;How you configure the data source depends on the method you used in Power BI Desktop.&lt;/p&gt;
 &lt;p&gt;&lt;strong&gt;Method 1 (DSN-based)&lt;/strong&gt;&lt;/p&gt;
 &lt;p&gt;Configure a System DSN on the gateway EC2 instance following the same ODBC DSN steps described in Method 1. When configuring, make sure that:&lt;/p&gt;
 &lt;ul&gt;
  &lt;li&gt;You use the &lt;strong&gt;System DSN&lt;/strong&gt; tab (not User DSN) because the gateway runs as a Windows service under a separate account.&lt;/li&gt;
  &lt;li&gt;The authentication type is set to &lt;strong&gt;SageMakerIam&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;The DSN name matches exactly the one configured on Power BI Desktop (for example, &lt;code&gt;pbi-iamdomain&lt;/code&gt;).&lt;/li&gt;
 &lt;/ul&gt;
 &lt;p&gt;&lt;strong&gt;Method 2 (DSN-less)&lt;/strong&gt;&lt;/p&gt;
 &lt;p&gt;No configuration is needed on the gateway machine itself. You configure the data source directly in Power BI Service.&lt;/p&gt;
 &lt;h3 id="configure-the-data-source-and-view-your-report-on-power-bi-service"&gt;Configure the data source and view your report on Power BI Service&lt;/h3&gt;
 &lt;p&gt;To view your report, complete the following steps:&lt;/p&gt;
 &lt;ol type="1"&gt;
  &lt;li&gt;Open the &lt;strong&gt;workspace&lt;/strong&gt; where you saved your report.&lt;/li&gt;
  &lt;li&gt;Search the &lt;strong&gt;Semantic Model&lt;/strong&gt; which has the same name as your report (for example, &lt;code&gt;generation-iamdomain&lt;/code&gt;) and choose the &lt;strong&gt;More options icon (three dots)&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;Choose &lt;strong&gt;Settings&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;Expand &lt;strong&gt;Gateway&lt;/strong&gt; and &lt;strong&gt;Cloud Connection&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;Choose &lt;strong&gt;View Datasources (play icon)&lt;/strong&gt; on your gateway.&lt;/li&gt;
  &lt;li&gt;Choose &lt;strong&gt;Manually add to gateway&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;Add a connection name (for example, &lt;code&gt;pbi-iamdomain&lt;/code&gt;).&lt;/li&gt;
 &lt;/ol&gt;
 &lt;p&gt;The next step depends on the method that you chose:&lt;/p&gt;
 &lt;h4 id="method-1-dsn-based"&gt;Method 1 (DSN-based)&lt;/h4&gt;
 &lt;ol start="8" type="1"&gt;
  &lt;li&gt;Add the &lt;strong&gt;DSN&lt;/strong&gt; (for example, &lt;code&gt;pbi-iamdomain&lt;/code&gt;) that matches exactly the one configured on Power BI Desktop.&lt;/li&gt;
 &lt;/ol&gt;
 &lt;h4 id="method-2-dsn-less"&gt;Method 2 (DSN-less)&lt;/h4&gt;
 &lt;ol start="8" type="1"&gt;
  &lt;li&gt;In the &lt;strong&gt;Connection string&lt;/strong&gt; field, enter the connection string that matches exactly the one used in Power BI Desktop.&lt;/li&gt;
 &lt;/ol&gt;
 &lt;p&gt;Next, continue with the configuration:&lt;/p&gt;
 &lt;ol start="9" type="1"&gt;
  &lt;li&gt;Select &lt;strong&gt;Anonymous&lt;/strong&gt; as &lt;strong&gt;Authentication Method&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;Choose &lt;strong&gt;Create&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;Expand again &lt;strong&gt;Gateway&lt;/strong&gt; and &lt;strong&gt;Cloud Connection&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;For &lt;strong&gt;Maps to&lt;/strong&gt;, choose the connection that you created (for example, &lt;code&gt;pbi-iamdomain&lt;/code&gt;).&lt;/li&gt;
  &lt;li&gt;Choose &lt;strong&gt;Apply&lt;/strong&gt;.&lt;/li&gt;
  &lt;li&gt;Return to the &lt;strong&gt;workspace&lt;/strong&gt; where you saved your report.&lt;/li&gt;
  &lt;li&gt;On the &lt;strong&gt;Content&lt;/strong&gt; section, choose your report (for example, &lt;code&gt;generation-iamdomain&lt;/code&gt;).&lt;/li&gt;
 &lt;/ol&gt;
 &lt;p&gt;The following screenshot shows a report on Power BI Service.&lt;/p&gt;
 &lt;div style="width: 810px" class="wp-caption alignnone"&gt;
  &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P2-13.jpg" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P2-13.jpg" alt="Published Power BI report rendering on Power BI Service" width="800"&gt;&lt;/a&gt;
  &lt;p class="wp-caption-text"&gt;Figure 13: Power BI report on Power BI Service&lt;/p&gt;
 &lt;/div&gt;
 &lt;p&gt;You can now see your report online with the data from your Amazon SageMaker Unified Studio project.&lt;/p&gt;
 &lt;h2 id="clean-up"&gt;Clean up&lt;/h2&gt;
 &lt;p&gt;To avoid additional charges after testing, delete the Amazon SageMaker Unified Studio domain and EC2 instances. Refer to &lt;a href="https://docs.aws.amazon.com/sagemaker-unified-studio/latest/adminguide/delete-domain.html" target="_blank" rel="noopener"&gt;Delete domains&lt;/a&gt; and &lt;a href="https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/terminating-instances.html" target="_blank" rel="noopener"&gt;Terminate Instances&lt;/a&gt; for instructions.&lt;/p&gt;
 &lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
 &lt;p&gt;In this two-part series, you connected Power BI to Amazon SageMaker Unified Studio through Amazon Athena. &lt;a href="https://aws.amazon.com/blogs/big-data/connect-amazon-sagemaker-unified-studio-to-microsoft-power-bi-part-1-iam-identity-center-idc-based-domains/" target="_blank" rel="noopener"&gt;Part 1&lt;/a&gt; covered IDC-based domains. This post covered IAM-based domains using SageMakerIam authentication. This provides a direct connection path, with no third-party licensing, while maintaining data governance and security.&lt;/p&gt;
 &lt;p&gt;You can automate many steps of this process. For information about automating DSN creation on the Power BI Gateway or Service, refer to &lt;a href="https://aws.amazon.com/blogs/big-data/how-engie-automates-the-deployment-of-amazon-athena-data-sources-on-microsoft-power-bi/" target="_blank" rel="noopener"&gt;How ENGIE automates the deployment of Amazon Athena data sources on Microsoft Power BI&lt;/a&gt;. If you don’t want users adding the gateway IAM role directly, you can create a custom blueprint as a self-service tool for gateway role addition. The blueprint uses a ProjectMembership resource with a configurable parameter that project owners can activate at project creation, automatically adding the gateway role as a project contributor.&lt;/p&gt;
 &lt;p&gt;For additional best practices, refer to the &lt;a href="https://docs.aws.amazon.com/whitepapers/latest/using-power-bi-with-aws-cloud/using-power-bi-with-aws-cloud.html" target="_blank" rel="noopener"&gt;Using Microsoft Power BI with the AWS Cloud Whitepaper&lt;/a&gt;. To learn more, visit &lt;a href="https://aws.amazon.com/sagemaker/unified-studio/" target="_blank" rel="noopener"&gt;Amazon SageMaker Unified Studio&lt;/a&gt; and &lt;a href="https://aws.amazon.com/athena/" target="_blank" rel="noopener"&gt;Amazon Athena&lt;/a&gt;.&lt;/p&gt;
 &lt;p style="clear: both"&gt;&lt;/p&gt;
 &lt;hr style="width: 100%"&gt;
 &lt;h2&gt;About the authors&lt;/h2&gt;
 &lt;footer&gt;
  &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
   &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
    &lt;img loading="lazy" class="alignnone size-full wp-image-94423" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/15/1756481575906.jpeg" alt="" width="100" height="100"&gt;
   &lt;/div&gt;
   &lt;h3 class="lb-h4"&gt;Ramesh H Singh&lt;/h3&gt;
   &lt;p style="overflow: hidden"&gt;Ramesh is a Senior Product Manager Technical at AWS in Seattle, focused on Amazon SageMaker. He’s passionate about building analytics and AI products that help enterprise customers unlock real value from their data. Away from work, he spends his time hiking with family and exploring spirituality. Connect with him on &lt;a href="http://www.linkedin.com/in/ramesh-harisaran-singh" target="_blank" rel="noopener"&gt;LinkedIn&lt;/a&gt;.&lt;/p&gt;
  &lt;/div&gt;
  &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
   &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
    &lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P2-15.jpg" alt="Armando Segnini" width="100" height="133"&gt;
   &lt;/div&gt;
   &lt;h3 class="lb-h4"&gt;Armando Segnini&lt;/h3&gt;
   &lt;p style="overflow: hidden"&gt;Armando is a Senior Analytics Specialist Solutions Architect at AWS, partnering with enterprise customers to architect scalable data, analytics, and AI platforms. He helps organizations turn complex data challenges into business value through expertise in streaming, BI integration, and generative AI. Outside of work, Armando enjoys traveling with his family, exploring new cultures, photography, and functional fitness competitions.&lt;/p&gt;
  &lt;/div&gt;
  &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
   &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
    &lt;img loading="lazy" class="alignnone size-full wp-image-94239" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/Screenshot-2026-09-10-at-1.21.59 PM.png" alt="" width="100" height="159"&gt;
   &lt;/div&gt;
   &lt;h3 class="lb-h4"&gt;Gaurav Sharma&lt;/h3&gt;
   &lt;p style="overflow: hidden"&gt;Gaurav is a Specialist Solutions Architect (Analytics) at AWS, supporting US public sector customers on their cloud journey. Outside of work, Gaurav enjoys spending time with his family and reading books.&lt;/p&gt;
  &lt;/div&gt;
  &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
   &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
    &lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P2-17.jpg" alt="Krishna Atluru" width="100" height="133"&gt;
   &lt;/div&gt;
   &lt;h3 class="lb-h4"&gt;Krishna Atluru&lt;/h3&gt;
   &lt;p style="overflow: hidden"&gt;Krishna is an Enterprise Support Lead TAM at AWS. He provides customers with in-depth guidance on improving security posture and operational excellence for their workloads, helping them build secure, resilient, and cost-effective solutions. His areas of expertise include building serverless architectures, and data and analytics solutions. Outside of work, Krishna enjoys cooking, swimming, and traveling.&lt;/p&gt;
  &lt;/div&gt;
  &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
   &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
    &lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-6011-P2-18.jpg" alt="Saushthav Saxena" width="100" height="133"&gt;
   &lt;/div&gt;
   &lt;h3 class="lb-h4"&gt;Saushthav Saxena&lt;/h3&gt;
   &lt;p style="overflow: hidden"&gt;Saushthav is a Software Development Engineer at AWS on the Amazon Athena team, where he has spent the past few years working on distributed systems and data analytics at scale. Based in the San Francisco Bay Area, his background spans full-stack development, high-performance computing, and large-scale infrastructure. Outside of work, he enjoys reading sci-fi novels, swimming, and traveling with family and friends.&lt;/p&gt;
  &lt;/div&gt;
 &lt;/footer&gt;
&lt;/div&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>How United Airlines uses Amazon Redshift and AWS Glue Data Catalog federation to query Databricks-managed data</title>
		<link>https://aws.amazon.com/blogs/big-data/how-united-airlines-uses-amazon-redshift-and-aws-glue-data-catalog-federation-to-query-databricks-managed-data/</link>
		
		<dc:creator><![CDATA[Vaibhav Agrawal]]></dc:creator>
		<pubDate>Tue, 15 Sep 2026 15:59:48 +0000</pubDate>
				<category><![CDATA[Advanced (300)]]></category>
		<category><![CDATA[Amazon Redshift]]></category>
		<category><![CDATA[AWS Glue]]></category>
		<category><![CDATA[Customer Solutions]]></category>
		<guid isPermaLink="false">e0f95e866482960f0d7ad6f8395513bc4a0c1d4c</guid>

					<description>Learn how United Airlines uses AWS Glue Data Catalog federation to query Databricks Unity Catalog data directly from Amazon Redshift Serverless without duplicating data, using resource links and AWS Lake Formation for governance.</description>
										<content:encoded>&lt;p&gt;&lt;em&gt;This post was co-written with Ankit Aggarwal and Raja Kalluri from United Airlines.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;United Airlines processes billions of events daily across its data platform, which spans Amazon Redshift and Databricks with Unity Catalog. To bridge these platforms without duplicating data, the team turned to &lt;a href="https://docs.aws.amazon.com/lake-formation/latest/dg/catalog-federation.html" target="_blank" rel="noopener"&gt;AWS Glue Data Catalog federation&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;In this post, we walk through how to configure AWS Glue Data Catalog federation to connect with Databricks Unity Catalog, so you can run live SQL queries from Amazon Redshift without moving or duplicating data.&lt;/p&gt;
&lt;h2 id="why-united-airlines-needed-catalog-federation"&gt;Why United Airlines needed catalog federation&lt;/h2&gt;
&lt;p&gt;United Airlines curates petabytes of data through a medallion architecture (bronze to silver to gold) on Amazon Simple Storage Service (Amazon S3). The airline user interaction data layer alone is several double-digit terabytes of near real-time streamed data. Teams use it to measure customer engagement patterns, feature adoption, and conversion behavior across web and mobile touchpoints. Analysts need to query this curated data through Amazon Redshift Serverless. As part of the existing data platform architecture these data tables are cataloged in Databricks Unity Catalog, not in the AWS Glue Data Catalog. As a result, Amazon Redshift has no native visibility into them. Without catalog federation, the only way to make this data queryable from Amazon Redshift would have been to duplicate it into Amazon Redshift Managed Storage (RMS) and build pipelines to keep it in sync.&lt;/p&gt;
&lt;p&gt;AWS Glue Data Catalog federation removed this need. Amazon Redshift users now query the gold layer stored in Amazon S3 directly, with Iceberg metadata resolved from Unity Catalog at query time and no data movement. AWS Glue Data Catalog federation connects Amazon Redshift to external catalogs like Unity Catalog, so analysts query cross-platform data without building sync pipelines or duplicating storage.&lt;/p&gt;
&lt;p&gt;Amazon Redshift Serverless is powered by the same Graviton-based query engine used in the new RG instance family, which delivers up to &lt;a href="https://aws.amazon.com/blogs/big-data/achieve-2x-faster-data-lake-query-performance-with-apache-iceberg-on-amazon-redshift/" target="_blank" rel="noopener"&gt;2x faster data lake query performance&lt;/a&gt; compared to prior generations. This engine is purpose-built for reading Apache Iceberg tables directly from Amazon S3, making it well-suited for such federated query workloads.&lt;/p&gt;
&lt;p&gt;United Airlines is taking a phased approach to adopting AWS Glue Data Catalog federation across its data platform. The initial focus is the most heavily used user interaction data tables, with 30 tables currently federated in production and 70 more in active rollout. Several hundred additional tables across different business domains are planned for production in the coming months.&lt;/p&gt;
&lt;h2 id="solution-overview"&gt;Solution overview&lt;/h2&gt;
&lt;p&gt;AWS Glue Data Catalog federation bridges these platforms at the metadata layer. Here’s how the architecture works.&lt;/p&gt;
&lt;p&gt;The architecture follows a four-layer federation chain:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;Databricks Unity Catalog exposes tables through its Iceberg REST API endpoint. For Delta tables, you can turn on UniForm format to make them Iceberg compatible.&lt;/li&gt;
 &lt;li&gt;AWS Glue Data Catalog creates a federated catalog that connects to Databricks Unity Catalog, making metadata visible within AWS without data movement.&lt;/li&gt;
 &lt;li&gt;A resource link database in the default AWS Glue catalog acts as a bridge, pointing to the federated catalog database. This is required for Amazon Redshift compute.&lt;/li&gt;
 &lt;li&gt;Amazon Redshift Serverless references the resource link database through an external schema. When a query runs, Amazon Redshift traverses the link, calls AWS Glue Federation, and reads the Iceberg data through the Databricks Unity Catalog REST API. AWS Lake Formation governs permissions throughout this chain.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Key services or service features used in this solution:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/glue/latest/dg/components-overview.html#data-catalog-intro" target="_blank" rel="noopener"&gt;AWS Glue Data Catalog&lt;/a&gt; – Federated catalog and resource link in the default catalog.&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/lake-formation/latest/dg/what-is-lake-formation.html" target="_blank" rel="noopener"&gt;AWS Lake Formation&lt;/a&gt; – Permission vending and fine-grained access control.&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/redshift/latest/mgmt/serverless-whatis.html" target="_blank" rel="noopener"&gt;Amazon Redshift Serverless&lt;/a&gt; – Queries through external schema.&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/IAM/latest/UserGuide/introduction.html" target="_blank" rel="noopener"&gt;AWS Identity and Access Management (IAM)&lt;/a&gt; – Role-based permissions for Amazon Redshift namespace and SAML-authenticated users.&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://docs.databricks.com/aws/en/data-governance/unity-catalog/" target="_blank" rel="noopener"&gt;Databricks Unity Catalog&lt;/a&gt; – Source catalog exposing Iceberg tables through a REST API.
  &lt;br&gt;
  &lt;figure&gt;
   &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-5928-1.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-5928-1.png" alt="Federation chain from Databricks Unity Catalog through AWS Glue and Lake Formation to Amazon Redshift Serverless" width="800"&gt;&lt;/a&gt;
   &lt;figcaption aria-hidden="true"&gt;Federation chain from Databricks Unity Catalog through AWS Glue and Lake Formation to Amazon Redshift Serverless&lt;/figcaption&gt;
  &lt;/figure&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;em&gt;Figure 1: Federation chain from Databricks Unity Catalog to Amazon Redshift Serverless through AWS Glue and Lake Formation&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;The architecture follows a six-step flow:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;A SQL analyst submits a query to Amazon Redshift Serverless.&lt;/li&gt;
 &lt;li&gt;Amazon Redshift resolves the external schema through the AWS Glue Data Catalog (resource link to federated catalog).&lt;/li&gt;
 &lt;li&gt;The AWS Glue federated catalog calls the Databricks Unity Catalog Iceberg REST API to retrieve current table metadata.&lt;/li&gt;
 &lt;li&gt;The namespace IAM role calls AWS Lake Formation GetDataAccess to obtain scoped, temporary S3 credentials.&lt;/li&gt;
 &lt;li&gt;Lake Formation evaluates fine-grained access policies and vends credentials for the authorized data files.&lt;/li&gt;
 &lt;li&gt;Amazon Redshift Serverless reads the Iceberg data files directly from S3 and returns results to the analyst.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="prerequisites"&gt;Prerequisites&lt;/h2&gt;
&lt;p&gt;Before you begin, make sure the following are in place:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;A Databricks workspace with Unity Catalog enabled and at least one catalog, schema, and table. Databricks uses UniForm to generate Iceberg metadata on Delta Lake tables on Amazon S3.&lt;/li&gt;
 &lt;li&gt;An AWS account with permissions to manage AWS Glue, AWS Lake Formation, Amazon Redshift Serverless, and IAM.&lt;/li&gt;
 &lt;li&gt;An Amazon Redshift Serverless workgroup and namespace already provisioned.&lt;/li&gt;
 &lt;li&gt;AWS Lake Formation set up with a data lake administrator.&lt;/li&gt;
 &lt;li&gt;AWS Command Line Interface (AWS CLI) configured with appropriate credentials.&lt;/li&gt;
 &lt;li&gt;Familiarity with Amazon Redshift Query Editor v2 or a SQL client.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; For setting up the Databricks Unity Catalog side (Phase 1), follow the steps in the AWS blog post &lt;a href="https://aws.amazon.com/blogs/big-data/access-databricks-unity-catalog-data-using-catalog-federation-in-the-aws-glue-data-catalog/" target="_blank" rel="noopener"&gt;Access Databricks Unity Catalog data using catalog federation in the AWS Glue Data Catalog.&lt;/a&gt; This walkthrough picks up after the federated catalog has been created in AWS Glue.&lt;/p&gt;
&lt;h2 id="solution-walkthrough"&gt;Solution walkthrough&lt;/h2&gt;
&lt;p&gt;The walkthrough is organized into six steps covering Lake Formation configuration, the resource link pattern, IAM role setup, and querying Databricks tables from Amazon Redshift.&lt;/p&gt;
&lt;h3 id="step-1-configure-aws-lake-formation"&gt;Step 1: Configure AWS Lake Formation&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;1a. Add a data lake administrator&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;In Lake Formation, choose &lt;strong&gt;Administration&lt;/strong&gt;, then choose&amp;nbsp;&lt;strong&gt;Administrators&lt;/strong&gt; and add your admin IAM user or role.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;1b. Confirm the federated catalog is registered&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;Choose &lt;strong&gt;Data Catalog&lt;/strong&gt;, then&amp;nbsp;&lt;strong&gt;Catalogs&lt;/strong&gt; and verify that databricks-federated-catalog is visible and registered.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="step-2-create-a-resource-link-in-the-default-aws-glue-catalog"&gt;Step 2: Create a resource link in the default AWS Glue catalog&lt;/h3&gt;
&lt;p&gt;This step is the key architectural detail in the walkthrough. Amazon Redshift resolves CREATE EXTERNAL SCHEMA only against the default AWS Glue Data Catalog. The federated catalog (databricks-federated-catalog) is a separate, non-default catalog object. To give Amazon Redshift a path to the federated data, you create a resource link database in the default catalog that points to the federated catalog’s database.&lt;/p&gt;
&lt;p&gt;A resource link does not copy data or metadata. It’s a pointer that Lake Formation resolves at query time.&lt;/p&gt;
&lt;p&gt;To create the resource link in the Lake Formation console:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;Choose &lt;strong&gt;Data Catalog&lt;/strong&gt;, &lt;strong&gt;Databases&lt;/strong&gt;,&amp;nbsp;&lt;strong&gt;Create database. &lt;/strong&gt;Then select &lt;strong&gt;Resource link&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;For Resource link name, enter databricks_federated_db_link.&lt;/li&gt;
 &lt;li&gt;For Target catalog, enter databricks-federated-catalog.&lt;/li&gt;
 &lt;li&gt;For Target database, enter the database name that was discovered by the AWS Glue crawler (for example, databricks_federated_db).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Alternatively, use the AWS CLI:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws glue create-database \
  --database-input '{
    "Name": "databricks_federated_db_link",
    "TargetDatabase": {
      "CatalogId": "&amp;lt;account-id&amp;gt;:databricks-federated-catalog",
      "DatabaseName": "databricks_federated_db"
    }
  }'&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h3 id="step-3-configure-the-amazon-redshift-serverless-namespace-iam-role"&gt;Step 3: Configure the Amazon Redshift Serverless namespace IAM role&lt;/h3&gt;
&lt;p&gt;When Amazon Redshift queries through the resource link, it uses the IAM role attached to the Amazon Redshift Serverless namespace to call the Lake Formation GetDataAccess API. Lake Formation permissions must be granted to this namespace role.&lt;/p&gt;
&lt;p&gt;Choose one of these two approaches:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;Option A – Update your existing namespace role by adding the following policy inline.&lt;/li&gt;
 &lt;li&gt;Option B – Create a new dedicated role (named RedshiftServerlessNamespaceRole) and attach it to the namespace alongside existing roles.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Attach the following IAM policy to the role:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-json"&gt;{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Action": [
        "glue:GetDatabase",
        "glue:GetDatabases",
        "glue:GetTable",
        "glue:GetTables",
        "glue:GetPartitions",
        "glue:GetCatalog",
        "glue:GetCatalogs"
      ],
      "Resource": "*"
    },
    {
      "Effect": "Allow",
      "Action": "lakeformation:GetDataAccess",
      "Resource": "*"
    }
  ]
}&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;Note: &lt;/strong&gt;&lt;/em&gt;The Resource: “*” in this policy is shown for simplicity. In production, scope resources to specific AWS Glue catalog ARNs, database ARNs, and table ARNs based on your use case.*&lt;/p&gt;
&lt;p&gt;After creating or updating the role, associate it with your Amazon Redshift Serverless namespace:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;In the Amazon Redshift Serverless console, choose &lt;strong&gt;Namespaces&lt;/strong&gt;, select [your namespace], then choose &lt;strong&gt;Security and encryption&lt;/strong&gt;, then&amp;nbsp;&lt;strong&gt;Manage IAM roles&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;If you use Option A, the existing role already has the new permissions, so no change is needed.&lt;/li&gt;
 &lt;li&gt;If you use Option B, add the new role alongside the existing roles.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="step-4-grant-lake-formation-permissions-to-the-amazon-redshift-namespace-role"&gt;Step 4: Grant Lake Formation permissions to the Amazon Redshift namespace role&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;4a. Grant DESCRIBE on the resource link database (default catalog)&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;In Lake Formation, choose &lt;strong&gt;Permissions,&lt;/strong&gt;&amp;nbsp;&lt;strong&gt;Data lake permissions&lt;/strong&gt;, then&amp;nbsp;&lt;strong&gt;Grant&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Principal: RedshiftServerlessNamespaceRole.&lt;/li&gt;
 &lt;li&gt;Resources: &lt;strong&gt;Named Data Catalog resources,&lt;/strong&gt;&amp;nbsp;&lt;strong&gt;Default catalog,&lt;/strong&gt;&amp;nbsp;&lt;strong&gt;databricks_federated_db_link (resouce link)&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Database permissions: DESCRIBE.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;4b. Grant SELECT and DESCRIBE on the target tables (Grant on Target)&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Resource links permit only DESCRIBE and DROP permissions on the link itself. To allow Amazon Redshift to actually read data, you must separately grant SELECT on the target tables in the federated catalog. This is the Lake Formation Grant on Target pattern.&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;Principal: RedshiftServerlessNamespaceRole.&lt;/li&gt;
 &lt;li&gt;Resources: &lt;strong&gt;Named Data Catalog resources,&lt;/strong&gt;&amp;nbsp;&lt;strong&gt;databricks-federated-catalog,&lt;/strong&gt;&amp;nbsp;&lt;strong&gt;databricks_federated_db&lt;/strong&gt;, then&amp;nbsp;&lt;strong&gt;Tables&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Table permissions: SELECT, DESCRIBE.&lt;/li&gt;
 &lt;li&gt;Catalog permission: DESCRIBE.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Important:&lt;/strong&gt; SELECT must be granted on the TARGET tables in the federated catalog, not on the resource link. Granting SELECT only on the resource link won’t work. This is a common configuration error.&lt;/p&gt;
&lt;h3 id="step-5-create-an-external-schema-in-amazon-redshift"&gt;Step 5: Create an external schema in Amazon Redshift&lt;/h3&gt;
&lt;p&gt;With the resource link in place and permissions granted, you can now create an external schema in Amazon Redshift that points to the resource link database. The external schema is the query interface. When a user runs SQL against it, Amazon Redshift traverses the link to the federated catalog and retrieves metadata and data from Databricks Unity Catalog.&lt;/p&gt;
&lt;p&gt;The DATABASE parameter must reference the resource link database name in the default AWS Glue catalog (databricks_federated_db_link), not the federated catalog name directly. The CATALOG_ARN parameter isn’t required here because the resource link lives in the default catalog and Amazon Redshift resolves it automatically.&lt;/p&gt;
&lt;p&gt;Connect to your Amazon Redshift cluster as a superuser (for example, using Amazon Redshift Query Editor v2) and run:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-sql"&gt;CREATE EXTERNAL SCHEMA databricks_schema
FROM DATA CATALOG
DATABASE 'databricks_federated_db_link'
IAM_ROLE '&amp;lt;iam-role-arn&amp;gt;'
REGION '&amp;lt;region&amp;gt;';&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;A key design principle in this architecture is the clear separation between data physically stored in Amazon Redshift and data accessed externally through federation. External schemas provide a transparent abstraction layer, so Amazon Redshift users can query data stored in S3 without ingestion. For consistency and clarity, United Airlines follows a standard naming convention for all federated schemas in Amazon Redshift: {domain}_iceberg. This convention makes it immediately clear that the data isn’t natively stored within Amazon Redshift but is accessed by using federation through AWS Glue and Lake Formation. This distinction is critical for analysts and engineers, because it improves discoverability, avoids ambiguity between storage layers, and reinforces architectural discipline when working across hybrid data environments.&lt;/p&gt;
&lt;p&gt;The User Interactions domain exposes curated datasets representing customer interaction activity, engagement behavior, and channel usage patterns. Operational datasets follow the same pattern, providing governed access to supporting business events and reference information through a common federation framework.&lt;/p&gt;
&lt;p&gt;You create a view layer over each external schema using WITH NO SCHEMA BINDING, so that analysts always resolve the freshest schema on each query execution. For example:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-sql"&gt;CREATE VIEW analytics.clickstream_events AS
SELECT * FROM {domain}_iceberg.interaction_events
WITH NO SCHEMA BINDING;&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h3 id="step-6-verify-and-query-databricks-tables-from-amazon-redshift"&gt;Step 6: Verify and query Databricks tables from Amazon Redshift&lt;/h3&gt;
&lt;p&gt;After creating the external schema, verify that the Databricks tables are visible and run a test query.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Verify table visibility&lt;/strong&gt;&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-sql"&gt;-- Confirm federated tables are visible in Redshift
SELECT * FROM SVV_EXTERNAL_TABLES
WHERE schemaname = 'databricks_schema';&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;Query a Databricks Unity Catalog table&lt;/strong&gt;&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-sql"&gt;-- Query a Databricks Unity Catalog table via the federated catalog
SELECT *
FROM databricks_schema.&amp;lt;table_name&amp;gt;
LIMIT 10;&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;When a query runs, Amazon Redshift calls Lake Formation GetDataAccess using the namespace IAM role to obtain temporary credentials. It then contacts the AWS Glue federated catalog, which in turn calls the Databricks Unity Catalog Iceberg REST API to retrieve metadata and read table data. The result is returned to the Amazon Redshift user transparently.&lt;/p&gt;
&lt;p&gt;For SAML-authenticated users, connect using your IdP JDBC plugin:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-text"&gt;jdbc:redshift:iam://&amp;lt;workgroup-name&amp;gt;.&amp;lt;account-id&amp;gt;.&amp;lt;region&amp;gt;.redshift-serverless.amazonaws.com:5439/&amp;lt;database&amp;gt;
?plugin_name=com.amazon.redshift.plugin.&amp;lt;YourIdPPlugin&amp;gt;
&amp;amp;idp_host=&amp;lt;your-idp-host&amp;gt;
&amp;amp;preferred_role=arn:aws:iam::&amp;lt;account-id&amp;gt;:role/RedshiftSAMLUserRole
&amp;amp;ssl=true&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;The Amazon Redshift JDBC driver handles authentication automatically. It authenticates with your IdP, receives a SAML assertion, and calls sts:AssumeRoleWithSAML for temporary IAM credentials. It then calls redshift-serverless:GetCredentials to connect as the mapped database user.&lt;/p&gt;
&lt;h2 id="business-impact"&gt;Business impact&lt;/h2&gt;
&lt;p&gt;AWS Glue Data Catalog federation delivered measurable architectural and operational improvements for United Airlines:&lt;/p&gt;
&lt;table border="1px" width="100%" cellpadding="10px"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;Area&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Before&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;After&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Impact&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Data access&lt;/td&gt;
   &lt;td&gt;Delta Lake and Amazon Redshift data were completely siloed, so Amazon Redshift users had no access to curated datasets on Databricks-managed S3 data&lt;/td&gt;
   &lt;td&gt;Amazon Redshift users get real-time access to Databricks-managed data through AWS Glue Data Catalog federation&lt;/td&gt;
   &lt;td&gt;~100 analysts gained access to user interaction data tables in the first phase without adding new pipelines.&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Disaster recovery&lt;/td&gt;
   &lt;td&gt;Cross-Region DR relied on Amazon Redshift snapshots every 3 hours (recovery point objective, or RPO, of 3 hours or more)&lt;/td&gt;
   &lt;td&gt;Amazon S3 cross-Region replication on the Delta Lake provides a near-continuous RPO. A new Amazon Redshift Serverless workgroup in the DR Region can federate to the same S3 data&lt;/td&gt;
   &lt;td&gt;More resilient architecture. Reduces cost for Amazon Redshift snapshot and copy maintenance across Regions&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Architecture simplification&lt;/td&gt;
   &lt;td&gt;Data processing happened in both Databricks and Amazon Redshift, requiring manual catalog synchronization between the two platforms which was operationally expensive and prone to drift&lt;/td&gt;
   &lt;td&gt;With the federated architecture, data processing is consolidated in Databricks, and Amazon Redshift acts solely as a query engine powering user queries and dashboards through catalog federation&lt;/td&gt;
   &lt;td&gt;Single processing platform, zero sync pipelines, single source of truth&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Infrastructure cost&lt;/td&gt;
   &lt;td&gt;Running dedicated Amazon Redshift ETL cluster with RMS storage, snapshots, and compute for data processing&lt;/td&gt;
   &lt;td&gt;For this use case with federation, Amazon Redshift is not needed for ETL but only as a query engine. No RMS storage duplication, no snapshot replication required&lt;/td&gt;
   &lt;td&gt;~$30K/month in redundant ETL infrastructure cost reduced&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="security-considerations"&gt;Security considerations&lt;/h2&gt;
&lt;p&gt;At United Airlines, identity governance is unified through Azure Active Directory groups. On the AWS consumption side, users authenticate to Amazon Redshift Serverless through SAML federation. AD group membership determines database-level access to federated schemas. On the Databricks side, the same AD groups govern access to Unity Catalog schemas. This single-identity model provides consistent access control across both platforms without requiring separate user provisioning. Lake Formation handles credential vending for S3 data access during federated queries, while schema-level access decisions are managed through the AD group mappings on each platform.&lt;/p&gt;
&lt;p&gt;The architecture also provides multiple layers of security controls built into the federation chain:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;AWS Lake Formation governs fine-grained access control throughout the federation chain, so that principals can only access authorized databases, tables, and columns.&lt;/li&gt;
 &lt;li&gt;IAM roles follow least-privilege principles. The Amazon Redshift namespace role is scoped only to AWS Glue metadata operations and Lake Formation GetDataAccess.&lt;/li&gt;
 &lt;li&gt;SAML-based authentication integrates enterprise identity providers, so that users authenticate through existing SSO infrastructure before accessing federated data.&lt;/li&gt;
 &lt;li&gt;All Amazon Redshift connections enforce TLS encryption (ssl=true), protecting data in transit between clients and the Amazon Redshift endpoint.&lt;/li&gt;
 &lt;li&gt;Lake Formation permission vending issues short-lived, scoped credentials for each query execution rather than long-lived static credentials.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="other-considerations"&gt;Other considerations&lt;/h2&gt;
&lt;p&gt;Review the &lt;a href="https://docs.aws.amazon.com/lake-formation/latest/dg/catalog-federation.html#catalog-federation-limitations" target="_blank" rel="noopener"&gt;catalog federation service limitations&lt;/a&gt; before deploying. Key requirements:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;Delta Lake tables must have UniForm enabled to expose Iceberg-compatible metadata.&lt;/li&gt;
 &lt;li&gt;We recommend that source tables be well-partitioned and regularly compacted, because the federated query performance reflects how efficiently the data is organized at write time.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="clean-up"&gt;Clean up&lt;/h2&gt;
&lt;p&gt;To avoid ongoing charges for resources created in this walkthrough, remove them in the following order. This teardown doesn’t affect Databricks metadata or your underlying data stored in Amazon S3.&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;Drop the external schema in Amazon Redshift: &lt;code&gt;DROP SCHEMA databricks_schema;&lt;/code&gt;.&lt;/li&gt;
 &lt;li&gt;Delete the resource link database in the default AWS Glue catalog (databricks_federated_db_link).&lt;/li&gt;
 &lt;li&gt;Revoke Lake Formation permissions granted to the Amazon Redshift namespace role on both the resource link database and the target tables in the federated catalog.&lt;/li&gt;
 &lt;li&gt;Delete the federated catalog in AWS Glue (databricks-federated-catalog).&lt;/li&gt;
 &lt;li&gt;Deregister the AWS Glue connection for the Databricks Unity Catalog if no longer needed.&lt;/li&gt;
 &lt;li&gt;Optionally, remove the IAM role (RedshiftServerlessNamespaceRole) if it was created solely for this walkthrough.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;In this post, we showed how United Airlines uses AWS Glue Data Catalog federation to give Amazon Redshift Serverless analysts real-time access to double-digit terabytes of curated user interaction data on Amazon S3, without duplicating a single byte or building sync pipelines.&lt;/p&gt;
&lt;p&gt;The architecture uses the Iceberg REST API, resource link databases, and Lake Formation credential vending to create a governed query path between Amazon Redshift and Unity Catalog. For United Airlines, this eliminated redundant ETL infrastructure costs, removed the need for catalog synchronization, and turned Amazon Redshift Serverless into a dedicated high-performance query engine for analysts and dashboards.&lt;/p&gt;
&lt;p&gt;For questions or feedback, leave a comment on this post.&lt;/p&gt;
&lt;p style="clear: both"&gt;&lt;/p&gt;
&lt;hr style="width: 100%"&gt;
&lt;h2&gt;About the authors&lt;/h2&gt;
&lt;footer&gt;
 &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
  &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
   &lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-5928-2.jpg" alt="Vaibhav Agrawal" width="100" height="133"&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Vaibhav Agrawal&lt;/h3&gt;
  &lt;p style="overflow: hidden"&gt;Vaibhav Agrawal is a Senior Analytics Specialist Solutions Architect at AWS, focused on helping enterprise customers design and implement modern data architectures using AWS Analytics services.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
  &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
   &lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/10/BDB-5928-3.jpg" alt="Ankit Aggarwal" width="100" height="133"&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Ankit Aggarwal&lt;/h3&gt;
  &lt;p style="overflow: hidden"&gt;Ankit Aggarwal is a Principal Enterprise Architect at United Airlines, where he leads the United Data Hub (UDH) platform architecture—a petabyte-scale data platform built on AWS and Databricks. He brings over 15 years of experience in data engineering and enterprise architecture.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
  &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
   &lt;img loading="lazy" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/15/raja-kalluri.png" alt="" width="100" height="109" class="alignnone size-full wp-image-94386"&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Raja Kalluri&lt;/h3&gt;
  &lt;p style="overflow: hidden"&gt;Raja Kalluri is a Principal Architect at United Airlines, where he leads enterprise-scale data architecture and modernization initiatives. He specializes in building cloud-native data platforms, enabling real-time analytics and AI, and transforming legacy ecosystems.&lt;/p&gt;
 &lt;/div&gt;
&lt;/footer&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>Scale down Kinesis Data Streams on-demand capacity with ODA warm throughput</title>
		<link>https://aws.amazon.com/blogs/big-data/scale-down-kinesis-data-streams-on-demand-capacity-with-oda-warm-throughput/</link>
		
		<dc:creator><![CDATA[Pratik Patel]]></dc:creator>
		<pubDate>Mon, 14 Sep 2026 15:36:36 +0000</pubDate>
				<category><![CDATA[Advanced (300)]]></category>
		<category><![CDATA[Announcements]]></category>
		<category><![CDATA[Kinesis Data Streams]]></category>
		<guid isPermaLink="false">1829c8cc55833a33b97ef1faeb95d4da5454dce4</guid>

					<description>Amazon Kinesis Data Streams now supports scaling down ingest capacity for on-demand Advantage streams with warm throughput. Learn how the scale-down works, how to monitor stream behavior with Amazon CloudWatch, and best practices for releasing excess capacity after transient traffic bursts.</description>
										<content:encoded>&lt;p&gt;Customers have been using Amazon Kinesis Data Streams to stream data at any scale. Some use On-demand Standard to let the service manage capacity, while others with predictable traffic patterns use &lt;span&gt;&lt;a href="https://aws.amazon.com/blogs/big-data/amazon-kinesis-data-streams-launches-on-demand-advantage-for-instant-throughput-increases-and-streaming-at-scale/" target="_blank" rel="noopener noreferrer"&gt;On-demand Advantage and warm throughput &lt;/a&gt;&lt;/span&gt; to ensure streams can handle instant throughput increases. Streaming workloads rarely run at peak volume all the time: flash sales end, batch migrations complete, and telemetry bursts subside. However, manual intervention is often required to scale back down after the burst subsides. &lt;u&gt;&lt;span&gt;&lt;a href="https://staging.prod.website.marketing.aws.dev/kinesis/data-streams/" target="_blank" rel="noopener noreferrer"&gt;Amazon Kinesis Data Streams&lt;/a&gt;&lt;/span&gt;&lt;/u&gt; now supports &lt;u&gt;&lt;span&gt;&lt;a href="https://staging.prod.website.marketing.aws.dev/about-aws/whats-new/2026/07/kinesis/on-demand-scale-down/" target="_blank" rel="noopener noreferrer"&gt;scaling down ingest capacity for on-demand Advantage streams with warm throughput&lt;/a&gt;&lt;/span&gt;&lt;/u&gt;, which optimizes downstream compute costs and performance by removing excess capacity. You configure this by turning on On-demand Advantage mode (ODA) and setting a new warm throughput value that is equal to or smaller than the existing amount.&lt;/p&gt;
&lt;p&gt;With this launch, you can now proactively reduce write throughput capacity, optimizing costs while maintaining performance and giving you more control over your stream’s provisioning.&lt;/p&gt;
&lt;p&gt;In this post, we explore the warm throughput scale-down capability. We cover the challenge it addresses, how it works, how to monitor stream behavior with Amazon CloudWatch metrics, and best practices for using it effectively.&lt;/p&gt;
&lt;h2 id="the-challenge-excess-capacity-after-traffic-spikes"&gt;The challenge: Excess capacity after traffic spikes&lt;/h2&gt;
&lt;p&gt;Amazon Kinesis Data Streams on-demand mode automatically scales to handle increases in data throughput. When your stream experiences a traffic spike, Kinesis Data Streams splits shards to accommodate the higher volume. This automatic scaling helps your applications keep pace with data during surges.&lt;/p&gt;
&lt;p&gt;However, many real-world workloads experience transient bursts that don’t represent sustained throughput needs. Consider a retail platform that processes a flash sale event, a healthcare system that ingests a large batch of patient records during a migration window, or an Internet of Things (IoT) fleet that transmits a high-volume firmware update telemetry burst. In each scenario, the stream scales up to accommodate the spike, but the elevated capacity remains long after the burst has subsided. Although Kinesis on-demand Advantage doesn’t charge for the elevated capacity, your consuming applications may see a higher cost and lower performance.&lt;/p&gt;
&lt;p&gt;Consider a Kinesis data stream running with 100 MB/s ingest throughput that requires 100 shards. A traffic spike of an additional 50 MB/s forces on-demand mode to scale streams to 150 shards. The spike subsides within minutes, but those 150 shards remain.&lt;/p&gt;
&lt;p&gt;If your AWS Lambda consumer uses a parallelization factor of 2, you go from 200 concurrent invocations (2 × 100 shards) to 300 (2 × 150 shards). This is a 50 percent jump in concurrent Lambda execution, even though ingest throughput has returned to 100 MB/s. Those extra 100 AWS Lambda invocations consume compute, count against your concurrent execution quota, and add cost while processing data with small batch sizes.&lt;/p&gt;
&lt;p&gt;Kinesis Client Library (KCL) consumers incur operational overhead. KCL tracks one lease per shard in Amazon DynamoDB, so 50 additional shards mean 50 more leases to scan, renew, and checkpoint every heartbeat cycle. The result is more Amazon DynamoDB overhead for lease management and reduced consumption performance overall.&lt;/p&gt;
&lt;p&gt;Before this launch, you had limited options to address this excess capacity:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;Switch to provisioned mode to manually set shard count, losing the benefits of automatic scaling.&lt;/li&gt;
 &lt;li&gt;Accept the higher capacity and associated costs until the stream self-adjusted.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These approaches either introduced operational overhead or resulted in paying for capacity that exceeded your workload’s actual requirements.&lt;/p&gt;
&lt;h2 id="the-solution-warm-throughput-scale-down"&gt;The solution: Warm throughput scale-down&lt;/h2&gt;
&lt;p&gt;With on-demand capacity reduction, you can now set a lower or equal warm throughput value on your on-demand stream to trigger a capacity reduction. The stream adjusts to the requested capacity or the amount needed to support peak data ingest usage within the last hour, whichever is higher. This safeguard helps your stream retain sufficient capacity for current traffic while releasing the excess you no longer need.&lt;/p&gt;
&lt;p&gt;This capability is available at no additional cost for all on-demand streams that have &lt;a href="https://docs.aws.amazon.com/streams/latest/dev/how-do-i-size-a-stream.html" target="_blank" rel="noopener"&gt;On-demand Advantage mode&lt;/a&gt; turned on.&lt;/p&gt;
&lt;h2 id="how-it-works"&gt;How it works&lt;/h2&gt;
&lt;p&gt;Warm throughput provides bidirectional capacity management for on-demand streams:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;Scale up (existing capability):&lt;/strong&gt; If you forecast an upcoming traffic event, you can configure warm throughput to a higher value to prepare the stream in advance so that capacity is available when data arrives without throttling.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Scale down (new capability):&lt;/strong&gt; If a transient burst has caused the stream to scale significantly beyond its steady-state needs, you can trigger a scale-down by setting warm throughput to a lower value.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;When you set a warm throughput value that is equal to or lower than the current value on an on-demand stream, Kinesis Data Streams evaluates the request against your stream’s recent traffic. The resulting capacity is the greater of:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;The warm throughput value you requested.&lt;/li&gt;
 &lt;li&gt;The capacity needed to support peak data ingest usage within the last hour.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This mechanism prevents you from accidentally reducing capacity below what your current workload demands. If data traffic increases after a scale-down has completed, on-demand mode can still expand stream ingest capacity through reactive scaling to avoid rate limiting.&lt;/p&gt;
&lt;h2 id="getting-started"&gt;Getting started&lt;/h2&gt;
&lt;h3 id="prerequisites"&gt;Prerequisites&lt;/h3&gt;
&lt;p&gt;To follow along, you need the following:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;An existing Kinesis data stream in on-demand mode.&lt;/li&gt;
 &lt;li&gt;On-demand Advantage mode turned on.&lt;/li&gt;
 &lt;li&gt;AWS Command Line Interface (AWS CLI) installed and configured.&lt;/li&gt;
 &lt;li&gt;AWS Identity and Access Management (IAM) permissions for kinesis:UpdateStreamMode.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;To trigger a scale-down, set a lower warm throughput value on your on-demand stream using the AWS CLI:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws kinesis update-stream-mode \
--stream-arn arn:aws:kinesis:us-east-1:111122223333:stream/my-stream/my-stream \
--warm-throughput-in-mb 50&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h2 id="monitoring-stream-behavior-with-amazon-cloudwatch"&gt;Monitoring stream behavior with Amazon CloudWatch&lt;/h2&gt;
&lt;p&gt;To observe the effects of a scale-down operation and understand your stream’s capacity and shard count, Amazon CloudWatch provides several key metrics. Monitoring these metrics helps you make informed decisions about when and how much to scale down.&lt;/p&gt;
&lt;h3 id="key-metrics-to-monitor"&gt;Key metrics to monitor&lt;/h3&gt;
&lt;p&gt;The following table summarizes the CloudWatch metrics most relevant to warm throughput scale-down:&lt;/p&gt;
&lt;table border="1px" width="100%" cellpadding="10px"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;Metric&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Namespace&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Description&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;code&gt;IncomingBytes&lt;/code&gt;&lt;/td&gt;
   &lt;td&gt;AWS/Kinesis&lt;/td&gt;
   &lt;td&gt;Total bytes ingested per period. Use the Sum statistic to see aggregate throughput across all shards.&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;code&gt;IncomingRecords&lt;/code&gt;&lt;/td&gt;
   &lt;td&gt;AWS/Kinesis&lt;/td&gt;
   &lt;td&gt;Total records ingested per period. Helps identify traffic patterns and burst frequency.&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;code&gt;WriteProvisionedThroughputExceeded&lt;/code&gt;&lt;/td&gt;
   &lt;td&gt;AWS/Kinesis&lt;/td&gt;
   &lt;td&gt;Number of records rejected because of throttling. A non-zero value after scale-down indicates capacity is set too low.&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;h3 id="observing-shard-count-behavior-during-scale-down"&gt;Observing shard count behavior during scale-down&lt;/h3&gt;
&lt;p&gt;To track shard count changes resulting from a scale-down, use the &lt;a href="https://docs.aws.amazon.com/kinesis/latest/APIReference/API_DescribeStreamSummary.html" target="_blank" rel="noopener"&gt;DescribeStreamSummary&lt;/a&gt; API, which returns the &lt;code&gt;OpenShardCount&lt;/code&gt; field in its response. Note that &lt;code&gt;OpenShardCount&lt;/code&gt; is not a CloudWatch metric. It’s available through the API and is also displayed on the Kinesis Data Streams console. You can poll this value periodically or build a custom CloudWatch metric using an AWS Lambda function to track shard count over time.&lt;/p&gt;
&lt;p&gt;Here is how you can expect the stream to behave:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;&lt;strong&gt;Before the burst:&lt;/strong&gt; Your stream operates at steady-state with a baseline shard count appropriate for your normal traffic. For example, a stream handling 20 MiB/s of write throughput might have approximately 67 open shards.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;During the burst:&lt;/strong&gt; As traffic spikes, Kinesis Data Streams automatically splits shards to accommodate the increased load. The &lt;code&gt;OpenShardCount&lt;/code&gt; rises, and &lt;code&gt;IncomingBytes&lt;/code&gt; increases correspondingly.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;After the burst (before scale-down):&lt;/strong&gt; Traffic returns to baseline, but the &lt;code&gt;OpenShardCount&lt;/code&gt; remains elevated because the stream retains capacity for up to double the recently observed peak.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;After triggering scale-down:&lt;/strong&gt; After you set a lower warm throughput, the &lt;code&gt;OpenShardCount&lt;/code&gt; decreases as Kinesis Data Streams merges shards to match the requested capacity (subject to the one-hour peak safeguard). You can observe this transition by polling DescribeStreamSummary or on the Kinesis console.&lt;/li&gt;
&lt;/ol&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/08/BDB-6001-1.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/08/BDB-6001-1.png" alt="Chart showing Kinesis Data Streams shard count rising during a traffic spike and decreasing after a warm throughput scale-down" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 1: Amazon Kinesis Data Streams shard count over time during a scale-down event, showing the incoming-data spike and the resulting change in shard count&lt;/p&gt;
&lt;/div&gt;
&lt;h2 id="best-practices"&gt;Best practices&lt;/h2&gt;
&lt;p&gt;When using warm throughput scale-down, consider the following recommendations:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;&lt;strong&gt;Analyze traffic patterns before scaling down.&lt;/strong&gt; Review at least 24 hours of &lt;code&gt;IncomingBytes&lt;/code&gt; and &lt;code&gt;IncomingRecords&lt;/code&gt; CloudWatch metrics to understand your baseline throughput before setting a lower warm throughput value. This helps you avoid setting capacity below your actual steady-state needs.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Set warm throughput above your observed steady-state peak.&lt;/strong&gt; Because on-demand streams accommodate up to double the observed peak, set your target warm throughput at or above your typical peak rather than your average. This maintains headroom for normal traffic variability without throttling.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Monitor throttling after scale-down.&lt;/strong&gt; Watch &lt;code&gt;WriteProvisionedThroughputExceeded&lt;/code&gt; closely in the hours following a scale-down. If throttling occurs, increase the warm throughput value. The stream will automatically scale back up, but proactive monitoring reduces the duration of any impact.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Use scale-down after known transient events.&lt;/strong&gt; The feature is most effective when you can identify that a traffic spike was temporary, for example, after a planned batch migration, marketing event, or scheduled data backfill. Avoid scaling down during periods of uncertain or growing traffic.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Use the one-hour safeguard.&lt;/strong&gt; The system won’t reduce capacity below what’s needed to serve peak ingest from the last hour. If you’re unsure about the right target, you can set a low warm throughput value and rely on this safeguard to prevent under-provisioning for active traffic.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;Amazon Kinesis Data Streams now supports scaling down ingest capacity with warm throughput, giving you elastic control over On-demand Advantage stream capacity. With this capability, you can release excess capacity after transient traffic bursts, improving cost efficiency while maintaining the automatic scaling benefits of on-demand mode.&lt;/p&gt;
&lt;p&gt;To get started, turn on &lt;a href="https://aws.amazon.com/blogs/big-data/amazon-kinesis-data-streams-launches-on-demand-advantage-for-instant-throughput-increases-and-streaming-at-scale/" target="_blank" rel="noopener"&gt;On-demand Advantage mode&lt;/a&gt; for your stream and use the warm throughput setting to manage capacity. Track shard count with DescribeStreamSummary to observe capacity changes and confirm your stream keeps appropriate headroom for your workload. Try warm throughput scale-down today in the &lt;a href="https://console.aws.amazon.com/kinesis/home" target="_blank" rel="noopener"&gt;Amazon Kinesis console&lt;/a&gt;, and to learn more, see &lt;a href="https://docs.aws.amazon.com/streams/latest/dev/how-do-i-size-a-stream.html" target="_blank" rel="noopener"&gt;Amazon Kinesis Data Streams on-demand capacity mode&lt;/a&gt; in the Developer Guide.&lt;/p&gt;
&lt;p style="clear: both"&gt;&lt;/p&gt;
&lt;hr style="width: 100%"&gt;
&lt;h2&gt;About the authors&lt;/h2&gt;
&lt;footer&gt;
 &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
  &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
   &lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/08/BDB-6001-2.jpg" alt="Pratik Patel" width="100" height="133"&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Pratik Patel&lt;/h3&gt;
  &lt;p style="overflow: hidden"&gt;Pratik is Sr Technical Account Manager and streaming analytics specialist. He works with AWS customers and provides ongoing support and technical guidance to help plan and build solutions using best practices and proactively helps in keeping customers’ AWS environments operationally healthy.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
  &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
   &lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/08/BDB-6001-3.jpg" alt="Priyanka Chaudhary" width="100" height="133"&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Priyanka Chaudhary&lt;/h3&gt;
  &lt;p style="overflow: hidden"&gt;Priyanka is Senior Solutions Architect at AWS. She is specialized in data lake and analytics services and helps many customers in this area. As a Solutions Architect, she plays a crucial role in guiding strategic customers through their cloud journey by designing scalable and secure cloud solutions. Outside of work, she loves spending time with friends and family, watching movies, and traveling.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
  &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
   &lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/08/BDB-6001-4.jpg" alt="Varsha Palepu" width="100" height="133"&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Varsha Palepu&lt;/h3&gt;
  &lt;p style="overflow: hidden"&gt;&lt;a href="https://www.linkedin.com/in/varsha-palepu/" target="_blank" rel="noopener"&gt;Varsha&lt;/a&gt; is a Solutions Architect and an analytics specialist on the AWS streaming team. She helps small and medium businesses innovate on AWS and creates technical streaming content to empower customers in their cloud journey.&lt;/p&gt;
 &lt;/div&gt;
&lt;/footer&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>How to migrate from Amazon CloudSearch to Amazon OpenSearch Serverless</title>
		<link>https://aws.amazon.com/blogs/big-data/how-to-migrate-from-amazon-cloudsearch-to-amazon-opensearch-serverless/</link>
		
		<dc:creator><![CDATA[Prasad Nadig]]></dc:creator>
		<pubDate>Thu, 10 Sep 2026 16:06:42 +0000</pubDate>
				<category><![CDATA[Advanced (300)]]></category>
		<category><![CDATA[Amazon CloudSearch]]></category>
		<category><![CDATA[Amazon OpenSearch Serverless]]></category>
		<category><![CDATA[Technical How-to]]></category>
		<guid isPermaLink="false">bf777f5bab542f56a24f05e05d9b0c2ad22ded30</guid>

					<description>Learn how to migrate an Amazon CloudSearch domain to Amazon OpenSearch Serverless: assess your configuration, create a collection with explicit index mappings, convert your documents and queries to the OpenSearch query DSL, configure security, load data with Amazon OpenSearch Ingestion, and validate before cutover.</description>
										<content:encoded>&lt;p&gt;If you run search on &lt;a href="https://aws.amazon.com/cloudsearch/" target="_blank" rel="noopener"&gt;Amazon CloudSearch&lt;/a&gt;, now is the time to plan your migration to &lt;a href="https://aws.amazon.com/opensearch-service/serverless/" target="_blank" rel="noopener"&gt;Amazon OpenSearch Serverless&lt;/a&gt;. Modern search has moved on to capabilities beyond what CloudSearch provides: semantic and hybrid search, Retrieval Augmented Generation (RAG), and agentic search. OpenSearch Serverless gives you all of these with automatic scaling on a pay-for-what-you-use basis. You don’t need to choose or maintain infrastructure. OpenSearch Serverless maintains the hands-off, operational simplicity of CloudSearch.&lt;/p&gt;
&lt;p&gt;This post shows you how to migrate your CloudSearch domain to an &lt;a href="https://aws.amazon.com/opensearch-service/" target="_blank" rel="noopener"&gt;Amazon OpenSearch Serverless&lt;/a&gt; collection. We walk you through assessing your CloudSearch configuration, creating an OpenSearch Serverless collection with explicit index mappings, converting your documents and queries, configuring security policies, loading your data with Amazon OpenSearch Ingestion, and validating the migration before cutting over.&lt;/p&gt;
&lt;h3 id="key-differences-to-note"&gt;Key differences to note&lt;/h3&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;Independent scaling&lt;/strong&gt; – OpenSearch Serverless scales indexing and search compute independently and can scale compute to zero when a collection is idle (you still pay for storage). For the cost structure, see &lt;a href="https://docs.aws.amazon.com/opensearch-service/latest/developerguide/serverless-scaling.html" target="_blank" rel="noopener"&gt;Managing capacity limits for Amazon OpenSearch Serverless&lt;/a&gt;.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Layered security model&lt;/strong&gt; – OpenSearch Serverless applies encryption, network, and data access policies at separate layers. For the security model, see &lt;a href="https://docs.aws.amazon.com/opensearch-service/latest/developerguide/serverless-security.html" target="_blank" rel="noopener"&gt;Security in Amazon OpenSearch Serverless&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="prerequisites"&gt;Prerequisites&lt;/h2&gt;
&lt;p&gt;To follow along with this post, you need the following:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;An AWS account.&lt;/li&gt;
 &lt;li&gt;An existing Amazon CloudSearch domain with indexed data.&lt;/li&gt;
 &lt;li&gt;Source data available in a durable store such as &lt;a href="https://aws.amazon.com/s3/" target="_blank" rel="noopener"&gt;Amazon Simple Storage Service (Amazon S3)&lt;/a&gt; or &lt;a href="https://aws.amazon.com/dynamodb/" target="_blank" rel="noopener"&gt;Amazon DynamoDB&lt;/a&gt; (CloudSearch doesn’t provide a built-in export or backup feature, so your original source data is required to re-ingest into OpenSearch).&lt;/li&gt;
 &lt;li&gt;AWS Identity and Access Management (IAM) permissions to create and manage Amazon OpenSearch Serverless collections, encryption policies, network policies, and data access policies.&lt;/li&gt;
 &lt;li&gt;An Amazon OpenSearch Ingestion pipeline (or alternative ingestion method) for loading data.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="plan-the-migration"&gt;Plan the migration&lt;/h2&gt;
&lt;p&gt;Planning is where you decide what success means: minimal downtime, no data loss, current functionality preserved, and custom configurations carried over. You don’t need to plan for infrastructure because OpenSearch Serverless provisions and scales compute for you. Your main planning task is to assess your current CloudSearch configuration so you can reproduce its behavior on the target.&lt;/p&gt;
&lt;p&gt;Document your existing setup from the Amazon CloudSearch console. Record the current instance type, the partition count, and the replication count. Capture the total document count and overall data size, and record every field definition, including field types and the search, facet, and sort settings for each field. Note any analyzers, synonyms, stopwords, or custom rank expressions. Note whether you use the 2011 or the 2013 CloudSearch API version, because the 2013 API added faceting and filtering features that change how you model the target.&lt;/p&gt;
&lt;p&gt;OpenSearch Serverless is the right target for most CloudSearch workloads, but not all of them. If your workload needs very low read-after-write latency (a short refresh interval), tight and predictable query response times, or direct control over instance configuration, choose an &lt;a href="https://docs.aws.amazon.com/opensearch-service/latest/developerguide/sizing-domains.html" target="_blank" rel="noopener"&gt;Amazon OpenSearch Service managed clusters deployment instead and size it from your workload profile&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The migration involves four main concerns: your source data format, your queries, your field definitions, and your access policies. Before you plan the details, it helps to see the whole migration at once. The following diagram maps the migration across four phases: your source CloudSearch environment, the migration pipeline that converts and moves your data, the OpenSearch Serverless target, and cutover and operations.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/09/BDB-5903-1.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/09/BDB-5903-1.png" alt="Migration workflow across four phases: source CloudSearch, migration pipeline, OpenSearch Serverless target, and cutover and operations" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 1: The migration workflow across four phases&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;In the source environment, you assess your CloudSearch configuration and back up your source data (Amazon S3, Amazon DynamoDB, or another store). Note the Source Data Format (SDF), the URL-based query syntax, and the IAM access policies you need to carry over. In the migration pipeline, you map field types, convert the data format from CloudSearch JSON to OpenSearch-compatible JSON, convert your queries to the OpenSearch query domain-specific language (DSL), configure security, bulk-ingest the data, and validate the result. The OpenSearch Serverless target holds the collection, index mappings, ingested documents, and the encryption, network, and data access policies, and it scales with your workload on a pay-per-use basis. In cutover and operations, you update your application to the new endpoint and clients, monitor with &lt;a href="https://aws.amazon.com/cloudwatch/" target="_blank" rel="noopener"&gt;Amazon CloudWatch&lt;/a&gt;, and decommission CloudSearch once no traffic remains.&lt;/p&gt;
&lt;h2 id="model-your-data-in-opensearch-service"&gt;Model your data in OpenSearch Service&lt;/h2&gt;
&lt;p&gt;OpenSearch Service uses &lt;a href="https://opensearch.org/docs/latest/field-types/" target="_blank" rel="noopener"&gt;index mappings&lt;/a&gt; to define the fields and data types in an index. Because you know your CloudSearch schema, define the target mapping explicitly when you create the index. Create the index and set its mapping in a single request, and set &lt;code&gt;dynamic&lt;/code&gt; to &lt;code&gt;strict&lt;/code&gt; so OpenSearch rejects any document that contains a field you did not define. Strict mapping catches schema drift at ingest time, avoiding the default OpenSearch behavior of creating new mappings for undefined fields.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-json"&gt;PUT /imdb_movies
{
  "mappings": {
    "dynamic": "strict",
    "properties": {
      "title": {
        "type": "text",
        "fields": {
          "keyword": { "type": "keyword" }
        }
      },
      "genres": { "type": "keyword" },
      "rating": { "type": "float" },
      "release_date": { "type": "date" }
      ...
    }
  }
}&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h3 id="field-type-mapping"&gt;Field type mapping&lt;/h3&gt;
&lt;p&gt;The following table maps CloudSearch field types to their OpenSearch Service equivalents.&lt;/p&gt;
&lt;table border="1px" width="100%" cellpadding="10px"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;CloudSearch&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;OpenSearch Service equivalent&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Notes&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;text&lt;/td&gt;
   &lt;td&gt;text&lt;/td&gt;
   &lt;td&gt;Text is tokenized. Stemming, synonyms, and stopwords apply. Good for matching user terms.&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;literal&lt;/td&gt;
   &lt;td&gt;keyword&lt;/td&gt;
   &lt;td&gt;Not tokenized. Good for exact-match search.&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;int&lt;/td&gt;
   &lt;td&gt;integer&lt;/td&gt;
   &lt;td&gt;Use for ranking, faceting, and narrowing.&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;double&lt;/td&gt;
   &lt;td&gt;float or double&lt;/td&gt;
   &lt;td&gt;&lt;span style="color: #ffffff"&gt;.&lt;/span&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;date&lt;/td&gt;
   &lt;td&gt;date&lt;/td&gt;
   &lt;td&gt;&lt;span style="color: #ffffff"&gt;.&lt;/span&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;boolean&lt;/td&gt;
   &lt;td&gt;boolean&lt;/td&gt;
   &lt;td&gt;&lt;span style="color: #ffffff"&gt;.&lt;/span&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;latlon&lt;/td&gt;
   &lt;td&gt;geo_point&lt;/td&gt;
   &lt;td&gt;&lt;span style="color: #ffffff"&gt;.&lt;/span&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;text-array&lt;/td&gt;
   &lt;td&gt;text&lt;/td&gt;
   &lt;td&gt;OpenSearch handles arrays natively, so map to the base text type.&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;literal-array&lt;/td&gt;
   &lt;td&gt;keyword&lt;/td&gt;
   &lt;td&gt;OpenSearch handles arrays natively, so map to the base keyword type.&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;multi-value&lt;/td&gt;
   &lt;td&gt;nested or object&lt;/td&gt;
   &lt;td&gt;&lt;span style="color: #ffffff"&gt;.&lt;/span&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;long&lt;/td&gt;
   &lt;td&gt;long&lt;/td&gt;
   &lt;td&gt;&lt;span style="color: #ffffff"&gt;.&lt;/span&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;binary&lt;/td&gt;
   &lt;td&gt;binary&lt;/td&gt;
   &lt;td&gt;&lt;span style="color: #ffffff"&gt;.&lt;/span&gt;&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Two mapping details deserve attention. First, pick the smallest numeric type that fits your data rather than copying the widths CloudSearch uses. CloudSearch stores integers as 64-bit values, but few datasets hold numbers that large. A &lt;code&gt;long&lt;/code&gt; or a &lt;code&gt;double&lt;/code&gt; consumes more disk than an &lt;code&gt;integer&lt;/code&gt;, a &lt;code&gt;short&lt;/code&gt;, or a &lt;code&gt;float&lt;/code&gt; with no benefit when the values are small. Evaluate the actual range of each field and choose the narrowest type that holds it. Reserve &lt;code&gt;long&lt;/code&gt; for values that genuinely exceed the roughly 2.1 billion ceiling of &lt;code&gt;integer&lt;/code&gt;, and use &lt;code&gt;float&lt;/code&gt; instead of &lt;code&gt;double&lt;/code&gt; unless you need double precision. Smaller types shrink your index and speed up queries.&lt;/p&gt;
&lt;p&gt;Second, if you sort or aggregate on a &lt;code&gt;text&lt;/code&gt; field, add a &lt;code&gt;keyword&lt;/code&gt; sub-field. The preceding example mapping has a &lt;code&gt;keyword&lt;/code&gt; subfield for the &lt;code&gt;title&lt;/code&gt; field. You access the field using dot notation: &lt;code&gt;title.keyword&lt;/code&gt;. OpenSearch doesn’t sort or aggregate analyzed &lt;code&gt;text&lt;/code&gt; fields by default.&lt;/p&gt;
&lt;p&gt;As noted earlier, if you run several CloudSearch domains, model each one as a separate index within a single OpenSearch Serverless collection to consolidate them.&lt;/p&gt;
&lt;h2 id="move-your-data"&gt;Move your data&lt;/h2&gt;
&lt;p&gt;Migrating to OpenSearch Service is a re-ingestion: you convert your source documents and index them into the collection you created. CloudSearch doesn’t provide a built-in backup or snapshot feature. It relies on the documents you send through the indexing process, so before you migrate, make sure your source data is available in a durable store such as Amazon S3, Amazon DynamoDB, or another database.&lt;/p&gt;
&lt;p&gt;The conversion is a format translation. CloudSearch accepts data in &lt;strong&gt;SDF&lt;/strong&gt; as JSON or XML, where a document batch is a collection of add and delete operations. The JSON that CloudSearch uses differs from the JSON that OpenSearch Service expects, so you must transform each source document into an OpenSearch document whose fields match the index mapping you defined earlier. Handle the same details the mapping calls out: emit each numeric value so it fits the narrow type you chose for its field rather than a wide &lt;code&gt;long&lt;/code&gt; or &lt;code&gt;double&lt;/code&gt;, format dates to match your &lt;code&gt;date&lt;/code&gt; mapping, and drop or rename any field that your strict mapping doesn’t define.&lt;/p&gt;
&lt;table style="border: 0;border-collapse: collapse;margin: 0 auto;width: 100%"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td style="border: 0;padding: 0 4px;text-align: center;width: 50%"&gt;&lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/09/BDB-5903-2.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/09/BDB-5903-2.png" alt="CloudSearch batch format showing add and delete operations in JSON" width="390"&gt;&lt;/a&gt;&lt;/td&gt;
   &lt;td style="border: 0;padding: 0 4px;text-align: center;width: 50%"&gt;&lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/09/BDB-5903-3.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/09/BDB-5903-3.png" alt="OpenSearch bulk batch format showing index operations in JSON" width="390"&gt;&lt;/a&gt;&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;em&gt;Figure 2: CloudSearch batch format (left) compared to OpenSearch batch format (right)&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;You can write a small conversion script. Have the script write its output to an Amazon S3 bucket so the converted documents live in a durable store you can re-ingest from as many times as you need.&lt;/p&gt;
&lt;p&gt;With your converted documents in Amazon S3, use &lt;a href="https://docs.aws.amazon.com/opensearch-service/latest/developerguide/ingestion.html" target="_blank" rel="noopener"&gt;Amazon OpenSearch Ingestion&lt;/a&gt; to load them. Amazon OpenSearch Ingestion is a feature of Amazon OpenSearch Service that you can use to ingest, filter, transform, enrich, and route data to an Amazon OpenSearch Service domain or an OpenSearch Serverless collection. Configure an OpenSearch Ingestion pipeline with an Amazon S3 source (you can use an &lt;a href="https://docs.aws.amazon.com/opensearch-service/latest/developerguide/pipeline-blueprint.html" target="_blank" rel="noopener"&gt;OpenSearch Ingestion blueprint to get started&lt;/a&gt;) that reads your converted documents. Let its built-in processors apply any final transformation before the pipeline writes to your collection. A managed pipeline reading from Amazon S3 gives you a repeatable, restartable load without operating ingestion infrastructure, which makes it the recommended path for most migrations.&lt;/p&gt;
&lt;p&gt;If you prefer to load data directly, OpenSearch Service exposes a REST API, so you can index documents with a standard client such as curl or with the OpenSearch &lt;a href="https://opensearch.org/docs/latest/clients/" target="_blank" rel="noopener"&gt;client libraries&lt;/a&gt; for many languages. Direct indexing is convenient for a small dataset or a quick test, but an Amazon S3 source with OpenSearch Ingestion is the better choice for a production migration.&lt;/p&gt;
&lt;h2 id="convert-your-queries"&gt;Convert your queries&lt;/h2&gt;
&lt;p&gt;CloudSearch uses a URL-based query format. You pass a query parameter in the URL and submit either a simple string search or a JSON-formatted query. OpenSearch Service uses a REST API and the OpenSearch query DSL in the request body, which gives you compound queries, function scoring, and richer relevance control. You can use generative AI coding assistants to help with this translation. Provide your CloudSearch query patterns, and the model generates the equivalent OpenSearch query DSL, which you then validate against your test cases.&lt;/p&gt;
&lt;h3 id="query-syntax-changes"&gt;Query syntax changes&lt;/h3&gt;
&lt;p&gt;CloudSearch appends parameters such as sort to the query URL, while OpenSearch expresses sorting, filtering, and boosting as explicit elements of the request body. For example, a title search for “shakespeare” in CloudSearch looks like the following.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-plaintext"&gt;https://my-cloudsearch-domain.us-east-1.cloudsearch.amazonaws.com/2013-01-01/search?q=shakespeare&amp;amp;size=10&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;The equivalent query in OpenSearch Service uses the query DSL.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-json"&gt;GET /imdb_movies/_search
{
  "query": {
    "match": { "title": "shakespeare" }
  }
}&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;To keep result sets consistent after migration, &lt;a href="https://docs.opensearch.org/latest/query-dsl/full-text/match/#operator" target="_blank" rel="noopener"&gt;set the default operator to AND&lt;/a&gt; in OpenSearch to match the default query behavior of CloudSearch. The following table shows common CloudSearch query patterns and their OpenSearch Service equivalents, using a sample IMDB movies dataset.&lt;/p&gt;
&lt;table border="1px" width="100%" cellpadding="10px"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;Query type&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;CloudSearch (Lucene syntax)&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;OpenSearch Service query DSL&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Compound AND&lt;/td&gt;
   &lt;td&gt;&lt;code&gt;title:"Inception" AND genres:"Sci-Fi"&lt;/code&gt;&lt;/td&gt;
   &lt;td&gt;&lt;code&gt;{"query":{"bool":{"must":[{"match":{"title":"Inception"}},{"match":{"genres":"Sci-Fi"}}]}}}&lt;/code&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Compound NOT&lt;/td&gt;
   &lt;td&gt;&lt;code&gt;title:"Star Wars" AND NOT genres:"Comedy"&lt;/code&gt;&lt;/td&gt;
   &lt;td&gt;&lt;code&gt;{"query":{"bool":{"must":[{"match":{"title":"Star Wars"}}],"must_not":[{"match":{"genres":"Comedy"}}]}}}&lt;/code&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Wildcard&lt;/td&gt;
   &lt;td&gt;&lt;code&gt;title:Batman*&lt;/code&gt;&lt;/td&gt;
   &lt;td&gt;&lt;code&gt;{"query":{"wildcard":{"title":{"value":"batman*"}}}}&lt;/code&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Numeric range&lt;/td&gt;
   &lt;td&gt;&lt;code&gt;rating:[7 TO 9]&lt;/code&gt;&lt;/td&gt;
   &lt;td&gt;&lt;code&gt;{"query":{"range":{"rating":{"gte":7,"lte":9}}}}&lt;/code&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Date range (after)&lt;/td&gt;
   &lt;td&gt;&lt;code&gt;release_date:[2015-01-01T00:00:00Z TO *]&lt;/code&gt;&lt;/td&gt;
   &lt;td&gt;&lt;code&gt;{"query":{"range":{"release_date":{"gte":"2015-01-01T00:00:00Z"}}}}&lt;/code&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Boosting&lt;/td&gt;
   &lt;td&gt;&lt;code&gt;title:"The Matrix"^6 OR genres:"Sci-Fi"^4&lt;/code&gt;&lt;/td&gt;
   &lt;td&gt;&lt;code&gt;{"query":{"bool":{"should":[{"query_string":{"query":"title": \"The Matrix\"^6","fields":["title"]}},{"query_string":{"query":"genres:\"Sci-Fi\"^4","fields":["genres"]}}]}}}&lt;/code&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Sorting&lt;/td&gt;
   &lt;td&gt;&lt;code&gt;title:"Batman" sort=release_date desc&lt;/code&gt;&lt;/td&gt;
   &lt;td&gt;&lt;code&gt;{"query":{"match":{"title":"Batman"}},"sort":[{"release_date":{"order":"desc"}}]}&lt;/code&gt;&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;h3 id="sorting-and-boosting"&gt;Sorting and boosting&lt;/h3&gt;
&lt;p&gt;Boosting is useful when you want certain fields or terms to carry more weight in relevance scoring. A higher boost value means the term contributes more to the score. OpenSearch also supports sorting by &lt;code&gt;_score&lt;/code&gt; (relevance), which is the default when you specify no sort. For the full query language, see the &lt;a href="https://opensearch.org/docs/latest/query-dsl/" target="_blank" rel="noopener"&gt;OpenSearch query DSL documentation&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="configure-security"&gt;Configure security&lt;/h2&gt;
&lt;p&gt;CloudSearch uses &lt;a href="https://aws.amazon.com/iam/" target="_blank" rel="noopener"&gt;AWS Identity and Access Management&lt;/a&gt; policies to control access to its configuration and domain service APIs. You attach user-based policies to an IAM role, user, or group, and the document, search, and suggest actions in those policies control access to the CloudSearch APIs.&lt;/p&gt;
&lt;p&gt;OpenSearch Serverless applies security through policies at several layers.&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;Collections&lt;/strong&gt;: Encrypted at rest by default, using either an AWS owned key or a customer managed key defined in an encryption policy.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Network policies&lt;/strong&gt;: Define whether a collection is reachable privately through a virtual private cloud (VPC) endpoint or over the internet.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Data access policies&lt;/strong&gt;: Control which IAM principals and Security Assertion Markup Language (SAML) identities can create indexes and read or write data in the collection.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Amazon OpenSearch Service provisioned domains also offer fine-grained access control, with role-based access control and security at the index, document, and field level. For OpenSearch Serverless, data access policies provide collection-level and index-level permissions, controlling which IAM principals and SAML identities can create, read, or write data within a collection.&lt;/p&gt;
&lt;h2 id="validate-the-migration"&gt;Validate the migration&lt;/h2&gt;
&lt;p&gt;Validation confirms that the migration is complete and correct before you send production traffic to OpenSearch Serverless. Work through five kinds of validation.&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;Documents&lt;/strong&gt;: Check your document count. Your OpenSearch Serverless indexes should have the same count as your CloudSearch indexes.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Queries&lt;/strong&gt;: Translate your most important queries and run them manually against your collection. Spot check the output for the presence of important results.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Ranking&lt;/strong&gt;: Check the order of results, especially for queries with custom rank functions or field weighting. Results might not match exactly, so look for anything that’s incorrect.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Latency&lt;/strong&gt;: Ideally you should tee your production traffic to your Serverless collection to get real latency metrics. Worst case, generate at least 100,000 synthetic queries across all your query types and run them. Monitor OpenSearch Compute Unit (OCU) consumption with &lt;a href="https://aws.amazon.com/cloudwatch/" target="_blank" rel="noopener"&gt;Amazon CloudWatch&lt;/a&gt; to understand your cost profile.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;To validate search functionality, run the same query against both systems and compare the results. Reuse the query pairs from the conversion step so you exercise the syntax differences directly. For example, to check a numeric range against the sample IMDB movies dataset, run the following query in CloudSearch.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-plaintext"&gt;https://my-cloudsearch-domain.us-east-1.cloudsearch.amazonaws.com/2013-01-01/search?q=rating: [7 TO 9]&amp;amp;size=10&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Run the equivalent query DSL against your OpenSearch Serverless collection.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-json"&gt;GET /imdb_movies/_search
{
  "query": {
    "range": { "rating": { "gte": 7, "lte": 9 } }
  }
}&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Confirm that both queries return the same set of movies. Then repeat the comparison for a query that exercises relevance, such as the boosted query from the conversion step, and confirm the top results appear in the same order.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-json"&gt;GET /imdb_movies/_search
{
  "query": {
    "bool": {
      "should": [
        { "match": { "title": { "query": "The Matrix", "boost": 6 } } },
        { "match": { "genres": { "query": "Sci-Fi", "boost": 4 } } }
      ]
    }
  }
}&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h2 id="cut-over-and-operate"&gt;Cut over and operate&lt;/h2&gt;
&lt;p&gt;When validation passes, update your application to use the OpenSearch Serverless endpoint and the query DSL, and switch from the CloudSearch SDK to the OpenSearch client libraries. After cutover, confirm that no application still points to a CloudSearch endpoint, retain your source data backups in Amazon S3 for rollback, and then delete the CloudSearch domain.&lt;/p&gt;
&lt;p&gt;Operating OpenSearch Serverless in production is lighter than operating a domain, because OpenSearch Serverless scales compute for you and you do not tune shards, instance types, or capacity. Your focus shifts to cost and search quality. Monitor OCU consumption and search latency with Amazon CloudWatch, and set alarms on the thresholds that matter to you. Review OCU usage patterns to understand cost and find optimization opportunities, and set capacity limits on the collection to cap the maximum OCUs it can consume. For guidance, see &lt;a href="https://docs.aws.amazon.com/opensearch-service/latest/developerguide/serverless-scaling.html" target="_blank" rel="noopener"&gt;Managing capacity limits for Amazon OpenSearch Serverless&lt;/a&gt; and &lt;a href="https://docs.aws.amazon.com/opensearch-service/latest/developerguide/serverless-monitoring.html" target="_blank" rel="noopener"&gt;Monitoring Amazon OpenSearch Serverless&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="cost-considerations"&gt;Cost considerations&lt;/h2&gt;
&lt;p&gt;With OpenSearch Serverless, you pay only for the compute and storage your workload consumes, and OpenSearch Serverless charges for compute and storage separately. OpenSearch Serverless scales indexing compute and search compute independently, so a write-heavy or a read-heavy workload scales only the dimension it needs, and compute can scale to zero when a collection is idle, in which case you pay only for storage. To share hardware across workloads, place collections in a collection group so they draw from the same compute rather than each provisioning its own. For pricing and unit details, see &lt;a href="https://aws.amazon.com/opensearch-service/pricing/" target="_blank" rel="noopener"&gt;Amazon OpenSearch Service pricing&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="clean-up"&gt;Clean up&lt;/h2&gt;
&lt;p&gt;Because you’re migrating to OpenSearch Serverless, the resources that you’ve created will likely become your production resources. If not, delete any OpenSearch Serverless collections and S3 buckets you created to avoid incurring ongoing cost.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;In this post, you saw how Amazon CloudSearch and Amazon OpenSearch Serverless compare, and how the concepts you rely on in CloudSearch (field types, query syntax, autoscaling, and access control) translate into OpenSearch Service. You assess your CloudSearch configuration, model your data with explicit OpenSearch mappings, move your converted documents into the collection with OpenSearch Ingestion, convert your URL-based queries into the OpenSearch query DSL, configure security, and validate before cutover. OpenSearch Serverless gives you the hands-off operational model you have with CloudSearch, and adds richer query capabilities, granular data access policies, and automatic scaling. To get started, create an OpenSearch Serverless collection on the AWS Management Console and follow the steps in this post.&lt;/p&gt;
&lt;p&gt;To learn more, see the following resources:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/opensearch-service/latest/developerguide/" target="_blank" rel="noopener"&gt;Amazon OpenSearch Service Developer Guide&lt;/a&gt;&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/opensearch-service/latest/developerguide/serverless.html" target="_blank" rel="noopener"&gt;Amazon OpenSearch Serverless&lt;/a&gt;&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://opensearch.org/docs/latest/" target="_blank" rel="noopener"&gt;OpenSearch documentation&lt;/a&gt;&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://aws.amazon.com/blogs/big-data/four-ways-to-build-with-amazon-opensearch-service/" target="_blank" rel="noopener"&gt;Four ways to build with Amazon OpenSearch Service&lt;/a&gt;&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://aws.amazon.com/blogs/big-data/amazon-opensearch-services-vector-database-capabilities-explained/" target="_blank" rel="noopener"&gt;Amazon OpenSearch Service vector database capabilities explained&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p style="clear: both"&gt;&lt;/p&gt;
&lt;hr style="width: 100%"&gt;
&lt;h2&gt;About the authors&lt;/h2&gt;
&lt;footer&gt;
 &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
  &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
   &lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/09/BDB-5903-4.jpg" alt="Prasad Nadig" width="100" height="133"&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Prasad Nadig&lt;/h3&gt;
  &lt;p style="overflow: hidden"&gt;&lt;a href="http://www.linkedin.com/in/prasad-nadig" target="_blank" rel="noopener"&gt;Prasad&lt;/a&gt; is a Senior Analytics Specialist Solutions Architect at Amazon Web Services (AWS), specializing in large-scale data analytics and AI. Prasad partners with customers to design, migrate, and modernize their analytics platforms on AWS into scalable, cost-effective solutions, with deep expertise in data lakes, data warehousing, distributed processing, and performance tuning at petabyte scale.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
  &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
   &lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/09/BDB-5903-5.jpg" alt="Jon Handler" width="100" height="133"&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Jon Handler&lt;/h3&gt;
  &lt;p style="overflow: hidden"&gt;&lt;a href="https://www.linkedin.com/in/jonhandler/" target="_blank" rel="noopener"&gt;Jon&lt;/a&gt; is a Senior Principal Solutions Architect for Search Services at Amazon Web Services. Jon works closely with OpenSearch and Amazon OpenSearch Service, providing help and guidance to a broad range of customers who have search and log analytics workloads. Prior to joining AWS, Jon’s career as a software developer included four years of coding a large-scale, eCommerce search engine.&lt;/p&gt;
 &lt;/div&gt;
&lt;/footer&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>Accelerating Spark queries with Iceberg materialized views</title>
		<link>https://aws.amazon.com/blogs/big-data/accelerating-spark-queries-with-iceberg-materialized-views/</link>
		
		<dc:creator><![CDATA[Yuzhou Sun]]></dc:creator>
		<pubDate>Thu, 10 Sep 2026 16:05:45 +0000</pubDate>
				<category><![CDATA[Advanced (300)]]></category>
		<category><![CDATA[Amazon EMR]]></category>
		<category><![CDATA[AWS Glue]]></category>
		<category><![CDATA[Technical How-to]]></category>
		<guid isPermaLink="false">f7504fce8cffe082d1e3efae9f5e5c99384ec13e</guid>

					<description>Accelerate slow, repetitive Apache Spark analytical queries on Apache Iceberg tables without rewriting any SQL. This post shows how automatic query rewrite in Amazon EMR and AWS Glue uses Iceberg materialized views in the AWS Glue Data Catalog to transparently substitute matching query plans, and how to design materialized views for the best speedup.</description>
										<content:encoded>&lt;p&gt;In this post, you learn how to reduce &lt;a href="https://spark.apache.org/" target="_blank" rel="noopener"&gt;Apache Spark&lt;/a&gt; query execution time with &lt;a href="https://iceberg.apache.org/" target="_blank" rel="noopener"&gt;Apache Iceberg&lt;/a&gt; materialized views without changing a single SQL query.&lt;/p&gt;
&lt;p&gt;Organizations running analytical workloads on their data lakes often hit a common wall: queries that are slow and costly, yet difficult to rewrite by hand. Multi-table joins, heavy aggregations, and window functions over large fact tables all drive up execution times, but the SQL behind them often can’t be changed. It might come from &lt;a href="https://aws.amazon.com/what-is/business-intelligence/" target="_blank" rel="noopener"&gt;business intelligence (BI)&lt;/a&gt; dashboards, packaged independent software vendor (ISV) applications, or legacy reports, where editing the source introduces regression risk that outweighs the performance gain.&lt;/p&gt;
&lt;p&gt;Starting with &lt;a href="https://aws.amazon.com/emr/" target="_blank" rel="noopener"&gt;Amazon EMR&lt;/a&gt; 7.12.0 and &lt;a href="https://aws.amazon.com/glue/" target="_blank" rel="noopener"&gt;AWS Glue&lt;/a&gt; 5.1, you can accelerate these queries without rewriting them. Automatic query rewrite analyzes the logical plan of each incoming query and compares it against a metadata cache of available MVs. When the optimizer finds a &lt;a href="https://aws.amazon.com/what-is/materialized-view/" target="_blank" rel="noopener"&gt;materialized view (MV)&lt;/a&gt; that satisfies all or part of a query, it rewrites the plan to read from that MV instead of the base tables. Matches can be structural (aggregations and joins) or exact (more complex patterns like window functions). If no MV matches, the original query runs unchanged with no impact on correctness.&lt;/p&gt;
&lt;p&gt;If you have previously tried to speed up slow analytical queries, you might have considered one of the following alternatives. Here is how automatic query rewrite compares:&lt;/p&gt;
&lt;table border="1px" width="100%" cellpadding="10px"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;Query modification approach&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Stored results&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Refreshes&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Modification to existing queries&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Standard views in AWS Glue&lt;/td&gt;
   &lt;td&gt;No (re-runs each time)&lt;/td&gt;
   &lt;td&gt;n/a&lt;/td&gt;
   &lt;td&gt;Required&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Custom ETL pipeline&lt;/td&gt;
   &lt;td&gt;Yes&lt;/td&gt;
   &lt;td&gt;Manual&lt;/td&gt;
   &lt;td&gt;Required&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Hand-rolled rewrite&lt;/td&gt;
   &lt;td&gt;Yes&lt;/td&gt;
   &lt;td&gt;Manual&lt;/td&gt;
   &lt;td&gt;Required&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Materialized views with automatic rewrite enabled&lt;/td&gt;
   &lt;td&gt;Yes&lt;/td&gt;
   &lt;td&gt;Automatically through &lt;a href="https://docs.aws.amazon.com/glue/latest/dg/catalog-and-crawler.html" target="_blank" rel="noopener"&gt;AWS Glue Data Catalog&lt;/a&gt; on a schedule when configured&lt;/td&gt;
   &lt;td&gt;Not required when supported&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;In this post, we:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;Give a high-level overview of how automatic query rewrite works in Apache Spark.&lt;/li&gt;
 &lt;li&gt;Walk through a concrete example, showing how the same query can benefit from MVs at different levels of coverage.&lt;/li&gt;
 &lt;li&gt;Discuss the trade-offs so you can choose the right MV shape for your workload.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="prerequisites"&gt;Prerequisites&lt;/h2&gt;
&lt;p&gt;To use automatic query rewrite with Iceberg materialized views, you need:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;Amazon EMR release 7.12.0 or later, or AWS Glue 5.1 or later.&lt;/li&gt;
 &lt;li&gt;Source tables in Apache Iceberg or Parquet format, registered in the AWS Glue Data Catalog, in the same AWS Region and account as the materialized view. Parquet source tables are supported for automatic query rewrite starting with Amazon EMR 7.14.0 and AWS Glue 8.1.&lt;/li&gt;
 &lt;li&gt;An Amazon Simple Storage Service (Amazon S3) Tables (a capability of Amazon S3) bucket, or an S3 general purpose bucket, for the materialized view data.&lt;/li&gt;
 &lt;li&gt;Permissions for the definer role. You can use AWS Identity and Access Management (IAM) policies or AWS Lake Formation.&lt;/li&gt;
 &lt;li&gt;Automatic query rewrite turned on in your Spark session: &lt;code&gt;--conf spark.sql.optimizer.answerQueriesWithMVs.enabled=true&lt;/code&gt;.&lt;/li&gt;
 &lt;li&gt;For Parquet source tables, set &lt;code&gt;spark.sql.materializedView.v1SourceTables.enabled=true&lt;/code&gt; and &lt;code&gt;spark.sql.materializedView.v1ETagVersioning.enabled=true&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For more Spark configurations, see &lt;a href="https://aws.amazon.com/blogs/big-data/introducing-apache-iceberg-materialized-views-in-aws-glue-data-catalog/" target="_blank" rel="noopener"&gt;Introducing Apache Iceberg materialized views in AWS Glue Data Catalog&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="how-it-works"&gt;How it works&lt;/h2&gt;
&lt;p&gt;Here is how MVs and automatic query rewrite work together:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;You define a SQL query &lt;/strong&gt;with aggregations, joins, or filters across your supported source tables.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;AWS Glue Data Catalog stores the precomputed results&lt;/strong&gt; as an Apache Iceberg table in your &lt;a href="https://aws.amazon.com/s3/" target="_blank" rel="noopener"&gt;Amazon S3&lt;/a&gt; bucket. You can store it in a general purpose S3 bucket or in &lt;a href="https://aws.amazon.com/s3/features/tables/" target="_blank" rel="noopener"&gt;Amazon S3 Tables&lt;/a&gt;. Any Apache Iceberg-compatible query engine can read the materialized view, including &lt;a href="https://aws.amazon.com/athena/" target="_blank" rel="noopener"&gt;Amazon Athena&lt;/a&gt;, Amazon EMR, AWS Glue, &lt;a href="https://aws.amazon.com/redshift/" target="_blank" rel="noopener"&gt;Amazon Redshift&lt;/a&gt;, and Iceberg-compatible third-party query engines. Automatic query rewrite is available on the AWS optimized Spark runtime in Amazon Athena, Amazon EMR, and AWS Glue. Other engines can query the materialized view directly, but they don’t rewrite queries to use it automatically.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Automatic refresh keeps the MV current&lt;/strong&gt; on a schedule that you define, for example &lt;code&gt;SCHEDULE REFRESH EVERY 1 DAY&lt;/code&gt;. You set it at creation time or later with &lt;code&gt;ALTER MATERIALIZED VIEW ... ADD SCHEDULE REFRESH&lt;/code&gt;. At that scheduled time, the refresh process checks the current Apache Iceberg snapshot ID or Parquet file ETags and refreshes the MV when it detects source-table changes.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Automatic query rewrite redirects matching queries to&lt;/strong&gt; the MV at query optimization time. Automatic query rewrite in Apache Spark uses two matching strategies:
  &lt;ul&gt;
   &lt;li&gt;&lt;strong&gt;Structural rewrite&lt;/strong&gt; (adapted from Amazon Redshift) handles an MV defined as a single SELECT-FROM-WHERE-GROUP-BY block over INNER joins. The optimizer can roll up an MV’s aggregates to a coarser grain and pull extra query predicates up onto the MV scan.&lt;/li&gt;
   &lt;li&gt;&lt;strong&gt;Exact-match rewrite&lt;/strong&gt; handles MVs defined as other shapes, such as window functions and outer joins, by matching a canonicalized form of the MV body against subtrees of the query plan.&lt;/li&gt;
  &lt;/ul&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;When the optimizer evaluates a query, it consults a metadata cache of MVs from the configured catalogs and chooses the best match. It also checks MV staleness during optimization. It skips stale MVs, so rewrite won’t return stale results. If no MV matches, the original query runs unchanged.&lt;/p&gt;
&lt;p&gt;Note that automatic query rewrite is opt-in: set &lt;code&gt;spark.sql.optimizer.answerQueriesWithMVs.enabled=true&lt;/code&gt; when creating the Apache Spark session.&lt;/p&gt;
&lt;h2 id="example-one-query-with-three-potential-mvs"&gt;Example: One query with three potential MVs&lt;/h2&gt;
&lt;p&gt;An MV doesn’t need to cover an entire query to help it. Automatic query rewrite in Apache Spark operates on subtrees: when an MV matches a portion of your query plan, the rewriter substitutes that subtree and lets the rest of the query run on the rewrite output unchanged. The same query can therefore be served by many possible MV designs, each making a different trade-off between per-query speedup, storage cost, and reuse across other queries.&lt;/p&gt;
&lt;p&gt;To make this concrete, consider a typical analytics query: &lt;strong&gt;“&lt;strong&gt;Top 100 preferred US customers by total store spending&lt;/strong&gt;.”&lt;/strong&gt; It joins fact and dimension tables, applies two selective filters on the customer dimension, aggregates per customer, ranks the result with a window function, and keeps only the top 100:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-sql"&gt;SELECT c_customer_id, total_revenue, num_transactions, avg_purchase, revenue_rank
FROM (
    SELECT cust.c_customer_id,
        SUM(sales.ss_quantity * sales.ss_sales_price) AS total_revenue,
        COUNT(*) AS num_transactions,
        AVG(sales.ss_quantity * sales.ss_sales_price) AS avg_purchase,
        RANK() OVER (ORDER BY SUM(sales.ss_quantity * sales.ss_sales_price) DESC) AS revenue_rank
    FROM base_catalog.base_db.store_sales sales
    INNER JOIN base_catalog.base_db.customer cust
        ON sales.ss_customer_sk = cust.c_customer_sk
    WHERE cust.c_birth_country = 'UNITED STATES'
        AND cust.c_preferred_cust_flag = 'Y'
    GROUP BY cust.c_customer_id
) ranked
WHERE revenue_rank &amp;lt;= 100
ORDER BY revenue_rank;&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;&lt;em&gt;Query 1: The original query. Top 100 preferred US customers by total store spending, before any materialized view.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Three MV designs cover progressively more of this query, from a single-table pre-aggregate to the full query body itself:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Tier 1: Pre-aggregate store_sales only, no join, no filter.&lt;/strong&gt; This tier is a single-table aggregate of &lt;code&gt;store_sales&lt;/code&gt; at customer-surrogate-key grain. The query still must join the &lt;code&gt;customer&lt;/code&gt; table, apply both filters, re-aggregate at &lt;code&gt;c_customer_id&lt;/code&gt; grain, and run the window function.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-sql"&gt;CREATE MATERIALIZED VIEW mv_catalog.mv_db.customer_tier_1 AS
SELECT ss_customer_sk,
    SUM(ss_quantity * ss_sales_price) AS sum_revenue,
    COUNT(ss_quantity * ss_sales_price) AS count_revenue,
    COUNT(*) AS num
FROM base_catalog.base_db.store_sales
GROUP BY ss_customer_sk;&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;&lt;em&gt;Tier 1 MV: Single-table pre-aggregate of store_sales by customer surrogate key (no join, no filter).&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;The following plans compare the original query plan to the rewritten plan:&lt;/p&gt;
&lt;pre class="text"&gt;&lt;code&gt;Window, filter, Sort
+- Aggregate by c_customer_id
:  total_revenue = SUM(ss_quantity * ss_sales_price)
:  num_transactions = COUNT(*)
:  avg_purchase = AVG(ss_quantity * ss_sales_price)
+- Project
   +- Join Inner ON ss_customer_sk = c_customer_sk
      :- BatchScan store_sales &lt;strong&gt;&amp;lt;-&lt;/strong&gt; reads the large store_sales table
      +- Filter c_birth_country='UNITED STATES' AND c_preferred_cust_flag='Y'
         +- BatchScan customer&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;em&gt;Plan 1: Original plan. Scans the large store_sales table.&lt;/em&gt;&lt;/p&gt;
&lt;pre class="text"&gt;&lt;code&gt;Window, filter, Sort
+- Aggregate by c_customer_id &lt;strong&gt;&amp;lt;-&lt;/strong&gt; rolls up pre-aggregated sums
:  total_revenue = SUM(sum_revenue) &lt;strong&gt;&amp;lt;-&lt;/strong&gt; sum of sum_revenue
:  num_transactions = SUM(num) &lt;strong&gt;&amp;lt;-&lt;/strong&gt; sum of num
:  avg_purchase = SUM(sum_revenue) / SUM(count_revenue) &lt;strong&gt;&amp;lt;-&lt;/strong&gt; sum of sum_revenue / sum of count_revenue
+- Project
   +- Join Inner ON ss_customer_sk = c_customer_sk
      :- BatchScan customer_tier_1 &lt;strong&gt;&amp;lt;-&lt;/strong&gt; reads pre-aggregated MV
      +- Filter c_birth_country='UNITED STATES' AND c_preferred_cust_flag='Y'
         +- BatchScan customer&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;em&gt;Plan 2: Rewritten plan (Tier 1). Reads the pre-aggregated customer_tier_1 MV.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Tier 2: Pre-join store_sales x customer, pre-apply one filter (c_preferred_cust_flag = ‘Y’).&lt;/strong&gt; The middle tier pre-joins both tables and bakes in the preferred-customer filter. The query still must apply the country filter as a residual on the MV scan and run the &lt;code&gt;RANK()&lt;/code&gt; window.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-sql"&gt;CREATE MATERIALIZED VIEW mv_catalog.mv_db.customer_tier_2 AS
SELECT cust.c_customer_id, cust.c_birth_country,
    SUM(sales.ss_quantity * sales.ss_sales_price) AS sum_revenue,
    COUNT(sales.ss_quantity * sales.ss_sales_price) AS count_revenue,
    COUNT(*) AS num
FROM base_catalog.base_db.store_sales sales
INNER JOIN base_catalog.base_db.customer cust
    ON sales.ss_customer_sk = cust.c_customer_sk
WHERE cust.c_preferred_cust_flag = 'Y'
GROUP BY cust.c_customer_id, cust.c_birth_country;&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;&lt;em&gt;Tier 2 MV: Pre-joins store_sales and customer, with the preferred-customer filter applied.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Rewritten query plan:&lt;/p&gt;
&lt;pre class="text"&gt;&lt;code&gt;Window, filter, Sort
+- Aggregate by c_customer_id &amp;lt;- rolls up pre-aggregated sums
:  total_revenue = SUM(sum_revenue) &amp;lt;- sum of sum_revenue
:  num_transactions = SUM(num) &amp;lt;- sum of num
:  avg_purchase = SUM(sum_revenue) / SUM(count_revenue) &amp;lt;- reads pre-aggregated MV
+- Filter c_birth_country='UNITED STATES' [residual filter on MV scan]
   +- BatchScan customer_tier_2 &amp;lt;- reads pre-aggregated MV&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;em&gt;Plan 3: Rewritten plan (Tier 2). Country filter applied as a residual on the MV scan.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Tier 3: Match the entire query, including the window function and top N filter.&lt;/strong&gt; This is the most specific tier. The MV body is the target query verbatim (minus the top-level &lt;code&gt;ORDER BY&lt;/code&gt;, which is meaningless for a stored set). The MV stores the top-ranked rows the query asks for (rank ≤ 100).&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-sql"&gt;CREATE MATERIALIZED VIEW mv_catalog.mv_db.customer_tier_3 AS
SELECT c_customer_id, total_revenue, num_transactions, avg_purchase, revenue_rank
FROM (
    SELECT cust.c_customer_id,
        SUM(sales.ss_quantity * sales.ss_sales_price) AS total_revenue,
        COUNT(*) AS num_transactions,
        AVG(sales.ss_quantity * sales.ss_sales_price) AS avg_purchase,
        RANK() OVER (ORDER BY SUM(sales.ss_quantity * sales.ss_sales_price) DESC) AS revenue_rank
    FROM base_catalog.base_db.store_sales sales
    INNER JOIN base_catalog.base_db.customer cust
        ON sales.ss_customer_sk = cust.c_customer_sk
    WHERE cust.c_birth_country = 'UNITED STATES'
        AND cust.c_preferred_cust_flag = 'Y'
    GROUP BY cust.c_customer_id
) ranked
WHERE revenue_rank &amp;lt;= 100;&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;&lt;em&gt;Tier 3 MV: Stores the exact ranked output of the query (exact-match path).&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;This tier exercises the &lt;strong&gt;exact-match rewrite path&lt;/strong&gt;: the rewriter canonicalizes the MV body and matches it against the query’s logical plan.&lt;/p&gt;
&lt;p&gt;Rewritten plan:&lt;/p&gt;
&lt;pre class="text"&gt;&lt;code&gt;Sort revenue_rank ASC
+- BatchScan customer_tier_3 &amp;lt;- reads around 100 stored rows&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;em&gt;Plan 4: Rewritten plan (Tier 3). Reads around 100 stored rows.&lt;/em&gt;&lt;/p&gt;
&lt;h3 id="the-trade-off"&gt;The trade-off&lt;/h3&gt;
&lt;p&gt;The three tiers trade per-query speedup against reuse and storage. In our testing on &lt;a href="https://www.tpc.org/tpcds/" target="_blank" rel="noopener"&gt;TPC-DS&lt;/a&gt; 3 TB, we observed the following:&lt;/p&gt;
&lt;table border="1px" width="100%" cellpadding="10px"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;MV design&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Pre-computed&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Reuse&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Per-query speedup&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;MV size&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Baseline (no MV)&lt;/td&gt;
   &lt;td&gt;nothing&lt;/td&gt;
   &lt;td&gt;n/a&lt;/td&gt;
   &lt;td&gt;1x&lt;/td&gt;
   &lt;td&gt;n/a&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Tier 1: store_sales agg by customer surrogate key&lt;/td&gt;
   &lt;td&gt;aggregate of all sales per customer&lt;/td&gt;
   &lt;td&gt;broadest: any per-customer aggregation&lt;/td&gt;
   &lt;td&gt;~5x faster&lt;/td&gt;
   &lt;td&gt;0.07% of store_sales for TPC-DS 3 TB&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Tier 2: store_sales x customer agg, one filter pre-applied&lt;/td&gt;
   &lt;td&gt;join + aggregate, preferred customers only&lt;/td&gt;
   &lt;td&gt;medium: any country filter, preferred customers&lt;/td&gt;
   &lt;td&gt;~10x faster&lt;/td&gt;
   &lt;td&gt;0.04% of store_sales for TPC-DS 3 TB&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Tier 3: entire query body verbatim (exact-match)&lt;/td&gt;
   &lt;td&gt;exact ranked output of this query&lt;/td&gt;
   &lt;td&gt;narrowest: only this exact query shape&lt;/td&gt;
   &lt;td&gt;20x+ faster&lt;/td&gt;
   &lt;td&gt;negligible (only 100 rows)&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Performance measured on TPC-DS 3 TB. Speedup is the ratio of baseline execution time to MV-accelerated execution time. Results might vary based on data characteristics, cluster size, and query complexity.&lt;/p&gt;
&lt;p&gt;In addition, MVs incur additional cost. Each one runs a query against your source tables once and stores the result. The more pre-computation it does (joining more tables, applying more filters), the more time it takes.&lt;/p&gt;
&lt;p&gt;The following chart plots per-query speedup and creation time for the three tiers in our testing on TPC-DS 3 TB. Per-query speedup rises steadily, from about 5x at Tier 1 to over 20x at Tier 3. Creation time doesn’t follow the same pattern: it peaks at Tier 2. Tier 2 pre-joins and aggregates all preferred customers across every country, so it materializes the most data work. Tier 3 applies both filters, so it processes far fewer rows and costs less to create.&lt;/p&gt;
&lt;p&gt;&lt;img loading="lazy" class="alignnone" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/17/BDB-5950-1.jpg" alt="Chart comparing three materialized view designs. In our testing with TPC-DS 3 TB, we observed per-query speedup rises from about 5x (Tier 1) to over 20x (Tier 3), while creation time peaks at Tier 2, which materializes the most data work. Stacked bars show creation time split into catalog setup, data work, and commit." width="800" height="829"&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Figure 1: Per-query speedup and creation time across the three materialized view tiers, measured on TPC-DS 3 TB&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Start by identifying one expensive query that runs repeatedly with stable filters. It is likely a good candidate for an exact-match MV.&lt;/p&gt;
&lt;h2 id="validating-automatic-query-rewrite"&gt;Validating automatic query rewrite&lt;/h2&gt;
&lt;p&gt;To confirm that your query benefited from automatic rewrite:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;&lt;strong&gt;Query plan inspection&lt;/strong&gt;: Check the query’s optimized logical plan or physical plan for a leaf scan node referencing the MV (for example, &lt;code&gt;BatchScan mv_catalog.mv_db.your_mv_name&lt;/code&gt;). If the MV appears as a scan source, rewrite succeeded.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Log confirmation (Amazon EMR 7.14.0+)&lt;/strong&gt;: Look for INFO-level log entries such as &lt;code&gt;AQMV outcome: rewritten=true, mvs=[mv_name], duration=12ms&lt;/code&gt;.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;No-rewrite diagnostics (Amazon EMR 7.14.0+)&lt;/strong&gt;: If rewrite didn’t occur, check the &lt;code&gt;MVRewriteMetricsEvent&lt;/code&gt; in the Apache Spark Event Log for the specific reason the optimizer skipped the MV.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;If you have set &lt;code&gt;spark.sql.optimizer.answerQueriesWithMVs.enabled=true&lt;/code&gt; but your query still runs against the base tables, check the following common causes:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;&lt;strong&gt;Write commands block rewrite by default.&lt;/strong&gt; INSERT and MERGE statements don’t trigger rewrite. Set &lt;code&gt;spark.sql.optimizer.answerQueriesWithMVs.commandBlockingEnabled=false&lt;/code&gt; to turn on rewrite within write command subqueries.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;The MV is stale.&lt;/strong&gt; Rewrite skips the MV when one or more source tables have changed since its last refresh. Wait for the next scheduled refresh, or force an immediate refresh with &lt;code&gt;REFRESH MATERIALIZED VIEW &amp;lt;mv_name&amp;gt;&lt;/code&gt;.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Heuristic candidate filtering.&lt;/strong&gt; The optimizer uses heuristic checks to narrow the set of MV candidates before attempting a full match. In some cases, an MV that could benefit the query might be filtered out early by these heuristics.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Spark version mismatch (Amazon EMR 7.13.0+).&lt;/strong&gt; Automatic query rewrite skips MVs whose stored &lt;code&gt;IMV_sparkVersion&lt;/code&gt; does not match the cluster’s current Apache Spark version. To bypass this check, set &lt;code&gt;spark.sql.materializedView.sparkVersionCompatibilityCheck.enabled=false&lt;/code&gt;.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;MV metadata cache not loaded.&lt;/strong&gt; The metadata cache loads lazily during optimization of the first rewritable query in a Spark session. If your critical query fires before the cache is warm, the MV will not be available. Run a small warm-up query (for example, &lt;code&gt;SELECT 1 FROM &amp;lt;some_iceberg_table&amp;gt;&lt;/code&gt;) at session start to pay this cost off the critical path.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;MV metadata cache memory limit reached.&lt;/strong&gt; If the cache was disabled or stopped loading MVs because of reaching its memory limit, increase &lt;code&gt;spark.driver.memory&lt;/code&gt;.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Too many tables in configured catalogs.&lt;/strong&gt; If there are many tables or MVs in the configured catalogs, the cache might not finish loading before your query starts. Place MVs in a dedicated catalog, add it to &lt;code&gt;spark.sql.materializedViews.additionalCatalogs&lt;/code&gt;, and set &lt;code&gt;spark.sql.materializedViews.scanCurrentCatalog=false&lt;/code&gt; to skip scanning the current catalog.&lt;/li&gt;
 &lt;li&gt;&lt;b&gt; Parquet base tables have additional limitations and configuration requirements&lt;/b&gt;. For automatic query rewrite with Parquet base tables, set &lt;code&gt;spark.sql.materializedView.v1SourceTables.enabled=true&lt;/code&gt; and &lt;code&gt;spark.sql.materializedView.v1ETagVersioning.enabled=true&lt;/code&gt;. Without ETag versioning, Spark can’t determine a usable source-table version and skips the MV. Partitioned Parquet base tables are also subject to additional validation limits.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="performance-considerations"&gt;Performance considerations&lt;/h2&gt;
&lt;p&gt;Turning on automatic query rewrite has overhead: it introduces trade-offs that might affect some queries negatively:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;&lt;strong&gt;Optimization overhead.&lt;/strong&gt; Enabling rewrite adds processing time during query optimization as the optimizer evaluates MV candidates against the query plan. This overhead applies to every query in the session, including those that ultimately don’t match any MV.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Reduced task parallelism.&lt;/strong&gt; Reading from an MV instead of the original base table might produce fewer tasks or introduce data skew, depending on the MV’s data layout. This reduces parallelism compared to a direct scan of the larger, more evenly distributed source table.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;In this post, we showed how automatic query rewrite can accelerate your existing Apache Spark workloads. It uses Apache Iceberg materialized views in the AWS Glue Data Catalog, without changing a single line of SQL. By storing precomputed results as managed Apache Iceberg tables, the AWS Glue Data Catalog lets the Apache Spark optimizer transparently substitute matching query plans. You get the performance benefit of pre-aggregation without the application-level rewiring. BI dashboards, ISV-generated reports, and legacy pipelines all benefit the moment a matching MV exists.&lt;/p&gt;
&lt;p&gt;We walked through three MV designs for the same analytical query, each striking a different balance between per-query speedup, storage footprint, and reuse across your workload. As the trade-off table shows, our testing found that a narrow, exact-match MV delivered 20x+ acceleration for a single query shape. A broader pre-aggregate served an entire family of queries at a more modest ~5x gain. The right choice depends on how many queries share the same join-and-aggregate pattern and how frequently your source data changes.&lt;/p&gt;
&lt;p&gt;To get started:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;Launch an Amazon EMR 7.12.0+ cluster or an AWS Glue 5.1+ job.&lt;/li&gt;
 &lt;li&gt;Create an MV over your most expensive repeating query using &lt;code&gt;CREATE MATERIALIZED VIEW&lt;/code&gt; in the AWS Glue Data Catalog.&lt;/li&gt;
 &lt;li&gt;Turn on automatic query rewrite by setting &lt;code&gt;spark.sql.optimizer.answerQueriesWithMVs.enabled=true&lt;/code&gt; in your Spark session configuration.&lt;/li&gt;
 &lt;li&gt;Verify the rewrite by inspecting the optimized query plan for an MV scan node, or by checking INFO-level logs on Amazon EMR 7.14.0+.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Queries with multi-table joins, heavy aggregations, or window functions over large fact tables are strong initial candidates. Start with one high-cost, frequently executed query. Validate the speedup, then expand to broader MVs as you identify shared patterns across your workload.&lt;/p&gt;
&lt;p&gt;Special thanks to everyone who contributed to the automatic query rewrite feature and this blog: Andre Hernich, Leon Lin, Yiyang Chen, Geeta Krishna Panda, Ashok Chintalapati, Muhammad Malik, Rishabh Bhatia, and Giovanni Fumarola.&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;p&gt;For more detail, see the following resources:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/emr/latest/ReleaseGuide/emr-spark-materialized-views.html" target="_blank" rel="noopener"&gt;Using materialized views with Amazon EMR&lt;/a&gt;.&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/glue/latest/dg/materialized-views.html" target="_blank" rel="noopener"&gt;Using materialized views with AWS Glue&lt;/a&gt;.&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/athena/latest/ug/querying-iceberg-gdc-mv.html" target="_blank" rel="noopener"&gt;Querying materialized views in Amazon Athena&lt;/a&gt;.&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://aws.amazon.com/blogs/big-data/introducing-apache-iceberg-materialized-views-in-aws-glue-data-catalog/" target="_blank" rel="noopener"&gt;Introducing Apache Iceberg materialized views in AWS Glue Data Catalog&lt;/a&gt;.&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/lake-formation/latest/dg/materialized-views.html" target="_blank" rel="noopener"&gt;Materialized views in AWS Lake Formation&lt;/a&gt;.&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://aws.amazon.com/blogs/big-data/how-to-use-streamlined-permissions-for-amazon-s3-tables-and-iceberg-materialized-views/" target="_blank" rel="noopener"&gt;How to use streamlined permissions for Amazon S3 Tables and Iceberg materialized views&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2&gt;About the authors&lt;/h2&gt;
&lt;footer&gt;
 &lt;div class="blog-author-box"&gt;
  &lt;div class="blog-author-image"&gt;
   &lt;p&gt;&lt;img loading="lazy" class="alignleft size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/17/BDB-5950-2.jpg" alt="Yuzhou Sun" width="100" height="100"&gt;&lt;/p&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Yuzhou Sun&lt;/h3&gt;
  &lt;p&gt;&lt;a href="https://www.linkedin.com/in/yuzhou-sun-b3242092/" target="_blank" rel="noopener"&gt;Yuzhou&lt;/a&gt; is a software development engineer for Open Data Analytics Engines at Amazon Web Services.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box"&gt;
  &lt;div class="blog-author-image"&gt;
   &lt;p&gt;&lt;img loading="lazy" class="alignleft size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/17/BDB-5950-3.jpg" alt="Srishti Mittal" width="100" height="100"&gt;&lt;/p&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Srishti Mittal&lt;/h3&gt;
  &lt;p&gt;&lt;a href="https://www.linkedin.com/in/srishti-mittal-8352bba0/" target="_blank" rel="noopener"&gt;Srishti&lt;/a&gt; is a product manager for Open Data Analytics Engines at Amazon Web Services.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box"&gt;
  &lt;div class="blog-author-image"&gt;
   &lt;p&gt;&lt;img loading="lazy" class="alignleft size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/17/BDB-5950-4.jpg" alt="Kinshuk Pahare" width="100" height="100"&gt;&lt;/p&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Kinshuk Pahare&lt;/h3&gt;
  &lt;p&gt;&lt;a href="https://www.linkedin.com/in/kinshukpahare/" target="_blank" rel="noopener"&gt;Kinshuk&lt;/a&gt; serves as Head of Product for Analytics Engines at AWS, where he leads the product teams responsible for Amazon Redshift, AWS Glue, Amazon EMR, and Amazon Athena. With over six years at AWS, he brings deep expertise in building and scaling cloud-native analytics platforms that help organizations unlock the value of their data at any scale.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box"&gt;
  &lt;div class="blog-author-image"&gt;
   &lt;p&gt;&lt;img loading="lazy" class="alignleft size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/17/BDB-5950-5.jpg" alt="Henry Laih" width="100" height="100"&gt;&lt;/p&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Henry Laih&lt;/h3&gt;
  &lt;p&gt;&lt;a href="https://www.linkedin.com/in/yen-li-laih-a975491a2/" target="_blank" rel="noopener"&gt;Henry&lt;/a&gt; is a software development engineer for Open Data Analytics Engines at Amazon Web Services.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box"&gt;
  &lt;div class="blog-author-image"&gt;
   &lt;p&gt;&lt;img loading="lazy" class="alignleft size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/17/BDB-5950-6.jpg" alt="Srikanth Kandula" width="100" height="100"&gt;&lt;/p&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Srikanth Kandula&lt;/h3&gt;
  &lt;p&gt;&lt;a href="https://www.linkedin.com/in/srikanthkandula/" target="_blank" rel="noopener"&gt;Srikanth&lt;/a&gt; is an engineer who works in analytics and distributed systems at Amazon Web Services.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box"&gt;
  &lt;div class="blog-author-image"&gt;
   &lt;p&gt;&lt;img loading="lazy" class="alignleft size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/17/BDB-5950-7.jpg" alt="Shahryar Baki" width="100" height="100"&gt;&lt;/p&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Shahryar Baki&lt;/h3&gt;
  &lt;p&gt;&lt;a href="https://www.linkedin.com/in/shahryar-baki-31a5b299/" target="_blank" rel="noopener"&gt;Shahryar&lt;/a&gt; is a software development engineer for Open Data Analytics Engines at Amazon Web Services.&lt;/p&gt;
 &lt;/div&gt;
&lt;/footer&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>Every team is a data team — bring Amazon Redshift analytics to ChatGPT Work</title>
		<link>https://aws.amazon.com/blogs/big-data/every-team-is-a-data-team-bring-amazon-redshift-analytics-to-chatgpt-work/</link>
		
		<dc:creator><![CDATA[Naresh Chainani]]></dc:creator>
		<pubDate>Thu, 10 Sep 2026 15:14:01 +0000</pubDate>
				<category><![CDATA[Amazon Redshift]]></category>
		<category><![CDATA[Announcements]]></category>
		<guid isPermaLink="false">bfcb19e5913c73719fbbf22751605789aa04153a</guid>

					<description>AWS is announcing the AWS Data Analytics plugin for the new Data agent in ChatGPT Work. Teams can ask questions in natural language, analyze governed data across their Amazon Redshift data warehouse and data lakes, and build shareable dashboards, all from a conversation in ChatGPT Work.</description>
										<content:encoded>&lt;p&gt;Today, AWS is announcing the &lt;a href="https://chatgpt.com/plugins/plugin_asdk_app_6a917708c64481918f7ddea0cef4b8e6?q=aws" target="_blank" rel="noopener"&gt;AWS Data Analytics plugin&lt;/a&gt; for the new &lt;a href="https://chatgpt.com/plugins/Plugin_fc9843a6fb34819195d6c7802398a8a7" target="_blank" rel="noopener"&gt;Data agent&lt;/a&gt; in ChatGPT Work. The plugin helps teams across an organization ask questions in natural language, analyze governed data across their Amazon Redshift data warehouse and data lakes, and create shareable dashboards. All this happens from a conversation in ChatGPT Work.&lt;/p&gt;
&lt;p&gt;Tens of thousands of customers choose Amazon Redshift every day to run their most demanding workloads, because it delivers analytics at scale with industry-leading price performance. They love how Amazon Redshift provides access to their data warehouses and data lakes together in one place. Teams can combine curated business data with the broader operational, historical, and third-party data stored in open formats like Apache Iceberg in their data lakes. This gives them a complete picture to make business-critical decisions across their data.&lt;/p&gt;
&lt;p&gt;Customers have asked AWS for a way to put that trusted data in the hands of more of their people. That means not only the analysts and engineers who write SQL, but also the sales leaders, operations managers, and finance teams who depend on the results. A sales leader wants to know how the customer pipeline has changed this quarter. An operations manager wants to understand why fulfillment times changed over the past month. That’s why we built the AWS Data Analytics plugin, bringing the power of Amazon Redshift and AWS analytics to ChatGPT Work.&lt;/p&gt;
&lt;blockquote&gt;
 &lt;p&gt;&lt;em&gt;“Business teams can make decisions faster when they can source their own analytics and build the dashboards they need. Our work with AWS gives more people that ability, helping them understand changes in performance and decide where to focus. The AWS Data Analytics plugin connects Amazon Redshift to the Data agent in ChatGPT Work, so employees can analyze trusted company data simply by asking, with their organization’s existing access controls in place.”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;— Arpan Shah, General Manager, Technology at OpenAI&lt;/p&gt;
&lt;p&gt;The new plugin helps shorten the path from question to decision for everyone. Using the Data agent in ChatGPT Work, employees can explore the data they are authorized to access in Amazon Redshift by asking questions in everyday language. They can then refine the analysis, investigate changes, and turn the results into a dashboard without leaving ChatGPT Work. The plugin works with both Amazon Redshift provisioned clusters and Serverless workgroups. Customers can integrate it into their existing multi-cluster or multi-workgroup environments and benefit from the cost and security controls they’ve already set up.&lt;/p&gt;
&lt;p&gt;Consider Maya, a business analyst supporting a revenue operations team. She wants to understand the revenue performance across various segments and regions.&lt;/p&gt;
&lt;p&gt;Maya starts by loading the AWS Data Analytics plugin in ChatGPT Work, and then asking:&lt;/p&gt;
&lt;blockquote&gt;
 &lt;p&gt;&lt;em&gt;What are the revenue metrics for the past 30 days compared to the previous 30-day period?&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/09/BDB-6266-1.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/09/BDB-6266-1.png" alt="ChatGPT Work conversation asking for revenue metrics over the past 30 days compared to the previous 30-day period" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 1: Asking for revenue metrics in ChatGPT Work using the AWS Data Analytics plugin&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;The plugin translates her question into SQL, or a sequence of queries if needed, and runs them against the relevant data in Amazon Redshift. It returns key revenue performance metrics based on the same curated revenue data that her analytics team maintains.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/09/BDB-6266-2.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/09/BDB-6266-2.png" alt="Table of revenue performance metrics the plugin returned from Amazon Redshift" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 2: Revenue performance metrics returned from Amazon Redshift&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;Maya notices that gross margin is declining and asks a follow-up question:&lt;/p&gt;
&lt;blockquote&gt;
 &lt;p&gt;&lt;em&gt;What is my revenue breakdown by product category and region for the past 90 days?&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/09/BDB-6266-3.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/09/BDB-6266-3.png" alt="Revenue results segmented by product category and region for the past 90 days in ChatGPT Work" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 3: Revenue breakdown by product category and region for the past 90 days&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;The plugin carries the context forward, segments the results, and helps Maya understand each segment’s performance for the past 90 days. She can inspect the analysis and ask additional questions to drill down even further to understand why certain regions are lagging or why certain segments are outperforming others.&lt;/p&gt;
&lt;p&gt;This conversational workflow doesn’t replace the data models, metric definitions, or governance practices that the analytics team has established. It helps more employees use that data directly, giving analysts more time for high-value work.&lt;/p&gt;
&lt;p&gt;The AWS Data Analytics plugin connects ChatGPT Work to Amazon Redshift and uses the context of the connected analytics environment to help answer questions with the Data agent. During a conversation, it can:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;Discover the schemas, tables, columns, and data types available to the user.&lt;/li&gt;
 &lt;li&gt;Translate a natural-language question into Amazon Redshift SQL.&lt;/li&gt;
 &lt;li&gt;Run the query against the customer’s Amazon Redshift environment.&lt;/li&gt;
 &lt;li&gt;Present the results in a table or concise explanation.&lt;/li&gt;
 &lt;li&gt;Use follow-up questions to filter, compare, or drill into the results.&lt;/li&gt;
 &lt;li&gt;Turn an analysis into an interactive dashboard that teams can share and explore.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Because the analysis runs against the customer’s existing data, teams can continue to use the curated datasets and business definitions they already maintain in Amazon Redshift. Customers whose Amazon Redshift environments query data in both a warehouse and a data lake can also make that data available through the governed datasets exposed to the plugin. The AWS Data Analytics plugin also supports our broader AWS data and analytics services. This includes the ability to work with AWS Glue Data Catalog, Amazon S3 Tables (a capability of Amazon Simple Storage Service (Amazon S3)), Amazon Athena, and vector search on AWS.&lt;/p&gt;
&lt;p&gt;Natural-language analytics requires more than passing a prompt to a database. The agent needs to understand SQL specific to Amazon Redshift, discover metadata, choose the right tables and columns, and construct queries that follow service best practices. The plugin was built using &lt;a href="https://github.com/aws/agent-toolkit-for-aws/tree/1799aae50f30a8e97bab4faf38f2eadbd665028b/skills/specialized-skills/analytics-skills/redshift-guide" target="_blank" rel="noopener"&gt;Amazon Redshift skills&lt;/a&gt; from the &lt;a href="https://aws.amazon.com/products/developer-tools/agent-toolkit-for-aws/" target="_blank" rel="noopener"&gt;Agent Toolkit for AWS&lt;/a&gt;. These skills provide tested procedures and service-specific guidance that agents can use when working with Amazon Redshift.&lt;/p&gt;
&lt;p&gt;To get started, install the AWS Data Analytics plugin in ChatGPT Work to connect it to Amazon Redshift. Give your teams a conversational path to governed insights across your data warehouse and data lake today.&lt;/p&gt;
&lt;p&gt;To learn more, see the following resources:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;a href="https://chatgpt.com/plugins/plugin_asdk_app_6a917708c64481918f7ddea0cef4b8e6?q=aws" target="_blank" rel="noopener"&gt;AWS Data Analytics plugin&lt;/a&gt;&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://chatgpt.com/plugins/Plugin_fc9843a6fb34819195d6c7802398a8a7" target="_blank" rel="noopener"&gt;Data agent in ChatGPT Work&lt;/a&gt;&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://openai.com/index/put-data-to-work/" target="_blank" rel="noopener"&gt;OpenAI’s announcement&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p style="clear: both"&gt;&lt;/p&gt;
&lt;hr style="width: 100%"&gt;
&lt;h2&gt;About the author&lt;/h2&gt;
&lt;footer&gt;
 &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
  &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
   &lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/09/BDB-6266-4.jpg" alt="Naresh Chainani" width="100" height="133"&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Naresh Chainani&lt;/h3&gt;
  &lt;p style="overflow: hidden"&gt;Naresh is a Director of Engineering at AWS, where he leads Amazon Redshift, one of the world’s most widely used cloud data warehouses. With over 20 years of experience across IBM and AWS, he is a recognized leader in high-performance database systems, holding more than a dozen patents and numerous publications at top venues including SIGMOD and VLDB. Naresh is passionate about advancing the state of the art in analytics and developing the next generation of engineering talent.&lt;/p&gt;
 &lt;/div&gt;
&lt;/footer&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>Build declarative ETL pipelines with AWS Glue 6.0</title>
		<link>https://aws.amazon.com/blogs/big-data/build-declarative-etl-pipelines-with-aws-glue-6-0/</link>
		
		<dc:creator><![CDATA[Syed Humair]]></dc:creator>
		<pubDate>Wed, 09 Sep 2026 17:38:22 +0000</pubDate>
				<category><![CDATA[Advanced (300)]]></category>
		<category><![CDATA[AWS Glue]]></category>
		<category><![CDATA[Technical How-to]]></category>
		<guid isPermaLink="false">fc8fc732ee4aa2580d6500a61a8577ae03e7fc35</guid>

					<description>AWS Glue 6.0 introduces Spark Declarative Pipelines. In this post, you build a single declarative AWS Glue 6.0 job that turns raw order records into validated, aggregated tables through a bronze, silver, and gold sequence, without writing any orchestration logic.</description>
										<content:encoded>&lt;p&gt;Data teams commonly build the extract, transform, and load (ETL) pipelines that turn raw order events into analyst-ready aggregates as a bronze, silver, and gold sequence, the medallion architecture. Bronze holds raw ingested records, silver holds cleaned and validated data, and gold holds the business-level aggregates that analysts query. Today you build this on &lt;a href="https://aws.amazon.com/glue/" target="_blank" rel="noopener"&gt;AWS Glue&lt;/a&gt; with an orchestrator such as &lt;a href="https://aws.amazon.com/managed-workflows-for-apache-airflow/" target="_blank" rel="noopener"&gt;Amazon Managed Workflows for Apache Airflow&lt;/a&gt; (Amazon MWAA) or &lt;a href="https://aws.amazon.com/step-functions/" target="_blank" rel="noopener"&gt;AWS Step Functions&lt;/a&gt; coordinating the stages. Many teams run production pipelines exactly this way. As a pipeline grows, the coordination work grows with it: you wire job dependencies, manage intermediate checkpoints, and add retry logic stage by stage.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://aws.amazon.com/blogs/aws/aws-glue-6-0-now-available-with-30-lower-price-and-full-apache-iceberg-v3-support/" target="_blank" rel="noopener"&gt;AWS Glue 6.0&lt;/a&gt;, powered by &lt;a href="https://spark.apache.org/docs/4.1.1/" target="_blank" rel="noopener"&gt;Apache Spark 4.1&lt;/a&gt;, introduces &lt;a href="https://docs.aws.amazon.com/glue/latest/dg/spark-declarative-pipelines.html" target="_blank" rel="noopener"&gt;Spark Declarative Pipelines&lt;/a&gt; (SDP), which simplifies this further. Instead of orchestrating jobs by hand, you declare what each dataset should contain and let the declarative framework resolve dependencies, manage checkpoints, and orchestrate execution order automatically. The result runs as a single declarative job, with no manual directed acyclic graph (DAG) wiring or imperative orchestration code.&lt;/p&gt;
&lt;p&gt;In this post, you build a single AWS Glue 6.0 job that turns raw order records into validated, aggregated, analytics-ready tables through the bronze, silver, and gold sequence. You do this without writing any orchestration logic. This walkthrough uses the AWS Command Line Interface (&lt;a href="https://aws.amazon.com/cli/" target="_blank" rel="noopener"&gt;AWS CLI&lt;/a&gt;), and the same operations are available through the &lt;a href="https://aws.amazon.com/developer/tools/" target="_blank" rel="noopener"&gt;AWS SDKs&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="solution-overview"&gt;Solution overview&lt;/h2&gt;
&lt;p&gt;You build a single AWS Glue 6.0 job that reads raw order records from a CSV file in Amazon Simple Storage Service (&lt;a href="https://aws.amazon.com/s3/" target="_blank" rel="noopener"&gt;Amazon S3&lt;/a&gt;). The job flows them through three declared datasets. These are a bronze materialized view (ingest as-is), a silver materialized view (type, validate, and classify), and a gold SQL materialized view (aggregate by region). With &lt;a href="https://docs.aws.amazon.com/glue/latest/dg/catalog-and-crawler.html" target="_blank" rel="noopener"&gt;AWS Glue Data Catalog&lt;/a&gt; integration turned on, all three land as Data Catalog tables, queryable with standard SQL tooling such as &lt;a href="https://aws.amazon.com/athena/" target="_blank" rel="noopener"&gt;Amazon Athena&lt;/a&gt;. SDP resolves the dependency order from the dataset references in your code, so you never orchestrate the steps yourself.&lt;/p&gt;
&lt;h2 id="before-and-after-imperative-compared-to-declarative"&gt;Two ways to build the pipeline&lt;/h2&gt;
&lt;p&gt;Before you build the pipeline, let’s understand this new way of writing ETL pipelines with a quick comparison of the imperative and declarative approaches.&lt;/p&gt;
&lt;p&gt;With the imperative approach, you need three AWS Glue jobs, plus an orchestrator to handle sequencing and error handling. A typical pipeline therefore has two layers: an orchestration layer and the ETL processing layer. The following diagram shows this two-layer imperative pipeline.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/03/BDB-6104-1.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/03/BDB-6104-1.png" alt="Two-layer imperative pipeline: three AWS Glue jobs coordinated by an orchestrator." width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 1: The two-layer imperative pipeline, with three AWS Glue jobs coordinated by an orchestrator.&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;Compared to that, the declarative approach runs as a single ETL job with SDP. The following diagram mirrors the previous one, but here it is a single AWS Glue ETL job instead of three jobs plus an orchestrator.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/03/BDB-6104-2.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/03/BDB-6104-2.png" alt="Declarative pipeline: a single AWS Glue job running the bronze, silver, and gold layers with Spark Declarative Pipelines." width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 2: The declarative pipeline, a single AWS Glue job running the bronze, silver, and gold layers with SDP.&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;The declarative approach reduces more than the number of jobs. It removes the boilerplate that surrounds them. An orchestrator such as Amazon MWAA or AWS Step Functions already handles retries and parallelism, but only at the granularity of a whole job. To get finer control, teams often split a pipeline into several jobs and then hand-wire the dependencies between them. With SDP, you no longer hand-wire a DAG, manage per-stage checkpoints, or split the pipeline into separate jobs for retries and parallelism. SDP derives the dependency graph from your table references and coordinates execution at the level of individual tables. You can still invoke an SDP job from an orchestrator when a broader workflow calls for it, but the pipeline’s internal coordination is no longer code you write and maintain.&lt;/p&gt;
&lt;p&gt;SDP separates the &lt;em&gt;what&lt;/em&gt; from the &lt;em&gt;how&lt;/em&gt;: you declare &lt;strong&gt;datasets&lt;/strong&gt; (the outputs you want), and SDP builds the &lt;strong&gt;flows&lt;/strong&gt; that produce them and runs them as one &lt;strong&gt;pipeline&lt;/strong&gt;, resolving dependencies and execution order automatically.&lt;/p&gt;
&lt;p&gt;You declare these abstractions through Python decorators. This post covers three of them, &lt;code&gt;@dp.table&lt;/code&gt;, &lt;code&gt;@dp.materialized_view&lt;/code&gt;, and &lt;code&gt;@dp.temporary_view&lt;/code&gt;, each with its own purpose:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;code&gt;@dp.table&lt;/code&gt; defines a streaming table, which processes new data incrementally on each run. Typical use cases are raw event ingestion and change data capture (CDC) feeds.&lt;/li&gt;
 &lt;li&gt;&lt;code&gt;@dp.materialized_view&lt;/code&gt; defines a materialized view for batch use cases. Today, this dataset type fully recomputes on each run. Common uses include parsing, aggregations, and machine learning (ML) feature engineering.&lt;/li&gt;
 &lt;li&gt;&lt;code&gt;@dp.temporary_view&lt;/code&gt; is for temporary computations and aggregations. It’s pipeline-scoped and isn’t persisted outside the pipeline. Use it for enrichment lookups and subqueries.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Streaming tables append only new arrivals. Materialized views fully recompute. This post uses &lt;code&gt;@dp.materialized_view&lt;/code&gt; for all three layers to keep the walkthrough focused. In production, you would typically use &lt;code&gt;@dp.table&lt;/code&gt; for the bronze layer to process only new files as they arrive rather than re-reading the full source each run.&lt;/p&gt;
&lt;h3 id="running-and-refreshing-the-pipeline"&gt;Running and refreshing the pipeline&lt;/h3&gt;
&lt;p&gt;When you rerun a pipeline, you don’t always want the same work to happen. Sometimes you only want to confirm the pipeline is well-formed before spending compute. Other times you want to run it but recompute only the datasets that changed rather than the entire graph. SDP handles both cases through two independent controls, and it helps to keep them separate:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;Execution mode&lt;/strong&gt; (the &lt;code&gt;spark.glue.sdp.jobMode&lt;/code&gt; key) answers &lt;em&gt;run or only validate?&lt;/em&gt;&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Refresh scope&lt;/strong&gt; (the &lt;code&gt;spark.glue.sdp.runMode&lt;/code&gt; key) answers &lt;em&gt;given that I’m running, what do I recompute?&lt;/em&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Execution mode.&lt;/strong&gt; &lt;code&gt;VALIDATE&lt;/code&gt; runs the pipeline in dry-run mode: SDP checks the YAML syntax, dependency resolution, and SQL and Python compilation without writing any data. Use it to verify your pipeline is well-formed before committing compute. &lt;code&gt;RUN&lt;/code&gt; (the default) executes the pipeline normally, resolving the dependency graph and materializing datasets.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;# Dry run: validate the graph, write nothing
aws glue start-job-run \
  --job-name "${JOB_NAME}" \
  --arguments '{"--conf":"spark.glue.sdp.jobMode=VALIDATE"}' \
  --region "${AWS_REGION}"

# Normal execution
aws glue start-job-run \
  --job-name "${JOB_NAME}" \
  --arguments '{"--conf":"spark.glue.sdp.jobMode=RUN"}' \
  --region "${AWS_REGION}"&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;Refresh scope.&lt;/strong&gt; By default, a &lt;code&gt;RUN&lt;/code&gt; recomputes every materialized view. You can narrow or widen that with &lt;code&gt;spark.glue.sdp.runMode&lt;/code&gt;:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;code&gt;--refresh &amp;lt;datasets&amp;gt;&lt;/code&gt; updates only the named datasets (comma-separated, no spaces).&lt;/li&gt;
 &lt;li&gt;&lt;code&gt;--full-refresh &amp;lt;datasets&amp;gt;&lt;/code&gt; resets and recomputes only the named datasets (for streaming tables, this also clears their checkpoints).&lt;/li&gt;
 &lt;li&gt;&lt;code&gt;--full-refresh-all&lt;/code&gt; resets and recomputes every dataset in the pipeline.&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;# Selective refresh of named datasets
aws glue start-job-run \
  --job-name "${JOB_NAME}" \
  --arguments '{"--conf":"spark.glue.sdp.jobMode=RUN --conf spark.glue.sdp.runMode=--refresh silver_orders,gold_sales_summary"}' \
  --region "${AWS_REGION}"

# Full reset and recompute of the entire pipeline
aws glue start-job-run \
  --job-name "${JOB_NAME}" \
  --arguments '{"--conf":"spark.glue.sdp.jobMode=RUN --conf spark.glue.sdp.runMode=--full-refresh-all"}' \
  --region "${AWS_REGION}"&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Selective refresh is useful during development, so you can iterate on a single layer without reprocessing the entire graph. Note that &lt;code&gt;--refresh&lt;/code&gt; and &lt;code&gt;--full-refresh&lt;/code&gt; each take an explicit list of datasets. To reset the whole pipeline, use &lt;code&gt;--full-refresh-all&lt;/code&gt;. Because materialized views hold no incremental state, resetting a materialized view and refreshing it both fully recompute it. The reset-versus-refresh distinction matters for streaming tables, where a refresh processes only new data and a reset clears the checkpoint and reprocesses from scratch.&lt;/p&gt;
&lt;p&gt;The multiple values are passed as a single &lt;code&gt;--conf&lt;/code&gt; argument string (&lt;code&gt;"spark.glue.sdp.jobMode=RUN --conf spark.glue.sdp.runMode=..."&lt;/code&gt;). This is the serialization the AWS Glue SDP mode expects for the run.&lt;/p&gt;
&lt;h3 id="materialized-views-batch-transforms-with-automatic-dependency-resolution"&gt;Materialized views: Batch transforms with automatic dependency resolution&lt;/h3&gt;
&lt;p&gt;Materialized views recompute their full result set on each run. SDP infers dependencies from table references: in this pipeline, &lt;code&gt;silver_orders&lt;/code&gt; references &lt;code&gt;bronze_orders&lt;/code&gt;, so SDP runs bronze first, as shown in the following diagram.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/03/BDB-6104-3.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/03/BDB-6104-3.png" alt="Dependency graph showing Spark Declarative Pipelines running the bronze layer before the silver layer." width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 3: SDP infers the dependency order from table references and runs bronze before silver.&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;The core pattern is a decorated function that returns a DataFrame:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-python"&gt;@dp.materialized_view(comment="Raw orders loaded from CSV")
def bronze_orders() -&amp;gt; DataFrame:
    return spark.read.schema(ORDERS_SCHEMA).option("header", "true").csv(ORDERS_PATH)&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;The silver layer references &lt;code&gt;bronze_orders&lt;/code&gt; through &lt;code&gt;spark.table("bronze_orders")&lt;/code&gt;, with no explicit dependency declaration. SDP builds the DAG by analyzing table references in your code and runs bronze first automatically.&lt;/p&gt;
&lt;p&gt;Bronze reads every column as a string by design: the bronze layer preserves raw source data without coercion. Type casting, validation, and filtering happen in the silver layer.&lt;/p&gt;
&lt;h3 id="sql-and-python-coexistence"&gt;SQL and Python coexistence&lt;/h3&gt;
&lt;p&gt;SDP supports both Python and SQL definitions in the same pipeline project. A SQL materialized view can reference a Python-defined table directly, for example the gold layer aggregating the silver table:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-sql"&gt;CREATE MATERIALIZED VIEW gold_sales_summary
COMMENT 'Completed-order metrics by region'
AS
SELECT
  region,
  COUNT(*) AS order_count,
  CAST(ROUND(SUM(amount), 2) AS DECIMAL(10, 2)) AS total_sales,
  CAST(ROUND(AVG(amount), 2) AS DECIMAL(10, 2)) AS average_order_value
FROM silver_orders
GROUP BY region;&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;In this post, Python files define ingestion and validation logic, and SQL files define reporting views and aggregations. SDP discovers both through the &lt;code&gt;libraries&lt;/code&gt; glob pattern in the pipeline specification and resolves the cross-language dependencies automatically. The complete source for all three layers follows in the step-by-step walkthrough.&lt;/p&gt;
&lt;h2 id="build-the-pipeline-step-by-step"&gt;Build the pipeline: Step by step&lt;/h2&gt;
&lt;p&gt;The rest of this post is a hands-on walkthrough. You build a single AWS Glue 6.0 job that reads &lt;code&gt;orders.csv&lt;/code&gt; and processes it through the bronze, silver, and gold layers. The steps are:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;&lt;strong&gt;Prerequisites&lt;/strong&gt;: AWS account, AWS Identity and Access Management (IAM) role, and S3 bucket.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Set up sample data&lt;/strong&gt;: create &lt;code&gt;orders.csv&lt;/code&gt; and upload it to Amazon S3.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Build the pipeline files&lt;/strong&gt; (the &lt;code&gt;spark-pipeline.yml&lt;/code&gt; specification plus the three transformation files).&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Package the pipeline&lt;/strong&gt; into a zip and upload it to Amazon S3.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Create the database&lt;/strong&gt;: a Data Catalog database with an S3 location.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Configure the job&lt;/strong&gt;: create the AWS Glue 6.0 job with the SDP flag.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Validate&lt;/strong&gt;: run in dry-run mode to verify the graph.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Run the pipeline&lt;/strong&gt; to materialize all datasets.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Query results&lt;/strong&gt;: inspect the tables with Amazon Athena.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Clean up&lt;/strong&gt;: delete the resources you created.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="step-1---prerequisites"&gt;Step 1 – Prerequisites&lt;/h2&gt;
&lt;p&gt;To follow along, you need:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;An AWS account with access to AWS Glue 6.0.&lt;/li&gt;
 &lt;li&gt;A dedicated IAM role trusted by &lt;code&gt;glue.amazonaws.com&lt;/code&gt; (set up in the following section).&lt;/li&gt;
 &lt;li&gt;A private, encrypted Amazon S3 bucket with Block Public Access enabled.&lt;/li&gt;
 &lt;li&gt;The AWS CLI configured with credentials for a non-production account.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="iam-role-for-the-pipeline"&gt;IAM role for the pipeline&lt;/h3&gt;
&lt;p&gt;Create a role that AWS Glue can assume, with the following trust policy:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-json"&gt;{
  "Version": "2012-10-17",
  "Statement": [{
    "Effect": "Allow",
    "Principal": { "Service": "glue.amazonaws.com" },
    "Action": "sts:AssumeRole"
  }]
}&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Attach the AWS managed policy &lt;strong&gt;AWSGlueServiceRole&lt;/strong&gt;, which grants the AWS Glue Data Catalog and &lt;a href="https://aws.amazon.com/cloudwatch/" target="_blank" rel="noopener"&gt;Amazon CloudWatch Logs&lt;/a&gt; access the job needs. Then add an inline policy that scopes Amazon S3 access to your bucket, covering the input data, the pipeline zip, the pipeline storage (state) path, and the warehouse location:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-json"&gt;{
  "Version": "2012-10-17",
  "Statement": [{
    "Effect": "Allow",
    "Action": ["s3:GetObject", "s3:PutObject", "s3:DeleteObject", "s3:ListBucket"],
    "Resource": [
      "arn:aws:s3:::amzn-s3-demo-bucket",
      "arn:aws:s3:::amzn-s3-demo-bucket/*"
    ]
  }]
}&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;For a full breakdown of the baseline permissions, see &lt;a href="https://docs.aws.amazon.com/glue/latest/dg/set-up-iam.html" target="_blank" rel="noopener"&gt;Setting up IAM permissions for AWS Glue&lt;/a&gt;.&lt;/p&gt;
&lt;h3 id="set-the-walkthrough-variables"&gt;Set the walkthrough variables&lt;/h3&gt;
&lt;p&gt;Set the following variables, replacing the example values (&lt;code&gt;us-east-1&lt;/code&gt;, &lt;code&gt;amzn-s3-demo-bucket&lt;/code&gt;, the account ID &lt;code&gt;111122223333&lt;/code&gt;, and the role name) with your own:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;export AWS_REGION="us-east-1"
export BUCKET="amzn-s3-demo-bucket"
export PREFIX="simple-sdp-demo"
export DATABASE="simple_sdp_demo_db"
export ROLE_ARN="arn:aws:iam::111122223333:role/AWSGlueServiceRole-sdp-demo"
export JOB_NAME="simple-sdp-demo"&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h2 id="step-2---set-up-sample-data"&gt;Step 2 – Set up sample data&lt;/h2&gt;
&lt;p&gt;The pipeline reads a CSV of order records. Save the following as &lt;code&gt;orders.csv&lt;/code&gt;:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-csv"&gt;order_id,customer_id,region,amount,status,order_ts
O-1001,C-101,EMEA,120.50,COMPLETE,2026-07-23T08:00:00Z
O-1002,C-102,AMER,750.00,COMPLETE,2026-07-23T08:15:00Z
O-1003,C-103,EMEA,-10.00,INVALID,2026-07-23T08:30:00Z
O-1004,C-104,APAC,320.25,COMPLETE,2026-07-23T09:00:00Z
O-1005,C-105,AMER,250.00,COMPLETE,2026-07-23T09:15:00Z
O-1006,C-106,EMEA,90.00,COMPLETE,2026-07-23T09:30:00Z&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Upload the file to the &lt;code&gt;input/&lt;/code&gt; location under your project prefix, which is where the bronze layer reads it (the &lt;code&gt;ORDERS_PATH&lt;/code&gt; in &lt;code&gt;01_bronze.py&lt;/code&gt;, shown in Step 3). Use the variables you exported in Step 1:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws s3 cp orders.csv \
  "s3://${BUCKET}/${PREFIX}/input/orders.csv" \
  --region "${AWS_REGION}"&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;The file includes one invalid order (O-1003, a negative amount), which the silver layer filters out to demonstrate the validation step. The AMER and EMEA regions each have two completed orders, so the gold layer’s &lt;code&gt;order_count&lt;/code&gt; and &lt;code&gt;average_order_value&lt;/code&gt; are meaningful aggregations rather than single-row passthroughs.&lt;/p&gt;
&lt;h2 id="step-3---build-the-pipeline-files"&gt;Step 3 – Build the pipeline files&lt;/h2&gt;
&lt;p&gt;The pipeline project uses the structure introduced earlier: a &lt;code&gt;transformations/&lt;/code&gt; folder holding the three layer definitions (&lt;code&gt;01_bronze.py&lt;/code&gt;, &lt;code&gt;02_silver.py&lt;/code&gt;, &lt;code&gt;03_gold.sql&lt;/code&gt;), plus the &lt;code&gt;spark-pipeline.yml&lt;/code&gt; specification. The following screenshot shows this layout in a code editor.&lt;/p&gt;
&lt;div style="width: 660px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/03/BDB-6104-4.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/03/BDB-6104-4.png" alt="Pipeline project layout in a code editor, showing the transformations folder and the spark-pipeline.yml file." width="650"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 4: The pipeline project layout in a code editor.&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;The complete contents of each file follow.&lt;/p&gt;
&lt;h3 id="a.-spark-pipeline.yml"&gt;3a. spark-pipeline.yml&lt;/h3&gt;
&lt;p&gt;The specification names the pipeline, points to the Data Catalog database, configures state storage, and discovers transformation files. As with the transformation files, it uses the &lt;code&gt;__DATABASE__&lt;/code&gt;, &lt;code&gt;__BUCKET__&lt;/code&gt;, and &lt;code&gt;__PREFIX__&lt;/code&gt; tokens, which you substitute at packaging time in Step 4:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-yaml"&gt;name: simple_sdp_demo
catalog: spark_catalog
database: __DATABASE__
storage: s3://__BUCKET__/__PREFIX__/state/
libraries:
  - glob:
      include: transformations/**
configuration:
  spark.sql.shuffle.partitions: "4"&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h3 id="b.-transformations01_bronze.py"&gt;3b. transformations/01_bronze.py&lt;/h3&gt;
&lt;p&gt;Bronze preserves the raw source as strings. No coercion, no filtering:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-python"&gt;"""Bronze layer: preserve source order records as strings."""
from pyspark import pipelines as dp
from pyspark.sql import DataFrame, SparkSession
from pyspark.sql.types import StringType, StructField, StructType

spark = SparkSession.active()

ORDERS_PATH = "s3://__BUCKET__/__PREFIX__/input/orders.csv"

ORDERS_SCHEMA = StructType([
    StructField("order_id", StringType(), True),
    StructField("customer_id", StringType(), True),
    StructField("region", StringType(), True),
    StructField("amount", StringType(), True),
    StructField("status", StringType(), True),
    StructField("order_ts", StringType(), True),
])


@dp.materialized_view(comment="Raw orders loaded from CSV")
def bronze_orders() -&amp;gt; DataFrame:
    return (
        spark.read
        .schema(ORDERS_SCHEMA)
        .option("header", "true")
        .csv(ORDERS_PATH)
    )&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;The path uses the tokens &lt;code&gt;__BUCKET__&lt;/code&gt; and &lt;code&gt;__PREFIX__&lt;/code&gt; rather than hardcoded values. AWS Glue reads these files from the packaged zip at runtime, so shell variables like &lt;code&gt;${BUCKET}&lt;/code&gt; are not expanded inside them. You substitute the tokens with your real values when you package the project in Step 4, which keeps every file consistent with the variables you exported in Step 1.&lt;/p&gt;
&lt;h3 id="c.-transformations02_silver.py"&gt;3c. transformations/02_silver.py&lt;/h3&gt;
&lt;p&gt;Silver casts types, filters to complete orders with positive amounts, and derives an &lt;code&gt;amount_band&lt;/code&gt; classification:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-python"&gt;"""Silver layer: type, validate, and classify complete orders."""
from pyspark import pipelines as dp
from pyspark.sql import DataFrame, SparkSession
from pyspark.sql.functions import col, to_timestamp, trim, when

spark = SparkSession.active()


@dp.materialized_view(comment="Validated complete orders with typed values")
def silver_orders() -&amp;gt; DataFrame:
    typed = (
        spark.table("bronze_orders")
        .select(
            trim(col("order_id")).alias("order_id"),
            trim(col("customer_id")).alias("customer_id"),
            trim(col("region")).alias("region"),
            col("amount").cast("double").alias("amount"),
            trim(col("status")).alias("status"),
            to_timestamp("order_ts", "yyyy-MM-dd'T'HH:mm:ss'Z'").alias("order_ts"),
        )
        .filter(
            col("order_id").isNotNull()
            &amp;amp; col("region").isNotNull()
            &amp;amp; col("order_ts").isNotNull()
            &amp;amp; (col("status") == "COMPLETE")
            &amp;amp; (col("amount") &amp;gt; 0)
        )
    )
    return typed.select(
        "*",
        when(col("amount") &amp;gt;= 500, "large")
        .when(col("amount") &amp;gt;= 100, "medium")
        .otherwise("small")
        .alias("amount_band"),
    )&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Silver reads bronze with &lt;code&gt;spark.table("bronze_orders")&lt;/code&gt;, so SDP infers the dependency and runs bronze first. Two details matter here:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;The &lt;code&gt;to_timestamp&lt;/code&gt; call passes an explicit format, &lt;code&gt;"yyyy-MM-dd'T'HH:mm:ss'Z'"&lt;/code&gt;. The source timestamps are ISO 8601 with a &lt;code&gt;Z&lt;/code&gt; suffix. Giving the format treats &lt;code&gt;Z&lt;/code&gt; as a literal and produces the same wall-clock value regardless of the job’s session time zone, which keeps the result deterministic.&lt;/li&gt;
 &lt;li&gt;The transformation runs in two projections: the first casts and filters, and the second derives &lt;code&gt;amount_band&lt;/code&gt; from the already-typed &lt;code&gt;amount&lt;/code&gt; column. Deriving columns with &lt;code&gt;.select(...)&lt;/code&gt; rather than a separate &lt;code&gt;.withColumn(...)&lt;/code&gt; step keeps SDP’s reference to &lt;code&gt;bronze_orders&lt;/code&gt; resolvable as a pipeline dependency. This way, SDP consistently orders the bronze layer before the silver layer. The order matters here too. Spark 4.1 enables ANSI mode by default, so comparing the raw string &lt;code&gt;amount&lt;/code&gt; against a number would fail. &lt;code&gt;amount_band&lt;/code&gt; therefore reads the already-cast &lt;code&gt;amount&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="d.-transformations03_gold.sql"&gt;3d. transformations/03_gold.sql&lt;/h3&gt;
&lt;p&gt;The gold layer aggregates order metrics by region using SQL:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-sql"&gt;CREATE MATERIALIZED VIEW gold_sales_summary
COMMENT 'Completed-order metrics by region'
AS
SELECT
  region,
  COUNT(*) AS order_count,
  CAST(ROUND(SUM(amount), 2) AS DECIMAL(10, 2)) AS total_sales,
  CAST(ROUND(AVG(amount), 2) AS DECIMAL(10, 2)) AS average_order_value
FROM silver_orders
GROUP BY region;&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h2 id="step-4---package-the-project"&gt;Step 4 – Package the project&lt;/h2&gt;
&lt;p&gt;Substitute the &lt;code&gt;__BUCKET__&lt;/code&gt;, &lt;code&gt;__PREFIX__&lt;/code&gt;, and &lt;code&gt;__DATABASE__&lt;/code&gt; tokens with the values you exported in Step 1. Then package &lt;code&gt;spark-pipeline.yml&lt;/code&gt; and the &lt;code&gt;transformations/&lt;/code&gt; folder into a zip with both at the zip root. Because AWS Glue reads these files from the zip at runtime, the substitution has to happen now, at packaging time, not through shell variables at run time:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;# Render the tokens into a build/ copy, leaving your source files untouched
rm -rf build/package &amp;amp;&amp;amp; mkdir -p build/package/transformations

sed -e "s|__BUCKET__|${BUCKET}|g" \
    -e "s|__PREFIX__|${PREFIX}|g" \
    -e "s|__DATABASE__|${DATABASE}|g" \
    spark-pipeline.yml &amp;gt; build/package/spark-pipeline.yml

sed -e "s|__BUCKET__|${BUCKET}|g" \
    -e "s|__PREFIX__|${PREFIX}|g" \
    transformations/01_bronze.py &amp;gt; build/package/transformations/01_bronze.py
cp transformations/02_silver.py transformations/03_gold.sql build/package/transformations/

# Zip with the spec and transformations at the zip root
(cd build/package &amp;amp;&amp;amp; zip -r -q ../simple-sdp-demo.zip spark-pipeline.yml transformations)

# Upload
aws s3 cp build/simple-sdp-demo.zip "s3://${BUCKET}/${PREFIX}/pipeline/simple-sdp-demo.zip" --region "${AWS_REGION}"&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Only &lt;code&gt;spark-pipeline.yml&lt;/code&gt; and &lt;code&gt;01_bronze.py&lt;/code&gt; carry tokens, so the other files are copied as-is. The uploaded object is named &lt;code&gt;simple-sdp-demo.zip&lt;/code&gt;, which is the same name the job references in Step 6.&lt;/p&gt;
&lt;h2 id="step-5---create-the-database"&gt;Step 5 – Create the database&lt;/h2&gt;
&lt;p&gt;The database named in &lt;code&gt;spark-pipeline.yml&lt;/code&gt; must already exist in the AWS Glue Data Catalog, with an S3 location URI, before the pipeline runs. SDP does not create it automatically:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws glue get-database --name "${DATABASE}" --region "${AWS_REGION}" &amp;gt;/dev/null 2&amp;gt;&amp;amp;1 \
|| aws glue create-database \
--database-input "{\"Name\":\"${DATABASE}\",\"LocationUri\":\"s3://${BUCKET}/${PREFIX}/warehouse/\"}" \
--region "${AWS_REGION}"&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h2 id="step-6---configure-the-job"&gt;Step 6 – Configure the job&lt;/h2&gt;
&lt;p&gt;Create an AWS Glue 6.0 job with the zip as &lt;code&gt;ScriptLocation&lt;/code&gt; and the SDP flag enabled:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="lang-bash"&gt;aws glue create-job \
--name "${JOB_NAME}" \
--role "${ROLE_ARN}" \
--command "{\"Name\":\"glueetl\",\"ScriptLocation\":\"s3://${BUCKET}/${PREFIX}/pipeline/simple-sdp-demo.zip\",\"PythonVersion\":\"3\"}" \
--glue-version "6.0" \
--worker-type "G.1X" \
--number-of-workers 2 \
--default-arguments "{\"--enable-spark-declarative-pipeline\":\"true\",\"--enable-glue-datacatalog\":\"true\"}" \
--region "${AWS_REGION}"&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Key arguments:&lt;/p&gt;
&lt;table border="1px" width="100%" cellpadding="10px"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;Argument&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Purpose&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;code&gt;--enable-spark-declarative-pipeline&lt;/code&gt;&lt;/td&gt;
   &lt;td&gt;Activates the SDP executor (required)&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;code&gt;--enable-glue-datacatalog&lt;/code&gt;&lt;/td&gt;
   &lt;td&gt;Uses the AWS Glue Data Catalog as the Spark Hive metastore, so the pipeline’s output tables register in the catalog&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;code&gt;ScriptLocation&lt;/code&gt;&lt;/td&gt;
   &lt;td&gt;Points to the pipeline zip, not a .py file&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;em&gt;Table 2: Key arguments for the create-job command.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;The &lt;code&gt;create-job&lt;/code&gt; command sets &lt;code&gt;ScriptLocation&lt;/code&gt; to the pipeline zip. You can also point it to an Amazon S3 prefix: upload the unzipped &lt;code&gt;spark-pipeline.yml&lt;/code&gt; and &lt;code&gt;transformations/&lt;/code&gt; to a prefix and set &lt;code&gt;ScriptLocation&lt;/code&gt; to that prefix (with a trailing &lt;code&gt;/&lt;/code&gt;). No other change is needed, and the &lt;code&gt;--enable-spark-declarative-pipeline&lt;/code&gt; flag stays the same. The zip keeps the upload to a single object.&lt;/p&gt;
&lt;h2 id="step-7---validate-dry-run"&gt;Step 7 – Validate (dry run)&lt;/h2&gt;
&lt;p&gt;Run the job in validation mode first to verify the dependency graph without materializing data:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws glue start-job-run \
  --job-name "${JOB_NAME}" \
  --arguments '{"--conf":"spark.glue.sdp.jobMode=VALIDATE"}' \
  --region "${AWS_REGION}"&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Validation analyzes the project structure, dependency graph, and SQL and Python compilation without creating tables, executing transforms, or writing data. Confirm that the database has no tables after validation completes.&lt;/p&gt;
&lt;p&gt;On AWS Glue, validation runs as a job (&lt;code&gt;jobMode=VALIDATE&lt;/code&gt;), so you create the job in Step 6 and then validate it here. If you develop locally with the open source &lt;code&gt;spark-pipelines&lt;/code&gt; CLI, you can run its &lt;code&gt;dry-run&lt;/code&gt; against the project before packaging and uploading.&lt;/p&gt;
&lt;h2 id="step-8---run-the-pipeline"&gt;Step 8 – Run the pipeline&lt;/h2&gt;
&lt;p&gt;Start the pipeline in normal execution mode:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws glue start-job-run \
  --job-name "${JOB_NAME}" \
  --arguments '{"--conf":"spark.glue.sdp.jobMode=RUN"}' \
  --region "${AWS_REGION}"&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;After the run completes, list the materialized tables:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws glue get-tables \
  --database-name "${DATABASE}" \
  --region "${AWS_REGION}" \
  --query 'TableList[].Name' \
  --output table&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Expected tables: &lt;code&gt;bronze_orders&lt;/code&gt;, &lt;code&gt;silver_orders&lt;/code&gt;, &lt;code&gt;gold_sales_summary&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;After the run, the AWS Glue console shows the three output tables in the &lt;code&gt;simple_sdp_demo_db&lt;/code&gt; database. The database’s &lt;strong&gt;Location&lt;/strong&gt; is the warehouse path you configured, &lt;code&gt;s3://amzn-s3-demo-bucket/simple-sdp-demo/warehouse/&lt;/code&gt;, and each table stores its data under that prefix. The following screenshot shows the database properties and the three tables (&lt;code&gt;bronze_orders&lt;/code&gt;, &lt;code&gt;silver_orders&lt;/code&gt;, and &lt;code&gt;gold_sales_summary&lt;/code&gt;), each registered in the AWS Glue Data Catalog.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/03/BDB-6104-5.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/03/BDB-6104-5.png" alt="The bronze_orders, silver_orders, and gold_sales_summary tables in the AWS Glue Data Catalog." width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 5: The three output tables in the AWS Glue Data Catalog.&lt;/p&gt;
&lt;/div&gt;
&lt;h2 id="step-9---query-results"&gt;Step 9 – Query results&lt;/h2&gt;
&lt;p&gt;Query the tables with Amazon Athena. If this is your first time using Athena in this Region, set an Amazon S3 query-results location for your workgroup first (Athena console, &lt;strong&gt;Settings&lt;/strong&gt;). Also make sure your identity can read the &lt;code&gt;simple_sdp_demo_db&lt;/code&gt; tables in the Data Catalog and the underlying S3 data.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-sql"&gt;-- Bronze preserves all 6 source rows
SELECT * FROM simple_sdp_demo_db.bronze_orders ORDER BY order_id;

-- Silver retains the 5 complete orders with positive amounts
SELECT * FROM simple_sdp_demo_db.silver_orders ORDER BY order_id;

-- Gold aggregates by region
SELECT * FROM simple_sdp_demo_db.gold_sales_summary ORDER BY region;&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Expected gold result:&lt;/p&gt;
&lt;table border="1px" width="100%" cellpadding="10px"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;region&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;order_count&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;total_sales&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;average_order_value&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;AMER&lt;/td&gt;
   &lt;td&gt;2&lt;/td&gt;
   &lt;td&gt;1000.00&lt;/td&gt;
   &lt;td&gt;500.00&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;APAC&lt;/td&gt;
   &lt;td&gt;1&lt;/td&gt;
   &lt;td&gt;320.25&lt;/td&gt;
   &lt;td&gt;320.25&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;EMEA&lt;/td&gt;
   &lt;td&gt;2&lt;/td&gt;
   &lt;td&gt;210.50&lt;/td&gt;
   &lt;td&gt;105.25&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;em&gt;Table 3: Gold layer aggregation results by region.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Running the query in the Amazon Athena console returns the aggregated result. The following screenshot shows the gold query and its three result rows (AMER, APAC, and EMEA), matching the values in the preceding table.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/03/BDB-6104-6.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/03/BDB-6104-6.png" alt="Amazon Athena console showing the gold query and its AMER, APAC, and EMEA result rows." width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 6: The gold table results in the Amazon Athena console.&lt;/p&gt;
&lt;/div&gt;
&lt;h2 id="cost-considerations"&gt;Cost considerations&lt;/h2&gt;
&lt;p&gt;AWS Glue 6.0 bills ETL jobs by the data processing unit (DPU)-hour, per second, with a 1-minute minimum per run. AWS Glue 6.0 is also priced 30 percent lower per DPU-hour than AWS Glue 5.1, with no change to your workload, so the same job costs less to run on 6.0. This walkthrough runs on 2 G.1X workers (2 DPUs), reads a 6-row CSV, and completes each run in about 2 minutes. It produces three tables in one AWS Glue Data Catalog database.&lt;/p&gt;
&lt;p&gt;To estimate the cost of a run, multiply the 2 DPUs by the run time in hours by your Region’s AWS Glue 6.0 DPU-hour rate. You can find that rate on the &lt;a href="https://aws.amazon.com/glue/pricing/" target="_blank" rel="noopener"&gt;AWS Glue pricing page&lt;/a&gt;, and rates differ by AWS Region. The Amazon S3 objects created are the 6-row CSV, the pipeline zip, and the three tables’ data. To stop further charges, delete the resources when you finish, as shown in the next step.&lt;/p&gt;
&lt;h2 id="step-10---clean-up"&gt;Step 10 – Clean up&lt;/h2&gt;
&lt;p&gt;To avoid ongoing charges, delete the resources you created:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;# Delete the AWS Glue job
aws glue delete-job --job-name "${JOB_NAME}" --region "${AWS_REGION}"

# Delete the Data Catalog database and its table metadata
aws glue delete-database --name "${DATABASE}" --region "${AWS_REGION}"

# Remove the S3 objects
aws s3 rm "s3://${BUCKET}/${PREFIX}/" --recursive --region "${AWS_REGION}"&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h2 id="whats-next"&gt;What’s next&lt;/h2&gt;
&lt;p&gt;You now have a single pipeline that turns raw order records into validated, aggregated analytics tables, without writing orchestration logic. From here you can:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;Extend&lt;/strong&gt;: Add transformation stages (additional &lt;code&gt;@dp.materialized_view&lt;/code&gt; functions) and connect them by referencing upstream tables. The pipeline picks up the new dependency automatically.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Scale&lt;/strong&gt;: This walkthrough uses materialized views throughout, so every layer fully recomputes on each run (materialized views don’t support incremental refresh). To process only new data as it arrives, convert the bronze layer to a streaming table, which maintains state across runs with checkpoints. For that cross-run state to persist, a streaming table’s data and checkpoint state must not be stored locally. Hive or AWS Glue managed tables require the database’s &lt;code&gt;LocationUri&lt;/code&gt; to point to an Amazon S3 path, while &lt;a href="https://iceberg.apache.org/" target="_blank" rel="noopener"&gt;Apache Iceberg&lt;/a&gt; tables manage their table metadata themselves.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Govern&lt;/strong&gt;: Protect the Data Catalog tables SDP produces with &lt;a href="https://docs.aws.amazon.com/glue/latest/dg/security-lf-enable.html" target="_blank" rel="noopener"&gt;AWS Lake Formation fine-grained access control&lt;/a&gt;. It enforces table-, row-, column-, and cell-level permissions on read queries in AWS Glue Spark jobs (Glue 5.0 and later, for Hive and Iceberg tables). Because this enforcement covers batch reads, it applies to SDP’s materialized views but not to streaming tables, which read through Spark Structured Streaming.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Automate&lt;/strong&gt;: Store the pipeline project in source control. Have your continuous integration and continuous delivery (CI/CD) pipeline package and upload it to Amazon S3 so each job run maps to a known build. Version the zip by object key, or upload the unzipped project to an S3 prefix and turn on Amazon S3 bucket versioning.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Monitor&lt;/strong&gt;: Use Amazon CloudWatch metrics and AWS Glue job run insights for pipeline observability, latency tracking, and failure alerting.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;In this post, you used Spark Declarative Pipelines, the declarative alternative to explicitly orchestrated ETL, now available in AWS Glue 6.0. Two decorated Python functions and one SQL file define the bronze, silver, and gold datasets, and SDP resolves the dependencies and manages execution order for you.&lt;/p&gt;
&lt;p&gt;With SDP, you declare what each dataset should contain and the declarative framework handles ordering and execution. A three-layer pipeline that would otherwise need separate transform and orchestration logic runs as one job that you can ship and maintain.&lt;/p&gt;
&lt;p&gt;To get started, open the &lt;a href="https://console.aws.amazon.com/glue/home" target="_blank" rel="noopener"&gt;AWS Glue console&lt;/a&gt; and build the walkthrough pipeline, or adapt the pattern to your own bronze, silver, and gold datasets. For the full set of features, see the &lt;a href="https://aws.amazon.com/blogs/aws/aws-glue-6-0-now-available-with-30-lower-price-and-full-apache-iceberg-v3-support/" target="_blank" rel="noopener"&gt;AWS Glue 6.0 launch announcement&lt;/a&gt;. To move existing jobs to the Spark 4.1 runtime, see &lt;a href="https://aws.amazon.com/blogs/big-data/upgrade-aws-glue-jobs-to-glue-6-0-with-ai-powered-spark-upgrades/" target="_blank" rel="noopener"&gt;Upgrade AWS Glue jobs to AWS Glue 6.0 with AI-powered Spark upgrades&lt;/a&gt;. For job configuration details, see the &lt;a href="https://docs.aws.amazon.com/glue/latest/dg/" target="_blank" rel="noopener"&gt;AWS Glue Developer Guide&lt;/a&gt;.&lt;/p&gt;
&lt;p style="clear: both"&gt;&lt;/p&gt;
&lt;hr style="width: 100%"&gt;
&lt;h2&gt;About the authors&lt;/h2&gt;
&lt;footer&gt;
 &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
  &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
   &lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/03/BDB-6104-7.jpg" alt="Syed Humair" width="100" height="133"&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Syed Humair&lt;/h3&gt;
  &lt;p style="overflow: hidden"&gt;Syed is a Senior Analytics Specialist Solutions Architect at Amazon Web Services, based in Dubai. He has nearly 20 years of experience in data strategy, data engineering, AI, and enterprise architecture across industries including financial services, retail, telecom, and healthcare. At AWS, he works with enterprise customers to build AI-ready data foundations, from lakehouse architectures and open data formats to real-time analytics and data governance. He is the co-author of the AWS Certified Data Engineer Study Guide (Wiley, 2025).&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
  &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
   &lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/03/BDB-6104-8.jpg" alt="Shrey Malpani" width="100" height="133"&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Shrey Malpani&lt;/h3&gt;
  &lt;p style="overflow: hidden"&gt;Shrey is a Senior Product Manager Technical at Amazon Web Services (AWS), where he works at the intersection of distributed data processing and data integration. He helps customers build AI-ready data platforms for analytics and machine learning. His focus is scaling data integration and data management across services like AWS Glue, Amazon EMR, and Amazon Redshift.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
  &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
   &lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/03/BDB-6104-9.jpg" alt="Bo Li" width="100" height="133"&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Bo Li&lt;/h3&gt;
  &lt;p style="overflow: hidden"&gt;Bo is a Senior Software Development Engineer on the AWS Glue team. He is devoted to designing and building end-to-end solutions to address customers’ data analytic and processing needs with cloud-based, data-intensive and generative AI technologies.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
  &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
   &lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/03/BDB-6104-10.jpg" alt="Kartik Panjabi" width="100" height="133"&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Kartik Panjabi&lt;/h3&gt;
  &lt;p style="overflow: hidden"&gt;Kartik is a Software Development Manager on the AWS Glue team. His team builds generative AI features for data integration and distributed systems for data integration.&lt;/p&gt;
 &lt;/div&gt;
&lt;/footer&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>AWS recognized as a Leader in the 2026 Gartner Magic Quadrant for Strategic Cloud Platform Services for the 16th consecutive year</title>
		<link>https://aws.amazon.com/blogs/big-data/aws-recognized-as-a-leader-in-the-2026-gartner-magic-quadrant-for-strategic-cloud-platform-services-for-the-16th-consecutive-year/</link>
					
		
		<dc:creator><![CDATA[Erika Ehrli]]></dc:creator>
		<pubDate>Tue, 08 Sep 2026 16:46:49 +0000</pubDate>
				<category><![CDATA[Thought Leadership]]></category>
		<guid isPermaLink="false">37151ac424ba52cd649c40924134a74925bb36b6</guid>

					<description>On September 1, Gartner published its Magic Quadrant for Strategic Cloud Platform Services (SCPS). Amazon Web Services (AWS) is the longest-running Leader in this Magic Quadrant, with Gartner naming AWS a Leader for the sixteenth consecutive year. In the report, Gartner once again placed AWS highest on the Ability to Execute axis. We believe this reflects our commitment to help customers innovate faster, operate more securely, and build at any scale, particularly as agentic AI drives the need for a data foundation that is production-ready.</description>
										<content:encoded>&lt;p&gt;On September 1, Gartner published its Magic Quadrant for Strategic Cloud Platform Services (SCPS). Amazon Web Services (AWS) is the longest-running Leader in this Magic Quadrant, with Gartner naming AWS a Leader for the sixteenth consecutive year.&lt;/p&gt;
&lt;p&gt;In the report, Gartner once again placed AWS highest on the Ability to Execute axis. We believe this reflects our commitment to help customers innovate faster, operate more securely, and build at any scale, particularly as agentic AI drives the need for a data foundation that is production-ready.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Here is the graphical representation of the 2026 Magic Quadrant for Strategic Cloud Platform Services.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;img loading="lazy" class="alignleft size-full wp-image-94140" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/08/global-mq-ardm-26-magic-quadrant-for-strategic-cloud-platform-services-graph.png" alt="" width="2893" height="3200"&gt;&lt;/p&gt;
&lt;p&gt;For the full evaluation and methodology, &lt;a href="https://aws.amazon.com/resources/analyst-reports/gartner/global-mq-ardm-26-magic-quadrant-for-strategic-cloud-platform-services/?trk=5f739cc3-c705-4f95-9178-13a84de030dc&amp;amp;sc_channel=el" target="_blank" rel="noopener"&gt;&lt;strong&gt;download the complete 2026 Gartner Magic Quadrant report&lt;/strong&gt;&lt;/a&gt; and read our &lt;a href="https://aws.amazon.com/blogs/migration-and-modernization/aws-named-a-leader-in-gartners-2026-strategic-cloud-platform-services-magic-quadrant-16-years-running/" target="_blank" rel="noopener"&gt;&lt;strong&gt;lead announcement post&lt;/strong&gt;&lt;/a&gt;.&lt;/p&gt;
&lt;h2&gt;&lt;strong&gt;Your AI strategy is only as good as your data strategy&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;Your agents are only as powerful as the data they rely upon. Agents need access to your data and shared context to reason accurately and deliver reliable responses.&lt;/p&gt;
&lt;p&gt;Today the knowledge agents need is scattered across databases, data lakes, warehouses and third-party applications with no shared context or governance. And the scale of the problem is new. Agents generate 10 to 100x more queries than humans. This means your data architecture must be agent-ready from day one. If it isn’t, your AI investments underperform.&lt;/p&gt;
&lt;p&gt;AWS gives your agents an &lt;a href="https://aws.amazon.com/data/" target="_blank" rel="noopener"&gt;open data foundation&lt;/a&gt; with governed context intelligence, built to scale while optimizing the cost of AI. Agentic data capabilities meet industry-specific compliance, security, and schematic requirements so you can move to production with confidence.&lt;/p&gt;
&lt;h2&gt;&lt;strong&gt;An open data architecture for your data and AI&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;Agents need to discover and access your data, wherever it is stored. That’s why AWS delivers an open architecture on Apache Iceberg so agents can use data across these silos. We offer the broadest native Iceberg support of any major cloud provider, with native Iceberg compatibility across every layer of the data stack – ingestion, storage, catalog, and analytics.&lt;/p&gt;
&lt;p&gt;Amazon Simple Storage Service (&lt;a href="https://aws.amazon.com/s3/" target="_blank" rel="noopener"&gt;Amazon S3&lt;/a&gt;) supports Apache Iceberg natively. &lt;a href="https://aws.amazon.com/s3/features/tables/" target="_blank" rel="noopener"&gt;S3 Tables&lt;/a&gt; delivers fully managed Apache Iceberg tables that automate compaction and maintenance as data grows. It works with any Iceberg-compatible engine, from Spark to Redshift, and supports natural language queries through MCP.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://aws.amazon.com/sagemaker/lakehouse/" target="_blank" rel="noopener"&gt;Amazon SageMaker lakehouse architecture&lt;/a&gt; is built with Apache Iceberg. It enables Amazon S3, &lt;a href="https://aws.amazon.com/pm/redshift/" target="_blank" rel="noopener"&gt;Amazon Redshift&lt;/a&gt;, &lt;a href="https://aws.amazon.com/opensearch-service/" target="_blank" rel="noopener"&gt;Amazon OpenSearch Service&lt;/a&gt;, &lt;a href="https://aws.amazon.com/emr/" target="_blank" rel="noopener"&gt;Amazon EMR&lt;/a&gt;, and &lt;a href="https://aws.amazon.com/athena/" target="_blank" rel="noopener"&gt;Amazon Athena&lt;/a&gt; to access the same Iceberg tables through a unified catalog, from a single governance layer. &lt;a href="https://docs.aws.amazon.com/glue/latest/dg/zero-etl-using.html"&gt;Zero-ETL integrations&lt;/a&gt; and federated querying remove remaining barriers across on-premises and third-party cloud sources.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://aws.amazon.com/blogs/aws/the-aws-mcp-server-is-now-generally-available/" target="_blank" rel="noopener"&gt;AWS MCP Server&lt;/a&gt;, part of the &lt;a href="https://aws.amazon.com/products/developer-tools/agent-toolkit-for-aws/" target="_blank" rel="noopener"&gt;Agent Toolkit for AWS&lt;/a&gt;, gives any tool (&lt;a href="https://aws.amazon.com/quick/" target="_blank" rel="noopener"&gt;Amazon Quick&lt;/a&gt;, a third-party agent, or a developer’s IDE) governed access to your data through a single path with inherited permissions. It standardizes tool discovery, authentication, and contextual data access for AI agents interacting with AWS services.&lt;/p&gt;
&lt;p&gt;AWS embraces open standards for flexibility and the best value. This includes PostgreSQL via &lt;a href="https://aws.amazon.com/rds/aurora/" target="_blank" rel="noopener"&gt;Amazon Aurora&lt;/a&gt; and &lt;a href="https://aws.amazon.com/rds/" target="_blank" rel="noopener"&gt;Amazon RDS&lt;/a&gt;, Apache Kafka via &lt;a href="https://aws.amazon.com/msk/" target="_blank" rel="noopener"&gt;Amazon MSK&lt;/a&gt;, OpenSearch via &lt;a href="https://aws.amazon.com/opensearch-service/" target="_blank" rel="noopener"&gt;Amazon OpenSearch Service&lt;/a&gt;, Apache Spark via &lt;a href="https://aws.amazon.com/emr/" target="_blank" rel="noopener"&gt;Amazon EMR&lt;/a&gt; and Trino via &lt;a href="https://aws.amazon.com/athena/" target="_blank" rel="noopener"&gt;Amazon Athena&lt;/a&gt;.&lt;/p&gt;
&lt;h2&gt;&lt;strong&gt;From data to contextual intelligence&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;Agents need more than data access to be accurate. They need contextual understanding of your data and the business rules governing how it should be used before they can make trusted decisions.&lt;/p&gt;
&lt;p&gt;This is why we introduced &lt;a href="https://aws.amazon.com/blogs/machine-learning/context-intelligence-for-your-data-and-ai-agents-at-scale/" target="_blank" rel="noopener"&gt;AWS Context&lt;/a&gt;, a new service that automatically maps the relationships across your existing data into a knowledge graph and provides agentic search so AI agents in the organization can access governed data relationships, business rules, and domain knowledge at runtime.&lt;/p&gt;
&lt;p&gt;For governance, &lt;a href="https://aws.amazon.com/glue/" target="_blank" rel="noopener"&gt;AWS Glue Data Catalog&lt;/a&gt; provides a single catalog for AWS and third-party Iceberg tables, while AWS Lake Formation enforces row-, column-, and cell-level access control so the right data reaches the right agent with the right permissions. AWS Glue Data Quality and SageMaker ML Lineage Tracking add the governance layer that production AI demands.&lt;/p&gt;
&lt;h2&gt;&lt;strong&gt;Foundational excellence at scale&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;Agentic AI workloads require resources that are always available, dynamically allocated, and optimized for price-performance. AWS delivers the most powerful combination of services and capabilities for automatic resource allocation, zero-tuning price performance, and the reliability that millions of customers have trusted for over 20 years.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://aws.amazon.com/products/databases/" target="_blank" rel="noopener"&gt;AWS Databases&lt;/a&gt; offer a high-performance, secure foundation to power agentic AI and data-driven applications at any scale. &lt;a href="https://aws.amazon.com/rds/aurora/" target="_blank" rel="noopener"&gt;Amazon Aurora&lt;/a&gt; delivers unparalleled high performance and availability at global scale for PostgreSQL, MySQL, and DSQL. &lt;a href="https://aws.amazon.com/dynamodb/" target="_blank" rel="noopener"&gt;Amazon DynamoDB&lt;/a&gt; and &lt;a href="https://aws.amazon.com/elasticache/" target="_blank" rel="noopener"&gt;Amazon ElastiCache&lt;/a&gt; serve up to tens of billions of requests per second at microsecond to single-digit millisecond latency at any scale, operating at agent speed. With native vector search built into Aurora PostgreSQL, DynamoDB, and ElastiCache, you can perform vector search — from billions to trillions of vectors — and integrate effortlessly across AWS services to build agentic applications.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://aws.amazon.com/s3/" target="_blank" rel="noopener"&gt;Amazon S3&lt;/a&gt; has evolved to support the demands of AI with purpose-built storage tiers. &lt;a href="https://aws.amazon.com/s3/features/files/" target="_blank" rel="noopener"&gt;S3 Files&lt;/a&gt; gives agents a shared file system directly on S3 data, so an entire agent fleet can read inputs, write outputs, and persist memory with no duplicated data and no new APIs to learn. &lt;a href="https://aws.amazon.com/s3/features/vectors/" target="_blank" rel="noopener"&gt;S3 Vectors&lt;/a&gt; is the first cloud object store with native support to store and query vectors. It cuts the cost of uploading, storing, and querying vector data by up to 90%, making it practical to build the large-scale vector datasets that give AI agents memory, context, and semantic search.&lt;/p&gt;
&lt;p&gt;For search and retrieval, AWS provides purpose-built &lt;a href="https://aws.amazon.com/blogs/machine-learning/aws-vector-solutions-build-agentic-ai-where-your-data-lives/" target="_blank" rel="noopener"&gt;vector engines&lt;/a&gt; that bring intelligent search to your data where it already lives. With &lt;a href="https://aws.amazon.com/opensearch-service/features/serverless/" target="_blank" rel="noopener"&gt;OpenSearch Service Serverless&lt;/a&gt;, your agents take advantage of lexical, vector, hybrid, and agentic search in a single system with high throughput, low latency, and relevant results at scale.&lt;/p&gt;
&lt;h3&gt;&lt;strong&gt;Key takeaways&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;The companies moving fastest with AI are the ones that treated data readiness as strategy from the start. We believe the Gartner recognition of AWS as a Leader for 16 consecutive years reflects the breadth and deepest set of core public cloud services and capabilities, including the data foundation that makes this possible.&lt;/p&gt;
&lt;p&gt;Ready to see the full evaluation? &lt;a href="https://aws.amazon.com/resources/analyst-reports/gartner/global-mq-ardm-26-magic-quadrant-for-strategic-cloud-platform-services/?trk=5f739cc3-c705-4f95-9178-13a84de030dc&amp;amp;sc_channel=el" target="_blank" rel="noopener"&gt;Download the 2026 Gartner Magic Quadrant for Strategic Cloud Platform Services&lt;/a&gt;.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;Gartner does not endorse any company, vendor, product or service depicted in its publications, and does not advise technology users to select only those vendors with the highest ratings or other designation. Gartner publications consist of the opinions of Gartner’s business and technology insights organization and should not be construed as statements of fact. Gartner disclaims all warranties, expressed or implied, with respect to this publication, including any warranties of merchantability or fitness for a particular purpose.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Gartner and Magic Quadrant are trademarks of Gartner, Inc., and/or its affiliates.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;This graphic was published by Gartner, Inc. as part of a larger research document and should be evaluated in the context of the entire document. The Gartner document is available to download: &lt;a href="https://aws.amazon.com/resources/analyst-reports/gartner/global-mq-ardm-26-magic-quadrant-for-strategic-cloud-platform-services/?trk=5f739cc3-c705-4f95-9178-13a84de030dc&amp;amp;sc_channel=el" target="_blank" rel="noopener"&gt;Complete 2026 Gartner Magic Quadrant report&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Gartner, Magic Quadrant for Strategic Cloud Platform Services, By&amp;nbsp;&lt;a href="https://www.gartner.com/analyst/b9c802bf7cae" target="_blank" rel="noopener"&gt;&lt;strong&gt;Alessandro Galimberti&lt;/strong&gt;&lt;/a&gt;,&amp;nbsp;&lt;a href="https://www.gartner.com/analyst/b9c905bd7ca1" target="_blank" rel="noopener"&gt;&lt;strong&gt;Carolin Zhou&lt;/strong&gt;&lt;/a&gt;,&amp;nbsp;&lt;a href="https://www.gartner.com/analyst/bcc807b978" target="_blank" rel="noopener"&gt;&lt;strong&gt;Douglas Toombs&lt;/strong&gt;&lt;/a&gt;,&amp;nbsp;&lt;a href="https://www.gartner.com/analyst/bdc901b97b" target="_blank" rel="noopener"&gt;&lt;strong&gt;Dennis Smith&lt;/strong&gt;&lt;/a&gt;,&amp;nbsp;&lt;a href="https://www.gartner.com/analyst/bcc803b97f" target="_blank" rel="noopener"&gt;&lt;strong&gt;Ed Anderson&lt;/strong&gt;&lt;/a&gt;,&amp;nbsp;&lt;a href="https://www.gartner.com/analyst/b1ca00b87e" target="_blank" rel="noopener"&gt;&lt;strong&gt;Tobi Bet&lt;/strong&gt;&lt;/a&gt;,&amp;nbsp;&lt;a href="https://www.gartner.com/analyst/b9cb05b87faf" target="_blank" rel="noopener"&gt;&lt;strong&gt;Chuck Lawton&lt;/strong&gt;&lt;/a&gt; , 1 September 2026&lt;/em&gt;&lt;/p&gt;
&lt;hr style="width: 100%"&gt;
&lt;h2&gt;About the author&lt;/h2&gt;
&lt;footer&gt;
 &lt;div class="blog-author-box"&gt;
  &lt;div class="blog-author-image"&gt;
   &lt;img loading="lazy" class="alignnone size-thumbnail wp-image-94104" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/03/erika-100x133.jpg" alt="" width="100" height="133"&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Erika Ehrli&lt;/h3&gt;
  &lt;p&gt;&lt;a href="https://www.linkedin.com/in/erikaehrli/" target="_blank" rel="noopener"&gt;Erika&lt;/a&gt; is Head of Product Marketing for Data for AI at Amazon Web Services, where she leads technical product marketing and go-to-market strategy across the AWS analytics, database, and storage portfolios. In this role she helps organizations build AI-ready data foundations for agentic AI and analytics workloads.&lt;/p&gt;
 &lt;/div&gt;
&lt;/footer&gt;</content:encoded>
					
					
			
		
		
			</item>
		<item>
		<title>How Sony LIV built real-time video streaming analytics with AWS</title>
		<link>https://aws.amazon.com/blogs/big-data/how-sony-liv-built-real-time-video-streaming-analytics-with-aws/</link>
		
		<dc:creator><![CDATA[Rahul Sureka]]></dc:creator>
		<pubDate>Tue, 08 Sep 2026 16:43:05 +0000</pubDate>
				<category><![CDATA[Amazon Kinesis]]></category>
		<category><![CDATA[Customer Solutions]]></category>
		<category><![CDATA[Media & Entertainment]]></category>
		<guid isPermaLink="false">0f1624b04e49f41793c895b8f3e582ded7e31e0c</guid>

					<description>Real-time data analytics is transforming how streaming platforms serve their audiences. Learn how Sony LIV built a comprehensive, real-time streaming analytics solution on AWS using Amazon Kinesis Data Streams, Amazon EMR, and Apache Iceberg.</description>
										<content:encoded>&lt;p&gt;&lt;em&gt;This guest post was co-written with Mukund Acharya from Sony LIV.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Real-time data analytics is transforming how streaming applications understand and serve their audiences. In this post, we share how Sony LIV used Amazon Kinesis Data Streams for sub-second processing and AWS services to build a comprehensive streaming analytics solution on AWS. In this post, we share how SonyLIV built a comprehensive streaming analytics solution on AWS using Amazon Kinesis Data Streams for sub-second event ingestion, Amazon Data Firehose for reliable delivery to storage, Amazon EMR with Apache Spark for scalable batch and micro-batch processing, and Apache Iceberg on Amazon S3 for ACID-compliant, queryable data lake tables. Together, these services enable SonyLIV to capture millions of concurrent viewer events, process them cost-efficiently at scale, and surface actionable insights from real-time engagement metrics to historical trend analysis — all within a fully managed, serverless-friendly architecture.&lt;/p&gt;
&lt;h2 id="about-sony-liv"&gt;About Sony LIV&lt;/h2&gt;
&lt;p&gt;Sony LIV, operated by Sony Pictures Networks India, is one of India’s leading over-the-top (OTT) streaming applications offering premium content including live sports, original series, movies, and TV shows to millions of users across mobile, web, and Smart TV devices.&lt;/p&gt;
&lt;p&gt;As part of the latest game broadcast rights for Asia Cup, Sony LIV built a streaming analytics solution using AWS services that successfully processed millions of concurrent sessions in real-time.&lt;/p&gt;
&lt;h2 id="the-challenge-of-batch-processing"&gt;The challenge of batch processing&lt;/h2&gt;
&lt;p&gt;As Sony LIV’s audience grew and live sports events attracted increasingly large viewership, the team identified an opportunity to move from batch-based analytics to a real-time data application. The existing architecture processed data reliably in scheduled batches, but the growing scale of live events called for faster, more granular insights. The team set out to address three key areas:&lt;/p&gt;
&lt;ol&gt;
 &lt;li&gt;&lt;strong&gt;Real-Time Visibility:&lt;/strong&gt; Track frontend user events, journeys, and conversion funnels in near real-time, especially during high-stakes live sports streams, to see peak concurrent viewership and respond to engagement patterns as they happen.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Comprehensive Quality Monitoring: &lt;/strong&gt;Implement end-to-end monitoring of video quality key performance indicators (KPIs) including buffering rates, playback failures, video start time, and rebuffering events to proactively enhance the viewer experience.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Unified Data Architecture: &lt;/strong&gt;Consolidate data from multiple sources into a unified architecture to enable comprehensive customer views, supporting future personalization and machine learning (ML)-driven recommendations.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Sony LIV needed to evolve from reactive batch processing to a real-time data application to deliver actionable engagement insights at scale and establish the foundation for unified customer profiles.&lt;/p&gt;
&lt;h2 id="streaming-analytics-architecture"&gt;Streaming analytics architecture&lt;/h2&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/26/BDB-5907-1.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/26/BDB-5907-1.png" alt="Streaming analytics architecture showing data from mobile, web, and Smart TV apps flowing through Amazon EKS into real-time and batch processing paths on AWS" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 1: Sony LIV streaming analytics architecture on AWS&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;Data from mobile, web, and Smart TV applications flows through Application Load Balancers to &lt;a href="https://aws.amazon.com/eks/"&gt;Amazon Elastic Kubernetes Service&lt;/a&gt; (Amazon EKS) pods for validation and preprocessing. The architecture implements parallel processing paths to balance speed and cost efficiency.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Real Time Path: &lt;/strong&gt;Priority events requiring immediate action such as playback failures, user interactions, and live sports engagement signals stream through &lt;a href="https://aws.amazon.com/kinesis/data-streams/"&gt;Amazon Kinesis Data Streams&lt;/a&gt; for sub-second processing. These events flow directly to ClickHouse using the ClickHouse connector, enabling real-time analytics with minimal latency.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Batch Path: &lt;/strong&gt;High-volume batch events are ingested via &lt;a href="https://aws.amazon.com/firehose/"&gt;Amazon Data Firehose&lt;/a&gt; into a raw landing zone on &lt;a href="https://aws.amazon.com/s3/"&gt;Amazon Simple Storage Service (Amazon S3)&lt;/a&gt;, partitioned by event type and time. &lt;a href="https://aws.amazon.com/emr/"&gt;Amazon EMR&lt;/a&gt; running Apache Spark then processes this raw data performing schema validation, deduplication, and transformations — and writes the curated output as &lt;a href="https://iceberg.apache.org/"&gt;Apache Iceberg&lt;/a&gt; tables back to S3&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Unified Data Layer: &lt;/strong&gt;AWS Glue catalogs data across streaming and batch sources, creating a unified schema that powers the customer data application. This unified data layer serves as the foundation for building comprehensive customer profiles and enabling personalized recommendations.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Analytics and Monitoring:&lt;/strong&gt; ClickHouse serves as the low-latency analytics engine, powering real-time dashboards and video quality monitoring. Operations teams track critical KPIs including peak concurrent viewership, buffering rates, video start time, and playback failures, enabling rapid response to issues during live events.&lt;/p&gt;
&lt;p&gt;Custom dashboards and Datadog provide visualization for business metrics and application performance monitoring, while Amazon CloudWatch tracks infrastructure health, delivering end-to-end visibility across the application.&lt;/p&gt;
&lt;h2 id="results-and-business-impact"&gt;Results and business impact&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;The results:&lt;/strong&gt; product, analytics and engineering teams now operate from a single unified dashboard, making real-time decisions with data that is seconds, not hours old. Auto scaling and serverless design have optimized costs while providing reliability at peak load. At the foundation, a centralized data catalog now serves as a single source of truth for business intelligence, machine learning workloads and operational monitoring.&lt;/p&gt;
&lt;p&gt;The transformation helped Sony LIV process millions of concurrent sessions in real time, particularly during high-profile events such as the Asia Cup 2025.&lt;/p&gt;
&lt;ol&gt;
 &lt;li&gt;&lt;strong&gt;Accelerated insights&lt;/strong&gt;: Data latency dropped from hours to seconds, enabling instant insights into viewer behavior, content performance, and streaming health.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Unified decision-making&lt;/strong&gt;: Product, analytics, and engineering teams now use unified dashboards for real-time decision-making.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Optimized operations&lt;/strong&gt;: The solution auto scaling and serverless design optimized costs while supporting reliability and fault tolerance.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Beyond performance gains, the new architecture established a strong foundation for personalized content recommendation and operational agility. By integrating a centralized data catalog and scalable analytics layer, the solution now provides a single source of truth for business intelligence, machine learning workloads, and monitoring.&lt;/p&gt;
&lt;h2 id="future-innovations"&gt;Future innovations&lt;/h2&gt;
&lt;p&gt;Sony LIV plans to expand its analytics capabilities by integrating advanced ML models for real-time recommendations, user churn prediction, and anomaly detection. The team will also focus on building unified customer profiles to enable hyper-personalized experiences.&lt;/p&gt;
&lt;p&gt;The AWS based architecture provides a strong foundation for future growth and innovation, enabling Sony LIV to deliver highly personalized experiences to an expanding viewer base.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;By adopting AWS, Sony LIV transformed its analytics architecture into a real-time, scalable, and insight-driven solution. The solution reduced data latency from hours to seconds, enabled processing of millions of concurrent sessions during peak events, and positioned Sony LIV as a leader in streaming analytics innovation in India.&lt;/p&gt;
&lt;p&gt;To learn more about how AWS can help your media organization implement real-time streaming analytics solutions, explore the following resources:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;a href="https://aws.amazon.com/media/" target="_blank" rel="noopener"&gt;AWS for Media &amp;amp; Entertainment&lt;/a&gt; – Discover how AWS powers content delivery, streaming, and analytics for media companies worldwide.&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://aws.amazon.com/kinesis/" target="_blank" rel="noopener"&gt;Amazon Kinesis Data Streams&lt;/a&gt; – Learn about the fully managed service for real-time data streaming used in this solution.&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://aws.amazon.com/blogs/media/category/analytics/" target="_blank" rel="noopener"&gt;Data Streaming Analytics on AWS (AWS Blog)&lt;/a&gt; – Read related posts on how customers are building streaming analytics pipelines on AWS.&lt;/li&gt;
&lt;/ul&gt;
&lt;p style="clear: both"&gt;&lt;/p&gt;
&lt;hr style="width: 100%"&gt;
&lt;h2&gt;About the authors&lt;/h2&gt;
&lt;footer&gt;
 &lt;div class="blog-author-box"&gt;
  &lt;div class="blog-author-image"&gt;
   &lt;p&gt;&lt;img loading="lazy" class="alignleft size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/26/BDB-5907-2.jpeg" alt="Rahul Sureka" width="100" height="100"&gt;&lt;/p&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Rahul Sureka&lt;/h3&gt;
  &lt;p&gt;&lt;a href="https://www.linkedin.com/in/rrsureka/" target="_blank" rel="noopener"&gt;Rahul&lt;/a&gt; is an Enterprise Solutions Architect at AWS, helping media and cross-industry customers design scalable, real-time streaming architectures on the cloud. With over 25 years of experience architecting and leading large-scale business transformation programs, Rahul specializes in streaming applications, data and analytics, AI/ML, and Agentic AI.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box"&gt;
  &lt;div class="blog-author-image"&gt;
   &lt;p&gt;&lt;img loading="lazy" class="alignleft size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/26/BDB-5907-3.jpeg" alt="Umesh Chaudhari" width="100" height="100"&gt;&lt;/p&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Umesh Chaudhari&lt;/h3&gt;
  &lt;p&gt;&lt;a href="https://www.linkedin.com/in/utc/" target="_blank" rel="noopener"&gt;Umesh&lt;/a&gt; is a Sr. Analytics Solutions Architect at AWS, helping Energy, Industrial, and Semiconductor customers shape their data and AI strategies. With over a decade of experience in distributed data systems and analytics Services, he turns complex data challenges into scalable, business-driving solutions. A recognized public speaker, he’s passionate about enabling innovation through modern, cloud-native architecture.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box"&gt;
  &lt;div class="blog-author-image"&gt;
   &lt;img loading="lazy" class="alignleft size-full wp-image-94080" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/03/mahesh-photo1.jpg" alt="" width="100" height="134"&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Maheshwaran G&lt;/h3&gt;
  &lt;p&gt;&lt;a href="https://www.linkedin.com/in/maheshwarang/" target="_blank" rel="noopener"&gt;Maheshwaran&lt;/a&gt; is a Principal Solution Architect in Media and Entertainment, dedicated to empowering Indian and SAARC media organizations to accelerate growth through cutting-edge cloud technologies that reimagine workflows, enhance scalability, and unlock new business opportunities — anchored in innovation with 17 granted patents spanning USPTO and IPO across diversified fields.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box"&gt;
  &lt;div class="blog-author-image"&gt;
   &lt;p&gt;&lt;img loading="lazy" class="alignleft size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/26/BDB-5907-5.jpeg" alt="Varsha Palepu" width="100" height="100"&gt;&lt;/p&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Varsha Palepu&lt;/h3&gt;
  &lt;p&gt;&lt;a href="https://www.linkedin.com/in/varsha-palepu/" target="_blank" rel="noopener"&gt;Varsha&lt;/a&gt; is a Solutions Architect at AWS and an analytics specialist on the AWS streaming team. She helps small and medium businesses innovate on AWS and creates technical streaming content to empower customers in their cloud journey.&lt;/p&gt;
 &lt;/div&gt;
&lt;/footer&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>From silos to insights: Federated data access patterns for AI agents</title>
		<link>https://aws.amazon.com/blogs/big-data/from-silos-to-insights-federated-data-access-patterns-for-ai-agents/</link>
		
		<dc:creator><![CDATA[James Wu]]></dc:creator>
		<pubDate>Fri, 04 Sep 2026 21:35:21 +0000</pubDate>
				<category><![CDATA[Advanced (300)]]></category>
		<category><![CDATA[Amazon Bedrock]]></category>
		<category><![CDATA[Technical How-to]]></category>
		<guid isPermaLink="false">7a4f3bacd2b1f0fcf76451c9d43e764b33bcd091</guid>

					<description>AI agents can reach enterprise data where it lives instead of routing every question through data engineers. This post presents three reference patterns for federated data access using Model Context Protocol (MCP) servers and Amazon Bedrock AgentCore: catalog-first, direct source, and hybrid access.</description>
										<content:encoded>&lt;p&gt;Enterprise data today is scattered across specialized systems, each with its own tools and expertise. Querying a database requires SQL. Accessing batch data on &lt;a href="https://aws.amazon.com/s3/" target="_blank" rel="noopener"&gt;Amazon Simple Storage Service (Amazon S3)&lt;/a&gt; requires compute engines such as &lt;a href="https://aws.amazon.com/athena/" target="_blank" rel="noopener"&gt;Amazon Athena&lt;/a&gt; and Trino. Consuming real-time streams from &lt;a href="https://aws.amazon.com/pm/kinesis/?trk=a3892bbf-ef62-45de-806e-8005a4c527b1&amp;amp;sc_channel=ps&amp;amp;ef_id=Cj0KCQjw2OnUBhC2ARIsACKyfaEyrI2sbnYx23H2oX1jJuYrd6FLICWBkJy6yGrjI-VvVS35yYTJgHsaAt8XEALw_wcB:G:s&amp;amp;gads_camp=23522747487&amp;amp;gads_ag=196433734047&amp;amp;gads_ad=795876995210&amp;amp;gads_kw=amazon%20kinesis&amp;amp;gads_matchtype=e&amp;amp;gads_network=g&amp;amp;gads_device=c&amp;amp;gads_geo=9032925&amp;amp;gad_campaignid=23522747487&amp;amp;gbraid=0AAAAADjHtp-V62uvc-Zq0dL718N9tnkEn&amp;amp;gclid=Cj0KCQjw2OnUBhC2ARIsACKyfaEyrI2sbnYx23H2oX1jJuYrd6FLICWBkJy6yGrjI-VvVS35yYTJgHsaAt8XEALw_wcB" target="_blank" rel="noopener"&gt;Amazon Kinesis&lt;/a&gt; requires streaming expertise. Each software as a service (SaaS) application has its own API, authentication model, and query language. Today, only data engineers can navigate this landscape, and business users file tickets, wait for reports, or rely on dashboards that answer yesterday’s questions. When a leader needs a one-time answer spanning multiple systems, they’re back in the ticket queue.&lt;/p&gt;
&lt;p&gt;Consider a streaming media company: customer profiles, content catalogs, and ad campaign performance are stored as batch data on Amazon S3. Viewership telemetry such as device type, stream quality, watch duration, and buffering events flows in real time through Amazon Kinesis. Subscriber management and support tickets live in a relational customer relationship management (CRM) database. Leaders routinely ask questions like:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;Which titles drove the most subscriber growth last quarter?&lt;/li&gt;
 &lt;li&gt;How does marketing spend correlate with viewing completion rates?&lt;/li&gt;
 &lt;li&gt;Is churn spiking among users who haven’t engaged with new content?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Answering these questions faces two challenges:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The data silo problem.&lt;/strong&gt; The data lives in multiple places with batch stores on S3, real-time streams in Kinesis, and an online transaction processing (OLTP) database, each with its own access patterns, query language, and authentication model. Organizations traditionally solve this by building data lakes or adopting a data mesh, but both require significant data engineering investment and ongoing maintenance.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The access gap.&lt;/strong&gt; The expertise to navigate the enterprise systems is concentrated in the hands of few data engineers, creating a bottleneck that no dashboard or business intelligence (BI) tool fully resolves. Every new one-time requirement means more engineering work, and it’s not self-service.&lt;/p&gt;
&lt;p&gt;A fundamentally different approach is emerging: instead of moving all data to one place or building bespoke integrations for each source, let AI agents talk directly to the systems where data lives. Model Context Protocol (MCP) makes this possible, an open protocol that standardizes how AI applications connect to external data sources and tools. MCP servers wrap diverse systems behind a uniform interface for tool discovery, invocation, and response handling. Any user can ask a question in natural language and the agent reaches the right data without knowing which system holds it, what API to use, or what query language is required.&lt;/p&gt;
&lt;p&gt;In this post, we propose reference architectures for accessing data stored in different systems and datastores using MCP and &lt;a href="https://aws.amazon.com/bedrock/agentcore/" target="_blank" rel="noopener"&gt;Amazon Bedrock AgentCore&lt;/a&gt;. The patterns apply to enterprises with mixed data sources, but we ground the narrative in our streaming media company example described earlier to make the problem concrete.&lt;/p&gt;
&lt;h2 id="solution-overview-our-approach"&gt;Solution overview&lt;/h2&gt;
&lt;p&gt;Our solution is a federated data foundation for a streaming media company. It supports real-time and batch analytics using MCP servers and Amazon Bedrock AgentCore, and it makes analytics accessible across the organization. The following reference architecture shows the complete picture from data ingestion through governance and compute layers to the generative AI layer where agents orchestrate across MCP servers. The demo uses synthetic data: batch datasets are generated with Python scripts, and streaming telemetry is produced by AWS Lambda. The complete source code is available in the accompanying GitHub repository, so you can deploy and try it yourself.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/04/BDB-5604-1.jpeg" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/04/BDB-5604-1.jpeg" alt="Reference architecture showing data ingestion, governance and compute layers, and the generative AI layer where agents orchestrate across MCP servers" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 1: Reference architecture for federated data access across batch, streaming, and relational sources&lt;/p&gt;
&lt;/div&gt;
&lt;h2 id="walkthrough"&gt;Walkthrough&lt;/h2&gt;
&lt;p&gt;This section covers the prerequisites and then walks through how a user request flows end to end through the reference architecture.&lt;/p&gt;
&lt;h3 id="prerequisites"&gt;Prerequisites&lt;/h3&gt;
&lt;ul&gt;
 &lt;li&gt;AWS account with access to &lt;a href="https://aws.amazon.com/bedrock/agentcore/?trk=f7596381-283b-40e2-90fe-22e22b18daf5&amp;amp;sc_channel=ps&amp;amp;ef_id=EAIaIQobChMIucuXyvPSlgMVsklHAR1iczz3EAAYASAAEgJGbvD_BwE&amp;amp;gads_camp=23527793912&amp;amp;gads_ag=197228065909&amp;amp;gads_ad=813189298265&amp;amp;gads_kw=agentcore&amp;amp;gads_matchtype=p&amp;amp;gads_network=g&amp;amp;gads_device=c&amp;amp;gads_geo=9003495&amp;amp;gad_campaignid=23527793912&amp;amp;gbraid=0AAAAADjHtp_Wn_Y6DYQbtRS27y3_2Ch4v&amp;amp;gclid=EAIaIQobChMIucuXyvPSlgMVsklHAR1iczz3EAAYASAAEgJGbvD_BwE" target="_blank" rel="noopener"&gt;Amazon Bedrock AgentCore&lt;/a&gt; and other AWS services. For more information, review &lt;a href="https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/runtime-permissions.html" target="_blank" rel="noopener"&gt;Permissions for AgentCore Runtime&lt;/a&gt; documentation.&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://aws.amazon.com/cli/" target="_blank" rel="noopener"&gt;AWS Command Line Interface (AWS CLI)&lt;/a&gt; configured.&lt;/li&gt;
 &lt;li&gt;Python 3.10 or later.&lt;/li&gt;
 &lt;li&gt;Docker or Finch installed.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="request-flow"&gt;Request flow&lt;/h3&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;&lt;strong&gt;User request:&lt;/strong&gt; A user submits a natural-language question through a React application served by Amazon CloudFront with static assets on Amazon S3.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Authentication:&lt;/strong&gt; Amazon Cognito authenticates the user and issues an identity token that travels with the request to the agent layer.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Agent orchestration:&lt;/strong&gt; The request reaches a Strands agent running on AgentCore runtime, a capability of Amazon Bedrock AgentCore. The agent reasons over the question and determines which data sources to query.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Gateway routing:&lt;/strong&gt; Amazon Bedrock AgentCore Gateway, a capability of Amazon Bedrock AgentCore, aggregates all three MCP servers behind a single endpoint, handling tool discovery, authentication, and routing.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;MCP server execution:&lt;/strong&gt; The agent routes the query to the appropriate MCP server(s), each running on Amazon Bedrock AgentCore runtime behind Amazon Bedrock AgentCore Gateway. The Data Processing MCP server queries AWS Glue Data Catalog and Amazon Athena for batch and streaming data on S3, the Amazon Aurora MCP server translates tool calls into SQL against the Amazon Aurora MySQL CRM database, and the AWS Documentation MCP server provides AWS service context.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Data sources:&lt;/strong&gt; The architecture deliberately spans multiple storage systems to reflect how enterprise data is typically fragmented across teams and technologies. Batch data (customer profiles, content titles, and ad campaigns) is generated by &lt;a href="https://aws.amazon.com/pm/lambda/?trk=2abe6167-e3db-40c4-a9fa-b283e7b4d7c8&amp;amp;sc_channel=ps&amp;amp;ef_id=Cj0KCQjw2OnUBhC2ARIsACKyfaEZYDVIosg2KZubmtmzq1BbkZbsmHk3VmS_yyj-eR3U86UAsAupfM8aAu9jEALw_wcB:G:s&amp;amp;gads_camp=23527793912&amp;amp;gads_ag=191938386622&amp;amp;gads_ad=802094701896&amp;amp;gads_kw=amazon%20lambda&amp;amp;gads_matchtype=e&amp;amp;gads_network=g&amp;amp;gads_device=c&amp;amp;gads_geo=9032925&amp;amp;gad_campaignid=23527793912&amp;amp;gbraid=0AAAAADjHtp8XHpQ1kaz37duFxmWQLTCuA&amp;amp;gclid=Cj0KCQjw2OnUBhC2ARIsACKyfaEZYDVIosg2KZubmtmzq1BbkZbsmHk3VmS_yyj-eR3U86UAsAupfM8aAu9jEALw_wcB" target="_blank" rel="noopener"&gt;AWS Lambda&lt;/a&gt; on an &lt;a href="https://aws.amazon.com/eventbridge/" target="_blank" rel="noopener"&gt;Amazon EventBridge&lt;/a&gt; schedule and lands as Parquet files on Amazon S3. Streaming viewership telemetry (what users watch, when they pause, where they drop off) flows through Amazon Kinesis Data Streams and Amazon Data Firehose to S3. CRM records (subscriber plans, support tickets, account status) live in an Amazon Aurora MySQL database. &lt;a href="https://docs.aws.amazon.com/prescriptive-guidance/latest/serverless-etl-aws-glue/aws-glue-data-catalog.html" target="_blank" rel="noopener"&gt;AWS Glue Data Catalog&lt;/a&gt; registers the S3-based sources under a unified metadata layer, and AWS Lake Formation enforces fine-grained access policies across the catalog. This mix of batch, streaming, and relational sources is what makes federated access essential. No single query engine can reach all datasets natively.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Response:&lt;/strong&gt; Results flow back through Amazon Bedrock AgentCore Gateway to the agent, which composes a natural-language answer and delivers it to the user through the front end.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;For deploying our reference architecture, follow the instructions in the &lt;a href="https://github.com/aws-samples/sample-aws-semantic-data-access-ai-agents" target="_blank" rel="noopener"&gt;code repository&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="design-patterns-for-federated-data-access"&gt;Design patterns for federated data access&lt;/h2&gt;
&lt;p&gt;Within our architecture, we propose three design patterns for federated data access, each on a spectrum between centralized governance and direct access flexibility.&lt;/p&gt;
&lt;h3 id="pattern-1-catalog-first-access"&gt;Pattern 1: Catalog-first access&lt;/h3&gt;
&lt;p&gt;AWS Glue Data Catalog registers all S3 sources under a unified metadata layer: schemas, business context, data quality metrics, and lineage. The AWS Data Processing MCP server, hosted on Amazon Bedrock AgentCore runtime, wraps AWS Glue Catalog metadata and Amazon Athena query capabilities behind standard MCP tool calls. So when a user asks &lt;em&gt;“Which ad campaigns drove the most subscriber activations last quarter?”&lt;/em&gt;, the agent discovers tables through catalog tools and resolves business terms from column metadata. It then executes the join through Athena without ever calling a Glue API directly.&lt;/p&gt;
&lt;p&gt;The following diagram traces how a single user request flows through the federated data access architecture: from the agent, through the MCP server, and down to the data in Amazon S3.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/04/BDB-5604-2.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/04/BDB-5604-2.png" alt="Request flow for the catalog-first access pattern, from the agent through the MCP server to data in Amazon S3" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 2: Request flow for the catalog-first access pattern&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;Internally, our agent built using Strands Agent framework has three components: a system prompt, a large language model (LLM), and a set of MCP tools. We use Claude Haiku 4.5 powered by Amazon Bedrock as the foundation LLM with tools discovered through the Amazon Bedrock AgentCore Gateway. The &lt;strong&gt;system prompt&lt;/strong&gt; teaches the agent how to use those tools not by listing every column in every table, but by providing intent-based routing rules and a mandatory schema discovery workflow. Here’s an extract from the system prompt:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-text"&gt;TOOL DISCOVERY &amp;amp; ROUTING:

You access tools via the MCP Gateway. Use x_amz_bedrock_agentcore_search
to find the right tool by keyword when unsure.

Routing by intent:
- Telemetry/streaming/viewing data → Glue catalog tools, then Athena query tools
- CRM/support tickets/ratings → MySQL tools (run_query, get_table_schema)
- AWS service questions → documentation search tools

SCHEMA DISCOVERY (MANDATORY before writing SQL):

Before writing any Athena query, retrieve the table schema:
→ Use manage_aws_glue_tables with operation='get-table',
database_name='acme_telemetry', table_name='&amp;lt;table&amp;gt;'

This returns all columns, data types, partition keys, and storage details.&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;To see this in action, consider what happens when a user asks “How many streaming events in February 2026 by event type?”:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;The agent’s routing rules match “streaming events” to the AWS Glue Catalog and Athena query path. If unsure which tool to use, the Gateway’s semantic search discovers tools by keyword rather than requiring exact names.&lt;/li&gt;
 &lt;li&gt;The agent calls &lt;code&gt;manage_aws_glue_tables&lt;/code&gt; exposed by the Data Processing MCP server to retrieve the full schema: column names and types, partition keys (year, month, day, hour), and storage format.&lt;/li&gt;
 &lt;li&gt;With the schema in hand, the agent writes Presto/Trino SQL with partition filters (&lt;code&gt;WHERE year='2026' AND month='02'&lt;/code&gt;).&lt;/li&gt;
 &lt;li&gt;The agent executes the query, retrieves results, and composes a natural-language answer. The user never sees SQL, Glue APIs, or partition strategies.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This discover-then-query workflow is what makes the pattern self-service. The Amazon Bedrock AgentCore Gateway provides unified tool discovery as new MCP servers appear without updating routing logic. The AWS Glue Data Catalog provides a live metadata layer for new tables and columns to appear immediately.&lt;/p&gt;
&lt;p&gt;This pattern isn’t unique to AWS. Other platforms adopt the same model. For example, Databricks offers managed MCP servers for Unity Catalog, letting agents discover and query governed datasets, AI models, and functions registered in Unity Catalog. The common trade-off across all of them: all data must be cataloged before agents can access it, which can bottleneck rapidly changing environments.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/04/BDB-5604-3.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/04/BDB-5604-3.png" alt="Catalog-first access where the agent uses AWS Glue Data Catalog and Amazon Athena to query governed data on Amazon S3" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 3: Catalog-first access with AWS Glue Data Catalog and Amazon Athena&lt;/p&gt;
&lt;/div&gt;
&lt;h3 id="pattern-2-direct-source-access"&gt;Pattern 2: Direct source access&lt;/h3&gt;
&lt;p&gt;Agents access source systems directly through dedicated MCP servers (no intermediate catalog). The Aurora MCP server, hosted on Amazon Bedrock AgentCore runtime, queries the Amazon Aurora CRM database directly. Therefore, a question like &lt;em&gt;“How many open support tickets from premium subscribers?”&lt;/em&gt; routes to the MCP server, which translates the tool call into SQL against Aurora. The agent never constructs a database connection or manages credentials. The MCP server handles authentication through &lt;a href="https://aws.amazon.com/secrets-manager/" target="_blank" rel="noopener"&gt;AWS Secrets Manager&lt;/a&gt; and exposes only two tools: &lt;code&gt;run_query&lt;/code&gt; for SQL execution and &lt;code&gt;get_table_schema&lt;/code&gt; for schema inspection.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/04/BDB-5604-4.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/04/BDB-5604-4.png" alt="Direct source access where the Aurora MCP server queries the Amazon Aurora CRM database without an intermediate catalog" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 4: Direct source access to the Amazon Aurora CRM database&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;Internally, the same agent architecture as Pattern 1 applies: a system prompt, an LLM, and a set of MCP tools. We use Claude Haiku 4.5 powered by Amazon Bedrock as the foundation LLM with tools discovered through the Amazon Bedrock AgentCore Gateway. There’s no catalog layer to query first. The system prompt provides lightweight schema hints: table names and key enum values needed for WHERE clauses so the agent can route correctly and write valid filters without a round trip:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-text"&gt;MYSQL CRM DATA (Aurora MySQL via RDS Data API):

Database: acme_crm

Tables:
- support_tickets: status (open|in_progress|resolved|closed),
  priority (low|medium|high|critical),
  category (billing|technical|content|account)
- content_ratings: rating (1-5), review_text

Use get_table_schema to verify full column details before complex queries.
Use run_query(sql='SELECT...') to execute. Default to read-only SELECT.
Use standard MySQL syntax (not Presto/Trino).&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;For straightforward queries, the agent writes SQL directly from these hints. For complex queries such as multi-table joins or unfamiliar columns, the agent calls &lt;code&gt;get_table_schema&lt;/code&gt; first to verify the full schema, mirroring the discover-then-query discipline from Pattern 1 but against the source database rather than a catalog. To see this in action, consider “Show me open critical support tickets by category”:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;The agent’s routing rules match “support tickets” to the MySQL CRM path and call &lt;code&gt;run_query&lt;/code&gt; with a &lt;code&gt;SELECT&lt;/code&gt; against &lt;code&gt;support_tickets&lt;/code&gt; filtered by &lt;code&gt;status='open'&lt;/code&gt; and &lt;code&gt;priority='critical'&lt;/code&gt;.&lt;/li&gt;
 &lt;li&gt;The Aurora MCP server translates this into a query against Amazon Aurora through the RDS Data API.&lt;/li&gt;
 &lt;li&gt;Results return through the AgentCore Gateway and the agent composes a formatted answer with ticket counts, categories, and so on.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The direct access pattern trades catalog governance for simplicity. There’s no metadata registration step. The MCP server queries the database as-is, which means schema changes in Aurora are immediately visible. This makes it ideal for operational databases where the schema is stable and well-understood, and where the overhead of cataloging every table would slow down access without adding value.&lt;/p&gt;
&lt;p&gt;Earlier this year, the &lt;a href="https://docs.aws.amazon.com/agent-toolkit/latest/userguide/mcp-server.html" target="_blank" rel="noopener"&gt;AWS MCP Server&lt;/a&gt; became generally available. It’s part of the &lt;a href="https://aws.amazon.com/products/developer-tools/agent-toolkit-for-aws/" target="_blank" rel="noopener"&gt;Agent Toolkit for AWS&lt;/a&gt;, a suite of tooling that includes the MCP Server, skills, and plugins that help coding agents build more effectively and efficiently on AWS. Rather than exposing a fixed set of per-service tools, the server provides generic AWS API access: &lt;code&gt;aws___run_script&lt;/code&gt; executes Python in a sandboxed environment with credentialed access to the AWS APIs, authenticated with SigV4 and authorized by your existing &lt;a href="https://aws.amazon.com/iam/?trk=7ad0b48b-3532-4cda-87b0-c63aef99e9cc&amp;amp;sc_channel=ps&amp;amp;ef_id=Cj0KCQjw2OnUBhC2ARIsACKyfaGBzrSr71HGBa4WI5HOshbbEOQ97opAHU7-5qLV8DBd85WALf9WDaEaAqDnEALw_wcB:G:s&amp;amp;gads_camp=23527793912&amp;amp;gads_ag=187898877250&amp;amp;gads_ad=795794010907&amp;amp;gads_kw=amazon%20iam&amp;amp;gads_matchtype=e&amp;amp;gads_network=g&amp;amp;gads_device=c&amp;amp;gads_geo=9032925&amp;amp;gad_campaignid=23527793912&amp;amp;gbraid=0AAAAADjHtp8XHpQ1kaz37duFxmWQLTCuA&amp;amp;gclid=Cj0KCQjw2OnUBhC2ARIsACKyfaGBzrSr71HGBa4WI5HOshbbEOQ97opAHU7-5qLV8DBd85WALf9WDaEaAqDnEALw_wcB" target="_blank" rel="noopener"&gt;AWS Identity and Access Management (IAM)&lt;/a&gt; policies. Because that reaches most of AWS APIs, you can connect your agents to relational data in Aurora through the RDS Data API or to real-time streaming data in Kinesis Data Streams, using &lt;code&gt;boto3&lt;/code&gt; calls such as &lt;code&gt;GetShardIterator&lt;/code&gt; and &lt;code&gt;GetRecords&lt;/code&gt;.&lt;/p&gt;
&lt;h3 id="pattern-3-hybrid-access"&gt;Pattern 3: Hybrid access&lt;/h3&gt;
&lt;p&gt;In practice, most organizations won’t pick only one pattern because the data landscape is too diverse. That’s exactly the case for our streaming media company: batch and streaming data on S3 benefits from catalog-first governance (Pattern 1), while the Aurora CRM database is better served by direct access (Pattern 2). Our reference architecture combines both patterns under a single orchestrator agent. Governed sources route through the catalog. Operational sources are accessed directly and both paths coexist behind the same agent. The key insight: &lt;strong&gt;both paths use the same protocol&lt;/strong&gt;. Amazon Bedrock AgentCore runtime hosts the MCP servers, and AgentCore Gateway handles tool discovery, authentication, and routing. Organizations can start with whichever pattern fits their current data maturity and grow into unified access as they onboard more sources.&lt;/p&gt;
&lt;h2 id="validate-the-deployment"&gt;Validate the deployment&lt;/h2&gt;
&lt;p&gt;Access the CloudFront URL from the stack outputs, log in with your test user credentials, and try these queries:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Query 1 – Customer analytics with visualization:&lt;/strong&gt;&lt;/p&gt;
&lt;blockquote&gt;
 &lt;p&gt;&lt;em&gt;“Build a chart on customer breakup by subscription type?”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The agent queries the &lt;code&gt;customers&lt;/code&gt; table in Athena and generates bar and pie charts showing the distribution across subscription tiers.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/04/BDB-5604-5.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/04/BDB-5604-5.png" alt="Bar and pie charts showing customer distribution across subscription tiers" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 5: Customer distribution across subscription tiers&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;Query 2 – CRM operational breakdown:&lt;/strong&gt;&lt;/p&gt;
&lt;blockquote&gt;
 &lt;p&gt;&lt;em&gt;“Show me the breakdown of support tickets by category and priority.”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This routes entirely to the MySQL MCP server, querying the Aurora CRM database for ticket distribution without touching S3 or Athena.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/04/BDB-5604-6.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/04/BDB-5604-6.png" alt="Support ticket breakdown by category and priority returned from the Aurora CRM database" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 6: Support ticket breakdown by category and priority&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;Query 3 – Federated cross-source query:&lt;/strong&gt;&lt;/p&gt;
&lt;blockquote&gt;
 &lt;p&gt;&lt;em&gt;“What are the top five highest-rated titles and how many streaming hours do they have?”&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This requires the agent to query &lt;code&gt;content_ratings&lt;/code&gt; from Aurora for ratings, then correlate with &lt;code&gt;streaming_events&lt;/code&gt; and &lt;code&gt;titles&lt;/code&gt; in Athena.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/04/BDB-5604-7.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/04/BDB-5604-7.png" alt="Query results listing the top five highest-rated titles alongside their streaming hours" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 7: Top five highest-rated titles and their streaming hours&lt;/p&gt;
&lt;/div&gt;
&lt;h2 id="things-to-consider"&gt;Things to consider&lt;/h2&gt;
&lt;p&gt;Consider these additional factors when you deploy the preceding architecture patterns to production:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;Application security:&lt;/strong&gt; Our architecture patterns use &lt;a href="https://aws.amazon.com/pm/cognito/?trk=36e1404e-1051-48b6-9dd0-51db40b9c756&amp;amp;sc_channel=ps&amp;amp;ef_id=EAIaIQobChMIosXv9cPrlQMVSkT_AR2RyQegEAAYASAAEgJWPPD_BwE&amp;amp;gads_camp=23527793912&amp;amp;gads_ag=187898877050&amp;amp;gads_ad=795794010901&amp;amp;gads_kw=amazon%20cognito&amp;amp;gads_matchtype=e&amp;amp;gads_network=g&amp;amp;gads_device=c&amp;amp;gads_geo=9003495&amp;amp;gad_campaignid=23527793912&amp;amp;gbraid=0AAAAADjHtp8PUlU1yblbfiwjYY5eYoSv6&amp;amp;gclid=EAIaIQobChMIosXv9cPrlQMVSkT_AR2RyQegEAAYASAAEgJWPPD_BwE" target="_blank" rel="noopener"&gt;Amazon Cognito&lt;/a&gt; for identity access and control. However, you should carefully review the identity used by the agent to interact with backend systems.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Data lineage and access control:&lt;/strong&gt; Consider using &lt;a href="https://aws.amazon.com/lake-formation/" target="_blank" rel="noopener"&gt;AWS Lake Formation&lt;/a&gt; for data governance, authentication, and authorization of data assets in the agentic AI application.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Semantic layer for agents:&lt;/strong&gt; Agentic response quality can be improved by providing agents with the right business context and building an independent semantic layer. AWS has recently announced support for &lt;a href="https://aws.amazon.com/about-aws/whats-new/2026/06/aws-glue-data-catalog/" target="_blank" rel="noopener"&gt;business context and semantic search&lt;/a&gt;. This can help the agent discover and understand data by semantic meaning, improve response quality and avoid hallucination, and many other issues.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="clean-up"&gt;Clean up&lt;/h2&gt;
&lt;p&gt;To avoid ongoing charges, destroy both AWS Cloud Development Kit (AWS CDK) stacks (agent stack first, then data stack) and remove any orphaned resources such as Kinesis streams and Amazon CloudWatch log groups. For detailed clean-up instructions, visit the repository’s README.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;Enterprise data stays locked behind &lt;strong&gt;silos&lt;/strong&gt; and an &lt;strong&gt;access gap&lt;/strong&gt;. Every one-time question routes through a handful of data engineers while the insight goes stale. MCP flips the model. Instead of centralizing data or wiring bespoke integrations, you deploy MCP servers that wrap each source behind a standardized protocol and let AI agents query them on behalf of the user. Whether you choose catalog-first access, direct access, or both unified behind a single agent, the agent navigates the complexity so the user doesn’t have to. Adding a new data source means deploying a new MCP server, not redesigning the pipeline.&lt;/p&gt;
&lt;p&gt;Open questions remain, for example, data lineage across agent-composed outputs, identity and authorization when agents are the primary data consumers, and audit trails that capture not only &lt;em&gt;what&lt;/em&gt; an agent accessed but &lt;em&gt;why&lt;/em&gt;. This landscape is growing fast: &lt;a href="https://aws.amazon.com/blogs/aws/the-aws-mcp-server-is-now-generally-available/" target="_blank" rel="noopener"&gt;AWS Labs MCP Servers&lt;/a&gt;, &lt;a href="https://docs.aws.amazon.com/agent-toolkit/latest/userguide/mcp-server.html" target="_blank" rel="noopener"&gt;AWS MCP documentation&lt;/a&gt;, and the &lt;a href="https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/registry.html" target="_blank" rel="noopener"&gt;MCP Gateway Registry&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Deploy the reference architecture, experiment with the patterns, and contribute back what you learn.&lt;/p&gt;
&lt;h3&gt;Acknowledgements&lt;/h3&gt;
&lt;p&gt;We would like to thank Yadgiri Pottabathini for his effort in testing the repository.&lt;/p&gt;
&lt;hr style="width: 100%"&gt;
&lt;h2&gt;About the authors&lt;/h2&gt;
&lt;footer&gt;
 &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
  &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
   &lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/04/BDB-5604-8.jpg" alt="James Wu" width="100" height="133"&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;James Wu&lt;/h3&gt;
  &lt;p style="overflow: hidden"&gt;James is a Principal GenAI/ML Specialist Solutions Architect at AWS, helping enterprises design and execute AI transformation strategies. Specializing in generative AI, agentic systems, and media supply chain automation, he is a featured conference speaker and technical author. Prior to AWS, he was an architect, developer, and technology leader for over 10 years, with experience spanning engineering and marketing industries.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
  &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
   &lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/04/BDB-5604-9.jpg" alt="Rahul Sharma" width="100" height="133"&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Rahul Sharma&lt;/h3&gt;
  &lt;p style="overflow: hidden"&gt;Rahul is a Sr.&amp;nbsp;Specialist Solutions Architect at Amazon Web Services. He is passionate about the data technologies that help leverage data as a strategic asset and is based out of New York.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
  &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
   &lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/04/BDB-5604-10.jpg" alt="Amit Kalawat" width="100" height="133"&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Amit Kalawat&lt;/h3&gt;
  &lt;p style="overflow: hidden"&gt;Amit is a Principal Solutions Architect at Amazon Web Services based out of New York. He works with enterprise customers as they transform their business and journey to the cloud.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
  &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
   &lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/04/BDB-5604-11.jpg" alt="Anirudha Joshi" width="100" height="133"&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Anirudha Joshi&lt;/h3&gt;
  &lt;p style="overflow: hidden"&gt;Anirudha is a Principal Customer Solutions Manager at AWS. A firm believer in &lt;a href="https://aws.amazon.com/video/watch/7a9dc2942e5/" target="_blank" rel="noopener"&gt;working backwards&lt;/a&gt; from customer problems, AJ partners with AWS Media &amp;amp; Entertainment (M&amp;amp;E) customers to guide them through their unique technology &lt;a href="https://aws.amazon.com/what-is/digital-transformation/" target="_blank" rel="noopener"&gt;transformation&lt;/a&gt; journeys. He is a member of the AWS Serverless and Machine Learning/Artificial Intelligence TFCs, with a focus on Agentic AI. Outside of work, AJ coaches and runs marathons, hits the trails hiking, and plays golf.&lt;/p&gt;
 &lt;/div&gt;
&lt;/footer&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>Network connectivity patterns for the next generation of Amazon OpenSearch Serverless</title>
		<link>https://aws.amazon.com/blogs/big-data/network-connectivity-patterns-for-the-next-generation-of-amazon-opensearch-serverless/</link>
		
		<dc:creator><![CDATA[Salman Ahmed]]></dc:creator>
		<pubDate>Thu, 03 Sep 2026 16:26:58 +0000</pubDate>
				<category><![CDATA[Advanced (300)]]></category>
		<category><![CDATA[Amazon OpenSearch Serverless]]></category>
		<category><![CDATA[Technical How-to]]></category>
		<guid isPermaLink="false">8e39757e00b66d46befcf4cc76c297f56946b5d7</guid>

					<description>The next generation of Amazon OpenSearch Serverless uses standard AWS PrivateLink endpoints on the on.aws domain. This post shows nine connectivity patterns for private access, from a single VPC to multiple VPCs, cross-account, on-premises, and cross-Region, with the DNS resolution and data path for each.</description>
										<content:encoded>&lt;p&gt;Network connectivity patterns for private access to &lt;a href="https://aws.amazon.com/opensearch-service/features/serverless/" target="_blank" rel="noopener"&gt;Amazon OpenSearch Serverless&lt;/a&gt; used to require considerable setup. You had to create virtual private cloud (VPC) endpoints in every consumer VPC and configure &lt;a href="https://docs.aws.amazon.com/Route53/latest/DeveloperGuide/profiles.html" target="_blank" rel="noopener"&gt;Amazon Route 53 Profiles&lt;/a&gt; for cross-account DNS. You also had to maintain custom private hosted zones with CNAME records and deploy resolver inbound endpoints for on-premises connectivity. &lt;a href="https://aws.amazon.com/blogs/big-data/the-next-generation-of-amazon-opensearch-serverless-built-from-the-ground-up-for-agents/" target="_blank" rel="noopener"&gt;The next generation of OpenSearch Serverless&lt;/a&gt; changes this. It uses standard &lt;a href="https://aws.amazon.com/privatelink/" target="_blank" rel="noopener"&gt;AWS PrivateLink&lt;/a&gt; interface endpoints with native private DNS support. Connectivity patterns that previously required multi-step DNS orchestration now work with the same endpoint mechanics you already use for other AWS services.&lt;/p&gt;
&lt;p&gt;Collections use &lt;a href="https://docs.aws.amazon.com/opensearch-service/latest/developerguide/serverless-collection-endpoints.html" target="_blank" rel="noopener"&gt;resource-based endpoints&lt;/a&gt; on the &lt;code&gt;on.aws&lt;/code&gt; domain in two formats. The per-collection endpoint (&lt;code&gt;&amp;lt;collectionId&amp;gt;.aoss.&amp;lt;region&amp;gt;.on.aws&lt;/code&gt;) reaches a single collection, and the hostname itself identifies which collection you want, so no additional routing information is needed. The per-account Regional endpoint (&lt;code&gt;&amp;lt;accountId&amp;gt;.aoss.&amp;lt;region&amp;gt;.on.aws&lt;/code&gt;) reaches any collection in your account through one hostname. Because the hostname alone does not identify a specific collection, you add the &lt;code&gt;x-amz-aoss-collection-name&lt;/code&gt; header (or &lt;code&gt;x-amz-aoss-collection-id&lt;/code&gt;) to each request to name the target collection. The AWS SDKs include this header automatically when they sign the request with &lt;a href="https://docs.aws.amazon.com/IAM/latest/UserGuide/reference_sigv.html" target="_blank" rel="noopener"&gt;Signature Version 4 (SigV4)&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Both formats use standard AWS PrivateLink. You create the VPC endpoint from the &lt;a href="https://aws.amazon.com/vpc/" target="_blank" rel="noopener"&gt;Amazon Virtual Private Cloud (Amazon VPC)&lt;/a&gt; console or the &lt;a href="https://aws.amazon.com/ec2/" target="_blank" rel="noopener"&gt;Amazon Elastic Compute Cloud (Amazon EC2)&lt;/a&gt; &lt;code&gt;CreateVpcEndpoint&lt;/code&gt; API, using the service name &lt;code&gt;com.amazonaws.&amp;lt;region&amp;gt;.aoss-data&lt;/code&gt;. It is the same interface endpoint you create for any other AWS service.&lt;/p&gt;
&lt;p&gt;In this post, each pattern shows the architecture, the DNS resolution flow, and the data traffic path. Patterns 1 through 8 operate within a single Region across one or more accounts, labeled Region A in the diagrams, so the repeated Region A boxes in a cross-account pattern are the same Region. Only Pattern 9 spans Regions, shown as Region A and Region B.&lt;/p&gt;
&lt;p&gt;These patterns apply to the collection (data) endpoint only. When you create a collection, you also receive an OpenSearch UI endpoint. That endpoint uses a separate PrivateLink mechanism today, with its own VPC endpoint and access policy, and is on a path to move to the standard PrivateLink model. OpenSearch UI connectivity is out of scope for this post.&lt;/p&gt;
&lt;h2 id="prerequisites"&gt;Prerequisites&lt;/h2&gt;
&lt;ul&gt;
 &lt;li&gt;A &lt;a href="https://docs.aws.amazon.com/opensearch-service/latest/developerguide/serverless-collection-groups-procedures.html" target="_blank" rel="noopener"&gt;collection group&lt;/a&gt; and at least one collection. Create it on the console with Express Create, or with the &lt;a href="https://aws.amazon.com/cli/" target="_blank" rel="noopener"&gt;AWS Command Line Interface (AWS CLI)&lt;/a&gt; using &lt;code&gt;--generation NEXTGEN&lt;/code&gt;.&lt;/li&gt;
 &lt;li&gt;A &lt;a href="https://docs.aws.amazon.com/opensearch-service/latest/developerguide/serverless-network.html" target="_blank" rel="noopener"&gt;network access policy&lt;/a&gt; that lists the VPC endpoints allowed to reach the collection.&lt;/li&gt;
 &lt;li&gt;A &lt;a href="https://docs.aws.amazon.com/opensearch-service/latest/developerguide/serverless-data-access.html" target="_blank" rel="noopener"&gt;data access policy&lt;/a&gt; that grants &lt;a href="https://aws.amazon.com/iam/" target="_blank" rel="noopener"&gt;AWS Identity and Access Management (IAM)&lt;/a&gt; principals permission to index and query.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="dns-resolution"&gt;DNS resolution&lt;/h2&gt;
&lt;p&gt;When you create a standard VPC endpoint for &lt;code&gt;com.amazonaws.&amp;lt;region&amp;gt;.aoss-data&lt;/code&gt; with private DNS enabled, AWS creates a private hosted zone for &lt;code&gt;*.aoss.&amp;lt;region&amp;gt;.on.aws&lt;/code&gt; and associates it with your VPC. This zone maps collection hostnames to the endpoint’s private elastic network interface (ENI) IP addresses. Your compute’s DNS query reaches the VPC’s Amazon Route 53 Resolver at VPC+2, which resolves the hostname to ENI IPs.&lt;/p&gt;
&lt;p&gt;One endpoint serves every collection hostname in the Region. The following AWS CLI command creates that interface endpoint, and the &lt;code&gt;--private-dns-enabled&lt;/code&gt; flag turns on the private DNS resolution described here.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws ec2 create-vpc-endpoint \
  --vpc-id vpc-abc123 \
  --service-name com.amazonaws.us-east-1.aoss-data \
  --vpc-endpoint-type Interface \
  --subnet-ids subnet-111 subnet-222 \
  --security-group-ids sg-xxx \
  --private-dns-enabled&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;In Regions that support &lt;a href="https://aws.amazon.com/compliance/fips/" target="_blank" rel="noopener"&gt;Federal Information Processing Standards (FIPS)&lt;/a&gt;, the same endpoint also resolves &lt;code&gt;*.aoss-fips.&amp;lt;region&amp;gt;.on.aws&lt;/code&gt; for FIPS-compliant access.&lt;/p&gt;
&lt;p&gt;OpenSearch Serverless has no per-collection Dashboards endpoint. Use &lt;a href="https://docs.aws.amazon.com/opensearch-service/latest/developerguide/application.html" target="_blank" rel="noopener"&gt;OpenSearch UI applications&lt;/a&gt; to explore and visualize collection data.&lt;/p&gt;
&lt;p&gt;The diagrams in the following patterns use an Amazon EC2 instance to represent the compute client. Any compute in the VPC reaches a collection the same way, including EC2 instances, &lt;a href="https://aws.amazon.com/lambda/" target="_blank" rel="noopener"&gt;AWS Lambda&lt;/a&gt; functions attached to the VPC, and containers on &lt;a href="https://aws.amazon.com/ecs/" target="_blank" rel="noopener"&gt;Amazon Elastic Container Service (Amazon ECS)&lt;/a&gt; or &lt;a href="https://aws.amazon.com/eks/" target="_blank" rel="noopener"&gt;Amazon Elastic Kubernetes Service (Amazon EKS)&lt;/a&gt;. The connectivity, DNS resolution, and access policies are the same regardless of the compute type.&lt;/p&gt;
&lt;h2 id="pattern-1.-private-access-from-a-single-vpc"&gt;Pattern 1: Private access from a single VPC&lt;/h2&gt;
&lt;p&gt;Compute in a VPC needs private access to collections in the same account. The following diagram shows the architecture for private access from a single VPC.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/28/BDB-6127-1.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/28/BDB-6127-1.png" alt="Compute in a single VPC reaches a collection through a VPC interface endpoint with private DNS enabled" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 1: Private access from a single VPC&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;&lt;a href="https://docs.aws.amazon.com/opensearch-service/latest/developerguide/serverless-vpc.html" target="_blank" rel="noopener"&gt;Create a standard VPC endpoint&lt;/a&gt; in the VPC where your compute runs, then reference its ID in the collection’s network policy.&lt;/p&gt;
&lt;p&gt;For the DNS resolution flow, (1) compute queries &lt;code&gt;&amp;lt;collectionId&amp;gt;.aoss.&amp;lt;region&amp;gt;.on.aws&lt;/code&gt;, and the VPC Route 53 Resolver at VPC+2 returns the endpoint ENI IP addresses because private DNS is enabled on the endpoint.&lt;/p&gt;
&lt;p&gt;For the data traffic path, (2) compute connects to the ENI, and (3) PrivateLink forwards the request to the service, which routes to the collection by hostname.&lt;/p&gt;
&lt;h2 id="pattern-2.-multiple-vpcs-in-the-same-account"&gt;Pattern 2: Multiple VPCs in the same account&lt;/h2&gt;
&lt;p&gt;Several VPCs, split by environment, tier, or team, need private access to the same collections. The following diagram shows how each VPC uses its own endpoint to reach the same collections.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/28/BDB-6127-2.png" target="_blank" rel="noopener"&gt;&lt;img loading="lazy" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/28/BDB-6127-2.png" alt="Three VPCs in one account, each with its own aoss-data interface endpoint reaching the same collections" width="800" height="1279"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 2: Multiple VPCs in the same account&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;Each VPC needs exactly one &lt;code&gt;aoss-data&lt;/code&gt; endpoint with private DNS enabled, and that single endpoint already reaches every collection in the Region. DNS resolves independently within each VPC, so there is no cross-VPC DNS dependency. Adding a new VPC takes two steps. Create the endpoint, then add its endpoint ID to the collection’s network policy. Do not create a second &lt;code&gt;aoss-data&lt;/code&gt; endpoint with private DNS enabled in the same VPC. Both endpoints share the same private hosted zone, which causes a conflict and the creation fails.&lt;/p&gt;
&lt;p&gt;For the DNS resolution flow, (1) compute in each VPC queries the collection hostname, and that VPC’s Route 53 Resolver at VPC+2 returns the endpoint ENI IP addresses because private DNS is enabled on the endpoint.&lt;/p&gt;
&lt;p&gt;For the data traffic path, (2) compute connects to its local ENI, and (3) PrivateLink forwards the request to the service, which routes to the collection by hostname.&lt;/p&gt;
&lt;h2 id="pattern-3.-on-premises-access-from-a-single-account"&gt;Pattern 3: On-premises access from a single account&lt;/h2&gt;
&lt;p&gt;On-premises clients reach collections over &lt;a href="https://aws.amazon.com/directconnect/" target="_blank" rel="noopener"&gt;AWS Direct Connect&lt;/a&gt; or &lt;a href="https://aws.amazon.com/vpn/site-to-site-vpn/" target="_blank" rel="noopener"&gt;AWS Site-to-Site VPN&lt;/a&gt;, which connect to the VPC through &lt;a href="https://aws.amazon.com/transit-gateway/" target="_blank" rel="noopener"&gt;AWS Transit Gateway&lt;/a&gt; or &lt;a href="https://aws.amazon.com/cloud-wan/" target="_blank" rel="noopener"&gt;AWS Cloud WAN&lt;/a&gt;. The following diagram shows the DNS and data path for on-premises access.&lt;/p&gt;
&lt;div id="attachment_94093" style="width: 3102px" class="wp-caption alignleft"&gt;
 &lt;img aria-describedby="caption-attachment-94093" loading="lazy" class="wp-image-94093 size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/03/BDB-6127-3-1.png" alt="" width="3092" height="1150"&gt;
 &lt;p id="caption-attachment-94093" class="wp-caption-text"&gt;Figure 3: On-premises access from a single account&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;On-premises DNS servers sit outside the VPC and cannot resolve PrivateLink private DNS names directly. Place an &lt;a href="https://docs.aws.amazon.com/Route53/latest/DeveloperGuide/resolver-forwarding-inbound-queries.html" target="_blank" rel="noopener"&gt;Amazon Route 53 Resolver inbound endpoint&lt;/a&gt; in the VPC that holds the &lt;code&gt;aoss-data&lt;/code&gt; VPC endpoint. On-premises DNS forwards queries for &lt;code&gt;aoss.&amp;lt;region&amp;gt;.on.aws&lt;/code&gt; to that inbound endpoint. The inbound endpoint resolves them against the private hosted zone. The inbound endpoint’s security group must allow TCP/UDP port 53 from your on-premises resolver ranges.&lt;/p&gt;
&lt;p&gt;For the DNS resolution flow, (1) the client queries the on-premises resolver. (2) The on-premises conditional forwarder for &lt;code&gt;*.aoss.&amp;lt;region&amp;gt;.on.aws&lt;/code&gt; sends the query over Direct Connect or VPN, through Transit Gateway or Cloud WAN, to the inbound endpoint. The inbound endpoint uses the VPC Route 53 Resolver to return the endpoint’s private ENI IPs.&lt;/p&gt;
&lt;p&gt;For the data traffic path, (3) the client sends an HTTPS request with the Transport Layer Security (TLS) Server Name Indication (SNI) header set to the collection hostname, over Direct Connect or VPN through Transit Gateway or Cloud WAN. (4) Traffic crosses the VPC’s attachment ENI, (5) reaches the VPC endpoint ENIs, and (6) PrivateLink forwards the request to the service.&lt;/p&gt;
&lt;h2 id="pattern-4.-cross-account-access-with-an-endpoint-in-each-consumer-vpc"&gt;Pattern 4: Cross-account access with an endpoint in each consumer VPC&lt;/h2&gt;
&lt;p&gt;A central account hosts collections, and compute in spoke accounts needs private access. Many enterprises start here. The following diagram shows the cross-account endpoint architecture.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/28/BDB-6127-4.png" target="_blank" rel="noopener"&gt;&lt;img loading="lazy" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/28/BDB-6127-4.png" alt="Spoke accounts each with their own interface endpoint reaching collections in a central account over PrivateLink" width="800" height="811"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 4: Cross-account access with an endpoint in each consumer VPC&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;Each spoke creates its own endpoint. The collection owner’s network policy references the spoke’s endpoint ID. The data access policy grants the spoke’s IAM role. PrivateLink carries the traffic end to end, with no Transit Gateway and no peering.&lt;/p&gt;
&lt;p&gt;The endpoint lives in the spoke account, not the collection account. The spoke team creates a standard interface VPC endpoint in the spoke VPC for the service name &lt;code&gt;com.amazonaws.&amp;lt;region&amp;gt;.aoss-data&lt;/code&gt; with private DNS enabled. The collection owner does not create this endpoint. After the endpoint is ready the spoke shares its endpoint ID with the collection owner, who adds that ID to the collection network policy under &lt;code&gt;SourceVPCEs&lt;/code&gt;. A network policy accepts endpoint IDs from accounts across your organization. Each spoke creates its own endpoint and shares the ID rather than peering VPCs or routing through another account’s endpoint.&lt;/p&gt;
&lt;p&gt;Network access and data access stay separate. The network policy authorizes the endpoint, and the data access policy authorizes the identity. A serverless data access policy grants principals from the collection’s own account. For a spoke in another account, you &lt;a href="https://docs.aws.amazon.com/IAM/latest/UserGuide/tutorial_cross-account-with-roles.html" target="_blank" rel="noopener"&gt;create an IAM role in the collection account&lt;/a&gt; and grant that role in the data access policy. The spoke role then assumes it to sign requests.&lt;/p&gt;
&lt;p&gt;The following network access policy lists the two spoke endpoint IDs under &lt;code&gt;SourceVPCEs&lt;/code&gt; and sets &lt;code&gt;AllowFromPublic&lt;/code&gt; to false, so only those endpoints reach the collection and the policy denies public access.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-json"&gt;[
  {
    "Description": "Cross-account access from spoke",
    "Rules": [
      {
        "ResourceType": "collection",
        "Resource": [
          "collection/my-collection"
        ]
      }
    ],
    "AllowFromPublic": false,
    "SourceVPCEs": [
      "vpce-spoke-b-id",
      "vpce-spoke-c-id"
    ]
  }
]&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;For the DNS resolution flow, (1) spoke compute queries the VPC Route 53 Resolver at VPC+2, which returns the local endpoint ENI IPs because private DNS is enabled on the endpoint.&lt;/p&gt;
&lt;p&gt;For the data traffic path, (2) compute connects to the local ENI. (3) PrivateLink forwards the request to the service, which checks the network policy for the endpoint ID and the data access policy for the IAM role before routing. Adding a spoke takes one API call and two policy edits.&lt;/p&gt;
&lt;h2 id="pattern-5.-centralized-shared-endpoint-with-amazon-route-53-profiles-over-transit-gateway"&gt;Pattern 5: Centralized shared endpoint with Amazon Route 53 Profiles over Transit Gateway&lt;/h2&gt;
&lt;p&gt;You want fewer PrivateLink endpoints, so you run one shared endpoint in a networking VPC and reach it from spoke accounts over Transit Gateway or AWS Cloud WAN, with no endpoint in each spoke. The following diagram shows this centralized architecture.&lt;/p&gt;
&lt;div id="attachment_94096" style="width: 4010px" class="wp-caption alignleft"&gt;
 &lt;img aria-describedby="caption-attachment-94096" loading="lazy" class="wp-image-94096 size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/03/BDB-6127-5-1.png" alt="" width="4000" height="1126"&gt;
 &lt;p id="caption-attachment-94096" class="wp-caption-text"&gt;Figure 5: Centralized shared endpoint with Amazon Route 53 Profiles over Transit Gateway&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;Pattern 5 consolidates access through a single shared endpoint in a central networking VPC rather than creating one per spoke. Because spoke VPCs have no local endpoint, they cannot resolve &lt;code&gt;*.aoss.&amp;lt;region&amp;gt;.on.aws&lt;/code&gt; on their own. You share the endpoint’s private DNS with spoke VPCs using Amazon Route 53 Profiles, shared through &lt;a href="https://aws.amazon.com/ram/" target="_blank" rel="noopener"&gt;AWS Resource Access Manager (AWS RAM)&lt;/a&gt;. This is the one pattern where you still manage DNS propagation.&lt;/p&gt;
&lt;p&gt;For the DNS resolution flow, (1) the spoke resolves the hostname through the shared Route 53 Profile, which returns the networking-VPC endpoint ENI IPs.&lt;/p&gt;
&lt;p&gt;For the data traffic path, (2) traffic leaves the compute through the spoke VPC’s attachment ENI, (3) crosses Transit Gateway or Cloud WAN into the networking VPC’s attachment ENI, (4) reaches the shared endpoint ENIs, and (5) PrivateLink forwards the request to the service.&lt;/p&gt;
&lt;h2 id="pattern-6.-cross-account-centralized-networking-with-on-premises"&gt;Pattern 6: Cross-account centralized networking with on-premises&lt;/h2&gt;
&lt;p&gt;A central account hosts collections. A separate networking account owns Direct Connect or VPN and Route 53. On-premises clients reach the collections through the networking account. The following diagram shows this architecture.&lt;/p&gt;
&lt;div id="attachment_94095" style="width: 3350px" class="wp-caption alignleft"&gt;
 &lt;img aria-describedby="caption-attachment-94095" loading="lazy" class="wp-image-94095 size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/03/BDB-6127-6.png" alt="" width="3340" height="1150"&gt;
 &lt;p id="caption-attachment-94095" class="wp-caption-text"&gt;Figure 6: Cross-account centralized networking with on-premises&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;The networking account runs the standard VPC endpoint and a Route 53 Resolver inbound endpoint. The collection owner’s network policy references the networking account’s endpoint ID.&lt;/p&gt;
&lt;p&gt;For the DNS resolution flow, (1) the on-premises client queries its resolver. (2) The on-premises conditional forwarder sends the query over Direct Connect or VPN, through Transit Gateway or Cloud WAN, to the inbound endpoint. The inbound endpoint uses the VPC Route 53 Resolver to return the endpoint’s private ENI IPs.&lt;/p&gt;
&lt;p&gt;For the data traffic path, (3) the client sends HTTPS over Direct Connect or VPN, through Transit Gateway or Cloud WAN. (4) Traffic crosses the networking VPC’s attachment ENI, (5) reaches the networking-VPC endpoint ENIs, and (6) PrivateLink forwards the request to the service in the central account. The two teams coordinate through one artifact, the endpoint ID.&lt;/p&gt;
&lt;h2 id="pattern-7.-distributed-multi-business-unit-with-spoke-account-access"&gt;Pattern 7: Distributed multi-business-unit with spoke-account access&lt;/h2&gt;
&lt;p&gt;Spoke accounts such as analytics or application teams need collections spread across several business unit accounts, and each unit manages its own collections. The following diagram shows the distributed multi-business-unit architecture.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/28/BDB-6127-7.png" target="_blank" rel="noopener"&gt;&lt;img loading="lazy" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/28/BDB-6127-7.png" alt="Spoke accounts reaching collections spread across several business unit accounts, each spoke with its own endpoint" width="800" height="1011"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 7: Distributed multi-business-unit with spoke-account access&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;Each spoke creates one standard endpoint, which resolves every collection hostname in the Region. Each business unit’s network policy lists the spoke endpoint IDs. Access control decides which collections a spoke reaches. DNS does not.&lt;/p&gt;
&lt;p&gt;For the DNS resolution flow, (1) spoke compute queries the VPC Route 53 Resolver at VPC+2, which returns the endpoint ENI IPs because private DNS is enabled on the endpoint.&lt;/p&gt;
&lt;p&gt;For the data traffic path, (2) compute connects to the local ENI, and (3) PrivateLink forwards the request to the service, which routes to the correct business unit collection by hostname.&lt;/p&gt;
&lt;table border="1px" width="100%" cellpadding="10px"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;Action&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Required change&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;New collection in any BU&lt;/td&gt;
   &lt;td&gt;No networking change is needed because in the network policy &lt;code&gt;collection/*&lt;/code&gt;&amp;nbsp; wildcard, already covers any new collection&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;New spoke account&lt;/td&gt;
   &lt;td&gt;Spoke creates an endpoint, and BUs add its ID to their policies&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Remove spoke access&lt;/td&gt;
   &lt;td&gt;BUs remove the endpoint ID and the IAM principal&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="pattern-8.-distributed-multi-business-unit-with-on-premises-access"&gt;Pattern 8: Distributed multi-business-unit with on-premises access&lt;/h2&gt;
&lt;p&gt;Several business units own collections in separate accounts. On-premises clients reach collections across all of those accounts through a central networking account. The following diagram shows this architecture.&lt;/p&gt;
&lt;div id="attachment_94094" style="width: 3390px" class="wp-caption alignleft"&gt;
 &lt;img aria-describedby="caption-attachment-94094" loading="lazy" class="wp-image-94094 size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/03/BDB-6127-8.png" alt="" width="3380" height="1112"&gt;
 &lt;p id="caption-attachment-94094" class="wp-caption-text"&gt;Figure 8: Distributed multi-business-unit with on-premises access&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;The networking account runs one standard endpoint that resolves &lt;code&gt;*.aoss.&amp;lt;region&amp;gt;.on.aws&lt;/code&gt; hostnames, regardless of which account owns the collection. Each business unit’s network policy includes the networking endpoint ID.&lt;/p&gt;
&lt;p&gt;For the DNS resolution flow, (1) the on-premises client queries its resolver. (2) The on-premises conditional forwarder for &lt;code&gt;*.aoss.&amp;lt;region&amp;gt;.on.aws&lt;/code&gt; sends the query over Direct Connect or VPN, through Transit Gateway or Cloud WAN, to the networking VPC’s inbound endpoint. The inbound endpoint uses the VPC Route 53 Resolver to return the shared endpoint’s private ENI IPs.&lt;/p&gt;
&lt;p&gt;For the data traffic path, (3) the client sends HTTPS over Direct Connect or VPN, through Transit Gateway or Cloud WAN. (4) Traffic crosses the networking VPC’s attachment ENI, (5) the request arrives at the shared endpoint ENIs, and (6) the service routes to business unit 1 or business unit 2 by hostname, as long as that business unit’s policy lists the networking endpoint ID. Adding a collection in any business unit needs no networking change if the network policy uses a &lt;code&gt;collection/*&lt;/code&gt; wildcard, since the wildcard already covers it.&lt;/p&gt;
&lt;h2 id="pattern-9.-cross-region-access-strategies"&gt;Pattern 9: Cross-Region access strategies&lt;/h2&gt;
&lt;p&gt;Consumers in Region B need data that lives in collections in Region A. The following diagram shows cross-Region access strategies.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/28/BDB-6127-9.png" target="_blank" rel="noopener"&gt;&lt;img loading="lazy" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/28/BDB-6127-9.png" alt="Independent collections in Region A and Region B, each with its own endpoint, with a replication arrow showing cross-Region data sync" width="800" height="780"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 9: Cross-Region access strategies&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;Collections are Regional. No built-in cross-Region endpoint or replication exists. Deploy independent collections in each Region, each with its own endpoint and policies, then synchronize data with one of these approaches.&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;Dual-write. The application writes to both Regions at ingestion time.&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://aws.amazon.com/opensearch-service/features/ingestion/" target="_blank" rel="noopener"&gt;Amazon OpenSearch Ingestion&lt;/a&gt; pipeline. A pipeline replicates index operations to the secondary Region with near real-time lag. The pipeline creates its own PrivateLink endpoint to the destination collection. It adds the endpoint to that collection’s network policy automatically. You only need to name the network policy and grant the pipeline role.&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://aws.amazon.com/s3/" target="_blank" rel="noopener"&gt;Amazon Simple Storage Service (Amazon S3)&lt;/a&gt; Cross-Region Replication with re-ingestion. Cross-Region Replication copies objects, and an OpenSearch Ingestion pipeline loads them into the local collection. Lag runs in minutes, at the lowest cost of these approaches.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For the DNS resolution flow, DNS resolves locally in each Region, the same as Pattern 1. Each collection hostname carries its Region, so a hostname in Region A resolves through Region A’s own endpoint and a hostname in Region B resolves through Region B’s own endpoint, with no cross-Region DNS.&lt;/p&gt;
&lt;p&gt;For the data traffic path, (1) compute in each Region uses that Region’s own endpoint to reach its local collection. Writes land in the primary Region and the sync approach you choose replicates them to the secondary Region, where local readers query the replica. The replicate arrow shows that cross-Region movement, such as an OpenSearch Ingestion pipeline that writes into the secondary-Region collection.&lt;/p&gt;
&lt;p&gt;Scale-to-zero changes the economics. An idle secondary-Region collection costs only storage until requests arrive.&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;table border="1px" width="100%" cellpadding="10px"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;Pattern&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Components&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;1. Same VPC&lt;/td&gt;
   &lt;td&gt;Standard endpoint and network policy&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;2. Multiple VPCs&lt;/td&gt;
   &lt;td&gt;Endpoint per VPC and a policy listing all IDs&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;3. On-premises&lt;/td&gt;
   &lt;td&gt;Endpoint, Route 53 inbound endpoint, on-premises forwarder, and Transit Gateway or Cloud WAN&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;4. Cross-account&lt;/td&gt;
   &lt;td&gt;Endpoint per consumer, network policy, and data policy&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;5. Centralized shared endpoint&lt;/td&gt;
   &lt;td&gt;Shared endpoint, Route 53 Profiles through RAM, and Transit Gateway or Cloud WAN&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;6. Central networking with on-premises&lt;/td&gt;
   &lt;td&gt;Networking endpoint, Route 53 inbound, forwarder, Transit Gateway or Cloud WAN, and policies&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;7. Multi-BU with spoke access&lt;/td&gt;
   &lt;td&gt;Endpoint per spoke, and each BU policy lists spoke IDs&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;8. Multi-BU with on-premises&lt;/td&gt;
   &lt;td&gt;One networking endpoint reached through Transit Gateway or Cloud WAN, and each BU policy lists its ID&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;9. Cross-Region&lt;/td&gt;
   &lt;td&gt;Independent collections per Region and a data-sync approach&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Across each private pattern, the VPC endpoint resolves all &lt;code&gt;*.aoss.&amp;lt;region&amp;gt;.on.aws&lt;/code&gt; hostnames through standard PrivateLink private DNS. Network policies control which endpoints reach a collection, and data access policies control which principals operate on the data. Only Pattern 5 asks you to manage DNS.&lt;/p&gt;
&lt;h2 id="cost-considerations"&gt;Cost considerations&lt;/h2&gt;
&lt;p&gt;The connectivity pattern you choose drives recurring cost, so match it to your scale instead of adding infrastructure you do not need. The two charges that come up most often, a Route 53 Resolver inbound endpoint and Route 53 Profiles, are both optional for access that stays inside AWS.&lt;/p&gt;
&lt;p&gt;A Route 53 Resolver inbound endpoint is needed only for the on-premises patterns (3, 6, and 8), where an on-premises resolver forwards queries into the VPC. Traffic that stays inside AWS never uses it. Route 53 Profiles apply only when a VPC has no endpoint of its own, as in Pattern 5, where the profile carries the shared endpoint’s private DNS to the spoke. When each VPC runs its own interface endpoint, DNS resolves locally through the VPC Route 53 Resolver at no extra charge, so neither the inbound endpoint nor a profile is required.&lt;/p&gt;
&lt;p&gt;For most multi-account and multi-Region deployments, an interface endpoint in each consumer VPC (Pattern 4) is the least complex and often the least expensive option. You pay for the interface endpoints you already need for private access, and local DNS resolution adds nothing. Because collections are Regional and each Region resolves on its own, this scales across Regions with no cross-Region DNS.&lt;/p&gt;
&lt;p&gt;Centralizing on one shared endpoint (Pattern 5) lowers the number of interface endpoints. However, it adds Transit Gateway or Cloud WAN data processing charges and the cost of sharing DNS. You share that DNS either through Route 53 Profiles or through a private hosted zone that you associate across accounts and maintain yourself. A smaller endpoint count is not automatically cheaper because transit data processing can exceed the savings. Compare both designs against your own traffic before you decide.&lt;/p&gt;
&lt;p&gt;Scale to zero also shapes cost. An idle collection, such as a secondary-Region replica in Pattern 9, releases its compute and bills only for storage until requests arrive. For current rates, see &lt;a href="https://aws.amazon.com/privatelink/pricing/" target="_blank" rel="noopener"&gt;AWS PrivateLink pricing&lt;/a&gt;, &lt;a href="https://aws.amazon.com/route53/pricing/" target="_blank" rel="noopener"&gt;Amazon Route 53 pricing&lt;/a&gt;, and &lt;a href="https://aws.amazon.com/opensearch-service/pricing/" target="_blank" rel="noopener"&gt;Amazon OpenSearch Service pricing&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;OpenSearch Serverless uses standard AWS PrivateLink for private connectivity. You create a VPC endpoint, enable private DNS, and reference the endpoint ID in your network policy. The model scales from single-VPC access to multi-account and multi-business-unit designs, and only Pattern 5 adds DNS infrastructure, where you share the endpoint’s private DNS with Route 53 Profiles. The per-account regional endpoint goes further and serves any collection in an account through one hostname and connection pool. To get started, create your first collection in the OpenSearch Serverless console, or explore the &lt;a href="https://docs.aws.amazon.com/opensearch-service/latest/developerguide/serverless.html" target="_blank" rel="noopener"&gt;OpenSearch Serverless documentation&lt;/a&gt; for detailed API references and tutorials.&lt;/p&gt;
&lt;hr style="width: 100%"&gt;
&lt;h2&gt;About the authors&lt;/h2&gt;
&lt;footer&gt;
 &lt;div class="blog-author-box"&gt;
  &lt;div class="blog-author-image"&gt;
   &lt;p&gt;&lt;img loading="lazy" class="alignleft size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/28/BDB-6127-10.png" alt="Salman Ahmed" width="100" height="100"&gt;&lt;/p&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Salman Ahmed&lt;/h3&gt;
  &lt;p&gt;Salman is a Senior Technical Account Manager at AWS, specializing in helping customers design, implement, and optimize their AWS environments. He combines deep networking expertise with a passion for exploring emerging technologies to help organizations get the most out of their cloud investments. Outside of work, he enjoys photography, traveling, and watching his favorite sports teams.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box"&gt;
  &lt;div class="blog-author-image"&gt;
   &lt;p&gt;&lt;img loading="lazy" class="alignleft size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/28/BDB-6127-11.png" alt="Ankush Goyal" width="100" height="100"&gt;&lt;/p&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Ankush Goyal&lt;/h3&gt;
  &lt;p&gt;Ankush is a Senior Technical Account Manager at AWS Enterprise Support, specializing in helping customers in the travel and hospitality industries optimize their cloud infrastructure. With over 20 years of IT experience, he focuses on using AWS networking services to drive operational efficiency and cloud adoption. Ankush is passionate about delivering impactful solutions and helping clients to streamline their cloud operations.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box"&gt;
  &lt;div class="blog-author-image"&gt;
   &lt;p&gt;&lt;img loading="lazy" class="alignleft size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/28/BDB-6127-12.png" alt="Ravi Bhatane" width="100" height="100"&gt;&lt;/p&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Ravi Bhatane&lt;/h3&gt;
  &lt;p&gt;Ravi is a Software Engineer at AWS working on Amazon OpenSearch Serverless. He builds the gateway layer that fronts the service, handling private connectivity, authentication, and request routing for customer traffic into collections. He’s drawn to distributed systems and the challenge of keeping them secure, highly available, and low latency as they grow. Outside of work, he enjoys photography and hiking.&lt;/p&gt;
 &lt;/div&gt;
&lt;/footer&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>How Moovit achieved 33% cost optimization through architectural modernization</title>
		<link>https://aws.amazon.com/blogs/big-data/how-moovit-achieved-33-cost-optimization-through-architectural-modernization/</link>
		
		<dc:creator><![CDATA[Saar Porat]]></dc:creator>
		<pubDate>Thu, 03 Sep 2026 16:23:46 +0000</pubDate>
				<category><![CDATA[Amazon EMR]]></category>
		<category><![CDATA[Amazon Redshift]]></category>
		<category><![CDATA[Customer Solutions]]></category>
		<guid isPermaLink="false">ae1009c12e15ea82d9ff5164d82c7638c5b8109c</guid>

					<description>Learn how Moovit modernized its data platform with a multi-engine lakehouse architecture: offloading heavy aggregation workloads from Amazon Redshift to Amazon EMR with Spark SQL, isolating workloads with Amazon Redshift Serverless, and cutting overall data pipeline cost by 33%.</description>
										<content:encoded>&lt;p&gt;&lt;a href="http://www.moovit.com/" target="_blank" rel="noopener"&gt;Moovit&lt;/a&gt;, part of Mobileye (Nasdaq: MBLY), is a leading Mobility-as-a-Service (MaaS) solutions provider and the creator of a leading urban mobility app. Moovit’s iOS, Android, and web apps offer users a smart mobility experience to get to their destination using any mode of public and shared transportation. Transit riders can benefit from mobile ticketing to plan, pay, and ride with transit services. Introduced in 2012, Moovit now serves over 1.7 billion users in more than 3,500 cities across 112 countries, in 45 languages.&lt;/p&gt;
&lt;p&gt;Behind these user-facing experiences is a data platform that processes large volumes of mobility, application, and operational data to support product analytics, business intelligence (BI), monitoring, and data science. As the platform grew, Moovit needed to keep analytical workloads reliable and cost-efficient without slowing down teams that depend on fresh data every day.&lt;/p&gt;
&lt;p&gt;Over several years, Moovit’s Amazon Redshift cluster grew continuously. It started with an expanding fleet of DC2 nodes, migrated to RA3 nodes, and scaled multiple times to keep pace with growing data demands, ultimately becoming the backbone of their entire data platform.&lt;/p&gt;
&lt;p&gt;To address this growth, Moovit transformed their data architecture by building an optimal multi-engine lakehouse architecture and assigning each workload to the most suitable option. This modernization reduced their Amazon Redshift cluster by 50 percent, while establishing a flexible, multi-engine architecture ready for future use cases.&lt;/p&gt;
&lt;p&gt;In this post, we share how Moovit gained visibility into workload patterns, cleaned up unnecessary load, selected candidates for offloading, and ran a successful proof of concept (POC) on Amazon EMR Serverless. Moovit ultimately divided the workload between multiple engines, building a modern and cost-optimized data platform that combines provisioned &lt;a href="https://docs.aws.amazon.com/redshift/latest/mgmt/overview.html" target="_blank" rel="noopener"&gt;Amazon Redshift&lt;/a&gt;, &lt;a href="https://docs.aws.amazon.com/redshift/latest/mgmt/working-with-serverless.html" target="_blank" rel="noopener"&gt;Amazon Redshift Serverless&lt;/a&gt;, and &lt;a href="https://docs.aws.amazon.com/emr/" target="_blank" rel="noopener"&gt;Amazon EMR&lt;/a&gt;.&lt;/p&gt;
&lt;h1&gt;The challenge: Outgrowing a single-engine data platform&lt;/h1&gt;
&lt;p&gt;The Amazon Redshift engine handled a wide variety of workloads, including:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;Heavy ETL processing:&lt;/strong&gt; Raw data ingestion from Amazon Simple Storage Service (Amazon S3) followed by complex aggregation pipelines (daily user-aggregation running once per day with a 3-day lookback, and weekly 10-day-lookback jobs).&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Near-real-time operational monitoring:&lt;/strong&gt; Queries executing every 20 minutes against raw data for system-health dashboards.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Business-intelligence reporting:&lt;/strong&gt; Tableau extracts and live dashboards.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Data-science workloads:&lt;/strong&gt; Exploratory analysis and model-feature engineering.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Ad-hoc analysis:&lt;/strong&gt; Non-recurring queries done by analysts and engineers.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;With business growth, storage grew by orders of magnitude over the past decade as the platform expanded. All these varied workloads competed for the same engine and pushed it to its limits. Jobs experienced increasing queue times, service level agreements (SLAs) were at risk, and adding nodes provided minimal performance gains, creating a need to isolate workloads.&lt;/p&gt;
&lt;h1&gt;Gaining visibility: Measuring workload impact&lt;/h1&gt;
&lt;p&gt;Moovit’s first modernization milestone was to create a trusted measurement foundation before changing any workloads. Instead of treating warehouse activity as a single opaque stream, the team implemented automated query attribution that continuously classified each query by workload owner and execution context. The classification combined multiple signals: who executed the query (user or service account), recognizable query-signature patterns, and metadata emitted by orchestration frameworks and scheduled processes.&lt;/p&gt;
&lt;p&gt;This produced a historical, query-level map of platform usage that answered three critical questions: &lt;em&gt;who&lt;/em&gt; is generating load, &lt;em&gt;what&lt;/em&gt; &lt;em&gt;kind&lt;/em&gt; of workload is running, and &lt;em&gt;how expensive&lt;/em&gt; each workload is in runtime and resource terms. With that baseline in place, the team made offload decisions from evidence rather than assumptions. This approach prioritized the largest and most stable optimization opportunities first and reduced the risk of moving business-critical workloads without visibility.&lt;/p&gt;
&lt;p&gt;These classifications and workload metrics were reflected in a Tableau report that aggregated query activity by classification label and execution context. The view exposed operational dimensions such as classification, time granularity, service class, execution-time bucket, unload flags, and sample-query context, supporting both trend monitoring and root-cause drill-down.&lt;/p&gt;
&lt;p&gt;The worksheet was parameterized to support multiple measurement modes over the same grouped workload population: total execution time, execution plus queue time, total CPU time, average execution time per query, and ratio-based efficiency views (execution/CPU and CPU/execution). This let the team compare “heavy by volume” workloads against “inefficient by behavior” workloads without creating separate artifacts.&lt;/p&gt;
&lt;p&gt;For decision-making, CPU time was used as the primary impact metric because it best represented sustained compute pressure. Execution time, queue time, query-count normalization, and workload-management segmentation were treated as secondary evidence to distinguish:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;compute-heavy but healthy workloads&lt;/li&gt;
 &lt;li&gt;queue-constrained workloads&lt;/li&gt;
 &lt;li&gt;high-frequency/low-cost workloads&lt;/li&gt;
 &lt;li&gt;noisy or weakly classified workloads that required attribution cleanup first&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Using this framework, prioritization became systematic: first improve classification coverage, then rank workloads by CPU contribution, then validate with queue and workload management (WLM) signals, and finally choose the action path per workload (optimize SQL, reschedule, isolate, retire, or move to another engine).&lt;/p&gt;
&lt;p&gt;The following figure shows an example of one of the dashboard widgets (CPU time by query).&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/01/BDB-5979-1.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/01/BDB-5979-1.png" alt="Dashboard widget showing CPU time consumed by each query" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 1: CPU time by query, highlighting the most resource-intensive queries and their usage patterns&lt;/p&gt;
&lt;/div&gt;
&lt;h1&gt;Cleanup: Reducing unnecessary data warehouse load&lt;/h1&gt;
&lt;p&gt;With a long-running data platform, in most cases the workloads will start accumulating, some of which become irrelevant at some point. For example, a report which was created and scheduled, yet it became irrelevant after a few years, but still running since no one disabled it. It’s important to indicate these workloads in general to reduce unnecessary load, yet even more critical before doing any significant architectural changes or migrations. Before migrating any workloads, Moovit first reduced unnecessary warehouse load.&lt;/p&gt;
&lt;p&gt;The team:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;Removed unused processes that were still consuming cluster resources.&lt;/li&gt;
 &lt;li&gt;Reduced unnecessary frequency where possible: some jobs ran more often than downstream consumers needed.&lt;/li&gt;
 &lt;li&gt;Reviewed workload-management guardrails to verify resource allocation matched actual priorities.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This cleanup phase was a prerequisite to migration. By removing waste first, the team verified that the workloads eventually selected for offloading were genuinely heavy rather than simply unoptimized or unnecessary.&lt;/p&gt;
&lt;p&gt;The no-longer-relevant processes consumed around 7 percent of overall CPU time and were removed before the optimization work began.&lt;/p&gt;
&lt;h1&gt;Workload selection: Choosing what to offload&lt;/h1&gt;
&lt;p&gt;With a clear picture of workload patterns, Moovit faced a common decision point: continue scaling the existing Redshift cluster, or re-architect towards a multi-engine approach. The team evaluated two main paths:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;Re-architect with Redshift multi-cluster and data sharing: Identify workloads that could benefit from resource isolation, then redistribute processing and queries between multiple Redshift clusters, combining both serverless and provisioned options. This would redistribute load across use-case-optimized clusters and potentially save costs through better resource use.&lt;/li&gt;
 &lt;li&gt;Re-architect with purpose-built engines: Identify workloads that could benefit from alternative processing frameworks and offload them to more suitable engines. This would reduce pressure on Amazon Redshift while building a more flexible, cost-efficient architecture.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Moovit decided to do both, because while some workloads benefited from being offloaded, others benefited from isolated Amazon Redshift compute.&lt;/p&gt;
&lt;p&gt;The measurement data revealed a primary candidate for offloading: raw-data aggregation pipelines. This workload loaded raw data into Amazon Redshift from Amazon S3, then performed heavy sessionization and aggregation transformations. Raw tables were still used for ad-hoc and exploratory analysis, but recurring production consumers primarily depended on aggregated outputs, making these transformations strong candidates for offloading.&lt;/p&gt;
&lt;h1&gt;Proof of concept: Offloading to EMR Serverless with Spark SQL&lt;/h1&gt;
&lt;p&gt;With target workload identified, Moovit initiated a POC using &lt;a href="https://docs.aws.amazon.com/emr/latest/EMR-Serverless-UserGuide/emr-serverless.html" target="_blank" rel="noopener"&gt;Amazon EMR Serverless&lt;/a&gt; with Spark SQL. The choice of EMR Serverless was driven by several factors:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;Spark SQL compatibility:&lt;/strong&gt; The existing Redshift SQL logic could be ported with minimal changes to Spark SQL syntax.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Serverless simplicity:&lt;/strong&gt; No cluster-management overhead during the evaluation phase.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Data-lake native:&lt;/strong&gt; Processing could occur directly on data in Amazon S3.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The POC defined quantified success criteria measured over five or more consecutive runs:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;Runtime reduction:&lt;/strong&gt; Greater than or equal to 40 percent reduction for the transform portion of selected pipelines.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Amazon Redshift cost reduction:&lt;/strong&gt; Greater than 30 percent reduction in Redshift RA3 compute with no performance degradation for remaining workloads.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Data-quality parity:&lt;/strong&gt; Exact match between Spark and Amazon Redshift outputs on row counts, distinct users, and all published metrics over a frozen parity window.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="overcoming-initial-performance-challenges"&gt;Overcoming initial performance challenges&lt;/h2&gt;
&lt;p&gt;The first POC attempts exposed significant challenges. Early Spark jobs with 100 executors took approximately 4 hours, far exceeding the 30–40-minute baseline on Amazon Redshift. Beyond raw performance, the team encountered memory pressure, data-parity gaps between Spark and Amazon Redshift outputs, and subtle SQL behavior differences between the two engines.&lt;/p&gt;
&lt;p&gt;The team systematically diagnosed and resolved these issues:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;Execution-plan analysis: Reviewing the Spark execution plan revealed suboptimal query patterns that generated excessive data shuffles.&lt;/li&gt;
 &lt;li&gt;Query rewrites: Rewriting specific SQL constructs to align with Spark’s distributed processing model, including splitting large monolithic logic into staged transformations.&lt;/li&gt;
 &lt;li&gt;Reducing or rewriting expensive DISTINCT patterns: Identifying and eliminating unnecessary DISTINCT operations that created heavy shuffle pressure.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;After applying these optimizations, execution time dropped from 4 hours to approximately 10 minutes, and the required executors dropped to fewer than 50, surpassing the original performance.&lt;/p&gt;
&lt;h1&gt;Validation: Ensuring data parity before cutover&lt;/h1&gt;
&lt;p&gt;Before transitioning any workload to production, Moovit implemented a rigorous validation process. The new Spark output was compared with the previous Amazon Redshift output using multiple dimensions:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;Row counts: ensuring no data was lost or duplicated.&lt;/li&gt;
 &lt;li&gt;Distinct users: verifying entity-level completeness.&lt;/li&gt;
 &lt;li&gt;Metric parity: all published business metrics matched.&lt;/li&gt;
 &lt;li&gt;Daily trends: time-series patterns remained consistent.&lt;/li&gt;
 &lt;li&gt;Row-level checks: spot-checking individual records for correctness.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Only after all validation checks passed consistently over multiple consecutive runs did the team proceed with cutover for each workload.&lt;/p&gt;
&lt;h1&gt;Moving to production: Expanding workload offloading&lt;/h1&gt;
&lt;p&gt;With a successful POC demonstrating both performance gains and cost savings, Moovit progressively moved additional workloads from Amazon Redshift to EMR:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;Heavy-aggregation jobs:&lt;/strong&gt; The primary daily and weekly aggregation pipelines transitioned fully to EMR.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Data-transformation stages:&lt;/strong&gt; Preprocessing steps that previously consumed Redshift compute moved to Spark, with only final aggregated results loaded back into Amazon Redshift for BI consumption.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Weekly batch workloads:&lt;/strong&gt; Large batch jobs that previously created resource contention during weekend processing windows.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The transition used a measured approach: each workload was migrated individually, with data-quality validation confirming parity before decommissioning the equivalent jobs which were running on Redshift.&lt;/p&gt;
&lt;h1&gt;Additional optimizations: Redshift Serverless, workload isolation, and Amazon EMR on Amazon EC2&lt;/h1&gt;
&lt;p&gt;Beyond EMR offloading, Moovit implemented further architectural improvements to isolate workloads and optimize costs.&lt;/p&gt;
&lt;h2 id="amazon-redshift-rightsizing-iterative-cluster-optimization"&gt;Amazon Redshift rightsizing: Iterative cluster optimization&lt;/h2&gt;
&lt;p&gt;With heavy workloads successfully offloaded and isolated, Moovit proceeded to right-size the Redshift cluster. Rather than a single resize, the team reduced the cluster incrementally, two nodes at a time, using &lt;a href="https://docs.aws.amazon.com/redshift/latest/mgmt/resizing-cluster.html#elastic-resize" target="_blank" rel="noopener"&gt;elastic resize&lt;/a&gt;. At each step, they validated that:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;Existing BI workloads maintained acceptable performance.&lt;/li&gt;
 &lt;li&gt;Queue wait times remained within SLA thresholds.&lt;/li&gt;
 &lt;li&gt;No workload degradation was observed under peak loads.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This iterative approach minimized risk and allowed the team to find the optimal cluster size with confidence.&lt;/p&gt;
&lt;h2 id="workload-isolation-with-redshift-serverless"&gt;Workload isolation with Redshift Serverless&lt;/h2&gt;
&lt;p&gt;Amazon Redshift persisted as the engine of choice for serving curated BI data. However, not all Amazon Redshift workloads needed provisioned capacity:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;Ad-hoc analyst queries:&lt;/strong&gt; Moved to Redshift Serverless, isolating unpredictable workloads from the provisioned cluster through data sharing.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Data-science workloads:&lt;/strong&gt; Transitioned to Redshift Serverless for flexible exploration without impacting production.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This workload isolation through Redshift Serverless provided resource separation without requiring additional provisioned capacity. The architecture now used data sharing to provide a unified view across provisioned and serverless clusters.&lt;/p&gt;
&lt;h2 id="operational-isolation-refinements"&gt;Operational isolation refinements&lt;/h2&gt;
&lt;p&gt;Moovit also refined workload isolation by rebalancing &lt;a href="https://docs.aws.amazon.com/redshift/latest/dg/c_workload_mngmt_classification.html" target="_blank" rel="noopener"&gt;WLM&lt;/a&gt; priorities on the provisioned cluster. Because the ETL queue mainly handled raw data loading from Amazon S3 (which was not the bottleneck after heavy aggregations moved to Spark), its priority was reduced. At the same time, with most human users moved to Redshift Serverless, Tableau serving workloads on provisioned Redshift were prioritized higher to keep dashboard performance predictable. The final result: a 50% reduction in provisioned Redshift capacity.&lt;/p&gt;
&lt;h2 id="transitioning-to-emr-on-ec2"&gt;Transitioning to EMR on EC2&lt;/h2&gt;
&lt;p&gt;EMR Serverless proved efficient for the POC phase: it allowed fast iteration without cluster management overhead. However, for longer-term recurring production workloads, Moovit moved to &lt;a href="https://docs.aws.amazon.com/emr/latest/ManagementGuide/emr-what-is-emr.html" target="_blank" rel="noopener"&gt;EMR on EC2&lt;/a&gt; to better fit their production cost and infrastructure model, using existing compute reservations.&lt;/p&gt;
&lt;p&gt;The transition between EMR deployment options required zero application code changes, demonstrating the flexibility of the EMR deployment options.&lt;/p&gt;
&lt;h2 id="ai-assisted-sql-translation"&gt;AI-assisted SQL translation&lt;/h2&gt;
&lt;p&gt;Additionally, Moovit used AI-assisted development tools, Claude Code and Cursor, to accelerate parts of the SQL transition process. These tools helped engineers identify Redshift SQL and Spark SQL syntax differences, suggest rewrites, and debug migration issues, while validation and production approval remained under engineer review.&lt;/p&gt;
&lt;h1&gt;Results: A modern multi-engine architecture&lt;/h1&gt;
&lt;p&gt;The architectural modernization delivered measurable outcomes:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;Cluster size reduction:&lt;/strong&gt; Redshift cluster size reduced to 50 percent of the initial capacity.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Performance improvement:&lt;/strong&gt; Key aggregation jobs ran faster and more consistently on EMR (50 percent execution time reduction for p90).&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Workload isolation:&lt;/strong&gt; No single workload type could impact others through resource contention.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;33 percent overall data pipeline cost reduction:&lt;/strong&gt; Combined savings from cluster reduction, transition to EMR, and efficient serverless usage.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Future flexibility:&lt;/strong&gt; The multi-engine architecture provided pathways for additional use cases without architectural changes.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The following figures compare aggregation-job performance before and after the transition.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/01/BDB-5979-2.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/01/BDB-5979-2.png" alt="Chart comparing aggregation-job execution times before and after the transition, with longer, inconsistent runtimes before and shorter, stable runtimes after" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 2: Aggregation-job execution times before and after the transition&lt;/p&gt;
&lt;/div&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/01/BDB-5979-3.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/01/BDB-5979-3.png" alt="Chart comparing wall-clock time for job executions across percentiles, with p90 at 5.48 hours before the transition and 2.77 hours after" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 3: Wall-clock time for job executions by percentile, before and after the transition&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;The resulting architecture assigned each workload to the engine that fits it best:&lt;/p&gt;
&lt;table border="1px" width="100%" cellpadding="10px"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;Workload type&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Engine&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Rationale&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Heavy ETL and aggregation&lt;/td&gt;
   &lt;td&gt;Amazon EMR (Spark SQL)&lt;/td&gt;
   &lt;td&gt;Distributed processing on Amazon S3. No data warehouse load required&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Ongoing processing and BI reporting&lt;/td&gt;
   &lt;td&gt;Amazon Redshift provisioned&lt;/td&gt;
   &lt;td&gt;24/7 running processes&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Ad-hoc queries&lt;/td&gt;
   &lt;td&gt;Amazon Redshift Serverless&lt;/td&gt;
   &lt;td&gt;Burst capacity with workload isolation&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Data science&lt;/td&gt;
   &lt;td&gt;Amazon Redshift Serverless&lt;/td&gt;
   &lt;td&gt;Flexible exploration without impacting production&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;h1&gt;Lessons learned&lt;/h1&gt;
&lt;p&gt;The Moovit modernization journey produced several key insights applicable to similar architectural transitions:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;Measure before you move: Establishing baseline metrics and automated classification was essential for identifying true offloading candidates. Without granular workload-level measurements, the team would not have identified which specific processes were exhausting the cluster.&lt;/li&gt;
 &lt;li&gt;Clean up before you migrate: Reducing unnecessary load first verified that migration efforts targeted genuinely heavy workloads rather than simply unoptimized or unused processes.&lt;/li&gt;
 &lt;li&gt;Small SQL changes, big impact: Moving from Redshift SQL to Spark SQL required relatively minor syntax adjustments. The core business logic remained intact, and most transformations translated directly with minimal refactoring.&lt;/li&gt;
 &lt;li&gt;Optimize for the engine: Porting SQL queries to Spark without optimization produced initially poor results for some workloads. Understanding Spark’s distributed execution model and optimizing for it was critical for achieving target performance.&lt;/li&gt;
 &lt;li&gt;Validate rigorously: Multi-dimensional data-parity checks (row counts, distinct users, metrics, daily trends, and row-level spot checks) gave the team confidence to cut over without data-quality regressions.&lt;/li&gt;
 &lt;li&gt;Moving between EMR options is straightforward: EMR Serverless proved very efficient for starting fast and evaluating Spark. When Moovit needed to move to EMR on EC2 to use existing reservations, the transition required no application code changes.&lt;/li&gt;
 &lt;li&gt;Iterative cluster rightsizing: Rather than a single resize, Moovit reduced the Redshift cluster incrementally (two nodes at a time) using elastic resize, validating performance at each step before proceeding further.&lt;/li&gt;
&lt;/ol&gt;
&lt;h1&gt;Conclusion&lt;/h1&gt;
&lt;p&gt;Looking ahead, as another potential optimization, Moovit will be evaluating the new &lt;a href="https://aws.amazon.com/redshift/features/rg/" target="_blank" rel="noopener"&gt;Amazon Redshift RG&lt;/a&gt; instances for provisioned clusters, providing up to 2.2x better price performance and priced 30% lower than RA3, powered by AWS Graviton.&lt;/p&gt;
&lt;p&gt;The broader takeaway is that AWS provides multiple purpose-built engines that can be used in a single data platform. In Moovit’s case, the biggest improvement came from assigning each workload to the engine that fit it best: &lt;a href="https://docs.aws.amazon.com/redshift/latest/mgmt/welcome.html" target="_blank" rel="noopener"&gt;Amazon Redshift&lt;/a&gt; for curated analytical serving, Redshift Serverless for isolated exploratory workloads, and &lt;a href="https://docs.aws.amazon.com/emr/latest/ManagementGuide/emr-what-is-emr.html" target="_blank" rel="noopener"&gt;Amazon EMR&lt;/a&gt; for large-scale transformations over data in Amazon S3. This architecture gives Moovit a foundation for future optimization and flexibility as data volumes grow and new analytical use cases emerge.&lt;/p&gt;
&lt;p&gt;&amp;nbsp;&lt;/p&gt;
&lt;p style="clear: both"&gt;&lt;/p&gt;
&lt;hr style="width: 100%"&gt;
&lt;h2&gt;About the authors&lt;/h2&gt;
&lt;footer&gt;
 &lt;div class="blog-author-box"&gt;
  &lt;div class="blog-author-image"&gt;
   &lt;img loading="lazy" class="alignleft size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/01/BDB-5979-4.jpeg" alt="Saar Porat" width="100" height="100"&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Saar Porat&lt;/h3&gt;
  &lt;p&gt;Saar is the Director of BI &amp;amp; Data Engineering at Moovit, where he has spent more than a decade building and scaling the company’s data engineering capabilities. With nearly 20 years of experience in BI, analytics, and data platforms, he focuses on designing reliable, maintainable, and cost-efficient systems that translate complex data into meaningful business impact. Saar led Moovit’s initiative to migrate major workloads from Amazon Redshift to Apache Spark, improving scalability, performance, and infrastructure efficiency while expanding the team’s engineering capabilities beyond SQL-based processing.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box"&gt;
  &lt;div class="blog-author-image"&gt;
   &lt;img loading="lazy" class="alignleft size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/01/BDB-5979-5.jpeg" alt="Vova Nevski" width="100" height="100"&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Vova Nevski&lt;/h3&gt;
  &lt;p&gt;Vova is a Senior Analytics Specialist Solutions Architect at AWS with more than 15 years of experience in the big data and analytics domain, including data lakes, batch and stream processing, both on premises and in the cloud. He partners with AWS customers to design and build solutions best suited to their unique needs.&lt;/p&gt;
 &lt;/div&gt;
&lt;/footer&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>Query Amazon S3 Tables from Amazon EMR Trino using the Iceberg REST endpoint</title>
		<link>https://aws.amazon.com/blogs/big-data/query-amazon-s3-tables-from-amazon-emr-trino-using-the-iceberg-rest-endpoint/</link>
		
		<dc:creator><![CDATA[Shubham Purwar]]></dc:creator>
		<pubDate>Wed, 02 Sep 2026 19:01:30 +0000</pubDate>
				<category><![CDATA[Advanced (300)]]></category>
		<category><![CDATA[Amazon EMR]]></category>
		<category><![CDATA[Amazon S3 Tables]]></category>
		<category><![CDATA[Technical How-to]]></category>
		<guid isPermaLink="false">f19d9c081172fd0038719ea9efb29d89f7624d00</guid>

					<description>Learn how to query Amazon S3 Tables from Trino on Amazon EMR using the Apache Iceberg REST catalog endpoint. This post shows how to deploy the integration with AWS CloudFormation, configure the Trino catalog, and run SQL to create, query, and manage Apache Iceberg tables.</description>
										<content:encoded>&lt;p&gt;Organizations running analytics on Amazon Simple Storage Service (Amazon S3) data lakes often struggle with the operational overhead of managing Apache Iceberg tables, including compaction, snapshot expiration, and metadata tracking, while still needing fast, interactive SQL access across large volumes of data. &lt;a href="https://aws.amazon.com/s3/features/tables/" target="_blank" rel="noopener"&gt;Amazon S3 Tables&lt;/a&gt;, a capability of Amazon S3, addresses this by providing a purpose-built storage layer with native Apache Iceberg support and automated table maintenance. When you query S3 Tables from &lt;a href="https://aws.amazon.com/emr/" target="_blank" rel="noopener"&gt;Amazon EMR&lt;/a&gt; using Trino and the Iceberg REST endpoint, you get a fully managed, open-standards-based analytics stack without the undifferentiated heavy lifting of table upkeep.&lt;/p&gt;
&lt;p&gt;When paired with Amazon EMR running Trino, organizations gain access to a high-performance distributed SQL query engine capable of processing large-scale datasets. Trino’s ability to query data across multiple sources, combined with the automated optimization features of S3 Tables, creates a flexible analytics platform. The integration uses Apache Iceberg’s REST catalog specification, providing a standardized interface that supports compatibility across different compute engines while maintaining full control over query execution and data processing logic.&lt;/p&gt;
&lt;p&gt;This architectural pattern is particularly valuable for organizations seeking to modernize their data platforms without vendor lock-in, as it relies on open standards and formats. The solution delivers high-throughput query performance with distributed SQL execution while significantly reducing the operational burden of managing table metadata, compaction, and snapshot lifecycle management. In this post, we show you how to create and query Amazon S3 Tables using Trino on Amazon EMR through the Apache Iceberg REST catalog endpoint.&lt;/p&gt;
&lt;h2 id="solution-overview"&gt;Solution overview&lt;/h2&gt;
&lt;p&gt;This implementation demonstrates a complete integration between the Trino distribution on Amazon EMR and Amazon S3 Tables through the Apache Iceberg REST catalog endpoint. The architecture uses several key AWS services working in concert:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Amazon EMR&lt;/strong&gt; serves as the managed compute layer, providing a scalable Hadoop framework that hosts the Trino query engine. Amazon EMR handles cluster provisioning, configuration management, and automatic scaling, allowing teams to focus on analytics rather than infrastructure management.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://docs.aws.amazon.com/emr/latest/ReleaseGuide/emr-trino.html" target="_blank" rel="noopener"&gt;&lt;strong&gt;Apache Trino&lt;/strong&gt;&lt;/a&gt; acts as the distributed SQL query engine, offering ANSI SQL compatibility and the ability to process queries across massive datasets with low latency for interactive workloads. Its connector architecture supports integration with various data sources, including the Iceberg REST catalog.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Amazon S3 Tables&lt;/strong&gt; provides the storage and catalog layer, managing Apache Iceberg tables with built-in optimization. The service automatically handles compaction, snapshot expiration, and metadata management, reducing operational overhead while maintaining query performance. S3 Tables exposes a REST API endpoint that conforms to the Apache Iceberg REST catalog specification, which provides standardized integration with any Iceberg-compatible engine.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Apache Iceberg REST endpoint&lt;/strong&gt; serves as the communication protocol between Trino and S3 Tables. This RESTful interface handles catalog operations including namespace management, table creation, metadata retrieval, and transaction coordination. The endpoint supports AWS Signature Version 4 authentication for secure access to table resources.&lt;/p&gt;
&lt;p&gt;The data flow follows this pattern: Users submit SQL queries through the Trino CLI or JDBC interface. Trino’s Iceberg connector communicates with the S3 Tables REST endpoint to retrieve table metadata and plan query execution. The query engine then reads data directly from S3 using optimized file formats (Parquet, ORC) while using Iceberg’s metadata layer for partition pruning and predicate pushdown. Write operations follow a similar path, with Trino coordinating with S3 Tables to commit new data files and update table metadata atomically.&lt;/p&gt;
&lt;p&gt;This architecture delivers several key benefits: separation of compute and storage for independent scaling, automated table maintenance reducing operational costs, open-source format compatibility preventing vendor lock-in, and fine-grained access control through AWS Identity and Access Management (IAM) and AWS Lake Formation integration.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/01/BDB-5695-1.jpg" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/01/BDB-5695-1.jpg" alt="Architecture diagram showing Trino on Amazon EMR querying Amazon S3 Tables through the Apache Iceberg REST catalog endpoint" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 1: Solution architecture for querying Amazon S3 Tables from Trino on Amazon EMR&lt;/p&gt;
&lt;/div&gt;
&lt;h2 id="prerequisites"&gt;Prerequisites&lt;/h2&gt;
&lt;p&gt;Before getting started, make sure that you have the following:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;An active AWS account with billing enabled.&lt;/li&gt;
 &lt;li&gt;An &lt;a href="https://aws.amazon.com/iam/" target="_blank" rel="noopener"&gt;AWS Identity and Access Management&lt;/a&gt; (IAM) user with specific permissions to create and manage resources, such as a virtual private cloud (VPC), subnet, security group, IAM roles, Amazon EMR, Interface VPC endpoints, S3 Tables bucket and S3 buckets.&lt;/li&gt;
 &lt;li&gt;Sufficient VPC capacity in your chosen &lt;a href="https://docs.aws.amazon.com/glossary/latest/reference/glos-chap.html#region" target="_blank" rel="noopener"&gt;AWS Region&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For this post, we create the solution resources in the US East (N. Virginia) Region (&lt;code&gt;us-east-1&lt;/code&gt;) using &lt;a href="http://aws.amazon.com/cloudformation" target="_blank" rel="noopener"&gt;AWS CloudFormation&lt;/a&gt; templates. In the following sections, we show you how to configure your resources and implement the solution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; Querying Amazon S3 Tables through Trino on Amazon EMR requires Trino version 475 or later, available in Amazon EMR 7.11 and later.&lt;/p&gt;
&lt;h2 id="part-a-configure-amazon-s3-tables-integration-with-trino-on-amazon-emr-using-aws-cloudformation"&gt;Part A: Configure Amazon S3 Tables integration with Trino on Amazon EMR using AWS CloudFormation&lt;/h2&gt;
&lt;p&gt;In this post, you use the CloudFormation template &lt;code&gt;emr-trino-s3tables.yaml&lt;/code&gt;.&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;This template deploys the following resources: a VPC with one private subnet, an S3 Tables interface VPC endpoint for private access, and an Amazon EMR cluster running Trino integrated with Amazon S3 Tables through the Apache Iceberg REST catalog endpoint.&lt;/li&gt;
 &lt;li&gt;It also creates an S3 Tables bucket, a general-purpose S3 bucket, IAM roles, and security groups.&lt;/li&gt;
 &lt;li&gt;At deploy time, it dynamically generates the Trino catalog configuration and bootstrap script.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;To create the solution resources, complete the following steps:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;Launch the stack &lt;code&gt;emr-trino-s3tables.yaml&lt;/code&gt; using the CloudFormation template.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;a href="https://us-east-1.console.aws.amazon.com/cloudformation/home?region=us-east-1#/stacks/create?templateURL=https://aws-blogs-artifacts-public.s3.us-east-1.amazonaws.com/artifacts/BDB-5260/trinos3tables/emr-trino-s3tables.yaml" target="_blank" rel="noopener"&gt;&lt;img loading="lazy" class="alignnone wp-image-54461 size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2023/09/29/cloudformation-launch-stack.png" alt="Launch Cloudformation Stack" width="144" height="27"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;ol type="2"&gt;
 &lt;li&gt;Provide the parameter values as listed in the following table.&lt;/li&gt;
&lt;/ol&gt;
&lt;table border="1px" width="100%" cellpadding="10px"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;Parameters&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Description&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Sample value&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Stack Name&lt;/td&gt;
   &lt;td&gt;Name of CloudFormation stack&lt;/td&gt;
   &lt;td&gt;emr-s3tables-trino&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;VPC CIDR block&lt;/td&gt;
   &lt;td&gt;IP range (CIDR notation) for this VPC.&lt;/td&gt;
   &lt;td&gt;10.0.0.0/16&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Private Subnet CIDR block&lt;/td&gt;
   &lt;td&gt;IP range (CIDR notation) for the private subnet in the second Availability Zone.&lt;/td&gt;
   &lt;td&gt;10.0.1.0/24&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Resource name Prefix&lt;/td&gt;
   &lt;td&gt;Short prefix applied to every resource name&lt;/td&gt;
   &lt;td&gt;emr-s3tables&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;S3 Tables bucket name&lt;/td&gt;
   &lt;td&gt;Name of S3 table Bucket&lt;/td&gt;
   &lt;td&gt;trinoemrs3tablebuck&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;EMR release&lt;/td&gt;
   &lt;td&gt;Release version of Amazon EMR&lt;/td&gt;
   &lt;td&gt;EMR 7.12&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The stack creation process can take approximately 15 minutes to complete. You can check the Outputs tab for the stack after the stack is created, as shown in the following screenshot.&lt;/p&gt;
&lt;div class="mceTemp"&gt;
 Figure 3: CloudFormation stack outputs
 &lt;p&gt;&lt;/p&gt;
 &lt;div id="attachment_94065" style="width: 881px" class="wp-caption alignnone"&gt;
  &lt;img aria-describedby="caption-attachment-94065" loading="lazy" class="size-full wp-image-94065" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/02/Emr1.jpg" alt="" width="871" height="822"&gt;
  &lt;p id="caption-attachment-94065" class="wp-caption-text"&gt;Figure 3: CloudFormation stack outputs&lt;/p&gt;
 &lt;/div&gt;
 &lt;h3 id="understanding-the-deployment"&gt;Understanding the deployment&lt;/h3&gt;
 &lt;p&gt;The CloudFormation template performs several key tasks:&lt;/p&gt;
 &lt;ol type="1"&gt;
  &lt;li&gt;&lt;strong&gt;Infrastructure provisioning&lt;/strong&gt;: Sets up the Amazon EMR cluster with Trino, VPC, subnet, security group, and S3 table bucket.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Configuration&lt;/strong&gt;: Creates necessary Trino configuration files.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Integration&lt;/strong&gt; configuration: Sets up the Iceberg REST connector for S3 Tables.&lt;/li&gt;
 &lt;/ol&gt;
 &lt;h2 id="part-b-connecting-trino-to-amazon-s3-tables-with-iceberg-rest-endpoint"&gt;Part B: Connecting Trino to Amazon S3 Tables with Iceberg REST endpoint&lt;/h2&gt;
 &lt;p&gt;The CloudFormation template automatically configures the S3 Tables catalog in Trino on Amazon EMR. In the next section, we examine the configuration that drives this integration.&lt;/p&gt;
 &lt;h3 id="catalog-configuration-details"&gt;1. Catalog configuration details&lt;/h3&gt;
 &lt;p&gt;A catalog in Trino on Amazon EMR is the configuration that grants access to a specific data source. Each Trino on Amazon EMR cluster can have multiple catalogs configured, allowing access to different data sources simultaneously.&lt;/p&gt;
 &lt;p&gt;As part of this setup, the CloudFormation template creates a catalog properties file at &lt;code&gt;/etc/trino/conf/catalog/s3tables_irc.properties&lt;/code&gt; with the following configuration:&lt;/p&gt;
 &lt;div class="hide-language"&gt;
  &lt;pre&gt;&lt;code class="language-properties"&gt;connector.name=iceberg
iceberg.catalog.type=rest
iceberg.rest-catalog.uri=https://s3tables.&amp;lt;REGION&amp;gt;.amazonaws.com/iceberg
iceberg.rest-catalog.warehouse=arn:aws:s3tables:AwsRegion:&amp;lt;ACCOUNT-ID&amp;gt;:bucket/&amp;lt;BUCKET-NAME&amp;gt;
iceberg.rest-catalog.sigv4-enabled=true
iceberg.rest-catalog.signing-name=s3tables
iceberg.rest-catalog.view-endpoints-enabled=false
fs.hadoop.enabled=false
fs.native-s3.enabled=true
s3.region=us-east-1
s3.iam-role=arn:aws:iam::&amp;lt;ACCOUNT-ID&amp;gt;:role/service-role/&amp;lt;ROLE-NAME&amp;gt;&lt;/code&gt;&lt;/pre&gt;
 &lt;/div&gt;
 &lt;h3 id="s3-tables-iceberg-rest-endpoint-configuration-properties"&gt;2. S3 Tables Iceberg REST endpoint configuration properties&lt;/h3&gt;
 &lt;p&gt;The following table lists the key properties in the catalog configuration on Trino:&lt;/p&gt;
 &lt;table border="1px" width="100%" cellpadding="10px"&gt;
  &lt;tbody&gt;
   &lt;tr&gt;
    &lt;td&gt;&lt;strong&gt;Property name&lt;/strong&gt;&lt;/td&gt;
    &lt;td&gt;&lt;strong&gt;Description&lt;/strong&gt;&lt;/td&gt;
   &lt;/tr&gt;
   &lt;tr&gt;
    &lt;td&gt;iceberg.rest-catalog.uri&lt;/td&gt;
    &lt;td&gt;REST server API endpoint URI (necessary).&lt;/td&gt;
   &lt;/tr&gt;
   &lt;tr&gt;
    &lt;td&gt;iceberg.rest-catalog.warehouse&lt;/td&gt;
    &lt;td&gt;Warehouse ID or location for the catalog (necessary). For S3 Tables, this is the ARN for the S3 table bucket as shown in the preceding properties example.&lt;/td&gt;
   &lt;/tr&gt;
   &lt;tr&gt;
    &lt;td&gt;iceberg.rest-catalog.sigv4-enabled&lt;/td&gt;
    &lt;td&gt;Must be set to ‘true’ (necessary)&lt;/td&gt;
   &lt;/tr&gt;
   &lt;tr&gt;
    &lt;td&gt;iceberg.rest-catalog.signing-name&lt;/td&gt;
    &lt;td&gt;Must be set to ‘s3tables’ (necessary)&lt;/td&gt;
   &lt;/tr&gt;
   &lt;tr&gt;
    &lt;td&gt;iceberg.rest-catalog.view-endpoints-enabled&lt;/td&gt;
    &lt;td&gt;Must be set to ‘false’ (necessary)&lt;/td&gt;
   &lt;/tr&gt;
   &lt;tr&gt;
    &lt;td&gt;fs.hadoop.enabled&lt;/td&gt;
    &lt;td&gt;Must be set to ‘false’&lt;/td&gt;
   &lt;/tr&gt;
   &lt;tr&gt;
    &lt;td&gt;fs.native-s3.enabled&lt;/td&gt;
    &lt;td&gt;Must be set to ‘true’&lt;/td&gt;
   &lt;/tr&gt;
   &lt;tr&gt;
    &lt;td&gt;s3.iam-role&lt;/td&gt;
    &lt;td&gt;&lt;a href="https://docs.aws.amazon.com/IAM/latest/UserGuide/reference-arns.html" target="_blank" rel="noopener"&gt;Amazon Resource Name (ARN)&lt;/a&gt; of the IAM role with permissions to S3 Tables. In this post, we use the same role, which is the service role for Amazon EMR.&lt;/td&gt;
   &lt;/tr&gt;
   &lt;tr&gt;
    &lt;td&gt;s3.region&lt;/td&gt;
    &lt;td&gt;AWS Region, for example us-east-1&lt;/td&gt;
   &lt;/tr&gt;
  &lt;/tbody&gt;
 &lt;/table&gt;
 &lt;p&gt;This configuration establishes a connection between Trino and the S3 Tables REST endpoint. You can have multiple catalogs registered, one per S3 table bucket, which is determined by the &lt;code&gt;iceberg.rest-catalog.warehouse&lt;/code&gt; property.&lt;/p&gt;
 &lt;h3 id="configure-amazon-emr-service-iam-role-trust-relationships"&gt;3. Configure Amazon EMR service IAM role trust relationships&lt;/h3&gt;
 &lt;p&gt;The Amazon EMR service role requires proper trust relationships to function correctly. Navigate to the IAM console and configure the trust policy for your Amazon EMR service role:&lt;/p&gt;
 &lt;div class="hide-language"&gt;
  &lt;pre&gt;&lt;code class="language-json"&gt;{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Effect": "Allow",
            "Principal": {
                "Service": "elasticmapreduce.amazonaws.com"
            },
            "Action": "sts:AssumeRole"
        },
        {
            "Effect": "Allow",
            "Principal": {
                "AWS": "arn:aws:iam::&amp;lt;ACCOUNT-ID&amp;gt;:role/service-role/AmazonEMR-InstanceProfile"
            },
            "Action": "sts:AssumeRole"
        }
    ]
}&lt;/code&gt;&lt;/pre&gt;
 &lt;/div&gt;
 &lt;p&gt;This trust policy establishes two critical relationships:&lt;/p&gt;
 &lt;ol type="1"&gt;
  &lt;li&gt;The Amazon EMR service can assume the role to manage cluster operations.&lt;/li&gt;
  &lt;li&gt;The EC2 instance profile can assume the role to access S3 Tables with elevated permissions.&lt;/li&gt;
 &lt;/ol&gt;
 &lt;h3 id="working-with-s3-tables-in-trino-on-amazon-emr"&gt;4. Working with S3 Tables in Trino on Amazon EMR&lt;/h3&gt;
 &lt;p&gt;Now that you have Trino on Amazon EMR set up and configured to work with S3 Tables, you can explore how to work with this integration.&lt;/p&gt;
 &lt;h3 id="connecting-to-trino-on-amazon-emr"&gt;4.1. Connecting to Trino on Amazon EMR&lt;/h3&gt;
 &lt;p&gt;Navigate to &lt;a href="https://us-east-1.console.aws.amazon.com/emr/" target="_blank" rel="noopener"&gt;Amazon EMR&lt;/a&gt; and select Connect to the primary node using AWS Systems Manager Session Manager for passwordless SSH.&lt;/p&gt;
 &lt;div id="attachment_94064" style="width: 1665px" class="wp-caption alignleft"&gt;
  &lt;img aria-describedby="caption-attachment-94064" loading="lazy" class="size-full wp-image-94064" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/02/EMR2.jpg" alt="" width="1655" height="411"&gt;
  &lt;p id="caption-attachment-94064" class="wp-caption-text"&gt;Figure 4: Connecting to the primary node with Session Manager&lt;/p&gt;
 &lt;/div&gt;
 &lt;p&gt;When you’re connected, you can use the Trino CLI with your S3 Tables catalog:&lt;/p&gt;
 &lt;div class="hide-language"&gt;
  &lt;pre&gt;&lt;code class="language-bash"&gt;sudo su - hadoop
trino-cli --catalog s3tables_irc&lt;/code&gt;&lt;/pre&gt;
 &lt;/div&gt;
 &lt;p&gt;This connects you to the Trino on Amazon EMR using the S3 Tables integration you configured.&lt;/p&gt;
 &lt;div style="width: 589px" class="wp-caption alignnone"&gt;
  &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/01/BDB-5695-5.jpg" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/01/BDB-5695-5.jpg" alt="Trino CLI connected to the s3tables_irc catalog on Amazon EMR" width="579"&gt;&lt;/a&gt;
  &lt;p class="wp-caption-text"&gt;Figure 5: Trino CLI connected to the S3 Tables catalog&lt;/p&gt;
 &lt;/div&gt;
 &lt;h3 id="examples-creating-and-querying-tables"&gt;4.2. Examples: Creating and querying tables&lt;/h3&gt;
 &lt;p&gt;In this section you run through some example queries to demonstrate the functionality.&lt;/p&gt;
 &lt;h3 id="creating-a-namespace"&gt;4.2.1 Creating a namespace&lt;/h3&gt;
 &lt;p&gt;First, you create a namespace (schema) in S3 Tables. A namespace in S3 Tables is a logical container or organizational unit that helps group related tables and objects together.&lt;/p&gt;
 &lt;div class="hide-language"&gt;
  &lt;pre&gt;&lt;code class="language-sql"&gt;CREATE SCHEMA blog_namespace;
USE blog_namespace;&lt;/code&gt;&lt;/pre&gt;
 &lt;/div&gt;
 &lt;h3 id="creating-a-table"&gt;4.2.2 Creating a table&lt;/h3&gt;
 &lt;p&gt;Create a table with various data types. You don’t need to specify the table type as Iceberg explicitly because you’re connecting to the Iceberg catalog. You can use all standard Iceberg capabilities, such as partitioning and sorting. Furthermore, some of the important Iceberg table properties that support table maintenance operations are configured with &lt;a href="https://docs.aws.amazon.com/AmazonS3/latest/userguide/s3-tables-considerations.html#s3-tables-maintenance-limits" target="_blank" rel="noopener"&gt;default values&lt;/a&gt;. You also have the option to edit the configurations using &lt;a href="https://docs.aws.amazon.com/AmazonS3/latest/userguide/s3-tables-maintenance.html" target="_blank" rel="noopener"&gt;S3 Tables maintenance APIs&lt;/a&gt;.&lt;/p&gt;
 &lt;div class="hide-language"&gt;
  &lt;pre&gt;&lt;code class="language-sql"&gt;CREATE TABLE IF NOT EXISTS customers (
customer_sk INT,
customer_id VARCHAR,
salutation VARCHAR,
first_name VARCHAR,
last_name VARCHAR,
preferred_cust_flag VARCHAR,
birth_day INT,
birth_month INT,
birth_year INT,
birth_country VARCHAR,
login VARCHAR
) WITH (
format = 'PARQUET',
sorted_by = ARRAY['customer_id']
);&lt;/code&gt;&lt;/pre&gt;
 &lt;/div&gt;
 &lt;p&gt;&lt;strong&gt;Table property explanation:&lt;/strong&gt;&lt;/p&gt;
 &lt;ul&gt;
  &lt;li&gt;&lt;code&gt;format = 'PARQUET'&lt;/code&gt;: Specifies Parquet as the file format for optimal compression and query performance.&lt;/li&gt;
  &lt;li&gt;&lt;code&gt;sorted_by = ARRAY['customer_id']&lt;/code&gt;: Defines sort order within data files, improving query performance for &lt;code&gt;customer_id&lt;/code&gt; filters.&lt;/li&gt;
 &lt;/ul&gt;
 &lt;p&gt;Verify the table creation:&lt;/p&gt;
 &lt;div class="hide-language"&gt;
  &lt;pre&gt;&lt;code class="language-sql"&gt;SHOW TABLES;&lt;/code&gt;&lt;/pre&gt;
 &lt;/div&gt;
 &lt;p&gt;You should see &lt;code&gt;customers&lt;/code&gt; in the output, confirming the table exists in the S3 Tables catalog.&lt;/p&gt;
 &lt;h3 id="inserting-data"&gt;4.2.3 Inserting data&lt;/h3&gt;
 &lt;p&gt;You can insert some sample data into your table. You can also use an existing table in any of the catalogs configured in Trino on Amazon EMR to read data and write into the S3 table with an &lt;code&gt;INSERT INTO ... SELECT&lt;/code&gt; statement.&lt;/p&gt;
 &lt;div class="hide-language"&gt;
  &lt;pre&gt;&lt;code class="language-sql"&gt;INSERT INTO customers VALUES
(1, 'AAAAA', 'Mrs', 'Martha', 'Rivera', 'Y', 8, 4, 1984, 'US', 'mrivera'),
(2, 'AAAAB', 'Mr', 'Mateo', 'Jackson', 'N', 22, 6, 2001, 'US', 'mjackson'),
(3, 'BAAAA', 'Ms', 'Mary', 'Major', 'Y', 16, 2, 1999, 'US', 'mmajor'),
(4, 'BBAAA', 'Mr', 'Paulo', 'Santos', 'N', 30, 3, 1973, 'US', 'psantos'),
(5, 'AACAA', 'Ms', 'Ana', 'Silva', 'N', 2, 6, 1982, 'CA', 'asilva'),
(6, 'ABAAA', 'Mr', 'Alejandro', 'Rosalez', 'N', 5, 12, 1988, 'US', 'arosalez'),
(7, 'BBAAA', 'Ms', 'Nikki', 'Wolf', 'N', 6, 1, 2006, 'MX', 'nwolf'),
(8, 'ACAAA', 'Mr', 'Arnav', 'Desai', 'N', 15, 7, 1976, 'US', 'adesai');&lt;/code&gt;&lt;/pre&gt;
 &lt;/div&gt;
 &lt;p&gt;This INSERT operation demonstrates Trino’s ability to write data to S3 Tables. Behind the scenes, Trino:&lt;/p&gt;
 &lt;ol type="1"&gt;
  &lt;li&gt;Writes data files in Parquet format to S3.&lt;/li&gt;
  &lt;li&gt;Communicates with the S3 Tables REST endpoint to register the new files.&lt;/li&gt;
  &lt;li&gt;Atomically commits the transaction, updating table metadata.&lt;/li&gt;
 &lt;/ol&gt;
 &lt;h3 id="querying-data"&gt;4.2.4 Querying data&lt;/h3&gt;
 &lt;p&gt;Execute a &lt;code&gt;SELECT&lt;/code&gt; query to retrieve and verify the inserted data:&lt;/p&gt;
 &lt;div class="hide-language"&gt;
  &lt;pre&gt;&lt;code class="language-sql"&gt;SELECT * FROM customers LIMIT 10;&lt;/code&gt;&lt;/pre&gt;
 &lt;/div&gt;
 &lt;p&gt;The query should return all eight customer records with proper formatting. You can also execute more complex analytical queries:&lt;/p&gt;
 &lt;div class="hide-language"&gt;
  &lt;pre&gt;&lt;code class="language-sql"&gt;-- Count customers by country
SELECT birth_country, COUNT(*) as customer_count
FROM customers
GROUP BY birth_country
ORDER BY customer_count DESC;

-- Find customers born after 1990
SELECT first_name, last_name, birth_year
FROM customers
WHERE birth_year &amp;gt; 1990
ORDER BY birth_year;&lt;/code&gt;&lt;/pre&gt;
 &lt;/div&gt;
 &lt;p&gt;These queries demonstrate Trino’s SQL capabilities and the integration with S3 Tables for both read and write operations.&lt;/p&gt;
 &lt;h3 id="explore-advanced-features"&gt;4.3 Explore advanced features&lt;/h3&gt;
 &lt;p&gt;S3 Tables with Iceberg provides several features for data management:&lt;/p&gt;
 &lt;h3 id="time-travel-queries"&gt;4.3.1 Time travel queries&lt;/h3&gt;
 &lt;p&gt;Step 1: Check available snapshots.&lt;/p&gt;
 &lt;div class="hide-language"&gt;
  &lt;pre&gt;&lt;code class="language-sql"&gt;-- Query table as of a specific timestamp. Check available snapshots
SELECT * FROM "customers$snapshots";&lt;/code&gt;&lt;/pre&gt;
 &lt;/div&gt;
 &lt;p&gt;Step 2: Query the table as of a specific snapshot.&lt;/p&gt;
 &lt;div class="hide-language"&gt;
  &lt;pre&gt;&lt;code class="language-sql"&gt;SELECT * FROM customers FOR VERSION AS OF &amp;lt;snapshot_id_from_step1&amp;gt;;&lt;/code&gt;&lt;/pre&gt;
 &lt;/div&gt;
 &lt;h3 id="schema-evolution"&gt;4.3.2 Schema evolution&lt;/h3&gt;
 &lt;div class="hide-language"&gt;
  &lt;pre&gt;&lt;code class="language-sql"&gt;-- Add a new column
ALTER TABLE customers ADD COLUMN email VARCHAR;

-- Rename a column
ALTER TABLE customers RENAME COLUMN login TO username;&lt;/code&gt;&lt;/pre&gt;
 &lt;/div&gt;
 &lt;h2 id="cleaning-up"&gt;Cleaning up&lt;/h2&gt;
 &lt;p&gt;To clean up the resources, navigate to &lt;a href="https://us-east-1.console.aws.amazon.com/cloudformation/home" target="_blank" rel="noopener"&gt;CloudFormation&lt;/a&gt; and delete the stack that you created.&lt;/p&gt;
 &lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
 &lt;p&gt;This solution demonstrates an integration between Amazon EMR Trino and Amazon S3 Tables using the Apache Iceberg REST catalog specification. In this post, we showed you how to create and query S3 Tables from Trino on Amazon EMR. The architecture delivers several advantages for modern data platforms:&lt;/p&gt;
 &lt;p&gt;&lt;strong&gt;Operational simplicity&lt;/strong&gt;: S3 Tables eliminates the complexity of managing Iceberg table metadata, compaction schedules, and snapshot lifecycle policies. The service handles these operations automatically, allowing data teams to focus on analytics rather than infrastructure maintenance.&lt;/p&gt;
 &lt;p&gt;&lt;strong&gt;Performance at scale&lt;/strong&gt;: The architecture is designed for large-scale workloads. Trino distributes query execution across the cluster while Iceberg’s metadata layer helps the engine locate only the relevant data files. Features like partition pruning, predicate pushdown, and columnar file formats can help improve performance for both interactive and batch workloads.&lt;/p&gt;
 &lt;p&gt;&lt;strong&gt;Cost efficiency&lt;/strong&gt;: This architecture separates compute and storage, so you can scale each independently based on workload requirements. S3 Tables automatically compacts small files to help reduce storage overhead, and Amazon EMR clusters can scale dynamically so you pay for compute only when needed.&lt;/p&gt;
 &lt;p&gt;&lt;strong&gt;Open standards and portability&lt;/strong&gt;: By using Apache Iceberg’s open table format and REST catalog specification, this solution avoids vendor lock-in. Other Iceberg-compatible engines can access tables created in S3 Tables including Apache Spark, Apache Flink, and Dremio, providing flexibility in tool selection.&lt;/p&gt;
 &lt;p&gt;&lt;strong&gt;Fine-grained access control&lt;/strong&gt;: Integration with IAM and resource-based policies provides access control at the table bucket, namespace, and table level. For fine-grained access at the column and row level, you can integrate with AWS Lake Formation. AWS Signature Version 4 authentication supports secure communication between Trino and S3 Tables.&lt;/p&gt;
 &lt;p&gt;&lt;strong&gt;ACID transactions&lt;/strong&gt;: Iceberg’s transaction model guarantees atomicity, consistency, isolation, and durability for all table operations. This supports reliable concurrent reads and writes, making the platform suitable for production workloads requiring data consistency.&lt;/p&gt;
 &lt;p&gt;This architectural pattern is particularly well-suited for organizations building modern data lakehouses, migrating from traditional data warehouses, or consolidating multiple analytics platforms. The combination of the managed compute of Amazon EMR, Trino’s versatile query engine, and the automated table management of S3 Tables creates a strong foundation for data-driven decision making.&lt;/p&gt;
 &lt;p&gt;To learn more about the services and features discussed in this post, see the following resources:&lt;/p&gt;
 &lt;ul&gt;
  &lt;li&gt;&lt;a href="https://aws.amazon.com/emr/" target="_blank" rel="noopener"&gt;Amazon EMR&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/AmazonS3/latest/userguide/s3-tables.html" target="_blank" rel="noopener"&gt;Getting started with Amazon S3 Tables&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/emr/latest/ReleaseGuide/emr-trino.html" target="_blank" rel="noopener"&gt;Trino on Amazon EMR&lt;/a&gt;&lt;/li&gt;
 &lt;/ul&gt;
 &lt;p style="clear: both"&gt;&lt;/p&gt;
 &lt;hr style="width: 100%"&gt;
 &lt;h2&gt;About the authors&lt;/h2&gt;
 &lt;footer&gt;
  &lt;div class="blog-author-box"&gt;
   &lt;div class="blog-author-image"&gt;
    &lt;img loading="lazy" class="alignleft size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/01/BDB-5695-6.jpeg" alt="Shubham Purwar" width="100" height="100"&gt;
   &lt;/div&gt;
   &lt;h3 class="lb-h4"&gt;Shubham Purwar&lt;/h3&gt;
   &lt;p&gt;&lt;a href="https://www.linkedin.com/in/shubham-purwar/" target="_blank" rel="noopener"&gt;Shubham&lt;/a&gt; is an AWS Analytics Specialist Solution Architect. He helps organizations unlock the full potential of their data by designing and implementing scalable, secure, and high-performance analytics solutions on AWS. In his free time, Shubham loves to spend time with his family and travel around the world.&lt;/p&gt;
  &lt;/div&gt;
  &lt;div class="blog-author-box"&gt;
   &lt;div class="blog-author-image"&gt;
    &lt;img loading="lazy" class="alignleft size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/01/BDB-5695-7.png" alt="Anirudh Chawla" width="100" height="100"&gt;
   &lt;/div&gt;
   &lt;h3 class="lb-h4"&gt;Anirudh Chawla&lt;/h3&gt;
   &lt;p&gt;&lt;a href="https://www.linkedin.com/in/chawla-anirudh/" target="_blank" rel="noopener"&gt;Anirudh&lt;/a&gt; is an AWS Analytics Specialist Solution Architect. He helps organizations empower businesses to harness their data effectively through the analytics services of AWS. His interest lies in building highly available distributed systems.&lt;/p&gt;
  &lt;/div&gt;
  &lt;div class="blog-author-box"&gt;
   &lt;div class="blog-author-image"&gt;
    &lt;img loading="lazy" class="alignleft size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/01/BDB-5695-8.png" alt="Nitin Kumar" width="100" height="100"&gt;
   &lt;/div&gt;
   &lt;h3 class="lb-h4"&gt;Nitin Kumar&lt;/h3&gt;
   &lt;p&gt;&lt;a href="https://www.linkedin.com/in/nitinkr91/" target="_blank" rel="noopener"&gt;Nitin&lt;/a&gt; is a Solutions Architect at AWS. He partners with customers to transform their cloud journey through innovative, scalable solutions. In his free time, he likes to watch movies and spend time with his family.&lt;/p&gt;
  &lt;/div&gt;
  &lt;div class="blog-author-box"&gt;
   &lt;div class="blog-author-image"&gt;
    &lt;img loading="lazy" class="alignleft size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/09/01/BDB-5695-9.png" alt="Prashanthi Chinthala" width="100" height="100"&gt;
   &lt;/div&gt;
   &lt;h3 class="lb-h4"&gt;Prashanthi Chinthala&lt;/h3&gt;
   &lt;p&gt;&lt;a href="https://www.linkedin.com/in/prashanthi-chinthala-22968b184/" target="_blank" rel="noopener"&gt;Prashanthi&lt;/a&gt; is a Cloud Engineer (DIST) at AWS. She helps customers overcome Amazon EMR challenges and develop scalable data processing and analytics pipelines on AWS.&lt;/p&gt;
  &lt;/div&gt;
 &lt;/footer&gt;
&lt;/div&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>Building medallion architecture with Iceberg materialized views in Amazon SageMaker</title>
		<link>https://aws.amazon.com/blogs/big-data/building-medallion-architecture-with-iceberg-materialized-views-in-amazon-sagemaker/</link>
		
		<dc:creator><![CDATA[Gaurav Sharma]]></dc:creator>
		<pubDate>Wed, 02 Sep 2026 18:34:32 +0000</pubDate>
				<category><![CDATA[Advanced (300)]]></category>
		<category><![CDATA[Amazon SageMaker]]></category>
		<category><![CDATA[Technical How-to]]></category>
		<guid isPermaLink="false">2e6460187e466160fe3ee7892e1c09bbef301318</guid>

					<description>With Apache Iceberg materialized views in Amazon SageMaker, you can build a Bronze, Silver, and Gold medallion architecture as three SQL statements. This declarative approach folds transformation, orchestration, and incremental processing into per-layer definitions, with no ETL jobs, orchestrators, or change-data-capture code to maintain.</description>
										<content:encoded>&lt;p&gt;Building a &lt;a href="https://dataengineering.wiki/Concepts/Data+Architecture/Medallion+Architecture" target="_blank" rel="noopener"&gt;Medallion Architecture&lt;/a&gt; today typically means that you must build three separate systems working in concert: extract, transform, and load (ETL) jobs to transform data between layers, an orchestrator (such as Apache Airflow or AWS Step Functions) to sequence those jobs in the correct order, and custom change-data-capture (CDC) logic to make sure that each job processes only new or modified records. Each component must be authored, tested, deployed, and maintained independently and when one breaks, the entire pipeline stalls.&lt;/p&gt;
&lt;p&gt;In this post, we show how Apache Iceberg materialized views in &lt;a href="https://aws.amazon.com/sagemaker/" target="_blank" rel="noopener"&gt;Amazon SageMaker&lt;/a&gt; collapse transformation, orchestration, and incremental processing into a single SQL definition per layer. You declare what each layer should contain, and the system handles when and how it refreshes based on your refresh configuration. With this approach, you can build a Bronze → Silver → Gold pipeline with three SQL statements. This reduces the complexity of maintaining separate orchestration code, CDC logic, and job artifacts.&lt;/p&gt;
&lt;h2 id="what-is-medallion-architecture"&gt;What is medallion architecture&lt;/h2&gt;
&lt;p&gt;The medallion architecture organizes data into three progressive layers:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;Bronze layer – Captures raw data as-is from source systems, preserving the original format for auditability and replay.&lt;/li&gt;
 &lt;li&gt;Silver layer – Applies cleaning, deduplication, type casting, and business logic to produce validated, query-ready datasets.&lt;/li&gt;
 &lt;li&gt;Gold layer – Aggregates Silver data into business-level metrics, key performance indicators (KPIs), and dimensional models optimized for analytics and reporting.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Each layer builds on the previous one, creating clear lineage from raw ingestion to business insight.&lt;/p&gt;
&lt;h2 id="traditional-versus-declarative-approach"&gt;Traditional versus declarative approach&lt;/h2&gt;
&lt;p&gt;The two approaches differ in how much infrastructure you build and maintain.&lt;/p&gt;
&lt;h3 id="traditional-approach"&gt;Traditional approach&lt;/h3&gt;
&lt;p&gt;You write an ETL job such as Apache Spark script for Bronze to Silver layer and another for Silver to Gold layer. You build a directed acyclic graph (DAG) in Apache Airflow or a Step Functions state machine to run them in order. You implement CDC logic like tracking high watermarks, comparing snapshots, or consuming change streams such that each job processes only new data.&lt;/p&gt;
&lt;h3 id="declarative-approach-with-iceberg-materialized-views"&gt;Declarative approach with Iceberg materialized views&lt;/h3&gt;
&lt;p&gt;You write one &lt;code&gt;CREATE MATERIALIZED VIEW&lt;/code&gt; statement per layer with a &lt;code&gt;SCHEDULE REFRESH EVERY N HOURS&lt;/code&gt; clause. The AWS Glue managed Spark compute executes the refresh, but you don’t author, version, or deploy a job artifact. Iceberg’s row-level change tracking (position-delete and equality-delete files) identifies which rows changed since the last refresh and AWS Glue processes only those rows. The dependency chain is implicit in the SQL definitions. The only code you maintain is the SQL transformation logic itself.&lt;/p&gt;
&lt;h2 id="apache-iceberg-and-materialized-views"&gt;Apache Iceberg and materialized views&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://iceberg.apache.org/" target="_blank" rel="noopener"&gt;Apache Iceberg&lt;/a&gt; is an open-source, high-performance table format designed for petabyte-scale analytic datasets in data lakes. It provides ACID transactions, time travel, schema evolution, and hidden partitioning.&lt;/p&gt;
&lt;p&gt;With an &lt;a href="https://aws.amazon.com/about-aws/whats-new/2025/11/aws-glue-apache-iceberg-based-materialized-views/" target="_blank" rel="noopener"&gt;Iceberg materialized view&lt;/a&gt;, you can define each layer of a medallion architecture as a SQL statement. Under the hood, AWS Glue uses Iceberg’s change-tracking metadata to identify which rows changed since the last refresh, then processes only those rows using managed Spark compute. You configure scheduling and incremental processing through SQL definitions, and the system executes atomic refreshes without requiring you to write pipeline code.&lt;/p&gt;
&lt;p&gt;When refreshed, the Gold materialized view reads incrementally from the Silver materialized view, which in turn reads from the Bronze table. This creates a declarative dependency chain: each layer’s definition points to the layer below it, and the system resolves which data to reprocess at each refresh.&lt;/p&gt;
&lt;h2 id="service-support-for-iceberg-materialized-views"&gt;Service support for Iceberg materialized views&lt;/h2&gt;
&lt;p&gt;At time of publication, the following services support creating and refreshing Iceberg materialized views:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/athena/latest/ug/notebooks-spark.html" target="_blank" rel="noopener"&gt;Amazon Athena Spark&lt;/a&gt;&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/glue/latest/dg/materialized-views.html" target="_blank" rel="noopener"&gt;AWS Glue 5.1&lt;/a&gt; and later&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/emr/latest/ReleaseGuide/emr-spark-materialized-views.html" target="_blank" rel="noopener"&gt;Amazon EMR 7.12&lt;/a&gt; and later&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For the latest version requirements, see the &lt;a href="https://docs.aws.amazon.com/glue/latest/dg/materialized-views.html" target="_blank" rel="noopener"&gt;AWS Glue materialized views documentation&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="technical-architecture"&gt;Technical architecture&lt;/h2&gt;
&lt;p&gt;The architecture uses Amazon S3 Tables, a capability of Amazon Simple Storage Service (Amazon S3), as the storage layer. Amazon S3 Tables is a managed Apache Iceberg offering that alleviates the administrative overhead of maintaining Iceberg tables. AWS Glue Data Catalog manages table metadata, and Amazon SageMaker Unified Studio provides the AI-powered notebook environment with AWS Glue 5.1 for authoring and executing materialized view definitions.&lt;/p&gt;
&lt;p&gt;The diagram illustrates a three-tier data lakehouse pipeline built on Apache Iceberg. The Bronze layer contains raw trip data (trips_bronze table on S3 Tables with fields: trip_id, city, vehicle_type, fare, status) that you ingest through INSERT/Append operations.&lt;/p&gt;
&lt;p&gt;An incremental REFRESH feeds the Silver layer, where a materialized view (mv_trips_silver) performs timestamp conversion, null filtering, and computes derived columns like revenue_per_mile and rating_category. It processes only new or changed rows.&lt;/p&gt;
&lt;p&gt;The Silver layer then refreshes two Gold layer materialized views on a daily schedule: mv_city_daily_metrics (city, date, trips, drivers, revenue, tips) and mv_vehicle_performance (vehicle_type, city, trips, revenue, distance). The Gold layer serves downstream consumers including Amazon Athena, Amazon Quick Sight, Amazon Redshift, and first-party (1P) or third-party (3P) compute engines supporting the Iceberg REST API.&lt;/p&gt;
&lt;p&gt;The pipeline flows as follows:&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-1.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-1.png" alt="Diagram of the medallion pipeline: a Bronze table feeds a Silver materialized view that feeds two Gold materialized views consumed by analytics engines" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 1: The three-tier medallion pipeline from the Bronze table through Silver and Gold materialized views to analytics consumers&lt;/p&gt;
&lt;/div&gt;
&lt;h2 id="prerequisites"&gt;Prerequisites&lt;/h2&gt;
&lt;p&gt;Before starting, verify that you have the following:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;An AWS account with permissions for Amazon SageMaker Unified Studio, AWS Glue, S3 Tables, and AWS Lake Formation.&lt;/li&gt;
 &lt;li&gt;An Amazon SageMaker Unified Studio domain.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="step-1-initialize-the-environment"&gt;Step 1: Initialize the environment&lt;/h2&gt;
&lt;p&gt;Open the AWS Management Console and navigate to &lt;strong&gt;Amazon SageMaker&lt;/strong&gt;.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-2.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-2.png" alt="Amazon SageMaker console landing page" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 2: The Amazon SageMaker console landing page&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;Choose &lt;strong&gt;Get Started&lt;/strong&gt; to set up Amazon SageMaker Unified Studio.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-3.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-3.png" alt="SageMaker Unified Studio Get Started setup page" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 3: The Get Started page for setting up SageMaker Unified Studio&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;Choose &lt;strong&gt;Open&lt;/strong&gt; to launch Amazon SageMaker Unified Studio.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-4.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-4.png" alt="Button to open and launch SageMaker Unified Studio" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 4: The option to open and launch SageMaker Unified Studio&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;After you’re in SageMaker Unified Studio, choose &lt;strong&gt;Data&lt;/strong&gt; in the left pane to create the S3 Tables bucket (a managed Apache Iceberg feature of Amazon S3) and a database. Choose &lt;strong&gt;Add&lt;/strong&gt;, then choose &lt;strong&gt;Create S3 Tables Catalog&lt;/strong&gt;, and provide a catalog and a database name. Finally, choose &lt;strong&gt;Create Catalog&lt;/strong&gt;.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-5.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-5.png" alt="Create S3 Tables Catalog dialog with catalog and database name fields" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 5: The Create S3 Tables Catalog dialog with catalog and database name fields&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;After the catalog creation is complete, in the left navigation pane, choose &lt;strong&gt;Notebooks&lt;/strong&gt;.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-6.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-6.png" alt="Notebooks option in the SageMaker Unified Studio left navigation pane" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 6: The Notebooks option in the SageMaker Unified Studio navigation pane&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;Choose &lt;strong&gt;Create Notebook&lt;/strong&gt;.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-7.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-7.png" alt="Create Notebook button in SageMaker Unified Studio" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 7: The Create Notebook button in SageMaker Unified Studio&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;Before using the notebook, select either &lt;strong&gt;Athena Spark&lt;/strong&gt; or &lt;strong&gt;Glue Spark&lt;/strong&gt; compute connection as the runtime engine for your notebook.&lt;/p&gt;
&lt;div style="width: 705px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-8.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-8.png" alt="Runtime engine selection showing Athena Spark and Glue Spark compute connections" width="695"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 8: Selecting Athena Spark or Glue Spark as the notebook runtime engine&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;Use the following code samples in individual notebook cells. You can also provide transformation requirements in natural language, and the &lt;a href="https://docs.aws.amazon.com/sagemaker-unified-studio/latest/userguide/sagemaker-data-agent.html" target="_blank" rel="noopener"&gt;SageMaker Data Agent&lt;/a&gt; will generate SQL code for you.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-9.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-9.png" alt="SageMaker Data Agent generating SQL from a natural language prompt" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 9: The SageMaker Data Agent generating SQL from a natural language request&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;Add each code block in a new cell by choosing the &lt;strong&gt;SQL&lt;/strong&gt; button:&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-10.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-10.png" alt="SQL cell-type button in the notebook toolbar" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 10: The SQL button for adding a code block to a notebook cell&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;Choose &lt;strong&gt;Athena Spark&lt;/strong&gt; or &lt;strong&gt;Glue Spark&lt;/strong&gt; as your compute from the cell menu.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-11.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-11.png" alt="Compute connection selection in the notebook cell menu" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 11: The compute selection in the notebook cell menu&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;If you encounter errors after cell execution, use the data agent chatbot or the &lt;strong&gt;Fix with AI&lt;/strong&gt; button to resolve them.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-12.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-12.png" alt="Fix with AI button and data agent chatbot for resolving cell errors" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 12: The Fix with AI button for resolving cell execution errors&lt;/p&gt;
&lt;/div&gt;
&lt;h2 id="step-2-ingest-data-into-bronze"&gt;Step 2: Ingest data into Bronze&lt;/h2&gt;
&lt;p&gt;Generate 300 realistic ride-sharing trips and insert them directly into the Bronze Iceberg table. This simulates a raw data ingestion layer. In production, you generally configure a streaming source or batch load based on your requirements.&lt;/p&gt;
&lt;p&gt;Copy the following code into the first notebook cell (use a Python cell type).&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-python"&gt;import random
from datetime import datetime, timedelta

CITIES = {
    "San Francisco": {"lat_range": (37.70, 37.82), "lon_range": (-122.52, -122.38), "surge_prob": 0.3},
    "Austin": {"lat_range": (30.22, 30.40), "lon_range": (-97.80, -97.68), "surge_prob": 0.15},
    "Chicago": {"lat_range": (41.85, 41.95), "lon_range": (-87.70, -87.60), "surge_prob": 0.2},
    "Seattle": {"lat_range": (47.55, 47.68), "lon_range": (-122.40, -122.28), "surge_prob": 0.25},
}
VEHICLE_TYPES = ["UberX", "Comfort", "XL", "Black"]
PAYMENT_METHODS = ["credit_card", "debit_card", "apple_pay", "google_pay", "cash"]
STATUSES = ["completed"] * 4 + ["cancelled_rider", "cancelled_driver"]
BASE_FARES = {"UberX": 2.50, "Comfort": 3.50, "XL": 4.00, "Black": 7.00}
PER_MILE = {"UberX": 1.75, "Comfort": 2.25, "XL": 2.50, "Black": 3.75}
PER_MIN = {"UberX": 0.35, "Comfort": 0.45, "XL": 0.50, "Black": 0.65}

rows = []
for i in range(300):
    city_name = random.choice(list(CITIES.keys()))
    city = CITIES[city_name]
    vehicle = random.choice(VEHICLE_TYPES)
    duration = random.randint(5, 45)
    distance = round(random.uniform(1.0, 20.0), 1)
    surge = round(random.uniform(1.0, 2.5), 1) if random.random() &amp;lt; city["surge_prob"] else 1.0
    base = BASE_FARES[vehicle]
    fare = round((base + distance * PER_MILE[vehicle] + duration * PER_MIN[vehicle]) * surge, 2)
    tip = round(fare * random.choice([0, 0, 0.1, 0.15, 0.2, 0.25]), 2)
    status = random.choice(STATUSES)
    day = random.randint(0, 2)
    hour = random.choices(range(24),
        weights=[1,1,1,1,1,2,4,8,10,8,6,5,6,5,5,5,6,8,10,8,6,4,2,1])[0]
    trip_time = datetime(2025, 12, 1) + timedelta(days=day, hours=hour, minutes=random.randint(0, 59))

    rows.append((
        f"TRIP-{i+1:06d}",
        f"DRV-{random.randint(1000, 5000)}",
        f"RDR-{random.randint(10000, 99999)}",
        city_name, vehicle,
        round(random.uniform(*city["lat_range"]), 6),
        round(random.uniform(*city["lon_range"]), 6),
        round(random.uniform(*city["lat_range"]), 6),
        round(random.uniform(*city["lon_range"]), 6),
        trip_time.isoformat(),
        (trip_time + timedelta(minutes=duration)).isoformat(),
        duration, distance, surge, base, fare, tip, round(fare + tip, 2),
        random.choice(PAYMENT_METHODS),
        random.choice([None, 3, 4, 4, 5, 5, 5]) if status == "completed" else None,
        status,
    ))

schema = ("trip_id STRING, driver_id STRING, rider_id STRING, city STRING, "
    "vehicle_type STRING, pickup_lat DOUBLE, pickup_lon DOUBLE, "
    "dropoff_lat DOUBLE, dropoff_lon DOUBLE, trip_start_time STRING, "
    "trip_end_time STRING, duration_minutes INT, distance_miles DOUBLE, "
    "surge_multiplier DOUBLE, base_fare DOUBLE, trip_fare DOUBLE, "
    "tip_amount DOUBLE, total_amount DOUBLE, payment_method STRING, "
    "rating INT, status STRING")

df = spark.createDataFrame(rows, schema)
df.writeTo("{CATALOG_NAME}.{NAMESPACE_NAME}.trips_bronze").createOrReplace()

print(f"Created Table and Inserted {len(rows)} trips into Bronze layer")&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h2 id="step-3-explore-bronze"&gt;Step 3: Explore Bronze&lt;/h2&gt;
&lt;p&gt;Run a preview on the bronze table. The output should look like the following screenshot:&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-13.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-13.png" alt="Preview of raw Bronze table trip records with string timestamps and nullable fields" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 13: A preview of raw trip records in the Bronze table&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;You should see raw, unprocessed trip records with string timestamps and nullable fields. This is exactly what the Silver layer will clean up.&lt;/p&gt;
&lt;p&gt;Now, verify the ingested data by querying the Bronze table for basic statistics.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-sql"&gt;SELECT COUNT(*) as total_trips, COUNT(DISTINCT city) as cities,
COUNT(DISTINCT vehicle_type) as vehicle_types,
MIN(trip_start_time) as earliest, MAX(trip_start_time) as latest
FROM ({CATALOG_NAME}.{NAMESPACE_NAME}.trips_bronze&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;The output should look like the following screenshot:&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-14.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-14.png" alt="Query results showing total trips, distinct cities, and vehicle types in the Bronze table" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 14: Bronze table statistics showing total trips, distinct cities, and vehicle types&lt;/p&gt;
&lt;/div&gt;
&lt;h2 id="step-4-create-the-silver-materialized-view"&gt;Step 4: Create the Silver materialized view&lt;/h2&gt;
&lt;p&gt;This SQL statement defines the Silver layer as a materialized view that cleans, transforms, and derives new columns from the Bronze table. Note that this is only a definition. The system processes the data at refresh time.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-sql"&gt;CREATE MATERIALIZED VIEW IF NOT EXISTS {CATALOG_NAME}.{DATABASE}.mv_trips_silver
COMMENT 'Silver layer: Cleaned trip data with proper types and derived columns'
SCHEDULE REFRESH EVERY 1 DAY
AS
SELECT
trip_id, driver_id, rider_id, city, vehicle_type,
pickup_lat, pickup_lon, dropoff_lat, dropoff_lon,
CAST(trip_start_time AS TIMESTAMP) as trip_start_timestamp,
CAST(trip_end_time AS TIMESTAMP) as trip_end_timestamp,
duration_minutes, distance_miles, surge_multiplier,
base_fare, trip_fare, tip_amount, total_amount,
payment_method, rating, status,
CASE WHEN distance_miles &amp;gt; 0 THEN total_amount / distance_miles ELSE 0 END as revenue_per_mile,
CASE WHEN rating &amp;gt;= 4 THEN 'High' WHEN rating &amp;gt;= 3 THEN 'Medium' ELSE 'Low' END as rating_category
FROM {CATALOG_NAME}.{DATABASE}.trips_bronze
WHERE trip_id IS NOT NULL AND driver_id IS NOT NULL AND rider_id IS NOT NULL
AND total_amount &amp;gt;= 0 AND distance_miles &amp;gt;= 0

print("Silver MV created: urbanride.mv_trips_silver")&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;Verify the Silver layer output:&lt;/strong&gt;&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-sql"&gt;SELECT trip_id, city, vehicle_type, total_amount,
ROUND(revenue_per_mile, 2) as rev_per_mile, rating_category
FROM {CATALOG_NAME}.{DATABASE}.mv_trips_silver LIMIT 5&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Notice how the Silver layer now has proper timestamps, derived revenue_per_mile, and rating categories: clean, typed, and ready for you to aggregate.&lt;/p&gt;
&lt;p&gt;The output should look like the following screenshot:&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-15.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-15.png" alt="Silver materialized view results with typed timestamps, revenue_per_mile, and rating_category columns" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 15: Silver materialized view results with typed timestamps and derived columns&lt;/p&gt;
&lt;/div&gt;
&lt;h2 id="step-5-create-gold-materialized-views"&gt;Step 5: Create Gold materialized views&lt;/h2&gt;
&lt;p&gt;Gold materialized views read incrementally from the Silver materialized view. This is a nested materialized view pattern: a materialized view built on top of another materialized view.&lt;/p&gt;
&lt;h3 id="gold-1-city-daily-metrics"&gt;Gold 1: City daily metrics&lt;/h3&gt;
&lt;p&gt;With this materialized view, you can aggregate trip data by city and date with a scheduled daily refresh.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-sql"&gt;CREATE MATERIALIZED VIEW IF NOT EXISTS {CATALOG_NAME}.urbanride.mv_city_daily_metrics
COMMENT 'Gold layer: Daily aggregated metrics by city'
SCHEDULE REFRESH EVERY 1 DAY
AS
SELECT
city, DATE(trip_start_timestamp) as trip_date,
COUNT(*) as total_trips,
COUNT(DISTINCT driver_id) as active_drivers,
COUNT(DISTINCT rider_id) as active_riders,
SUM(total_amount) as total_revenue,
SUM(distance_miles) as total_distance,
SUM(tip_amount) as total_tips
FROM {CATALOG_NAME}.{DATABASE}.mv_trips_silver
WHERE status = 'completed'
GROUP BY city, DATE(trip_start_timestamp)

print("Gold MV created: mv_city_daily_metrics (reads from Silver MV, refreshes daily)")&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h3 id="gold-2-vehicle-performance"&gt;Gold 2: Vehicle performance&lt;/h3&gt;
&lt;p&gt;With this materialized view, you can aggregate performance metrics by vehicle type and city.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-sql"&gt;CREATE MATERIALIZED VIEW IF NOT EXISTS {CATALOG_NAME}.{DATABASE}.mv_vehicle_performance
COMMENT 'Gold layer: Vehicle type performance metrics'
SCHEDULE REFRESH EVERY 1 DAY
AS
SELECT
vehicle_type, city,
COUNT(*) as trip_count,
SUM(total_amount) as total_revenue,
SUM(distance_miles) as total_distance,
SUM(tip_amount) as total_tips
FROM {CATALOG_NAME}.{DATABASE}.mv_trips_silver
WHERE status = 'completed'
GROUP BY vehicle_type, city

print("Gold MV created: mv_vehicle_performance (reads from Silver MV, refreshes daily)")&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h3 id="dependency-chain"&gt;Dependency chain&lt;/h3&gt;
&lt;p&gt;The complete pipeline dependency is:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-text"&gt;trips_bronze (table)
└── mv_trips_silver (materialized view)
    ├── mv_city_daily_metrics (MV on MV, daily schedule)
    └── mv_vehicle_performance (MV on MV, daily schedule)&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Each layer is defined by a single SQL statement. There are no DAGs to maintain, no job definitions to deploy, and no watermark tracking to implement.&lt;/p&gt;
&lt;h2 id="step-6-query-the-gold-layer"&gt;Step 6: Query the Gold layer&lt;/h2&gt;
&lt;p&gt;Query the Gold materialized views to see aggregated business metrics.&lt;/p&gt;
&lt;h3 id="city-daily-metrics-gold-table"&gt;City daily metrics Gold table&lt;/h3&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-sql"&gt;SELECT city, trip_date, total_trips, active_drivers,
ROUND(total_revenue, 2) as revenue,
ROUND(total_revenue / total_trips, 2) as avg_per_trip
FROM {CATALOG_NAME}.{DATABASE}.mv_city_daily_metrics
ORDER BY trip_date DESC, revenue DESC LIMIT 15&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;The output should look like the following screenshot:&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-16.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-16.png" alt="City daily metrics results with trips, active drivers, and revenue per city" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 16: City daily metrics from the Gold materialized view&lt;/p&gt;
&lt;/div&gt;
&lt;h3 id="vehicle-performance-gold-table"&gt;Vehicle performance Gold table&lt;/h3&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-sql"&gt;SELECT vehicle_type, city, trip_count,
ROUND(total_revenue, 2) as revenue,
ROUND(total_revenue / trip_count, 2) as avg_per_trip
FROM {CATALOG_NAME}.{DATABASE}.mv_vehicle_performance
ORDER BY revenue DESC&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;The output should look like the following screenshot:&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-17.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-17.png" alt="Vehicle performance results with trip counts and revenue by vehicle type and city" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 17: Vehicle performance metrics from the Gold materialized view&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;The Gold layer gives you pre-aggregated, business-ready metrics without writing aggregation jobs.&lt;/p&gt;
&lt;h2 id="step-7-data-propagation-demo"&gt;Step 7: Data propagation demo&lt;/h2&gt;
&lt;p&gt;This section demonstrates how changes propagate through the layers using INSERT, UPDATE (MERGE), and DELETE operations followed by incremental refresh. In production, the scheduled refresh handles this automatically. We trigger it manually here for demonstration purposes.&lt;/p&gt;
&lt;h3 id="insert-new-records"&gt;INSERT new records&lt;/h3&gt;
&lt;p&gt;Insert new trip records into the Bronze table.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-sql"&gt;INSERT INTO {CATALOG_NAME}.{DATABASE}.trips_bronze VALUES
('DEMO_TRIP_001', 'DRIVER_999', 'RIDER_888', 'Seattle', 'UberX',
47.6062, -122.3321, 47.6205, -122.3493,
'2024-12-15 14:30:00', '2024-12-15 14:50:00',
20, 5.2, 1.0, 10.0, 15.0, 3.0, 18.0, 'credit_card', 5, 'completed'),
('DEMO_TRIP_002', 'DRIVER_888', 'RIDER_777', 'Seattle', 'XL',
47.6101, -122.3300, 47.6550, -122.3080,
'2024-12-15 15:00:00', '2024-12-15 15:35:00',
35, 8.5, 1.5, 15.0, 30.0, 5.0, 35.0, 'cash', 4, 'completed'),
('DEMO_TRIP_003', 'DRIVER_777', 'RIDER_666', Portland, 'Comfort',
30.2672, -97.7431, 30.2800, -97.7400,
'2024-12-15 16:00:00', '2024-12-15 16:15:00',
15, 3.0, 1.0, 8.0, 12.0, 2.0, 14.0, 'credit_card', 5, 'completed')

print("Inserted 3 new trips into Bronze")&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h3 id="refresh-silver-incremental"&gt;Refresh Silver (incremental)&lt;/h3&gt;
&lt;p&gt;Refresh the Silver materialized view. Iceberg materialized view processes only three new records.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-sql"&gt;REFRESH MATERIALIZED VIEW {CATALOG_NAME}.{DATABASE}.mv_trips_silver"&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h3 id="verify-the-new-records-propagated"&gt;Verify the new records propagated&lt;/h3&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-sql"&gt;SELECT trip_id, city, total_amount, ROUND(revenue_per_mile, 2) as rev_per_mile, rating_category
FROM {CATALOG_NAME}.{DATABASE}.mv_trips_silver
WHERE trip_id LIKE 'DEMO_TRIP_%' ORDER BY trip_id&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;The output should look like the following screenshot:&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-18.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-18.png" alt="Silver materialized view showing three newly inserted demo trips" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 18: The Silver materialized view showing the three newly inserted demo trips&lt;/p&gt;
&lt;/div&gt;
&lt;h3 id="refresh-gold-cascading-from-the-silver-materialized-view"&gt;Refresh Gold (cascading from the Silver materialized view)&lt;/h3&gt;
&lt;p&gt;Refresh the Gold materialized view. It reads from the refreshed Silver materialized view and processes only the incremental changes.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-sql"&gt;REFRESH MATERIALIZED VIEW {CATALOG_NAME}.{DATABASE}.mv_city_daily_metrics&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;Verify the Gold layer reflects the new trips&lt;/strong&gt;&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-sql"&gt;SELECT city, trip_date, total_trips, ROUND(total_revenue, 2) as revenue
FROM {CATALOG_NAME}.{DATABASE}.mv_city_daily_metrics
WHERE trip_date = '2024-12-15' ORDER BY city&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;The output should look like the following screenshot:&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-19.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-19.png" alt="City daily metrics reflecting the newly added trips for December 15, 2024" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 19: City daily metrics reflecting the new trips for 2024-12-15&lt;/p&gt;
&lt;/div&gt;
&lt;h3 id="update-through-merge"&gt;UPDATE through MERGE&lt;/h3&gt;
&lt;p&gt;Use MERGE to update existing records in Bronze, then refresh incrementally.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-sql"&gt;MERGE INTO {CATALOG_NAME}.{DATABASE}.trips_bronze AS target
USING (SELECT 'DEMO_TRIP_002' as trip_id, 5 as new_rating, 20.0 as new_tip) AS source
ON target.trip_id = source.trip_id
WHEN MATCHED THEN UPDATE SET
target.rating = source.new_rating,
target.tip_amount = source.new_tip,
target.total_amount = target.trip_fare + source.new_tip&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;Refresh Silver and verify&lt;/strong&gt;&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-sql"&gt;REFRESH MATERIALIZED VIEW {CATALOG_NAME}.{DATABASE}.mv_trips_silver")

SELECT trip_id, rating, rating_category, tip_amount, total_amount,
ROUND(revenue_per_mile, 2) as rev_per_mile
FROM {CATALOG_NAME}.{DATABASE}.mv_trips_silver WHERE trip_id = 'DEMO_TRIP_002'

print("UPDATE propagated: rating 4-&amp;gt;5, tip $5-&amp;gt;$20, total $35-&amp;gt;$50")&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;The output should look like the following screenshot:&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-20.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-20.png" alt="Silver materialized view showing the updated rating and tip for DEMO_TRIP_002" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 20: The Silver materialized view showing the updated rating and tip for the demo trip&lt;/p&gt;
&lt;/div&gt;
&lt;h2 id="step-8-cleanup"&gt;Step 8: Cleanup&lt;/h2&gt;
&lt;p&gt;Drop materialized views, tables, the namespace, and delete the S3 Tables bucket to fully clean up resources.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-python"&gt;# Drop MVs (Gold first, then Silver, due to dependency order)
spark.sql(f"DROP MATERIALIZED VIEW IF EXISTS {CATALOG_NAME}.{DATABASE}.mv_city_daily_metrics")
spark.sql(f"DROP MATERIALIZED VIEW IF EXISTS {CATALOG_NAME}.{DATABASE}.mv_vehicle_performance")
spark.sql(f"DROP MATERIALIZED VIEW IF EXISTS {CATALOG_NAME}.{DATABASE}.mv_trips_silver")
print("All materialized views dropped")

# Drop base table
spark.sql(f"DROP TABLE IF EXISTS {CATALOG_NAME}.{DATABASE}.trips_bronze")
print("Base table dropped")

# Drop the namespace
spark.sql(f"DROP NAMESPACE IF EXISTS {CATALOG_NAME}.{DATABASE} ")
print("Namespace dropped")

# Delete the S3 table bucket
import boto3
s3tables_client = boto3.client("s3tables")

# List and delete all remaining tables in the bucket
tables_response = s3tables_client.list_tables(
    tableBucketARN=TABLE_BUCKET_ARN, namespace="{DATABASE}"
)
for table in tables_response.get("tables", []):
    s3tables_client.delete_table(
        tableBucketARN=TABLE_BUCKET_ARN, namespace="{DATABASE}", name=table['name']
    )
    print(f" Deleted table: {table['name']}")

# Delete the namespace and bucket
s3tables_client.delete_namespace(tableBucketARN=TABLE_BUCKET_ARN, namespace="urbanride")
s3tables_client.delete_table_bucket(tableBucketARN=TABLE_BUCKET_ARN)
print(f"S3 table bucket deleted: {TABLE_BUCKET_NAME}")&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h2 id="limitations-and-considerations"&gt;Limitations and considerations&lt;/h2&gt;
&lt;p&gt;While materialized views remove most orchestration code, note the following:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;No sub-hour freshness. The minimum schedule granularity is one hour (&lt;code&gt;SCHEDULE REFRESH EVERY 1 HOUR&lt;/code&gt;).&lt;/li&gt;
 &lt;li&gt;Cascading refresh isn’t automatic. Refreshing Silver doesn’t trigger Gold in the same operation. Each layer refreshes on its own schedule or must be triggered sequentially.&lt;/li&gt;
 &lt;li&gt;Deletes require a FULL refresh. An incremental REFRESH that feeds the Silver layer detects inserts and updates through Iceberg metadata but cannot detect row removals. Use &lt;code&gt;REFRESH ... FULL&lt;/code&gt; when delete propagation is needed.&lt;/li&gt;
 &lt;li&gt;SQL subset only. Some window functions, user-defined functions (UDFs), and complex expressions might not be supported in materialized view definitions.&lt;/li&gt;
 &lt;li&gt;Schema evolution requires recreation. If the source schema changes in a way that affects the materialized view definition, you must drop and recreate it.&lt;/li&gt;
 &lt;li&gt;AWS-specific extension. Iceberg materialized views are not part of the open-source Apache Iceberg specification. They aren’t portable to non-AWS environments.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="pricing"&gt;Pricing&lt;/h2&gt;
&lt;p&gt;AWS bills materialized view auto-refresh at USD $0.44 per DPU-hour (4 vCPU, 16 GB memory), billed per second with a 1-minute minimum. When you configure scheduled refresh, the AWS Glue Data Catalog uses managed Spark compute to incrementally update the materialized view. You pay only for the compute time of each refresh run.&lt;/p&gt;
&lt;p&gt;There are no separate charges for storing materialized view metadata in the Data Catalog (covered under standard catalog pricing: first million objects at no additional cost, then $1.00 per 100K objects/month). The materialized view data itself is stored as Iceberg files in S3 Tables or Amazon S3, charged at standard Amazon S3 storage rates.&lt;/p&gt;
&lt;p&gt;Manual refreshes triggered from Spark (through Amazon Athena, Amazon EMR, or AWS Glue notebooks) are billed under those services’ respective compute pricing rather than the materialized view auto-refresh rate. For the latest pricing details, see the &lt;a href="https://aws.amazon.com/glue/pricing/" target="_blank" rel="noopener"&gt;AWS Glue pricing&lt;/a&gt; page.&lt;/p&gt;
&lt;p&gt;Estimated cost for this tutorial: Running through all steps once with 300 records typically consumes less than 0.5 DPU-hours total (~$0.22 in AWS Glue compute plus negligible Amazon S3 storage).&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;In this post, you built a Bronze → Silver → Gold medallion architecture using three SQL statements with nested materialized views and no orchestration code. The full pipeline creation took under 2 minutes, and incremental refreshes processed only changed data with no watermarks, no DAGs, no CDC plumbing.&lt;/p&gt;
&lt;p&gt;To get started with your own data, create an Amazon SageMaker Unified Studio project, define your Bronze table, and express your transformation logic as Iceberg materialized views. For more information, see the Apache Iceberg materialized views documentation in the &lt;a href="https://docs.aws.amazon.com/glue/latest/dg/materialized-views.html" target="_blank" rel="noopener"&gt;AWS Glue Developer Guide&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://docs.aws.amazon.com/glue/latest/dg/materialized-views.html" target="_blank" rel="noopener"&gt;Using materialized views with AWS Glue&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://docs.aws.amazon.com/athena/latest/ug/querying-iceberg-gdc-mv.html" target="_blank" rel="noopener"&gt;Query AWS Glue Data Catalog materialized views&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://docs.aws.amazon.com/emr/latest/ReleaseGuide/emr-spark-materialized-views.html" target="_blank" rel="noopener"&gt;Using materialized views with Amazon EMR&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://docs.aws.amazon.com/AmazonS3/latest/userguide/s3-tables.html" target="_blank" rel="noopener"&gt;Working with Amazon S3 Tables and table buckets&lt;/a&gt;&lt;/p&gt;
&lt;hr style="width: 100%"&gt;
&lt;h2&gt;About the authors&lt;/h2&gt;
&lt;footer&gt;
 &lt;div class="blog-author-box"&gt;
  &lt;div class="blog-author-image"&gt;
   &lt;p&gt;&lt;img loading="lazy" class="alignleft size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-21.png" alt="Gaurav Sharma" width="100" height="100"&gt;&lt;/p&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Gaurav Sharma&lt;/h3&gt;
  &lt;p&gt;Gaurav is a Specialist Solutions Architect (Analytics) at AWS, supporting US public sector customers on their cloud journey. Outside of work, Gaurav enjoys spending time with his family and staying informed on technology, politics, and history through books, videos, and podcasts.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box"&gt;
  &lt;div class="blog-author-image"&gt;
   &lt;p&gt;&lt;img loading="lazy" class="alignleft size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-5987-22.png" alt="Matt David" width="100" height="100"&gt;&lt;/p&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Matt David&lt;/h3&gt;
  &lt;p&gt;Matt is a Product Marketing Manager at AWS, specializing in helping data teams with AI-powered analytics. His areas of interest include self-service analytics, data democratization, and preparing organizations for the age of AI agents. He brings extensive experience from his roles at Atlassian, Hex, and DataCamp.&lt;/p&gt;
 &lt;/div&gt;
&lt;/footer&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>Build a dynamic streaming data lake with Apache Iceberg and Apache Flink</title>
		<link>https://aws.amazon.com/blogs/big-data/build-a-dynamic-streaming-data-lake-with-apache-iceberg-and-apache-flink/</link>
		
		<dc:creator><![CDATA[Francisco Morillo]]></dc:creator>
		<pubDate>Wed, 02 Sep 2026 18:28:38 +0000</pubDate>
				<category><![CDATA[Advanced (300)]]></category>
		<category><![CDATA[Amazon Managed Service for Apache Flink]]></category>
		<category><![CDATA[Kinesis Data Streams]]></category>
		<category><![CDATA[Technical How-to]]></category>
		<guid isPermaLink="false">04edd6bf40eaf258061255f25d4b1c3964fba481</guid>

					<description>Learn how to build a dynamic streaming data lake on Amazon Managed Service for Apache Flink that adapts to new event types and schema changes without stopping the pipeline, using Apache Iceberg's Dynamic Iceberg Sink for per-record table routing and automatic schema evolution.</description>
										<content:encoded>&lt;p&gt;Handling upstream schema changes is a common operational challenge in streaming data pipelines that write to a data lake. When a source schema changes, teams often face a difficult choice: restart the pipeline or perform a manual migration. A restart can pause ingestion and delay or lose in-flight data. A manual migration consumes engineering time and introduces the risk of schema inconsistencies while the data lake falls behind the source.&lt;/p&gt;
&lt;p&gt;For example, consider an &lt;a href="https://flink.apache.org/" target="_blank" rel="noopener"&gt;Apache Flink&lt;/a&gt; job that ingests &lt;code&gt;order_events&lt;/code&gt; and writes to an Iceberg table. On Monday, the pipeline runs normally. By Wednesday, the upstream team adds a new &lt;code&gt;loyalty_tier&lt;/code&gt; field and introduces a new &lt;code&gt;interaction_events&lt;/code&gt; event type. Traditionally, you would need to stop the Flink job, update your schema definitions, and redeploy. With &lt;a href="https://iceberg.apache.org/" target="_blank" rel="noopener"&gt;Apache Iceberg’s&lt;/a&gt; &lt;a href="https://iceberg.apache.org/docs/nightly/flink-writes/#flink-dynamic-iceberg-sink" target="_blank" rel="noopener"&gt;Dynamic Iceberg Sink&lt;/a&gt; on &lt;a href="https://aws.amazon.com/managed-service-apache-flink/" target="_blank" rel="noopener"&gt;Amazon Managed Service for Apache Flink&lt;/a&gt;, the pipeline can handle both changes at the record level without disruption. The DynamicSink routes each event to the right Iceberg table and evolves table schemas as new columns appear, with no operator intervention.&lt;/p&gt;
&lt;p&gt;Managed Service for Apache Flink is a fully managed AWS service that you can use to build and deploy streaming applications without setting up infrastructure and managing resources. Apache Flink’s distributed processing engine with exactly once processing guarantees through &lt;a href="https://nightlies.apache.org/flink/flink-docs-stable/docs/dev/datastream/fault-tolerance/checkpointing/" target="_blank" rel="noopener"&gt;checkpointing&lt;/a&gt; paired with Apache Iceberg’s two-phase commit provides end-to-end consistency without duplications or data loss.&lt;/p&gt;
&lt;p&gt;In this post, we show you how to build a dynamic streaming data lake that adapts to new event types and schema changes without stopping the pipeline. Using Apache Flink 2.3 and Apache Iceberg 1.11.0 on Managed Service for Apache Flink, we walk through the DataStream API patterns for per-record table routing and automatic schema evolution. The complete implementation is available in this &lt;a href="https://github.com/aws-samples/sample-streaming-data-lake-with-apache-iceberg-and-apache-flink/tree/main" target="_blank" rel="noopener"&gt;GitHub repository&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="apache-iceberg-dynamic-sink"&gt;Apache Iceberg dynamic sink&lt;/h2&gt;
&lt;p&gt;The &lt;a href="https://iceberg.apache.org/docs/nightly/flink-writes/#flink-dynamic-iceberg-sink" target="_blank" rel="noopener"&gt;Dynamic Iceberg Sink&lt;/a&gt; allows Flink to dynamically route records to multiple Iceberg tables based on user-defined logic. It also creates and updates tables on the fly and evolves both table schemas and partition specs during streaming execution, controlled through the &lt;code&gt;DynamicRecord&lt;/code&gt; class, which eliminates the need for Flink job restarts when requirements change.&lt;/p&gt;
&lt;h3 id="per-record-table-routing-with-dynamicicebergsink"&gt;Per-record table routing with DynamicIcebergSink&lt;/h3&gt;
&lt;p&gt;The DynamicIcebergSink resolves the target table at the record level rather than at pipeline configuration time. Records flow through a DynamicRecordGenerator that, for each input, emits one or more DynamicRecord values. Each DynamicRecord carries its own target table ID, schema, partition spec, and row payload, so the sink knows where to write and how the table should look:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-java"&gt;DynamicIcebergSink.forInput(events)
    .generator(generator)
    .catalogLoader(catalogLoader)
    .immediateTableUpdate(true)
    .cacheMaxSize(cacheMaxSize)
    .cacheRefreshMs(cacheRefreshMs)
    .append();&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;The generator receives each record and emits a DynamicRecord targeting a resolved table that looks as follows:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-java"&gt;return new DynamicRecord(
    tableId,
    tableBranch,
    icebergSchema,
    rowData,
    partitionSpec,
    distributionMode,
    1);&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;The sink creates the table if it does not exist and evolves its schema when a record carries new columns. cacheMaxSize and cacheRefreshMs bound the sink’s per-table metadata cache, so a job that writes to many tables does not reload metadata on every record. immediateTableUpdate(true) controls how those catalog changes are applied, which the following section on automatic schema evolution explains. A single Flink job can ingest and route order_events, interaction_events, user_events, and future event types without additional sink definitions.&lt;/p&gt;
&lt;p&gt;However, the sink also needs to know what the table looks like. That is why every DynamicRecord also carries the Iceberg schema so that DynamicIcebergSink can create the table on first sight and evolve it as new fields appear. The schema information can be inferred from the data or read from a schema registry.&lt;/p&gt;
&lt;h3 id="automatic-schema-evolution"&gt;Automatic schema evolution&lt;/h3&gt;
&lt;p&gt;Streaming sources add new fields over time, and DynamicIcebergSink handles them without a restart. Before writing each record, it compares the record’s schema against the target table. If the record has a new field, Iceberg adds it as an optional column and commits the change with the next data file. Existing files stay valid and no table rewrite is needed. When you query older files, the new column returns null.&lt;/p&gt;
&lt;p&gt;The immediateTableUpdate setting controls where the catalog change happens. The GitHub sample repository sets immediateTableUpdate=true, so the writer subtask that sees the new schema applies the create or alter inline, before it emits the record. This gives the lowest latency but makes more concurrent calls to the catalog. When set to false, records that require a table change take a detour. Records whose table, schema, and partition spec already match the sink’s cached metadata go straight to the writers. Records that do need a change are routed, keyed by table name, to an update operator, so updates for the same table apply one at a time. Once the update commits and the cache refreshes, subsequent records match again and skip the detour. In steady state, with no schema changes arriving, this path adds no extra shuffle. Either way, the schema comparison and the resulting table change are the same.&lt;/p&gt;
&lt;p&gt;Schema changes are non-destructive by default. The sink can add new columns, widen existing types (for example, int to long or float to double), relax a required column to optional, and drop columns. Importantly, DynamicIcebergSink does not support renaming columns at the time of writing.&lt;/p&gt;
&lt;p&gt;Source schemas are identified in two ways: inferring the schema from source records (for example, JSON inference) and reading serialized records from a schema registry (for example, AWS Glue Schema Registry (GSR)). Schema evolution behavior for the Iceberg sink table depends on the schema source. JSON inference adds any new field it sees, with no contract. For example, this allows the job to initially infer a schema as an integer, and later expand to a long when larger values are detected. Schema registry serialized records define the policy using the registry’s compatibility rules (for example, BACKWARD). This means that incompatible producer changes are rejected when the schema is registered rather than at write time.&lt;/p&gt;
&lt;p&gt;The partition spec travels on each DynamicRecord, so the sink applies it when it creates or updates the table. How our sample derives that spec is covered in the partitioning section.&lt;/p&gt;
&lt;h2 id="solution-overview"&gt;Solution overview&lt;/h2&gt;
&lt;p&gt;The following diagram illustrates the solution architecture. A data generator (a local Java application) writes events to an Amazon Kinesis Data Stream. In Avro mode it also registers each event schema in the AWS Glue Schema Registry. A Managed Service for Apache Flink application consumes the stream, resolves a target Iceberg table for each record, and writes to Iceberg tables in Amazon S3, cataloged either in the AWS Glue Data Catalog or, for fully managed tables, in Amazon S3 Tables, a capability of Amazon S3.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/28/BDB-5973-1.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/28/BDB-5973-1.png" alt="Data generator sends events to Amazon Kinesis Data Streams, and Managed Service for Apache Flink routes each record to an Iceberg table in Amazon S3" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 1: Solution architecture for routing streaming records to per-event Iceberg tables on Managed Service for Apache Flink&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;At a high level, a single Managed Service for Apache Flink application reads raw records from Kinesis and resolves a target Iceberg table for each record. It uses the DynamicIcebergSink to create and evolve tables on demand. The same job handles many event types because the destination is decided per record, not per sink.&lt;/p&gt;
&lt;p&gt;A note on stream topology: the examples assume one Kinesis stream carrying multiple event types, which keeps the walkthrough focused. This is not a requirement for the pattern. If your events arrive on separate streams (for example, one stream per producer or per domain), create one KinesisStreamsSource per stream and union them into a single DataStream before the sink. The routing generator chooses the destination table from the record itself, so many sources can fan into one DynamicIcebergSink and still land in the correct tables.&lt;/p&gt;
&lt;p&gt;Unioning does not add shuffle cost. The sink always re-distributes records by an internal per-table writer key, so a unioned stream and N separate pipelines incur the same per-record exchange. The distribution mode each DynamicRecord carries only changes which writer subtask a row lands on, not whether a shuffle occurs. The real tradeoff is isolation. All tables share one writer pool, one commit aggregator, and one committer. A hot stream’s backpressure and checkpoint alignment therefore couple to every other stream, and writer parallelism is a single job-wide setting. Prefer one unioned pipeline when you have many small-to-medium event types that should pool capacity. Split into separate applications when one stream is high-volume enough to need its own writer parallelism and failure isolation.&lt;/p&gt;
&lt;p&gt;DynamicIcebergSink needs a schema for every record. The sample provides two interchangeable ways to obtain it, implemented as two generator variants: Option 1 infers the schema from each JSON record at runtime. Option 2 reads the registered schema from AWS Glue Schema Registry. Everything downstream (routing, table creation, and schema evolution) is identical, and only the generator changes.&lt;/p&gt;
&lt;h3 id="option-1-infer-the-schema-from-the-json-record"&gt;Option 1: Infer the schema from the JSON record&lt;/h3&gt;
&lt;p&gt;SchemaAgnosticRoutingGenerator implements Iceberg’s DynamicRecordGenerator. Its generate method maps the routing field to a table name, infers the schema, derives a partition spec, and emits a DynamicRecord through the collector:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-java"&gt;@Override
public void generate(JsonNode json, Collector&amp;lt;DynamicRecord&amp;gt; out) {
    String tableName = determineTableName(json); // routing field -&amp;gt; table name
    TableIdentifier tableId = TableIdentifier.of(database, tableName);
    Schema schema = inferSchemaFromJson(json); // cached by schema signature
    RowData rowData = convertJsonToRowData(json, schema);
    PartitionSpec spec = buildPartitionSpec(schema); // cached per schema
    out.collect(new DynamicRecord(
        tableId, "main", schema, rowData, spec, DistributionMode.NONE, 4));
}&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;The table name comes from an explicit table-name field when present, otherwise from the routing field (event_type by default).&lt;/p&gt;
&lt;p&gt;For schemaless or semi-structured JSON, the generator infers an Iceberg schema directly from each record. This is convenient, but inference is fundamentally lossy because JSON does not carry type information. The generator therefore applies deliberately conservative rules and selects a stable type rather than the narrowest one:&lt;/p&gt;
&lt;table border="1px" width="100%" cellpadding="10px"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;JSON value&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Iceberg type&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Integer&lt;/td&gt;
   &lt;td&gt;LongType (all integral values are widened to long)&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;String&lt;/td&gt;
   &lt;td&gt;StringType&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Floating-point values&lt;/td&gt;
   &lt;td&gt;DoubleType&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Boolean&lt;/td&gt;
   &lt;td&gt;BooleanType&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;ISO-8601 timestamps&lt;/td&gt;
   &lt;td&gt;TimestampType (microseconds)&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Nested JSON object&lt;/td&gt;
   &lt;td&gt;StructType (with fields inferred recursively)&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;JSON array&lt;/td&gt;
   &lt;td&gt;ListType (with element type inferred from array contents)&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;h4 id="partitioning-the-routed-tables"&gt;Partitioning the routed tables&lt;/h4&gt;
&lt;p&gt;Partitioning is decided by our generator, not by the sink, and the same mechanism applies to both schema options: the JSON-inference and schema-registry generators share the partition-candidate logic. The open source DynamicIcebergSink applies whatever PartitionSpec each DynamicRecord carries. Our sample’s SchemaAgnosticRoutingGenerator builds that spec at runtime: it reads a list of candidate partition fields from the partition.candidates application property and derives a per-table spec from the fields it observes. For each table, buildPartitionSpec walks that list and keeps only the candidates present in the table’s schema.&lt;/p&gt;
&lt;p&gt;The same list adapts to each table. A table with event_date and region is partitioned by identity(event_date) and identity(region). A table with none of the candidates is created unpartitioned. The resulting spec travels on each DynamicRecord, so the sink applies it when it first creates the table.&lt;/p&gt;
&lt;p&gt;For example, with partition.candidates = event_time,region,product: a table whose schema has event_time and product is created partitioned by those two. A table with only event_time gets identity(event_time). A table with none of the candidates is created unpartitioned. Partition specs are not frozen at creation time either: the sink evolves them through Iceberg partition-spec evolution, adding a candidate field when it later appears in the table’s schema and removing one that disappears. This is a metadata-only change, so existing data files keep the spec they were written with.&lt;/p&gt;
&lt;p&gt;Two operational practices follow. First, always include your event-time field among the candidates so every table is at least time-partitioned, and monitor for unpartitioned tables through the table’s $partitions metadata or its spec in the catalog: a producer that emits create_timestamp instead of event_time will silently create unpartitioned tables until the candidate list is updated. Second, be deliberate with generic fields like region. If a source produces high-cardinality values for a candidate field, you can correct the spec later. Evolution applies to newly written files only, so the small files already written remain until compaction rewrites them.&lt;/p&gt;
&lt;p&gt;Note that the candidate list is global, not per table. It tracks every field you might partition on, and each table takes only the ones it has.&lt;/p&gt;
&lt;h3 id="option-2-read-the-schema-from-a-schema-registry"&gt;Option 2: Read the schema from a schema registry&lt;/h3&gt;
&lt;p&gt;Inference is convenient but lossy, and it offers no contract: nothing stops a producer from silently changing a field’s type or meaning. The second option removes the guesswork by reading the schema from a registry instead of the data. Many production streaming platforms standardize on strongly typed Avro schemas managed through AWS Glue Schema Registry. With GSR, producers register schemas explicitly, each record on Kinesis is Avro-encoded and prefixed with a schema-version ID, and the consumer decodes against the exact registered schema. That gives you three things JSON inference cannot: precise types (a long stays a long, a timestamp-micros stays a timestamp-micros), a governed evolution policy enforced at registration, and a single source of truth shared across producers and consumers.&lt;/p&gt;
&lt;p&gt;The pattern works with any schema registry that gives consumers the writer’s schema per record. The sample implements it with AWS Glue Schema Registry, but the same generator shape applies to other registries.&lt;/p&gt;
&lt;p&gt;The dynamic-sink-avro-sample module applies GSR-managed Avro schemas to the same dynamic routing and schema evolution pattern. For each record, AvroToDynamicRecordGenerator reads the schema-version ID and fetches the writer schema from GSR, caching it after the first lookup. It then converts that schema to an Iceberg schema, decodes the payload into RowData, and emits a DynamicRecord, exactly as the JSON generator does:&lt;/p&gt;
&lt;p&gt;The sink wiring is identical to option 1. Only the generator changes, and because the source carries raw Avro bytes the input stream is byte[] rather than parsed JSON:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-java"&gt;AvroToDynamicRecordGenerator generator = new AvroToDynamicRecordGenerator(
    awsRegion, registryName, database, partitionCandidates, branch);
DynamicIcebergSink.forInput(eventBytes)
    .generator(generator)
    // identical catalogLoader, immediateTableUpdate(true), cache, and write settings as option 1
    .append();&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Because the schema comes from GSR rather than from inspecting bytes, the Avro-to-Iceberg type mapping is exact:&lt;/p&gt;
&lt;table border="1px" width="100%" cellpadding="10px"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;Category&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Avro type&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Iceberg type&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Primitive&lt;/td&gt;
   &lt;td&gt;int&lt;/td&gt;
   &lt;td&gt;IntegerType&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Primitive&lt;/td&gt;
   &lt;td&gt;long&lt;/td&gt;
   &lt;td&gt;LongType&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Primitive&lt;/td&gt;
   &lt;td&gt;float&lt;/td&gt;
   &lt;td&gt;FloatType&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Primitive&lt;/td&gt;
   &lt;td&gt;double&lt;/td&gt;
   &lt;td&gt;DoubleType&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Primitive&lt;/td&gt;
   &lt;td&gt;string&lt;/td&gt;
   &lt;td&gt;StringType&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Primitive&lt;/td&gt;
   &lt;td&gt;boolean&lt;/td&gt;
   &lt;td&gt;BooleanType&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Logical&lt;/td&gt;
   &lt;td&gt;timestamp-millis&lt;/td&gt;
   &lt;td&gt;TimestampType (preserves millisecond precision)&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Logical&lt;/td&gt;
   &lt;td&gt;timestamp-micros&lt;/td&gt;
   &lt;td&gt;TimestampType (preserves microsecond precision)&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Logical&lt;/td&gt;
   &lt;td&gt;decimal&lt;/td&gt;
   &lt;td&gt;DecimalType&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Complex&lt;/td&gt;
   &lt;td&gt;record&lt;/td&gt;
   &lt;td&gt;StructType (nested fields mapped recursively)&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Complex&lt;/td&gt;
   &lt;td&gt;array&lt;/td&gt;
   &lt;td&gt;ListType (element type inferred from items schema)&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Complex&lt;/td&gt;
   &lt;td&gt;map&lt;/td&gt;
   &lt;td&gt;MapType (keys are always StringType)&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The GSR integration handles schema versioning transparently. As soon as a producer registers a new schema version containing additional fields, the Flink consumer deserializes the updated payload and evolves the Iceberg table to match, with no job restart.&lt;/p&gt;
&lt;h2 id="prerequisites"&gt;Prerequisites&lt;/h2&gt;
&lt;p&gt;To follow along, you need the following:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;An AWS account with permissions to create Amazon Kinesis Data Streams, Managed Service for Apache Flink applications, AWS Glue resources, and Amazon S3 buckets (plus Amazon S3 Tables if you choose that catalog).&lt;/li&gt;
 &lt;li&gt;The AWS Command Line Interface (AWS CLI) configured with credentials.&lt;/li&gt;
 &lt;li&gt;&lt;code&gt;Node.js&lt;/code&gt; 18 or later and the AWS Cloud Development Kit (AWS CDK) CLI.&lt;/li&gt;
 &lt;li&gt;Java 17 or later and Apache Maven 3.9 or later, to build the data generator.&lt;/li&gt;
 &lt;li&gt;Docker running locally. The CDK build bundles the Flink application jars inside a Maven image.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="deploy-and-test-the-solution"&gt;Deploy and test the solution&lt;/h2&gt;
&lt;p&gt;The accompanying repository provisions everything through a single parameterized AWS CDK stack.&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;Install the CDK dependencies and bootstrap your environment (first time only):
  &lt;div class="hide-language"&gt;
   &lt;pre&gt;&lt;code class="language-bash"&gt;cd cdk-infrastructure &amp;amp;&amp;amp; npm install
npx cdk bootstrap aws://&amp;lt;account&amp;gt;/&amp;lt;region&amp;gt;&lt;/code&gt;&lt;/pre&gt;
  &lt;/div&gt;&lt;/li&gt;
 &lt;li&gt;Deploy the variant you want to try:
  &lt;div class="hide-language"&gt;
   &lt;pre&gt;&lt;code class="language-bash"&gt;npx cdk deploy -c appType=dynamic -c tableFormatVersion=2 # JSON inference variant
npx cdk deploy -c appType=dynamic-avro -c tableFormatVersion=2 # GSR Avro variant&lt;/code&gt;&lt;/pre&gt;
  &lt;/div&gt;
  &lt;p&gt;Add -c catalogType=s3tables to either command to use Amazon S3 Tables instead of the AWS Glue Data Catalog. The walkthrough sets tableFormatVersion=2 so you can query the results with a broad range of engines. Omit it to use the default, Iceberg format version 3, when you query with a v3-aware engine such as Spark on Amazon EMR 7.12+ or AWS Glue ETL.&lt;/p&gt;&lt;/li&gt;
 &lt;li&gt;Start the application using the ApplicationName value from the stack outputs:
  &lt;div class="hide-language"&gt;
   &lt;pre&gt;&lt;code class="language-bash"&gt;aws kinesisanalyticsv2 start-application --application-name &amp;lt;ApplicationName&amp;gt; --run-configuration 'ApplicationRestoreConfiguration={ApplicationRestoreType=SKIP_RESTORE_FROM_SNAPSHOT}'&lt;/code&gt;&lt;/pre&gt;
  &lt;/div&gt;&lt;/li&gt;
 &lt;li&gt;Send test events with the included data generator. Start with the v1 payloads, which create the tables without the optional fields:
  &lt;div class="hide-language"&gt;
   &lt;pre&gt;&lt;code class="language-bash"&gt;java -jar data-generator/target/data-generator-1.0-SNAPSHOT.jar &amp;lt;stream-name&amp;gt; &amp;lt;region&amp;gt; 100 60 v1&lt;/code&gt;&lt;/pre&gt;
  &lt;/div&gt;
  &lt;p&gt;Then send v2 payloads, which add the userAgent and scrollDepth fields. This second run is the schema evolution you observe in the next step:&lt;/p&gt;
  &lt;div class="hide-language"&gt;
   &lt;pre&gt;&lt;code class="language-bash"&gt;java -jar data-generator/target/data-generator-1.0-SNAPSHOT.jar &amp;lt;stream-name&amp;gt; &amp;lt;region&amp;gt; 100 60 v2&lt;/code&gt;&lt;/pre&gt;
  &lt;/div&gt;
  &lt;p&gt;For the Avro variant, the generator registers each schema version in the AWS Glue Schema Registry as it sends:&lt;/p&gt;
  &lt;div class="hide-language"&gt;
   &lt;pre&gt;&lt;code class="language-bash"&gt;java -jar data-generator/target/data-generator-1.0-SNAPSHOT.jar avro &amp;lt;stream-name&amp;gt; &amp;lt;region&amp;gt; &amp;lt;registry-name&amp;gt; 100 60&lt;/code&gt;&lt;/pre&gt;
  &lt;/div&gt;&lt;/li&gt;
 &lt;li&gt;Query the routed tables in Amazon Athena. You should see one Iceberg table per event type appear in the database within a checkpoint interval, and after sending v2 events, the new fields (userAgent, scrollDepth) show up as optional columns on the same tables. The Iceberg metadata tables (for example, &lt;code&gt;SELECT * FROM "db"."table$snapshots"&lt;/code&gt;) show each commit the sink makes.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="clean-up"&gt;Clean up&lt;/h2&gt;
&lt;p&gt;When you finish testing, delete the resources to stop incurring charges:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;cd cdk-infrastructure &amp;amp;&amp;amp; npx cdk destroy&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;CDK removes the Kinesis Data Stream, the Managed Service for Apache Flink application, and the stack-created AWS Identity and Access Management (IAM) roles. Additionally, empty and delete the S3 warehouse bucket to remove the Iceberg data and metadata files, delete any schemas the Avro variant registered in the AWS Glue Schema Registry, and delete the table bucket contents if you used the S3 Tables catalog.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;With Apache Iceberg 1.11.0 and Flink 2.3, you can build streaming data lake architectures that adapt to change without stopping the pipeline. With per-record routing, a single Flink application can write multiple event types to separate Iceberg tables, while automatic schema evolution keeps table definitions aligned with changing source data. Choosing AWS Glue Schema Registry over runtime JSON inference adds precise types and a governed evolution contract, and a configurable partition-candidate list keeps each routed table partitioned correctly without pre-declaring its schema.&lt;/p&gt;
&lt;p&gt;The result is fewer pipeline redeployments, reduced operational overhead, and a data lake that remains synchronized with evolving application schemas.&lt;/p&gt;
&lt;p&gt;To get started, follow the deploy and test section, then adapt the routing field and partition candidates to your own event types.&lt;/p&gt;
&lt;p&gt;The full sample code is available in the &lt;a href="https://github.com/aws-samples/sample-streaming-data-lake-with-apache-iceberg-and-apache-flink" target="_blank" rel="noopener"&gt;accompanying GitHub repository&lt;/a&gt;.&lt;/p&gt;
&lt;hr style="width: 100%"&gt;
&lt;h2&gt;About the authors&lt;/h2&gt;
&lt;footer&gt;
 &lt;div class="blog-author-box"&gt;
  &lt;div class="blog-author-image"&gt;
   &lt;p&gt;&lt;img loading="lazy" class="alignleft size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/28/BDB-5973-2.jpg" alt="Francisco Morillo" width="100" height="100"&gt;&lt;/p&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Francisco Morillo&lt;/h3&gt;
  &lt;p&gt;Francisco is a Sr.&amp;nbsp;Streaming Solutions Architect at AWS, specializing in real-time analytics architectures. With over five years in the streaming data space, Francisco has worked as a data analyst for startups and as a big data engineer for consultancies, building streaming data pipelines. He has deep expertise in Amazon Managed Streaming for Apache Kafka (Amazon MSK) and Amazon Managed Service for Apache Flink.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box"&gt;
  &lt;div class="blog-author-image"&gt;
   &lt;p&gt;&lt;img loading="lazy" class="alignleft size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/28/BDB-5973-3.jpg" alt="Felix John" width="100" height="100"&gt;&lt;/p&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Felix John&lt;/h3&gt;
  &lt;p&gt;&lt;a href="https://www.linkedin.com/in/fel-john/" target="_blank" rel="noopener"&gt;Felix&lt;/a&gt;&amp;nbsp;is a Global Solutions Architect and data &amp;amp; AI expert at AWS, based out of Germany. He focuses on supporting AWS’ strategic global automotive &amp;amp; manufacturing customers on their data &amp;amp; AI transformation journey.&lt;/p&gt;
 &lt;/div&gt;
&lt;/footer&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>Observing and evaluating production agents using OpenSearch Agent Health</title>
		<link>https://aws.amazon.com/blogs/big-data/observing-and-evaluating-production-agents-using-opensearch-agent-health/</link>
		
		<dc:creator><![CDATA[Ulrich Hinze]]></dc:creator>
		<pubDate>Tue, 01 Sep 2026 16:19:02 +0000</pubDate>
				<category><![CDATA[Advanced (300)]]></category>
		<category><![CDATA[Amazon OpenSearch Service]]></category>
		<category><![CDATA[Technical How-to]]></category>
		<guid isPermaLink="false">b0b6f1eac15ab79405c85629aa2ce9e8420cb617</guid>

					<description>Learn how to observe and evaluate production AI agents by combining an agent running on AWS with OpenSearch Agent Health. This post walks through deploying an agent and its observability pipeline to AWS, then using Agent Health to explore traces and run evaluations that measure and improve agent quality over time.</description>
										<content:encoded>&lt;p&gt;As AI agents are moving from experimental prototypes to production workloads, teams need visibility into what agents are doing and a systematic way to measure whether they’re doing it well. Traditional testing methodologies like unit and integration tests fall short for this task, as measuring an agent’s quality isn’t a straightforward true/false decision. Instead, agent observability and evaluations (evals for short) provide a two-legged solution to this problem. Agent observability captures the details of an agent’s behavior, and evals compare this behavior to the behavior that you want. With this approach, teams can monitor their agent’s quality over time and introduce agent-specific quality gates in their software development lifecycle.&lt;/p&gt;
&lt;p&gt;In this post, we show how to combine an AI agent running on AWS with &lt;a href="https://github.com/opensearch-project/agent-health" target="_blank" rel="noopener"&gt;OpenSearch Agent Health&lt;/a&gt; for observability and evals. You will deploy an agent and its observability data pipeline to AWS, then use Agent Health as a local development tool connecting to your cloud resources.&lt;/p&gt;
&lt;h2 id="overview-of-solution"&gt;Overview of solution&lt;/h2&gt;
&lt;p&gt;Agent observability and evaluations rely on &lt;a href="https://opentelemetry.io/docs/concepts/signals/traces/" target="_blank" rel="noopener"&gt;OpenTelemetry traces&lt;/a&gt; to understand agent behavior. Traces describe the flow of a request through components of a system. OpenSearch Agent Health is a purpose-built tool for analyzing agent traces and running evaluations against an agent for quality control. Although Agent Health works with any open source OpenSearch installation, many AWS customers choose &lt;a href="https://docs.aws.amazon.com/opensearch-service/latest/developerguide/ingestion.html" target="_blank" rel="noopener"&gt;Amazon OpenSearch Ingestion&lt;/a&gt; and &lt;a href="https://aws.amazon.com/opensearch-service/" target="_blank" rel="noopener"&gt;Amazon OpenSearch Service&lt;/a&gt; for ingesting and storing their OpenTelemetry data. You can connect OpenSearch Agent Health to these AWS resources to fetch live data and store its own configuration and evaluation history.&lt;/p&gt;
&lt;p&gt;The following diagram shows the overall architecture of the solution presented in this post: &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/26/BDB-5986-1.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/26/BDB-5986-1.png" alt="Architecture diagram showing the agent, Amazon OpenSearch Ingestion, Amazon OpenSearch Service, and OpenSearch Agent Health observability and evaluation flow" width="800"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Figure 1: Solution overview&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;The individual parts are:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;&lt;a href="https://aws.amazon.com/amplify/" target="_blank" rel="noopener"&gt;AWS Amplify&lt;/a&gt; for hosting an &lt;a href="https://github.com/assistant-ui/assistant-ui" target="_blank" rel="noopener"&gt;assistant-ui&lt;/a&gt; chat interface. Connects to the agent backend using the &lt;a href="https://docs.ag-ui.com/introduction" target="_blank" rel="noopener"&gt;Agent-User Interaction (AG-UI)&lt;/a&gt; protocol.&lt;/li&gt;
 &lt;li&gt;Sample ecommerce AI agent using Strands Agents SDK, deployed to &lt;a href="https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/agents-tools-runtime.html" target="_blank" rel="noopener"&gt;Amazon Bedrock AgentCore runtime&lt;/a&gt;, exposing an &lt;a href="https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/runtime-agui.html" target="_blank" rel="noopener"&gt;AG-UI Server-Sent Events (SSE) endpoint&lt;/a&gt;. This agent has access to multiple tools, such as product search and shopping basket operations. For this sample project, the tool calls are all simulated within the agent runtime rather than including API calls to other systems. The agent emits messages, reasoning steps, and tool calls as OpenTelemetry traces.&lt;/li&gt;
 &lt;li&gt;Large language models (LLMs) on Amazon Bedrock. One model (&lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-amazon-nova-2-lite.html" target="_blank" rel="noopener"&gt;Amazon Nova 2 Lite&lt;/a&gt;) is used to power the agent, the other model (&lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-anthropic-claude-opus-4-6.html" target="_blank" rel="noopener"&gt;Anthropic Claude Opus 4.6&lt;/a&gt;) is used to evaluate the agent behavior.&lt;/li&gt;
 &lt;li&gt;Amazon OpenSearch Ingestion for collecting and transforming the raw agent traces and loading them into an Amazon OpenSearch Service domain. Agent traces have the same structure as regular OpenTelemetry traces, with the addition of &lt;a href="https://opentelemetry.io/docs/specs/semconv/gen-ai/" target="_blank" rel="noopener"&gt;generative AI semantics&lt;/a&gt; (for example, tool calls and token usage). This means a &lt;a href="https://docs.opensearch.org/latest/data-prepper/pipelines/configuration/sources/otlp-source/" target="_blank" rel="noopener"&gt;regular OpenTelemetry pipeline configuration&lt;/a&gt; can be used to process agent traces.&lt;/li&gt;
 &lt;li&gt;OpenSearch Agent Health for analyzing traces and running evaluation test cases and benchmarks against the agent. Agent Health uses the same AG-UI endpoint as the front-end application. It authenticates to the application, to Amazon Bedrock for model functionality, and to Amazon OpenSearch Service using AWS SigV4 authentication.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="walkthrough"&gt;Walkthrough&lt;/h2&gt;
&lt;p&gt;In this walkthrough, we showcase how you can use Agent Health and Strands to measure and improve your agent’s quality over time.&lt;/p&gt;
&lt;p&gt;We follow these steps:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;Deploy solution to AWS and test the application.&lt;/li&gt;
 &lt;li&gt;Start Agent Health locally and connect it to cloud resources.&lt;/li&gt;
 &lt;li&gt;Explore agent traces and run evaluations.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;We have created a &lt;a href="https://github.com/aws-samples/sample-agent-health-with-amazon-opensearch-service" target="_blank" rel="noopener"&gt;GitHub repository&lt;/a&gt; for you to follow along.&lt;/p&gt;
&lt;h3 id="prerequisites"&gt;Prerequisites&lt;/h3&gt;
&lt;p&gt;For this walkthrough, you should have the following prerequisites:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;An &lt;a href="https://signin.aws.amazon.com/signin?redirect_uri=https%3A%2F%2Fportal.aws.amazon.com%2Fbilling%2Fsignup%2Fresume&amp;amp;client_id=signup" target="_blank" rel="noopener"&gt;AWS account&lt;/a&gt;&lt;/li&gt;
 &lt;li&gt;Git&lt;/li&gt;
 &lt;li&gt;Node.js&lt;/li&gt;
 &lt;li&gt;AWS Cloud Development Kit (AWS CDK)&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="deploy-solution-to-aws-and-test-the-application"&gt;Deploy solution to AWS and test the application&lt;/h3&gt;
&lt;p&gt;In this section, you check out the repository and deploy the infrastructure to AWS. Be aware that these steps create AWS resources that incur cost. We cover cleanup steps at the end of this post.&lt;/p&gt;
&lt;p&gt;First, clone the repository to a local directory:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;git clone https://github.com/aws-samples/sample-agent-health-with-amazon-opensearch-service &amp;amp;&amp;amp; cd sample-agent-health-with-amazon-opensearch-service&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Switch to the infra folder and install dependencies:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;cd infra &amp;amp;&amp;amp; npm install&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Before you can start the deployment, determine the AWS Identity and Access Management (IAM) user or role that you will use to start Agent Health later on. In many cases, this will be the same role that you use to deploy the infrastructure. Set this ARN in your environment by issuing the following command:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;export AGENT_HEALTH_READER_ARN=arn:aws:iam::&amp;lt;YOUR_ACCOUNT_ID&amp;gt;:role/&amp;lt;YOUR_ROLE_NAME&amp;gt;&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Bootstrap your AWS account for use with AWS CDK:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;cdk bootstrap -c agentHealthReaderArn=$AGENT_HEALTH_READER_ARN&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Run the infrastructure deployment. Review and acknowledge IAM statement changes when prompted. This takes around 25 minutes to complete:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;cdk deploy -R -c agentHealthReaderArn=$AGENT_HEALTH_READER_ARN&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;The &lt;code&gt;-R&lt;/code&gt; parameter defines that if something fails during this deployment, the successfully provisioned resources are retained. Be aware that this command creates AWS resources, incurring cost. Review the cleanup section at the end of this post for removing all created resources.&lt;/p&gt;
&lt;p&gt;When the deploy command finishes successfully, you should see an output like the following:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-plaintext"&gt;...
AgentObservabilityStack

✨ Deployment time: 1402.42 s

Outputs:
AgentObservabilityStack.AgentEndpoint = https://bedrock-agentcore.us-east-1.amazonaws.com/runtimes/arn%3Aaws%3Abedrock-agentcore%3Aus-east-1%3A123456789012%3Aruntime%2Fretail_agent-abcdefghij/invocations?qualifier=AgentObservabilityStAgentRuntimeEndpointABCDEFGH
AgentObservabilityStack.ChatUrl = https://main.abcdefghijklmn.amplifyapp.com
...&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Next, create a user for your application. Retrieve the CDK output value for &lt;code&gt;AgentObservabilityStack.UserPoolId&lt;/code&gt;. Create a user for the application using the user pool ID, an email address, and a strong password (minimum eight characters including uppercase, lowercase, letter, and digit):&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;export COGNITO_EMAIL=&amp;lt;YOUR_EMAIL&amp;gt;
export COGNITO_PASSWORD=&amp;lt;YOUR_PASSWORD&amp;gt;
export USER_POOL=&amp;lt;YOUR_USER_POOL_ID&amp;gt;
aws cognito-idp admin-create-user --user-pool-id $USER_POOL --username $COGNITO_EMAIL --message-action SUPPRESS --user-attributes Name=email_verified,Value=true
aws cognito-idp admin-set-user-password --user-pool-id $USER_POOL --username $COGNITO_EMAIL --password "$COGNITO_PASSWORD" --permanent&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;You can now access the retail agent application. From the CDK output values, retrieve the value for &lt;code&gt;AgentObservabilityStack.ChatUrl&lt;/code&gt;. Copy and paste this URL into your browser. Log in with your email and password. You should now see the agent interface:&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/26/BDB-5986-2.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/26/BDB-5986-2.png" alt="Sample retail agent chat interface showing the ecommerce assistant ready for queries" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 2: Sample retail agent user interface&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;Experiment with the application. Here is an example sequence of queries you can put in:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;Do you have books on Python?&lt;/li&gt;
 &lt;li&gt;Is this in stock?&lt;/li&gt;
 &lt;li&gt;Put it into my basket.&lt;/li&gt;
 &lt;li&gt;What else can you do for me?&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="start-agent-health"&gt;Start Agent Health&lt;/h3&gt;
&lt;p&gt;Now that you have the infrastructure running, you can start OpenSearch Agent Health locally and connect it to your cloud resources.&lt;/p&gt;
&lt;p&gt;The CDK infrastructure deployment created a file &lt;code&gt;cdk-output.json&lt;/code&gt;, which contains all relevant configuration values for Agent Health. We’ve already created a file &lt;code&gt;agent-health/agent-health.config.ts&lt;/code&gt; that pulls these values dynamically in your environment, so you can start Agent Health without any further configuration.&lt;/p&gt;
&lt;p&gt;Open a terminal and start Agent Health by running the following command:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;cd ../agent-health &amp;amp;&amp;amp; npm install
npx @opensearch-project/agent-health&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;Open &lt;a class="uri" href="http://localhost:4001" target="_blank" rel="noopener"&gt;http://localhost:4001&lt;/a&gt; in your browser to access Agent Health UI. Choose &lt;strong&gt;Agent Traces&lt;/strong&gt; in the sidebar menu to access your agent’s traces. You should see traces from your previous interactions:&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/26/BDB-5986-3.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/26/BDB-5986-3.png" alt="Agent Health Traces view listing agent traces captured from previous interactions" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 3: Agent traces. As Agent Health is in active development, this interface might have changed since the time of writing&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;Expand the trace and explore the information it contains, such as token count and agent trajectory (sequence of messages, reasoning steps, and tool calls).&lt;/p&gt;
&lt;p&gt;If you’re unable to access the application or see any traces, verify the following:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;Check Agent Health logs in your terminal for any errors. Also check whether Agent Health is running on an alternative port, like 4002 instead of 4001.&lt;/li&gt;
 &lt;li&gt;If there are permission errors when accessing traces from OpenSearch, verify that the AWS credentials in your terminal match the principal (user or role) that you specified under the &lt;code&gt;agentHealthReaderArn&lt;/code&gt; CDK parameter during &lt;code&gt;cdk deploy&lt;/code&gt;. This principal must have &lt;code&gt;ESHttpGet:*&lt;/code&gt; IAM permissions. Agent Health uses your current AWS credentials to access the OpenSearch API for querying traces. The OpenSearch API is guarded by both IAM and &lt;a href="https://docs.aws.amazon.com/opensearch-service/latest/developerguide/fgac.html" target="_blank" rel="noopener"&gt;OpenSearch fine-grained access control&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="create-and-run-a-test"&gt;Create and run a test&lt;/h3&gt;
&lt;p&gt;Choose &lt;strong&gt;Test Cases&lt;/strong&gt; and &lt;strong&gt;New Test Case&lt;/strong&gt;. Fill out the required fields with the following information:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;Name: &lt;strong&gt;Should add to cart&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Initial Prompt: &lt;strong&gt;Add some wireless headphones to my cart. Take any that you have in stock.&lt;/strong&gt;&lt;/li&gt;
 &lt;li&gt;Expected Outcomes: &lt;strong&gt;PROD-001 added to cart&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Back in the test cases overview, select the created test case and choose &lt;strong&gt;Run Test&lt;/strong&gt;. In the &lt;strong&gt;Configure Run&lt;/strong&gt; dialog, choose &lt;strong&gt;Retail Assistant (production)&lt;/strong&gt; for &lt;strong&gt;Agent&lt;/strong&gt;, &lt;strong&gt;Tool Usage Efficiency&lt;/strong&gt; for &lt;strong&gt;Evaluator&lt;/strong&gt;, &lt;strong&gt;Claude Opus 4.8&lt;/strong&gt; for &lt;strong&gt;Judge Model&lt;/strong&gt;, and choose &lt;strong&gt;Start Run&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Agent Health now runs the configured initial prompt against the agent. The agent completes the task and sends execution traces to OpenSearch. Agent Health uses an evaluation model to check both agent responses and traces on successful execution, according to the defined expected outcomes. After the test is completed, go through the different tabs to check the test results.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/26/BDB-5986-4.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/26/BDB-5986-4.png" alt="Agent Health evaluation report showing test results across multiple tabs" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 4: Agent Health evaluation report&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;If you’re unable to run the test, check the following:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;Agent Health automatically creates an Amazon Cognito token for your user upon start, but this token can expire. Restarting Agent Health creates a new token. Verify that both the &lt;code&gt;COGNITO_EMAIL&lt;/code&gt; and &lt;code&gt;COGNITO_PASSWORD&lt;/code&gt; variables are still set in your terminal environment.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="beyond-test-cases"&gt;Beyond test cases&lt;/h3&gt;
&lt;p&gt;After running a single test case, choose &lt;strong&gt;Benchmarks&lt;/strong&gt; in the sidebar menu. With Benchmarks, you can run multiple test cases in parallel and summarize their results. You can compare benchmark runs by choosing &lt;strong&gt;Evaluation Runs&lt;/strong&gt; in the sidebar, where you can analyze trends in pass rate, cost, and duration over time. Lastly, choose &lt;strong&gt;Evaluators&lt;/strong&gt; to define your own evaluation logic beyond the predefined ones.&lt;/p&gt;
&lt;p&gt;You can also run Agent Health tests with its command-line interface, which is handy for automation and continuous integration (CI). The equivalent command of running the preceding test is:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;npx @opensearch-project/agent-health run -t &amp;lt;TEST_CASE_ID&amp;gt; -a "Retail Assistant" -e system-tool-usage --judge-model claude-opus-4.8 -e system-tool-usage&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;where &lt;code&gt;TEST_CASE_ID&lt;/code&gt; can be retrieved from the browser URL when you visit the Agent Health UI (test case IDs start with &lt;strong&gt;tc-&lt;/strong&gt;).&lt;/p&gt;
&lt;p&gt;Agent Health stores all test cases, other configuration, and reports locally on disk in the &lt;code&gt;agent-health/agent-health-data&lt;/code&gt; directory.&lt;/p&gt;
&lt;h3 id="cleaning-up"&gt;Cleaning up&lt;/h3&gt;
&lt;p&gt;To avoid incurring future charges, delete the resources:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;cd ../infra &amp;amp;&amp;amp; cdk destroy -c agentHealthReaderArn=$AGENT_HEALTH_READER_ARN&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;In this post, you learned to set up and use OpenSearch Agent Health for production agent observability and evaluations. To discover more features, see the Agent Health &lt;a href="https://observability.opensearch.org/docs/agent-health/" target="_blank" rel="noopener"&gt;documentation pages&lt;/a&gt;. You can discuss and request additional features, and get help with setup, through the issues in the &lt;a href="https://github.com/opensearch-project/agent-health/" target="_blank" rel="noopener"&gt;GitHub project&lt;/a&gt;. For more information, see the &lt;a href="https://docs.aws.amazon.com/opensearch-service/latest/developerguide/observability.html" target="_blank" rel="noopener"&gt;observability documentation&lt;/a&gt; for Amazon OpenSearch Service, where you can learn about the features available to build observability for both agents and traditional systems using OpenSearch. To investigate issues in production AI agents, see the recent post &lt;a href="https://aws.amazon.com/blogs/big-data/unified-observability-in-amazon-opensearch-service-metrics-traces-and-ai-agent-debugging-in-a-single-interface/" target="_blank" rel="noopener"&gt;Unified observability in Amazon OpenSearch Service&lt;/a&gt;.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;About the authors&lt;/h2&gt;
&lt;footer&gt;
 &lt;div class="blog-author-box"&gt;
  &lt;div class="blog-author-image"&gt;
   &lt;img loading="lazy" class="alignleft size-full wp-image-93941" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/28/BDB-5986-5.jpg" alt="" width="100" height="126"&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Ulli Hinze&lt;/h3&gt;
  &lt;p&gt;Ulli is a Solutions Architect based in Berlin, Germany. He focuses on SaaS, agentic AI, and OpenSearch, and helps customers build and modernize their solutions on AWS. His previous roles included software development, platform engineering, and architecture.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box"&gt;
  &lt;div class="blog-author-image"&gt;
   &lt;img loading="lazy" class="alignleft size-full wp-image-93942" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/28/BDB-5986-6-1.jpeg" alt="" width="100" height="133"&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Megha Goyal&lt;/h3&gt;
  &lt;p&gt;Megha is a Senior Software Engineer at AWS OpenSearch. For the past year she has focused on AI agent observability and evaluations, building Agent Health — an open-source developer tool for agents. Previously, she worked on data integrations with Amazon CloudWatch and Amazon Security Lake for the observability and security space. When she’s not building software, she enjoys designing and 3D-printing models at home.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box"&gt;
  &lt;div class="blog-author-image"&gt;
   &lt;p&gt;&lt;img loading="lazy" class="alignleft size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/26/BDB-5986-7.jpg" alt="Rekha Thottan" width="100" height="100"&gt;&lt;/p&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Rekha Thottan&lt;/h3&gt;
  &lt;p&gt;Rekha is a Senior Product Manager Technical on the Amazon OpenSearch Service team.&lt;/p&gt;
 &lt;/div&gt;
&lt;/footer&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>Accelerate Apache Spark debugging on Amazon EMR with AWS DevOps Agent</title>
		<link>https://aws.amazon.com/blogs/big-data/accelerate-apache-spark-debugging-on-amazon-emr-with-aws-devops-agent/</link>
		
		<dc:creator><![CDATA[Kalyan Janaki]]></dc:creator>
		<pubDate>Tue, 01 Sep 2026 16:06:44 +0000</pubDate>
				<category><![CDATA[Advanced (300)]]></category>
		<category><![CDATA[Amazon EMR]]></category>
		<category><![CDATA[Technical How-to]]></category>
		<guid isPermaLink="false">aca2a06510bb5601101dc9ed433507ac09b5ecf5</guid>

					<description>Extend AWS DevOps Agent to investigate Apache Spark failures on Amazon EMR. This post shows how to register the Apache Spark Troubleshooting Agent for Amazon EMR as a custom MCP capability provider over AWS PrivateLink, so a single agent chat session diagnoses a failing Spark job from an Amazon CloudWatch alarm to a line-numbered root cause.</description>
										<content:encoded>&lt;p&gt;When an Apache Spark job fails on Amazon EMR, the root cause can hide in executor logs, memory profiles, or application code. As data pipelines grow in complexity, correlating logs, metrics, and traces across multiple services requires significant operational effort. &lt;a href="https://aws.amazon.com/devops-agent/" target="_blank" rel="noopener"&gt;AWS DevOps Agent&lt;/a&gt; handles this investigation autonomously while keeping operators in the loop to review findings and approve fixes. From a single chat prompt, it produces a root cause and mitigation plan, often without any human involvement beyond the initial question.&lt;/p&gt;
&lt;p&gt;The native AWS API tools in AWS DevOps Agent don’t extend into Spark-internal artifacts. Sometimes those tools can’t reach the evidence that pins down the root cause: a Spark History Server event log, executor Python worker memory, or a line of code that allocated too much. In these cases, AWS DevOps Agent can describe symptoms (“the executor exited with code 1”) but can’t identify the actual antipattern that caused them.&lt;/p&gt;
&lt;p&gt;This post shows how to extend AWS DevOps Agent to investigate failures in Apache Spark workloads on Amazon EMR. You register the &lt;a href="https://docs.aws.amazon.com/emr/latest/ReleaseGuide/spark-troubleshoot.html" target="_blank" rel="noopener"&gt;&lt;strong&gt;Apache Spark Troubleshooting Agent for Amazon EMR&lt;/strong&gt;&lt;/a&gt;, a managed Model Context Protocol (MCP) server hosted by AWS, as a custom capability provider in your AWS DevOps Agent space. You route the traffic over AWS PrivateLink so MCP calls never traverse the public internet. Then you watch a single agent chat session investigate a deliberately failing Spark job, from Amazon CloudWatch alarm to line-numbered root cause, in about two minutes.&lt;/p&gt;
&lt;h2 id="prerequisites"&gt;Prerequisites&lt;/h2&gt;
&lt;p&gt;Before you begin, make sure you have the following:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;An AWS account with permissions to deploy &lt;a href="https://aws.amazon.com/cloudformation/" target="_blank" rel="noopener"&gt;AWS CloudFormation&lt;/a&gt; stacks that create &lt;a href="https://aws.amazon.com/iam/" target="_blank" rel="noopener"&gt;AWS Identity and Access Management&lt;/a&gt; (IAM) roles and &lt;a href="https://aws.amazon.com/vpc/" target="_blank" rel="noopener"&gt;Amazon Virtual Private Cloud&lt;/a&gt; (Amazon VPC) resources.&lt;/li&gt;
 &lt;li&gt;The latest version of the &lt;a href="https://docs.aws.amazon.com/cli/latest/userguide/getting-started-install.html" target="_blank" rel="noopener"&gt;AWS Command Line Interface (AWS CLI)&lt;/a&gt;, Boto3, and Botocore, installed and configured with credentials for the same account and AWS Region.&lt;/li&gt;
 &lt;li&gt;This walkthrough assumes you are familiar with AWS DevOps Agent and know how to trigger an investigation. For an introduction, see Getting started with AWS DevOps Agent.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="how-aws-devops-agent-discovers-custom-tools-through-mcp"&gt;How AWS DevOps Agent discovers custom tools through MCP&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://modelcontextprotocol.io/" target="_blank" rel="noopener"&gt;Model Context Protocol (MCP)&lt;/a&gt; is an open standard that defines how AI agents discover and invoke external tools. AWS DevOps Agent supports connecting to custom MCP servers, which means you can expose new capabilities to it without modifying the agent itself. When you connect an MCP server to AWS DevOps Agent, the agent automatically discovers the available tools, understands their schemas, and calls them as part of its investigation workflow. You build and connect the MCP server, and the agent handles the rest.&lt;/p&gt;
&lt;p&gt;MCP tools sit alongside the agent’s built-in AWS API tools. During a single investigation, the agent can interleave calls to &lt;code&gt;cloudwatch.describe-alarms&lt;/code&gt;, &lt;code&gt;emr-serverless.get-job-run&lt;/code&gt;, and a custom MCP tool such as &lt;code&gt;analyze_spark_workload&lt;/code&gt;. The agent picks the right one for each subtask. You augment the agent’s reach without replacing what it already does.&lt;/p&gt;
&lt;p&gt;For this integration, you don’t build an MCP server. The Apache Spark Troubleshooting Agent for Amazon EMR is itself a managed MCP server, hosted by AWS at a regional endpoint. Your job is to register that endpoint with AWS DevOps Agent and authorize the agent to call it. This requires a network path from the agent to the endpoint, plus an IAM role for AWS Signature Version 4 request signing.&lt;/p&gt;
&lt;h2 id="why-spark-internals-visibility-matters"&gt;Why Spark internals visibility matters&lt;/h2&gt;
&lt;p&gt;The actual root cause for a Spark failure usually lives somewhere none of those APIs (such as Amazon CloudWatch Logs Insights, AWS CloudTrail, or Amazon EMR step-status calls) can reach:&lt;/p&gt;
&lt;p&gt;The Apache Spark Troubleshooting Agent for Amazon EMR reads the following sources.&lt;/p&gt;
&lt;p&gt;The Spark History Server event log is a per-job archive in Amazon Simple Storage Service (Amazon S3) with stage timings, task-level metrics, executor utilization, shuffle read/write volumes, and garbage-collection pauses. Amazon EMR exposes this data through the Spark UI on Amazon EMR Serverless, Amazon EMR on Amazon Elastic Compute Cloud (Amazon EC2), and Amazon EMR on Amazon Elastic Kubernetes Service (Amazon EKS), but interpreting signals like data skew, executor memory pressure, or stages that take significantly longer than expected requires familiarity with Spark internals.&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;The Spark query plan&lt;/strong&gt; — the logical and physical plan the driver compiled. Without it, you can’t identify antipatterns such as unnecessary data repartitioning or missing broadcast hints that trigger expensive shuffles.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;The application source code&lt;/strong&gt; in Amazon S3 — the &lt;code&gt;.py&lt;/code&gt; or &lt;code&gt;.jar&lt;/code&gt; code artifact the job ran. Without it, you can’t quote the offending line of a &lt;code&gt;mapPartitions&lt;/code&gt; user-defined function or an inefficient &lt;code&gt;collect()&lt;/code&gt;.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;The Python worker process telemetry&lt;/strong&gt; — the PySpark worker is a separate Python subprocess outside the Java Virtual Machine’s (JVM) managed memory. When it crashes from &lt;code&gt;spark.executor.pyspark.memory&lt;/code&gt; exhaustion, the JVM driver sees a generic “executor exited unexpectedly” message. The actual cause is invisible to standard JVM-level logs.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;When the agent invokes &lt;code&gt;analyze_spark_workload&lt;/code&gt; during an investigation, it returns a structured analysis with the antipattern identified at the line level, the offending stage isolated, and a concrete fix: both code changes and configuration changes.&lt;/p&gt;
&lt;h2 id="integrating-aws-devops-agent-with-apache-spark-troubleshooting-mcp"&gt;Integrating AWS DevOps Agent with Apache Spark Troubleshooting MCP&lt;/h2&gt;
&lt;p&gt;This section explains how AWS DevOps Agent connects to the Apache Spark Troubleshooting Agent through a private MCP endpoint and orchestrates the investigation workflow.&lt;/p&gt;
&lt;h3 id="how-it-works"&gt;How it works&lt;/h3&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-6062-1.jpg" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-6062-1.jpg" alt="Architecture diagram showing AWS DevOps Agent connecting to the Apache Spark Troubleshooting Agent over AWS PrivateLink" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 1: Integration architecture between AWS DevOps Agent and the Apache Spark Troubleshooting Agent for Amazon EMR over AWS PrivateLink&lt;/p&gt;
&lt;/div&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;You submit an investigation prompt in AWS DevOps Agent.&lt;/li&gt;
 &lt;li&gt;AWS DevOps Agent sends a SigV4-signed MCP call into your Amazon VPC through the AWS DevOps Agent private connection.&lt;/li&gt;
 &lt;li&gt;The private connection forwards the request to the Interface VPC Endpoint.&lt;/li&gt;
 &lt;li&gt;The endpoint routes the request over AWS PrivateLink to the Apache Spark Troubleshooting Agent for Amazon EMR, which AWS manages.&lt;/li&gt;
 &lt;li&gt;The MCP service reads from your data sources (Amazon EMR, Amazon S3, Amazon CloudWatch Logs) using the same IAM role AWS DevOps Agent assumed for the call.&lt;/li&gt;
 &lt;li&gt;When a CloudWatch alarm transitions to ALARM state (for example, a failed-jobs alarm for your Amazon EMR Serverless application), AWS DevOps Agent automatically triggers an investigation without manual intervention.&lt;/li&gt;
 &lt;li&gt;AWS DevOps Agent decides which tools to call based on the prompt. For a Spark failure, that includes the Apache Spark Troubleshooting MCP server &lt;a href="https://docs.aws.amazon.com/devopsagent/latest/userguide/configuring-integrations-and-knowledge-connecting-mcp-servers.html" target="_blank" rel="noopener"&gt;you registered as a capability provider&lt;/a&gt;.&lt;/li&gt;
 &lt;li&gt;Each MCP request is signed with AWS Signature Version 4 using the IAM role assigned to the capability provider. The request travels from AWS DevOps Agent into your Amazon VPC through the private connection. This private connection is a managed VPC Lattice resource gateway you created during setup.&lt;/li&gt;
 &lt;li&gt;From the resource gateway, the request flows to the Interface VPC Endpoint for the &lt;a href="https://docs.aws.amazon.com/emr/latest/ReleaseGuide/spark-troubleshooting-agent-setup.html" target="_blank" rel="noopener"&gt;Amazon SageMaker Unified Studio MCP service&lt;/a&gt;, then on to the Apache Spark Troubleshooting Agent. The traffic stays entirely on the AWS network.&lt;/li&gt;
 &lt;li&gt;The MCP server reads the inputs it needs from your AWS account using the IAM role that you assigned to the capability provider during MCP server registration. This role grants access to the Spark History Server event log and application source code in Amazon S3, the driver and executor stdout streams in Amazon CloudWatch Logs, and the job-run metadata from Amazon EMR Serverless.&lt;/li&gt;
 &lt;li&gt;The MCP server returns its diagnostic findings to AWS DevOps Agent. The agent then analyzes the results, identifies the root cause, and presents recommended fixes both code-level and configuration-level in your chat.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="setting-up-the-demo"&gt;Setting up the demo&lt;/h2&gt;
&lt;p&gt;As part of this demo, this post includes a sample AWS CloudFormation template, tested in the us-east-1 Region, that provisions the following resources for the walkthrough:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;A dedicated Amazon Virtual Private Cloud (Amazon VPC) with two private subnets in Availability Zones supported by the Apache Spark Troubleshooting Agent for Amazon EMR.&lt;/li&gt;
 &lt;li&gt;An Interface VPC Endpoint for the Apache Spark Troubleshooting Agent for Amazon EMR.&lt;/li&gt;
 &lt;li&gt;An IAM role that AWS DevOps Agent assumes to invoke the Apache Spark Troubleshooting MCP server with AWS Signature Version 4.&lt;/li&gt;
 &lt;li&gt;A deliberately failing PySpark workload running on Amazon EMR Serverless, including the Amazon EMR Serverless application, the Spark execution role, and the demo logs stored in Amazon S3 bucket.&lt;/li&gt;
 &lt;li&gt;An Amazon CloudWatch alarm that fires when the demo job fails. This alarm is used as the trigger for the agent investigation later in this section.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="step-1-clone-the-repository"&gt;Step 1: Clone the repository&lt;/h3&gt;
&lt;p&gt;Clone the git &lt;a href="https://github.com/aws-samples/sample-aws-data-processing-and-analytics.git" target="_blank" rel="noopener"&gt;repository&lt;/a&gt; for the CloudFormation template, PySpark script, and Parquet data.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;git clone https://github.com/aws-samples/sample-aws-data-processing-and-analytics.git&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h3 id="step-2-deploy-the-aws-cloudformation-stack"&gt;Step 2: Deploy the AWS CloudFormation stack&lt;/h3&gt;
&lt;p&gt;Deploy the template using the following AWS CLI command.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;cd sample-aws-data-processing-and-analytics/blogs/devops-agent-spark-mcp-integration

aws cloudformation create-stack \
  --stack-name spark-troubleshooting-demo \
  --template-body file://cloudformation/spark-troubleshooting-devops-agent-blog.yaml \
  --capabilities CAPABILITY_NAMED_IAM \
  --region us-east-1&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;The stack reaches &lt;code&gt;CREATE_COMPLETE&lt;/code&gt; in approximately 4–6 minutes. When it does, capture the following stack outputs, which you paste into the AWS DevOps Agent console in the next two steps:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;code&gt;DemoVpcId&lt;/code&gt; — the VPC ID for the AWS DevOps Agent private connection.&lt;/li&gt;
 &lt;li&gt;&lt;code&gt;DemoSubnetIds&lt;/code&gt; — the two subnet IDs for the AWS DevOps Agent private connection.&lt;/li&gt;
 &lt;li&gt;&lt;code&gt;SMUSVpcEndpointSecurityGroupId&lt;/code&gt; — the security group ID.&lt;/li&gt;
 &lt;li&gt;&lt;code&gt;TroubleshootingRoleArn&lt;/code&gt; — the IAM role Amazon Resource Name (ARN).&lt;/li&gt;
 &lt;li&gt;&lt;code&gt;MCPEndpointURL&lt;/code&gt; — the MCP endpoint URL to register.&lt;/li&gt;
 &lt;li&gt;&lt;code&gt;FailedJobsAlarmName&lt;/code&gt; — the CloudWatch alarm name to reference in your investigation prompt.&lt;/li&gt;
 &lt;li&gt;&lt;code&gt;DemoBucket&lt;/code&gt; — the S3 bucket name where you copy the demo script and Parquet data.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;To retrieve all outputs at once, use the following AWS CLI command.&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws cloudformation describe-stacks \
  --region us-east-1 \
  --stack-name spark-troubleshooting-demo \
  --query "Stacks[0].Outputs" --output table&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;# Get your bucket name from the stack outputs
DEMO_BUCKET=$(aws cloudformation describe-stacks --stack-name spark-troubleshooting-demo --region us-east-1 --query 'Stacks[0].Outputs[?OutputKey==`DemoBucket`].OutputValue' --output text)

# Copy the script
cd sample-aws-data-processing-and-analytics/blogs/devops-agent-spark-mcp-integration
aws s3 cp scripts/customer_events_aggregator.py s3://$DEMO_BUCKET/customer_events_aggregator.py

# Copy the Parquet data
aws s3 cp data/ s3://$DEMO_BUCKET/data/ --recursive&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h3 id="step-3-create-an-agent-space"&gt;Step 3: Create an agent space&lt;/h3&gt;
&lt;p&gt;The agent space defines which AWS account and Region the agent monitors, which IAM role it assumes, and which capability providers, including MCP servers, it can call.&lt;/p&gt;
&lt;p&gt;Follow the steps in &lt;a href="https://docs.aws.amazon.com/devopsagent/latest/userguide/getting-started-with-aws-devops-agent-creating-an-agent-space.html#creating-an-agent-space" target="_blank" rel="noopener"&gt;Creating an Agent Space&lt;/a&gt; in the AWS DevOps Agent User Guide. When completing those steps, use the following values:&lt;/p&gt;
&lt;table border="1px" width="100%" cellpadding="10px"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;Parameter&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Value&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Name&lt;/td&gt;
   &lt;td&gt;data-pipeline-troubleshooting&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Region&lt;/td&gt;
   &lt;td&gt;us-east-1&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Agent Space role&lt;/td&gt;
   &lt;td&gt;Choose Auto-create a new DevOps Agent role — the console generates a DevOpsAgentRole-AgentSpace* role with AIOpsAssistantPolicy attached&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Optional integrations&lt;/td&gt;
   &lt;td&gt;Not required&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;After the agent space reaches &lt;strong&gt;Active&lt;/strong&gt; status, proceed to create the private connection.&lt;/p&gt;
&lt;h3 id="step-4-create-the-aws-devops-agent-private-connection"&gt;Step 4: Create the AWS DevOps Agent private connection&lt;/h3&gt;
&lt;p&gt;AWS DevOps Agent uses the private connection to reach into your Amazon VPC. Follow the steps in &lt;a href="https://docs.aws.amazon.com/devopsagent/latest/userguide/configuring-integrations-and-knowledge-connecting-to-privately-hosted-tools.html" target="_blank" rel="noopener"&gt;Connecting to privately hosted tools&lt;/a&gt; in the AWS DevOps Agent User Guide. You can use either the console or the AWS CLI command documented under &lt;a href="https://docs.aws.amazon.com/devopsagent/latest/userguide/configuring-integrations-and-knowledge-connecting-to-privately-hosted-tools.html#create-a-private-connection" target="_blank" rel="noopener"&gt;Create a private connection&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;When completing those steps, use the following values from your CloudFormation stack outputs:&lt;/p&gt;
&lt;table border="1px" width="100%" cellpadding="10px"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;Parameter&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Value&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Name&lt;/td&gt;
   &lt;td&gt;A descriptive name (for example, spark-private)&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;VPC&lt;/td&gt;
   &lt;td&gt;DemoVpcId from your stack outputs&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Subnets&lt;/td&gt;
   &lt;td&gt;Both subnet IDs from DemoSubnetIds&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Security group&lt;/td&gt;
   &lt;td&gt;SMUSVpcEndpointSecurityGroupId&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;TCP port ranges (Advanced configuration)&lt;/td&gt;
   &lt;td&gt;443&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Host address (Service target details)&lt;/td&gt;
   &lt;td&gt;sagemaker-unified-studio-mcp.us-east-1.api.aws&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;DNS resolution&lt;/td&gt;
   &lt;td&gt;In VPC (private DNS)&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Certificate public key&lt;/td&gt;
   &lt;td&gt;None&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;After the connection reaches &lt;strong&gt;Active&lt;/strong&gt; status, proceed to Step 5.&lt;/p&gt;
&lt;h3 id="step-5-register-the-apache-spark-troubleshooting-mcp-server-as-a-capability-provider"&gt;Step 5: Register the Apache Spark Troubleshooting MCP server as a capability provider&lt;/h3&gt;
&lt;p&gt;With the private connection in place, register the MCP server as a capability provider. Follow the steps in &lt;a href="https://docs.aws.amazon.com/devopsagent/latest/userguide/configuring-integrations-and-knowledge-connecting-mcp-servers.html#registering-an-mcp-server-account-level" target="_blank" rel="noopener"&gt;Registering an MCP server at the account level&lt;/a&gt; in the AWS DevOps Agent User Guide.&lt;/p&gt;
&lt;p&gt;When completing those steps, use the following values:&lt;/p&gt;
&lt;table border="1px" width="100%" cellpadding="10px"&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;strong&gt;Parameter&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;&lt;strong&gt;Value&lt;/strong&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Name&lt;/td&gt;
   &lt;td&gt;spark-troubleshooting&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Endpoint URL&lt;/td&gt;
   &lt;td&gt;MCPEndpointURL from your stack outputs&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Connect to endpoint using a private connection&lt;/td&gt;
   &lt;td&gt;Selected&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;h3 id="step-6-add-the-mcp-server-to-the-agent-space"&gt;Step 6: Add the MCP server to the agent space&lt;/h3&gt;
&lt;p&gt;With the MCP server registered, you need a workspace where investigations run. The agent space defines which AWS account and Region the agent monitors, which IAM role it assumes, and which capability providers, including MCP servers, it can call.&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;In the &lt;strong&gt;MCP Server&lt;/strong&gt; section, choose &lt;strong&gt;Add&lt;/strong&gt;.&lt;/li&gt;
&lt;/ol&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-6062-2.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-6062-2.png" alt="MCP Server section of the agent space detail page with an Add button" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 2: MCP Server section of the agent space detail page&lt;/p&gt;
&lt;/div&gt;
&lt;ol start="2" type="1"&gt;
 &lt;li&gt;In the &lt;strong&gt;Add a capability&lt;/strong&gt; dialog, locate &lt;code&gt;spark-troubleshooting&lt;/code&gt; in the list of registered MCP servers and choose &lt;strong&gt;Add&lt;/strong&gt;.&lt;/li&gt;
&lt;/ol&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-6062-3.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-6062-3.png" alt="Add a capability dialog listing the spark-troubleshooting MCP server" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 3: Add a capability with the spark-troubleshooting MCP server listed&lt;/p&gt;
&lt;/div&gt;
&lt;ol start="3" type="1"&gt;
 &lt;li&gt;On the &lt;strong&gt;Select MCP server tools&lt;/strong&gt; page, both tools that the Apache Spark Troubleshooting Agent for Amazon EMR publishes are listed: &lt;code&gt;analyze_spark_workload&lt;/code&gt; and &lt;code&gt;analyze_spark_history_server_endpoint&lt;/code&gt;. Select both checkboxes, then choose &lt;strong&gt;Save&lt;/strong&gt;.&lt;/li&gt;
&lt;/ol&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-6062-4.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-6062-4.png" alt="Select MCP server tools page with both Spark troubleshooting tools checked" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 4: The Select MCP server tools page with both Spark troubleshooting tools selected&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;The agent space connects to the MCP server, lists its tools, and displays &lt;strong&gt;2 Available / 2 Connected&lt;/strong&gt;. Both tools are now part of your agent’s catalog.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-6062-5.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-6062-5.png" alt="MCP Server section showing spark-troubleshooting connected with two available and two connected tools" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 5: MCP Server section showing spark-troubleshooting connected with both tools available&lt;/p&gt;
&lt;/div&gt;
&lt;h2 id="seeing-it-in-action"&gt;Seeing it in action&lt;/h2&gt;
&lt;p&gt;To see the integration end to end, you submit a PySpark job, watch the CloudWatch alarm move to &lt;code&gt;ALARM&lt;/code&gt;, and then ask AWS DevOps Agent to investigate using the alarm name.&lt;/p&gt;
&lt;h3 id="the-failing-workload"&gt;The failing workload&lt;/h3&gt;
&lt;p&gt;The CloudFormation template provisioned an Amazon EMR Serverless application called &lt;code&gt;analytics-events-platform&lt;/code&gt; and configured a sample PySpark job, &lt;code&gt;customer_events_aggregator.py&lt;/code&gt;. The script simulates a common Python-side memory bug: a &lt;code&gt;mapPartitions&lt;/code&gt; user-defined function accumulates 11 copies of every input row in an in-memory Python list before yielding results, while the job runs with &lt;code&gt;spark.executor.pyspark.memory=256m&lt;/code&gt;. The Python worker process exceeds the 256 MB cap, the kernel kills it, Spark retries four times, and the stage is marked failed.&lt;/p&gt;
&lt;h3 id="submit-the-failing-job"&gt;Submit the failing job&lt;/h3&gt;
&lt;p&gt;Run the &lt;code&gt;DemoSubmitJobCommand&lt;/code&gt; from your stack outputs in your terminal. It looks like this:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws emr-serverless start-job-run \
  --region us-east-1 \
  --application-id &amp;lt;DemoApplicationId&amp;gt; \
  --execution-role-arn &amp;lt;DemoExecutionRoleArn&amp;gt; \
  --name daily-customer-events-rollup \
  --job-driver '{"sparkSubmit":{"entryPoint":"s3://&amp;lt;DemoBucket&amp;gt;/customer_events_aggregator.py","entryPointArguments":["&amp;lt;DemoBucket&amp;gt;"],"sparkSubmitParameters":"--conf spark.executor.cores=2 --conf spark.executor.memory=1g --conf spark.executor.pyspark.memory=256m --conf spark.executor.instances=2"}}' \
  --configuration-overrides '{"monitoringConfiguration":{"s3MonitoringConfiguration":{"logUri":"s3://&amp;lt;DemoBucket&amp;gt;/logs/"}}}'&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;The command returns a &lt;code&gt;jobRunId&lt;/code&gt;. Note it down. You will see it later in the agent’s investigation.&lt;/p&gt;
&lt;p&gt;The job goes through PENDING to SCHEDULED to RUNNING to FAILED and reaches FAILED state in roughly four minutes.&lt;/p&gt;
&lt;h3 id="watch-the-cloudwatch-alarm-fire"&gt;Watch the CloudWatch alarm fire&lt;/h3&gt;
&lt;p&gt;The CloudFormation template also created a CloudWatch alarm named &lt;code&gt;&amp;lt;DemoApplicationId&amp;gt;-FailedJobs&lt;/code&gt; (the exact name is in the &lt;code&gt;FailedJobsAlarmName&lt;/code&gt; stack output). The alarm watches the &lt;code&gt;FailedJobs&lt;/code&gt; metric in the &lt;code&gt;AWS/EMRServerless&lt;/code&gt; namespace, scoped to your demo application, and flips to &lt;code&gt;ALARM&lt;/code&gt; within a minute or two of the job failing.&lt;/p&gt;
&lt;p&gt;Open the Amazon CloudWatch console, choose &lt;strong&gt;Alarms&lt;/strong&gt; in the left navigation pane, and confirm the alarm is in &lt;code&gt;In alarm&lt;/code&gt; state.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-6062-6.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-6062-6.png" alt="Amazon CloudWatch console alarm detail page showing the FailedJobs alarm in alarm state" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 6: The Amazon CloudWatch alarm detail page showing the FailedJobs alarm in the In alarm state&lt;/p&gt;
&lt;/div&gt;
&lt;h3 id="ask-aws-devops-agent-to-investigate"&gt;Ask AWS DevOps Agent to investigate&lt;/h3&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;Open your AWS DevOps Agent space.&lt;/li&gt;
 &lt;li&gt;In the left navigation pane, choose &lt;strong&gt;Operator Access&lt;/strong&gt;, then choose &lt;strong&gt;Incidents&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Choose &lt;strong&gt;Start an investigation&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Paste the following prompt, replacing &lt;code&gt;&amp;lt;FailedJobsAlarmName&amp;gt;&lt;/code&gt; with the value from your stack outputs:&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;em&gt;CloudWatch alarm in us-east-1 just went into ALARM state. Investigate why and recommend a fix&lt;/em&gt;&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-6062-7.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-6062-7.png" alt="AWS DevOps Agent Start an investigation panel with the alarm prompt entered" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 7: AWS DevOps Agent Start an investigation panel with the Amazon CloudWatch alarm investigation prompt&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;The agent’s investigation chains together native AWS API tools and the Apache Spark Troubleshooting MCP tool you registered:&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;&lt;strong&gt;&lt;code&gt;use_aws cloudwatch describe-alarms&lt;/code&gt;&lt;/strong&gt; — fetches the alarm definition and reads its metric dimensions, identifying that the alarm is scoped to Amazon EMR Serverless application &lt;code&gt;&amp;lt;DemoApplicationId&amp;gt;&lt;/code&gt;.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;&lt;code&gt;use_aws emr-serverless list-job-runs&lt;/code&gt;&lt;/strong&gt; — finds the most recent FAILED job run on that application.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;&lt;code&gt;use_aws emr-serverless get-job-run&lt;/code&gt;&lt;/strong&gt; — pulls the FAILED run’s metadata and last-known error.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;&lt;code&gt;spark-troubleshooting analyze_spark_workload&lt;/code&gt;&lt;/strong&gt; — invokes the Apache Spark Troubleshooting Agent for Amazon EMR through the MCP capability provider, passing the application ID and job run ID. This is where the deep analysis happens.&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id="review-the-root-cause-and-fix"&gt;Review the root cause and fix&lt;/h3&gt;
&lt;p&gt;When the investigation completes, AWS DevOps Agent presents the results across two tabs: &lt;strong&gt;Investigation timeline&lt;/strong&gt; and &lt;strong&gt;Root cause&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;The &lt;strong&gt;Investigation timeline&lt;/strong&gt; shows every step the agent took: skills loaded, native AWS API calls made, and the moment it called the &lt;code&gt;analyze_spark_workload&lt;/code&gt; MCP tool to analyze the failed Spark job. Each entry is expandable so you can audit the inputs and outputs.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-6062-8.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-6062-8.png" alt="Investigation timeline listing the agent tool calls and the MCP invocation" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 8: Investigation timeline tab showing the sequence of agent tool calls and the spark-troubleshooting MCP invocation&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;The &lt;strong&gt;Root cause&lt;/strong&gt; tab is where the answer lands. It is organized into three sections that mirror what an experienced engineer would write in an incident report:&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-6062-9.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-6062-9.png" alt="Root cause tab showing impact, root causes, and key findings for the memory failure" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 9: The Root cause tab showing the impact summary, identified root causes, and key findings for the Spark memory exhaustion failure&lt;/p&gt;
&lt;/div&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;Impact&lt;/strong&gt; — what failed, when, and for how long. For our demo, this calls out that the &lt;code&gt;daily-customer-events-rollup&lt;/code&gt; job on the &lt;code&gt;analytics-events-platform&lt;/code&gt; application failed with a &lt;code&gt;MemoryError&lt;/code&gt; and that the alarm transitioned to ALARM state at the time of the failure.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Root causes&lt;/strong&gt; — the actual antipattern. The agent identifies that &lt;code&gt;customer_events_aggregator.py&lt;/code&gt; combines three compounding issues: an &lt;code&gt;expand_event&lt;/code&gt; function (line 23) that amplifies each input row 11×, a &lt;code&gt;repartition(1)&lt;/code&gt; that funnels all data into a single partition on a single executor, and a &lt;code&gt;collect()&lt;/code&gt; (line 31) that pulls the amplified dataset back to the driver. All three run with only 1 GB of executor memory.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;Key findings&lt;/strong&gt; — supporting facts behind the diagnosis, including the executor memory configuration, the application’s maximum capacity, and how the agent confirmed each fact from the analyzed artifacts.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Both the antipattern identification and the supporting evidence come from artifacts the agent could only reach through the MCP tool: the application source code in Amazon S3, the Spark History Server event log, and the query plan. Without the Apache Spark Troubleshooting Agent for Amazon EMR plugged in, AWS DevOps Agent would have stopped at “the executor exited with a memory error.”&lt;/p&gt;
&lt;h2 id="clean-up"&gt;Clean up&lt;/h2&gt;
&lt;p&gt;To avoid ongoing charges, delete the resources you created. Some resources are managed by the AWS DevOps Agent console and must be removed there first. Otherwise, the CloudFormation stack deletion fails.&lt;/p&gt;
&lt;ol type="1"&gt;
 &lt;li&gt;In the &lt;strong&gt;AWS DevOps Agent console&lt;/strong&gt;, open your &lt;code&gt;data-pipeline-troubleshooting&lt;/code&gt; agent space, choose the &lt;strong&gt;MCP Server&lt;/strong&gt; section, select &lt;code&gt;spark-troubleshooting&lt;/code&gt;, and choose &lt;strong&gt;Remove&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;From the &lt;strong&gt;Agent spaces&lt;/strong&gt; list, select &lt;code&gt;data-pipeline-troubleshooting&lt;/code&gt; and choose &lt;strong&gt;Delete&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;In &lt;strong&gt;Capability Providers&lt;/strong&gt;, select &lt;code&gt;spark-troubleshooting&lt;/code&gt; and choose &lt;strong&gt;Deregister&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;In &lt;strong&gt;Capability Providers&lt;/strong&gt; → &lt;strong&gt;Private connections&lt;/strong&gt;, select &lt;code&gt;smus-spark-private&lt;/code&gt; and choose &lt;strong&gt;Delete&lt;/strong&gt;.&lt;/li&gt;
 &lt;li&gt;Delete the AWS CloudFormation stack. This removes the Amazon VPC, the Interface VPC Endpoint, the security group, the IAM role, the Amazon EMR Serverless application, the Spark execution role, the Amazon CloudWatch alarm, and the demo logs bucket.&lt;/li&gt;
&lt;/ol&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws cloudformation delete-stack \
  --region us-east-1 \
  --stack-name spark-troubleshooting-demo&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;In this post, you connected the Apache Spark Troubleshooting Agent for Amazon EMR to AWS DevOps Agent as a custom MCP capability provider. You kept the traffic on the AWS network with AWS PrivateLink, and ran a failing PySpark job to see the integration end to end. A CloudWatch alarm fired, you asked the agent to investigate, and a single chat session returned the root cause along with code and configuration fixes.&lt;/p&gt;
&lt;p&gt;You can extend this pattern beyond the demo scenario. Consider connecting the MCP server to agent spaces that monitor your production Amazon EMR environment. Any Spark job that writes a History Server event log becomes diagnosable through the same workflow.&lt;/p&gt;
&lt;p&gt;To continue learning, explore the following resources:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/devopsagent/latest/userguide/" target="_blank" rel="noopener"&gt;AWS DevOps Agent documentation&lt;/a&gt; — learn how to create agent spaces, configure integrations, and manage investigations.&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/emr/latest/ReleaseGuide/spark-troubleshooting-agent-setup.html" target="_blank" rel="noopener"&gt;Apache Spark Troubleshooting Agent for Amazon EMR setup guide&lt;/a&gt; — detailed prerequisites and configuration options for the MCP server.&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/devopsagent/latest/userguide/configuring-integrations-and-knowledge-connecting-mcp-servers.html" target="_blank" rel="noopener"&gt;Connecting MCP servers to AWS DevOps Agent&lt;/a&gt; — register additional MCP servers to expand your agent’s capabilities.&lt;/li&gt;
 &lt;li&gt;&lt;a href="https://github.com/aws-samples/sample-aws-data-processing-and-analytics/tree/main/blogs/devops-agent-spark-mcp-integration" target="_blank" rel="noopener"&gt;Sample code on GitHub&lt;/a&gt; — clone the CloudFormation template, PySpark script, and Parquet data used in this walkthrough.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If you’ve already integrated the Apache Spark Troubleshooting Agent into your operational workflow, or if you’re exploring other MCP-based extensions for AWS DevOps Agent, we want to hear about your experience. Share your thoughts and questions in the comments.&lt;/p&gt;
&lt;hr style="width: 100%"&gt;
&lt;h2&gt;About the authors&lt;/h2&gt;
&lt;footer&gt;
 &lt;div class="blog-author-box"&gt;
  &lt;div class="blog-author-image"&gt;
   &lt;p&gt;&lt;img loading="lazy" class="alignleft size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-6062-10.png" alt="Kalyan Janaki" width="100" height="100"&gt;&lt;/p&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Kalyan Janaki&lt;/h3&gt;
  &lt;p&gt;Kalyan is Senior Big Data &amp;amp; Analytics Specialist with Amazon Web Services. He helps customers architect and build highly scalable, performant, and secure cloud-based solutions on AWS.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box"&gt;
  &lt;div class="blog-author-image"&gt;
   &lt;p&gt;&lt;img loading="lazy" class="alignleft size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/31/BDB-6062-11.jpg" alt="Aneesh Varghese" width="100" height="100"&gt;&lt;/p&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Aneesh Varghese&lt;/h3&gt;
  &lt;p&gt;Aneesh is a Senior Technical Account Manager at AWS with more than 20 years of Information Technology industry experience. Aneesh supports enterprise customers in cost optimization strategies, Cloud operations, MLOps, providing advocacy and strategic technical guidance to help plan and build solutions using AWS best practices. Outside of work, Aneesh likes to spend time with family, play Basketball and Badminton&lt;/p&gt;
 &lt;/div&gt;
&lt;/footer&gt;</content:encoded>
					
		
		
			</item>
		<item>
		<title>Deliver real-time data to streaming tables for Apache Iceberg with Amazon Kinesis Data Streams</title>
		<link>https://aws.amazon.com/blogs/big-data/deliver-real-time-data-to-streaming-tables-for-apache-iceberg-with-amazon-kinesis-data-streams/</link>
		
		<dc:creator><![CDATA[Nikit Pednekar]]></dc:creator>
		<pubDate>Mon, 31 Aug 2026 23:03:25 +0000</pubDate>
				<category><![CDATA[Amazon Kinesis]]></category>
		<category><![CDATA[Announcements]]></category>
		<category><![CDATA[Intermediate (200)]]></category>
		<guid isPermaLink="false">dcfc2043b085b9273944141d7e341354d1c31e28</guid>

					<description>Amazon Kinesis Data Streams now supports streaming tables, a fully managed capability that continuously delivers your streaming data as queryable Apache Iceberg tables on Amazon S3 Tables. Streaming tables reduce data delivery costs to S3 Tables by up to 50% compared to self-managed alternatives and reduce downstream query costs by up to 30% through intelligent inline compaction that eliminates the small file problem. You need no custom applications, no self-managed compute, and no operational overhead.</description>
										<content:encoded>&lt;p&gt;&lt;a href="https://aws.amazon.com/kinesis/data-streams/" target="_blank" rel="noopener"&gt;Amazon Kinesis Data Streams&lt;/a&gt; now supports streaming tables, a fully managed capability that continuously delivers your streaming data as queryable &lt;a href="https://iceberg.apache.org/" target="_blank" rel="noopener"&gt;Apache Iceberg&lt;/a&gt; tables on &lt;a href="https://aws.amazon.com/s3/features/tables/" target="_blank" rel="noopener"&gt;Amazon S3 Tables&lt;/a&gt;. Amazon S3 Tables is a capability of &lt;a href="https://aws.amazon.com/s3/" target="_blank" rel="noopener"&gt;Amazon Simple Storage Service (Amazon S3)&lt;/a&gt;. Streaming tables reduce data delivery costs to S3 Tables by up to 50% compared to self-managed alternatives and reduce downstream query costs by up to 30% through intelligent inline compaction that eliminates the small file problem. You need no custom applications, no self-managed compute, and no operational overhead.&lt;/p&gt;
&lt;p&gt;Customers increasingly want to unify streaming data with Apache Iceberg for near-real-time analytics, fraud detection, personalization, and artificial intelligence and machine learning (AI/ML) feature pipelines. But integrating the two has meant operating complex custom connectors, managing format conversions, and contending with the performance impact of many small Parquet files that slow queries and increase costs. Streaming tables solve this: configure delivery in a few steps from the console or through APIs, and your data becomes queryable from &lt;a href="https://aws.amazon.com/athena/" target="_blank" rel="noopener"&gt;Amazon Athena&lt;/a&gt;, &lt;a href="https://aws.amazon.com/redshift/" target="_blank" rel="noopener"&gt;Amazon Redshift&lt;/a&gt;, and &lt;a href="https://spark.apache.org/" target="_blank" rel="noopener"&gt;Apache Spark&lt;/a&gt; within minutes. Tables are automatically registered in &lt;a href="https://aws.amazon.com/glue/" target="_blank" rel="noopener"&gt;AWS Glue Data Catalog&lt;/a&gt;, making them immediately discoverable for analytics engines and AI agents.&lt;/p&gt;
&lt;p&gt;For workloads that don’t require Iceberg table format, you can also deliver streaming data to Amazon S3 general purpose buckets. Delivery is in the source data format, ideal for archival, backup, and ML training data pipelines, with the same serverless, fully managed delivery and no infrastructure to operate.&lt;/p&gt;
&lt;h2 id="challenges-with-delivering-streaming-data-to-apache-iceberg"&gt;Challenges with delivering streaming data to Apache Iceberg&lt;/h2&gt;
&lt;p&gt;Customers today face three challenges when integrating streaming data with Apache Iceberg.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Operational complexity:&lt;/strong&gt; Connecting Kinesis Data Streams to Iceberg tables today requires deploying and maintaining custom connectors, Apache Flink jobs, or consumer applications. Teams must manage pipeline failures, handle format conversions, scale infrastructure, and monitor delivery reliability. These operational tasks consume significant engineering time and introduce ongoing risk of downtime.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Resiliency and the small file problem:&lt;/strong&gt; Without proper coordination, simultaneous writes from multiple high-throughput shards can conflict, leading to failed commits, data freshness delays, and degraded performance. Streaming ingestion of high-volume data creates large numbers of small Parquet files in Iceberg tables, forcing a difficult trade-off between data freshness and query efficiency.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cost:&lt;/strong&gt; Customers typically spend up to $28/TB operating streaming extract, transform, and load (ETL) pipelines from Kinesis Data Streams using self-managed alternatives based on internal analysis. This creates a high price barrier to getting streaming data into queryable formats and makes cost unpredictable as volume grows.&lt;/p&gt;
&lt;h2 id="how-delivery-to-streaming-tables-solves-these-challenges"&gt;How delivery to streaming tables solves these challenges&lt;/h2&gt;
&lt;p&gt;Streaming tables are a native capability built directly into Amazon Kinesis Data Streams. There is no separate service to deploy, no connector to version, and no consumer application to maintain. You enable delivery in a few steps from the console or through APIs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Zero operational overhead:&lt;/strong&gt; Streaming tables remove the need to build and operate custom consumer applications for data delivery. No pipeline infrastructure to provision, no scaling logic to write, no failure handling to implement. The capability automatically scales to process gigabytes per second of throughput.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Built-in resiliency:&lt;/strong&gt; Streaming tables provide write coordination and exactly once delivery semantics across all shards in your stream, resolving concurrent writer conflicts and ensuring data integrity without manual intervention.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intelligent compaction, no trade-offs:&lt;/strong&gt; During ingestion, streaming tables perform inline compaction that produces query-optimized Parquet files, eliminating the small file problem while maintaining minute-level data freshness. This reduces downstream query costs by up to 30 percent compared to uncompacted delivery.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Consumption-based pricing:&lt;/strong&gt; You pay only for data delivered: $14/TB for Iceberg delivery to S3 Tables in US East (N. Virginia) Region (us-east-1) (50% savings compared to self-managed alternatives) and $11/TB for general purpose S3 delivery (60% savings compared to self-managed alternatives). When your stream is idle, you pay nothing for delivery. Combined with Kinesis Data Streams On-Demand Advantage pricing, which eliminates per-shard charges and scales automatically, the entire path from ingestion to queryable Iceberg tables operates on a pure consumption model.&lt;/p&gt;
&lt;h2 id="end-to-end-managed-streaming-analytics-architecture"&gt;End-to-end managed streaming analytics architecture&lt;/h2&gt;
&lt;p&gt;With delivery to streaming tables, you now have a fully managed end-to-end real-time data architecture from data ingestion through storage to analytics. Your producers publish events to a Kinesis Data Stream, which continuously delivers data as optimized Iceberg read-only tables in S3 Tables. From there, you can query your streaming data using analytics engines like Amazon Athena, Amazon Redshift, &lt;a href="https://aws.amazon.com/emr/" target="_blank" rel="noopener"&gt;Amazon EMR&lt;/a&gt; (Apache Spark), or &lt;a href="https://aws.amazon.com/managed-service-apache-flink/" target="_blank" rel="noopener"&gt;Apache Flink&lt;/a&gt;. You can also let AI agents discover and reason over your data through Glue Data Catalog semantic search. This managed experience removes the intermediate infrastructure that customers previously assembled: separate connector clusters, compaction jobs, and custom consumers. It replaces them with a single, serverless pipeline from stream to insight.&lt;/p&gt;
&lt;p&gt;The following diagram illustrates this end-to-end architecture.&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/25/BDB-6152-1.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/25/BDB-6152-1.png" alt="End-to-end architecture from Kinesis Data Streams producers to Iceberg tables in S3 Tables, queryable by Athena, Redshift, EMR, and Flink" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 1: End-to-end managed streaming analytics architecture from ingestion to query&lt;/p&gt;
&lt;/div&gt;
&lt;h2 id="getting-started"&gt;Getting started&lt;/h2&gt;
&lt;p&gt;To get started, sign in to the Amazon Kinesis Data Streams console, navigate to your streams, and enable delivery to streaming tables in a few steps. Specify the stream you want to deliver, configure your schema settings using AWS Glue Schema Registry, and choose your destination S3 Tables location. After you enable it, delivery to streaming tables immediately begins materializing your streaming data as queryable Iceberg tables in S3 with no further intervention required. There’s no infrastructure to provision and no minimum commitment. You pay only for data delivered.&lt;/p&gt;
&lt;p&gt;Additionally, you can use Amazon Kinesis Data Streams APIs to programmatically set up, update, or delete delivery to streaming tables configurations for your data streams. With these APIs, teams can build agentic workflows and infrastructure-as-code patterns to manage configurations across multiple data streams at scale.&lt;/p&gt;
&lt;h2 id="getting-started-with-the-kinesis-data-streams-agent-skill"&gt;Getting started with the Kinesis Data Streams Agent Skill&lt;/h2&gt;
&lt;p&gt;The Kinesis Data Streams Agent Skill provides AI-assisted guidance for setting up streaming tables integrations for your existing or new data streams. The skill helps you configure delivery to S3 Tables (Iceberg) or S3, including schema registry setup, AWS Identity and Access Management (IAM) role configuration, and validation.&lt;/p&gt;
&lt;h3 id="installing-as-an-agent-skill"&gt;Installing as an Agent Skill&lt;/h3&gt;
&lt;p&gt;&lt;a href="https://docs.aws.amazon.com/agent-toolkit/latest/userguide/skills.html" target="_blank" rel="noopener"&gt;Agent Skills&lt;/a&gt; are discovered automatically by compatible tools through the SKILL.md file. Refer to the &lt;a href="https://docs.aws.amazon.com/agent-toolkit/latest/userguide/aws-cli.html" target="_blank" rel="noopener"&gt;Agent Toolkit for AWS Skill Installation Guide&lt;/a&gt; to install the managing-amazon-kinesis-data-streams Agent Skill. We also recommend you install the AWS MCP Server in your developer tool of choice, which exposes tools for searching AWS documentation, blogs, and Skills dynamically at runtime. These capabilities make agents more accurate and powerful for AWS related development and operational tasks, and make skill discovery and installation more flexible. Refer to &lt;a href="https://docs.aws.amazon.com/agent-toolkit/latest/userguide/getting-started-aws-mcp-server.html" target="_blank" rel="noopener"&gt;Setting up the AWS MCP Server&lt;/a&gt; for guidance on installing the AWS MCP Server in your environment.&lt;/p&gt;
&lt;p&gt;For example:&lt;/p&gt;
&lt;div class="hide-language"&gt;
 &lt;pre&gt;&lt;code class="language-bash"&gt;aws configure agent-toolkit
aws agent-toolkit add-skill --skill-name managing-amazon-kinesis-data-streams&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
&lt;p&gt;To verify the installation, interact with the skill in your preferred tool.&lt;/p&gt;
&lt;p&gt;To start delivering data from your data streams to Apache Iceberg tables in real time, prompt “Create me a streaming table on my events data stream” to your agent of choice:&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/25/BDB-6152-2.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/25/BDB-6152-2.png" alt="Agent chat showing a prompt to create a streaming table on the events data stream" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 2: Prompting the agent to create a streaming table&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;The agent dynamically loads the managing-amazon-kinesis-data-streams skill and starts by gathering the available resources in your AWS account for the streaming tables integration. After it gathers that data, it confirms the resources to use or create, and creates the integration:&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/25/BDB-6152-3.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/25/BDB-6152-3.png" alt="Agent confirming the AWS resources to use or create for the streaming tables integration" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 3: The agent confirming resources before creating the integration&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;After creating the integration, the agent summarizes the status and can then help with any other operational tasks with your data. For example, the agent can help you set up &lt;a href="https://aws.amazon.com/lake-formation/" target="_blank" rel="noopener"&gt;AWS Lake Formation&lt;/a&gt; permissions to query the data in S3 Tables with Athena, or configure your table maintenance behavior in S3 Tables:&lt;/p&gt;
&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
 &lt;a href="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/25/BDB-6152-4.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/25/BDB-6152-4.png" alt="Agent summarizing integration status and offering Lake Formation permissions or table maintenance setup" width="800"&gt;&lt;/a&gt;
 &lt;p class="wp-caption-text"&gt;Figure 4: The agent offering follow-up operational tasks&lt;/p&gt;
&lt;/div&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;Streaming tables are available in all AWS Regions where Amazon Kinesis Data Streams is offered. Pricing is $14/TB for delivery to S3 Tables (Apache Iceberg) and $11/TB for delivery to general purpose S3 buckets. To learn more, visit the &lt;a href="https://docs.aws.amazon.com/streams/latest/dev/introduction.html" target="_blank" rel="noopener"&gt;documentation&lt;/a&gt; and &lt;a href="https://aws.amazon.com/kinesis/data-streams/pricing/" target="_blank" rel="noopener"&gt;pricing&lt;/a&gt; pages.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;About the authors&lt;/h2&gt;
&lt;footer&gt;
 &lt;div class="blog-author-box"&gt;
  &lt;div class="blog-author-image"&gt;
   &lt;p&gt;&lt;img loading="lazy" class="alignleft size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/25/BDB-6152-5.jpg" alt="Nikit Pednekar" width="100" height="100"&gt;&lt;/p&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Nikit Pednekar&lt;/h3&gt;
  &lt;p&gt;Nikit is Principal Product Manager for Amazon Kinesis Data Streams. He leads product vision, strategy, and the P&amp;amp;L for AWS’s real-time data streaming portfolio- Amazon Kinesis Data Streams and related services. Working backwards from customer needs, he drives the streaming roadmap to help AWS customers build scalable, low-latency, real time data architectures.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box"&gt;
  &lt;div class="blog-author-image"&gt;
   &lt;p&gt;&lt;img loading="lazy" class="alignleft size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/25/BDB-6152-6.png" alt="Mazrim Mehrtens" width="100" height="100"&gt;&lt;/p&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Mazrim Mehrtens&lt;/h3&gt;
  &lt;p&gt;&lt;a href="https://www.linkedin.com/in/mmehrtens/" target="_blank" rel="noopener"&gt;Mazrim&lt;/a&gt; is a Sr.&amp;nbsp;Specialist Solutions Architect for messaging and streaming workloads. Mazrim works with customers to build and support systems that process and analyze terabytes of streaming data in real time, run enterprise machine learning (ML) pipelines, and create systems to share data across teams seamlessly with varying data toolsets and software stacks.&lt;/p&gt;
 &lt;/div&gt;
 &lt;div class="blog-author-box"&gt;
  &lt;div class="blog-author-image"&gt;
   &lt;p&gt;&lt;img loading="lazy" class="alignleft size-full" src="https://d2908q01vomqb2.cloudfront.net/b6692ea5df920cad691c20319a6fffd7a4a766b8/2026/08/25/BDB-6152-7.jpg" alt="Ren Liu" width="100" height="100"&gt;&lt;/p&gt;
  &lt;/div&gt;
  &lt;h3 class="lb-h4"&gt;Ren Liu&lt;/h3&gt;
  &lt;p&gt;&lt;a href="https://www.linkedin.com/in/renliu11/" target="_blank" rel="noopener"&gt;Ren&lt;/a&gt; is a Solutions Architect at AWS in Seattle, working across the full stack from landing zone design and cloud governance to real-time streaming and ML inference. He works with ISV customers in cybersecurity, FinOps, and healthcare to architect secure, scalable solutions powered by generative AI.&lt;/p&gt;
 &lt;/div&gt;
&lt;/footer&gt;</content:encoded>
					
		
		
			</item>
	</channel>
</rss>