A Glue job that had run every morning for two years failed on a Monday. Access denied on GetTable, thrown by Lake Formation. Nothing about the job had changed. Nothing about the IAM role had changed. What had changed was that someone spent Friday afternoon “moving us onto the lakehouse” and, as part of that, revoked a permission from a virtual group most people have never heard of.
That is the shape of the risk. When people ask whether a SageMaker Lakehouse migration is worth it, they picture data moving. Almost none of it does. What moves is the permission model, and that is the part nobody tests.
This post covers what the migration actually changes underneath, the failure mode that bites hardest, where the lakehouse earns its place, where staying on Glue plus Athena is the better call, how the cost levers shift, and a decision procedure you can run in an afternoon.
First, understand what you are actually migrating
The framing “SageMaker Lakehouse versus Glue and Athena” is misleading, and it drives bad decisions.
SageMaker Lakehouse is built on the AWS Glue Data Catalog and AWS Lake Formation. It is not a replacement metastore. It is a reorganisation of the catalog you already have, plus connectors, plus an Apache Iceberg REST interface, plus a console to look at all of it through.
The unit it introduces is the catalog: a logical container of databases, tables and views, nestable to mirror the shape of the source system. A managed catalog holds data you own in S3 or Redshift Managed Storage. A federated catalog points at something external such as Redshift, DynamoDB, PostgreSQL, Oracle, BigQuery or Snowflake, and lets you query it in place.
Your existing databases and tables do not disappear. They sit under the account catalog and keep working. AWS states in its own FAQ that you do not have to migrate, and that Athena, EMR, Glue and MWAA continue to work as they do today. That is an unusual thing for a vendor to publish about a flagship product, and it is the most useful sentence in the documentation set.
So the question is not “should we move our data”. It is “should we move our governance”.
The failure mode that actually bites
Lake Formation grants Super permission to a virtual group called IAMAllowedPrincipals on every existing Data Catalog database and table by default. That group effectively means “anyone whose IAM policy already lets them in”. It exists purely for backward compatibility. While it holds Super on a resource, access is decided by IAM and S3 policies exactly as it was before Lake Formation existed. Your Lake Formation grants are recorded, but they are not what gets enforced.
So the first time someone writes a careful column-level policy and tests it, the policy appears not to work. The analyst can still see the column. The natural next move is to revoke IAMAllowedPrincipals so the grants take effect.
That revoke is the moment everything relying on IAM-only access stops working. Glue jobs. Notebooks. Scheduled dbt runs. The QuickSight dataset nobody has looked at since the person who built it left. None of them had a Lake Formation grant, because none of them ever needed one.
It is invisible in advance because nothing warns you. Lake Formation cannot know which principals are silently depending on the compatibility group. There is no dry run.
Hybrid access mode is the way through
Hybrid access mode exists exactly for this. It supports two permission pathways to the same catalog objects at once. Principals you explicitly opt in are governed by Lake Formation grants. Everyone else keeps reaching the data through their IAM and S3 policies, untouched.
The detail worth internalising: when an S3 location is registered in hybrid access mode, Lake Formation enforces only CREATE_TABLE, CREATE_PARTITION and UPDATE_TABLE by default. The write path gets governed first. The read path stays on IAM until you opt a principal in. That asymmetry is what lets you migrate one team at a time instead of one account at a time.
The safe sequence: register the location in hybrid mode, opt in one low-stakes principal, verify its queries respect the fine-grained grants, verify a non-opted-in principal still works, repeat. Only after every consumer is opted in and verified do you consider revoking the compatibility group. Plenty of teams never need that last step.
Where the lakehouse genuinely wins
- One permission surface across engines. If governance today is “Lake Formation for the lake, Redshift GRANTs for the warehouse, IAM for DynamoDB”, you maintain three access models that disagree with each other. Catalogs let you express column and row-level policy once and have Athena, Redshift, EMR and Spark honour it.
- Tag-based access control across federated sources. LF-TBAC extends to federated catalogs covering S3 Tables, Redshift and external sources. If you hand-maintain named-resource grants across dozens of databases, tag expressions are a real operational win.
- Querying across boundaries without copying. A federated catalog lets one statement touch a Redshift table and an S3 Iceberg table, with no ETL job whose only purpose is moving rows so they can sit next to each other. Every pipeline you delete is a pipeline that cannot break at 3am.
- The Iceberg REST interface. Exposing the catalog through Iceberg REST APIs lets engines outside AWS read your tables over a standard protocol. If you run Trino, Spark or DuckDB elsewhere, this is what makes the catalog genuinely portable rather than nominally open.
- Cross-account sharing without replication. For data mesh setups, sharing a catalog instead of shipping a copy per consumer removes a whole category of staleness bug. Cross-account grants on federated catalog objects require the cross-account data sharing version set to 4 or higher.
Notice what these share. Every one is a governance or interoperability problem. None is a query performance problem, and none is a cost problem.
Where Glue plus Athena is still the right answer
- Single account, S3 only, one team. Catalogs solve a problem you do not have. You would be adding a governance layer to govern one group.
- Your access control is coarse and that is correct. Not every dataset needs column masking. If bucket-level IAM boundaries match your actual risk model, Lake Formation enforcement adds a failure mode without adding safety.
- You have heavy IAM-only automation. Every Glue job, Lambda, Airflow DAG and third-party tool touching the catalog is a principal you must eventually opt in and re-verify. That inventory is the real cost of migration, and it is always larger than the first estimate.
- Your bottleneck is query cost, not governance. Partitioning, file sizes, compaction and columnar layout drive Athena spend. A lakehouse migration touches none of them. Fix the small-files problem first; it pays back faster and it is reversible.
- You depend on a specific engine behaviour today. Fine-grained access control support varies by engine and table format. Confirm your exact combination is supported rather than assuming parity across Athena, Redshift, EMR and Glue Spark.
The cost levers that move, and the ones that don’t
Pricing pages change, so here are the billing mechanisms rather than numbers.
Unchanged: metadata storage and API requests follow existing Glue Data Catalog pricing, including free tier allowances. Reorganising into catalogs does not put you on a new metadata meter. Athena still bills on data scanned. S3 storage and request charges are what they were.
New line items:
- Lambda invocations behind federated catalogs. When a query hits a federated table, Lake Formation vends credentials that invoke a Lambda function defined in the Glue connection to retrieve metadata from the source. That is compute you did not previously pay for, and it is latency on the metadata path. Chatty BI tools that re-resolve schemas on every refresh are the ones to watch.
- Service-managed Redshift Serverless compute. Queries against a federated catalog pointing at an existing Redshift warehouse execute through a service-managed Redshift Serverless workgroup, which stages results before returning them. That compute is billed, and it does not appear where you expect it.
- Redshift Managed Storage. Managed catalogs backed by RMS rather than S3 put you on Redshift storage and Redshift Serverless compute economics, which behave nothing like S3 plus Athena.
The service-managed workgroup is the one that catches people, because the spend lands under Redshift Serverless with no obvious link to the analyst who ran the query. AWS publishes a generated cost allocation tag for exactly this, which you must activate in the billing console before it produces anything:
# Cost allocation tag on lakehouse-managed Redshift Serverless compute.
# Activate under Billing -> Cost allocation tags before it appears in reports.
Key: aws:redshift-serverless:LakehouseManagedWorkgroup
Value: "True"
Activate it on day one, not after the first surprising invoice. Tags only apply to usage recorded after activation, so a late start leaves a blind period you cannot reconstruct. If you already run a cost visibility layer such as Vantage, CloudZero or Datadog’s cloud cost views, build a view isolating this tag before you create your first federated catalog. If you do not, the AWS Cost and Usage Report grouped by that tag key is enough to start.
Is a SageMaker Lakehouse migration worth it? A decision procedure
Run these in order. Stop at the first one that gives a clear answer.
- Count your non-S3 governed sources. Redshift clusters, DynamoDB tables, operational databases and external warehouses that analysts need and that each carry their own access model today. Zero or one means the federated catalog story does not apply to you yet. Three or more and the unified permission surface starts paying for itself.
- Inventory every principal touching the Data Catalog. Pull CloudTrail for Glue and Lake Formation API calls across a full billing cycle, long enough to catch monthly jobs. Every distinct role in that list is work. This number, not your data volume, is the migration estimate.
- Ask whether column or row-level control has actually been requested. If the answer is a compliance requirement with a name attached, proceed. If it is “it would be nice”, you are buying a permission model to solve a hypothetical.
- Check engine coverage against your real workloads. Confirm fine-grained access control support for each production engine in current documentation. One unsupported critical path makes hybrid mode permanent rather than transitional, which is a legitimate outcome but should be a decision, not a discovery.
- Price the federated path against a copy. For your highest-frequency cross-source query, estimate metadata Lambda invocations plus any service-managed compute, then compare against a nightly job that lands a copy in S3. Federation wins on freshness and on one less pipeline. It does not automatically win on cost.
- Decide the destination, not just the direction. “Hybrid access mode indefinitely, Lake Formation enforcement on the two regulated datasets” is a good end state. Write it down first, because migrations without a finish line become permanent half-migrations.
Three things to run before you commit
Enumerate what Athena can see today. This is your before picture, and it exposes a tooling gap worth knowing up front:
aws athena list-data-catalogs
# Returns catalogs registered directly with Athena, for example:
# {
# "DataCatalogsSummary": [
# { "CatalogName": "AwsDataCatalog", "Type": "GLUE" },
# { "CatalogName": "cw_logs_catalog", "Type": "LAMBDA" }
# ]
# }
S3 Tables catalogs use a s3tablescatalog/bucket-name naming convention and are not returned by this call, even though you can query them by naming the catalog explicitly. The console shows them through mechanisms the public API does not expose. Any inventory or drift-detection tooling built on list-data-catalogs will quietly under-report, so enumerate S3 Tables catalogs separately.
Next, confirm the role you intend to use as data lake administrator can create catalogs at all. If it is not already a data lake admin, it needs an explicit grant:
aws lakeformation grant-permissions
--cli-input-json
'{
"Principal": {
"DataLakePrincipalIdentifier": "arn:aws:iam::123456789012:role/Admin"
},
"Resource": { "Catalog": {} },
"Permissions": [ "CREATE_CATALOG", "DESCRIBE" ]
}'
This avoids the common false start where catalog creation works in the console for an admin user and fails for the automation role you actually meant to use.
Finally, if you read lakehouse tables from a Glue Spark job, fine-grained access control is not on by default. The job parameters that enable it:
--datalake-formats iceberg
--enable-lakeformation-fine-grained-access true
Without the second one, the job reads through its IAM role and your column-level grants are not applied. That is a governance gap that looks exactly like a working pipeline, which is the worst kind.
Arguments that don’t survive contact
- “It’s the same Glue Data Catalog, so migrating is low risk.” True about the metadata, false about enforcement. The catalog is the same object. The thing deciding whether a query returns rows is not.
- “We’ll migrate to avoid falling behind.” AWS states plainly that you can keep using Athena, EMR, Glue and MWAA as you do today. Nothing here is a deprecation. Migrating on an invented schedule is how you end up owning a permission model nobody needed.
- “Federated catalogs mean we can retire our ETL.” Some of it, sometimes. Federation removes copies that existed only for co-location. It does not remove transformation, deduplication, late-arriving data handling or historical snapshots. Those jobs stay.
- “Iceberg REST support means no lock-in.” The table format is portable and that is real. The Lake Formation grants, LF-Tags, hybrid mode opt-ins and catalog hierarchy are not, and they become the hardest part of your platform to reproduce elsewhere. Portability moved down a layer; it did not appear.
- “We can test this in a sandbox.” You can test the mechanics. You cannot test the thing that actually breaks, which is the set of production principals depending on IAM-only access. That inventory does not exist in a sandbox.
How I’d decide
S3-only, one account, small team: I would not migrate. The same effort spent on partitioning and compaction returns more, and it keeps returning.
Multiple governed sources, several consuming teams, a named compliance requirement for fine-grained control: migrate, and treat hybrid access mode as the destination rather than a waypoint. Enforce Lake Formation on the datasets that need it, leave the rest on IAM. The pressure to reach a “clean” fully-enforced state is aesthetic, not operational.
Either way, start with a federated catalog over one source that is currently painful to reach. Prove the query path and the cost, and let the rest follow evidence rather than a plan drawn before anyone had numbers.
Frequently asked questions
Does SageMaker Lakehouse replace AWS Glue and Athena?
No. It is built on the Glue Data Catalog and Lake Formation, and Athena remains one of the engines you query it with. AWS documentation states you can continue using Athena, EMR, Glue and MWAA without migrating. It is a governance and connectivity layer, not a replacement stack.
Will a SageMaker Lakehouse migration break my existing Athena queries?
Reorganising into catalogs does not break queries on its own. What breaks them is revoking Super from IAMAllowedPrincipals before every consuming principal has a Lake Formation grant. Use hybrid access mode, opt principals in one at a time, and existing queries keep working throughout.
Do I have to move data into Redshift Managed Storage?
No. Managed catalogs can be backed by S3, including S3 Tables, and existing S3 data lakes integrate directly. RMS is a storage choice with different cost and performance behaviour, not a requirement of the architecture.
How much extra does it cost?
The catalog layer itself is not a new meter: metadata storage and API requests follow existing Glue Data Catalog pricing including free tier allowances. New costs come from what you build on it, mainly Lambda invocations for federated metadata retrieval, service-managed Redshift Serverless compute for federated Redshift queries, and RMS storage if you choose it.
Is SageMaker Unified Studio required?
No. Unified Studio is the integrated console experience. The catalogs are addressable from Athena, Redshift, EMR, Glue and Iceberg-compatible engines directly. Adopting the catalog structure without adopting the studio is a valid and common choice.
What is the smallest useful first step?
Create one federated catalog over a single non-S3 source that analysts currently reach through an awkward workaround. Grant read access to one team. Measure query latency and the resulting Lambda and compute charges. That gives you real numbers on your own workload, and it is trivially reversible.
Conclusion
A SageMaker Lakehouse migration is not a data migration. The Glue Data Catalog underneath is the same catalog, the tables are the same tables, Athena is still Athena. What changes is which system decides whether a query returns rows.
That makes the call simpler than the marketing implies. Several governed sources, multiple consuming teams and a real fine-grained control requirement, and the unified permission surface is worth the work. Without them, you are taking on a permission model and a Lambda-backed metadata path in exchange for a nicer console.
Whichever way you go: register in hybrid access mode, opt principals in one at a time, and do not revoke IAMAllowedPrincipals until you can name every principal that will notice.
Need help deciding, or unpicking a migration already underway?
An outside read on the actual inventory usually saves more time than any amount of architecture discussion. Things I help with:
- Building the principal inventory from CloudTrail so you know exactly who depends on IAM-only Data Catalog access before anything is revoked
- Designing a hybrid access mode rollout that opts teams in incrementally, with a verification step and a rollback position at every stage
- Modelling federated catalog costs against the copy-the-data alternative for your query patterns, including the service-managed compute nobody budgets for
- Diagnosing Lake Formation access denied errors across Athena, Glue Spark and Redshift, including the fine-grained access flags that silently bypass your grants
- Cost allocation tagging so lakehouse spend is attributable per catalog and per team from the first query rather than retrofitted later
- Honest second opinions on whether a migration is worth doing at all, including the answer that it isn’t
Send me a failing query, a Lake Formation grant listing or a CloudTrail export and I’ll tell you what I’d do with it.