{"id":855,"date":"2026-09-20T09:00:00","date_gmt":"2026-09-20T06:00:00","guid":{"rendered":"https:\/\/john-nessime.com\/blog\/?p=855"},"modified":"2026-09-20T11:56:52","modified_gmt":"2026-09-20T08:56:52","slug":"iceberg-table-maintenance-storage-costs","status":"publish","type":"post","link":"https:\/\/john-nessime.com\/blog\/technical-guides\/iceberg-table-maintenance-storage-costs\/","title":{"rendered":"Iceberg Table Maintenance: Why Your Storage Bill Keeps Growing"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">The question usually comes from finance, not engineering. Someone points at the bucket holding the analytics lake and asks why it is twice the size it was two quarters ago. Nobody added a pipeline. The tables are the same tables. One of them had a large chunk of history deleted on purpose, and the bucket still grew.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">That gap between what a table logically contains and what the object store physically holds is what Iceberg table maintenance exists to close. Iceberg never overwrites a file in place. Deletes and updates write new files and a new pointer, and the old files sit where they are until something explicitly removes them. If nothing does, the bill only moves one way.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This post covers where those bytes hide, why expiring snapshots often frees nothing, how orphan files accumulate without showing up in any metadata you would think to check, and the order to run cleanup so you do not corrupt a table while shrinking it.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Why deleting rows does not shrink anything<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Every write produces a snapshot: an immutable view of the table at a point in time, pointing at a manifest list, which points at manifests, which point at data files. Nothing in that chain is mutated. A new commit builds a new chain that reuses most of the old files.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">That design is what gives you time travel, rollback and safe concurrent writes. It is also why storage grows silently. A data file only becomes a deletion candidate once <em>no<\/em> retained snapshot references it. Delete a year of rows and you free nothing, because the snapshot from before the delete still points at every one of those files.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">So maintenance is not an optimization. It is the mechanism by which deletes become real.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">The four places your storage is actually sitting<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">When a lakehouse bucket balloons, the bytes are in one of four categories. Each needs a different tool, and confusing them is how teams end up running the same job for weeks with no effect.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">1. Data files held alive by snapshots you never expired<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">This is the one everyone knows about and usually the largest share. Snapshots accumulate until something expires them, and nothing expires them by default.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Before changing anything, look at the table&#8217;s own metadata. Iceberg exposes it as queryable tables:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>-- How many snapshots exist, and how far back do they go?\nSELECT committed_at, snapshot_id, operation\nFROM db.orders.snapshots\nORDER BY committed_at;\n\n-- How much data does the CURRENT snapshot reference?\nSELECT count(*) AS file_count,\n       sum(file_size_in_bytes) AS bytes\nFROM db.orders.files;<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Compare that second number against what your object store reports for the table prefix. The difference is your maintenance debt. On a neglected table the referenced set is often a minority of the physical bytes.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Expiring snapshots is a Spark stored procedure:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>CALL my_catalog.system.expire_snapshots(\n  table =&gt; 'db.orders',\n  older_than =&gt; TIMESTAMP '2000-01-01 00:00:00',\n  retain_last =&gt; 10,\n  stream_results =&gt; true\n);<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Three arguments matter. <code>older_than<\/code> sets the age cutoff. <code>retain_last<\/code> is the floor: it keeps that many ancestor snapshots regardless of age, which is your rollback safety net. <code>stream_results<\/code> streams the delete list back to the driver instead of collecting it at once, which on a table with long history is the difference between the job finishing and the driver running out of memory.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The trap is that Iceberg treats these conditions as guidance rather than a strict contract, leaning toward retention when they conflict. An aggressive <code>older_than<\/code> plus a generous <code>retain_last<\/code> gives you the generous answer. Safe, but it surprises people who expected the timestamp to win.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Branches and tags are separate: they are not subject to snapshot age at all. A tag created for a regulatory export a year ago pins its entire snapshot chain, and no amount of expiring will touch it. Reference age is governed by <code>history.expire.max-ref-age-ms<\/code>.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">2. Orphan files that no metadata points at<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">This is the invisible one, and the reason many teams conclude that maintenance &#8220;does not work.&#8221;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">When a Spark task writes a data file and the job then fails before the commit lands, that file is already on disk. It is in no manifest, no snapshot, no metadata file. Expiry will never find it, because expiry compares the file sets of expired snapshots against retained ones, and a file that was never in a snapshot is in neither set.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Streaming writers, speculative execution, retried stages and killed compaction jobs all produce these. The only way to find them is to list the storage prefix and diff it against the metadata:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>-- Always dry run first. This lists candidates without deleting.\nCALL my_catalog.system.remove_orphan_files(\n  table =&gt; 'db.orders',\n  dry_run =&gt; true\n);<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Read that output before you run it live. Recent timestamps or unfamiliar paths mean stop and investigate.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The safety mechanism is <code>older_than<\/code>, which defaults to three days. That default is not arbitrary caution. Orphan removal cannot distinguish a file abandoned by a dead job from one being written right now by a live one. Set the interval shorter than your longest-running write and the procedure can delete a file a job is about to commit, corrupting the table. Shorten it only when you know your write durations, and never below your longest backfill.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The second hazard is less obvious. Iceberg compares file paths as strings. If your metadata records <code>s3a:\/\/bucket\/...<\/code> and your listing returns <code>s3:\/\/bucket\/...<\/code>, or you migrated between HDFS namenodes and the authority changed, every live file looks like an orphan. Running that deletes the table. The procedure defends against it by erroring on a prefix mismatch by default, via <code>prefix_mismatch_mode<\/code> with <code>ERROR<\/code>, <code>IGNORE<\/code> and <code>DELETE<\/code> as options. Resolve a mismatch with <code>equal_schemes<\/code> or <code>equal_authorities<\/code>. Do not reach for <code>DELETE<\/code> to silence it.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">3. Metadata JSON files and manifests<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Every commit writes a new table metadata JSON file. On a batch table that is a handful a day. On a streaming table committing every minute it is a lot of small objects, and those cost you in request charges and planning time as much as in bytes.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>ALTER TABLE db.orders SET TBLPROPERTIES (\n  'write.metadata.delete-after-commit.enabled' = 'true',\n  'write.metadata.previous-versions-max' = '20'\n);<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">The first property is off by default. The second defaults to 100 and controls how many previous metadata files are <em>tracked<\/em> in the metadata log. The catch is that cleanup only deletes tracked files. If you have been running with deletion disabled, files that fell out of the log are already untracked, and turning the property on later will not reach back for them. They are orphans now, and only orphan file removal clears them.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">You can see the tracked set with <code>SELECT timestamp, file FROM db.orders.metadata_log_entries<\/code>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Manifests rarely dominate the bill, but a table with thousands of tiny ones plans queries slowly. <code>rewrite_manifests<\/code> consolidates them. Treat it as a query performance job and run it after the deletion work, so it is not reorganizing metadata about to be thrown away.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">4. The object store&#8217;s own hidden inventory<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Most Iceberg guides skip this category, and it often explains a bill nobody can account for. Iceberg has no visibility into any of it.<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Incomplete multipart uploads.<\/strong> Large Parquet files upload in parts. When a writer dies partway, the completed parts sit in the bucket, billed as storage, and they do not appear in a normal object listing. Nothing removes them automatically.<\/li>\n\n\n\n<li><strong>Noncurrent object versions.<\/strong> With versioning enabled, every file your maintenance job &#8220;deletes&#8221; becomes a noncurrent version you keep paying for, indefinitely, unless a rule expires it.<\/li>\n\n\n\n<li><strong>Delete markers.<\/strong> Deletes on a versioned bucket leave markers behind. Small individually, but on a table with millions of expired files they add up and slow listings down.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">A lifecycle configuration covers all three. On S3 the shape looks like this:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>{\n  \"Rules\": [\n    {\n      \"ID\": \"abort-stale-multipart\",\n      \"Status\": \"Enabled\",\n      \"Filter\": { \"Prefix\": \"warehouse\/\" },\n      \"AbortIncompleteMultipartUpload\": { \"DaysAfterInitiation\": 7 }\n    },\n    {\n      \"ID\": \"expire-noncurrent-versions\",\n      \"Status\": \"Enabled\",\n      \"Filter\": { \"Prefix\": \"warehouse\/\" },\n      \"NoncurrentVersionExpiration\": { \"NoncurrentDays\": 14 },\n      \"Expiration\": { \"ExpiredObjectDeleteMarker\": true }\n    }\n  ]\n}<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Now the warning that matters most here. <strong>Never put an age-based expiration or storage class transition rule on the prefix holding live Iceberg data files.<\/strong> A rule deleting objects older than ninety days will happily delete data files a current snapshot still references, and the table breaks with missing-file errors on the next read. Transitioning them to an archive class is just as bad, because reads then fail or stall on retrieval.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The rules above act only on things the table does not reference: abandoned upload parts, superseded versions, tombstones. Iceberg decides what live data to delete; lifecycle rules clean up after that decision.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This applies equally on S3-compatible storage. If you run a lakehouse on Cloudflare R2, Backblaze B2, Wasabi or Contabo Object Storage to dodge egress charges, check which lifecycle actions the provider actually implements. Support for aborting incomplete multipart uploads varies, and a provider that silently ignores the rule leaves you paying for parts you cannot see.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Running Iceberg table maintenance in a safe order<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Order matters, and the reason is not stylistic. These procedures read each other&#8217;s inputs.<\/p>\n\n\n\n<ol class=\"wp-block-list\">\n<li><strong>Compact first, if you are compacting.<\/strong> <code>rewrite_data_files<\/code> writes new files and leaves the old ones referenced by older snapshots. Doing it after expiry means the space you just freed is immediately replaced.<\/li>\n\n\n\n<li><strong>Expire snapshots.<\/strong> This is what turns unreferenced data files into deleted ones. It must happen before orphan removal, because expiry needs to read the manifests belonging to the snapshots it is dropping.<\/li>\n\n\n\n<li><strong>Remove orphan files.<\/strong> Only after expiry has committed. Running the two concurrently is a genuine race: orphan removal can delete files the in-flight expiry job still needs to read, and expiry fails.<\/li>\n\n\n\n<li><strong>Rewrite manifests, if planning is slow.<\/strong> Optional, and worth measuring before and after rather than running on faith.<\/li>\n<\/ol>\n\n\n\n<p class=\"wp-block-paragraph\">On cadence: expiry is cheap and should run often, daily or better on high-commit tables. Orphan removal is expensive because it lists the entire table prefix, and the official guidance is to run it periodically but not frequently. Weekly on a busy table, monthly on a quiet one. Where supported, the <code>prefix_listing<\/code> option makes that scan substantially cheaper on object stores, where recursive directory listing is the slow part.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Letting table properties do the boring part<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Procedures are what you run. Properties are what you set once so the procedures behave sensibly without a config file per table.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>ALTER TABLE db.orders SET TBLPROPERTIES (\n  'history.expire.max-snapshot-age-ms' = '604800000',   -- 7 days\n  'history.expire.min-snapshots-to-keep' = '10'\n);<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">The snapshot age default is five days and the minimum to keep defaults to one. Setting both means a generic maintenance job can call <code>expire_snapshots<\/code> with no tuning and each table gets its own policy. A regulated reporting table keeps ninety days; a staging table keeps an hour. That beats maintaining per-table retention in your scheduler.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The other property worth knowing is <code>gc.enabled<\/code>. Setting it false blocks physical deletion entirely. It exists for tables whose files are shared with something else, and it is correct for those. It is also a common accidental cause of &#8220;maintenance runs successfully and frees nothing.&#8221; Check it before debugging anything else.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Troubleshooting: the job succeeded and nothing shrank<\/h2>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>The counters were zero.<\/strong> <code>expire_snapshots<\/code> returns counts of deleted data files, delete files, manifests and manifest lists. All zero means nothing qualified. Check whether <code>retain_last<\/code> or the table&#8217;s minimum-to-keep property is holding everything.<\/li>\n\n\n\n<li><strong>A branch or tag is pinning history.<\/strong> Refs do not age out with snapshots. Look at the table&#8217;s refs metadata table and see what is anchored.<\/li>\n\n\n\n<li><strong>Garbage collection is off.<\/strong> Either <code>gc.enabled<\/code> is false, or the table was created by the <code>snapshot<\/code> procedure, which is barred from expiry precisely because it does not own its data files.<\/li>\n\n\n\n<li><strong>Files were deleted but the bucket did not shrink.<\/strong> Versioning. The deletes created noncurrent versions. This is the most common cause of healthy counters and a flat bill.<\/li>\n\n\n\n<li><strong>Orphan removal hit a file limit.<\/strong> Some managed implementations cap deletions per run. On a backlogged bucket you may need several passes before the number stops dropping.<\/li>\n\n\n\n<li><strong>The driver died.<\/strong> On a table with very long history, add <code>stream_results<\/code> and raise <code>max_concurrent_deletes<\/code> so deletes spread across a thread pool instead of serializing.<\/li>\n<\/ul>\n\n\n\n<h2 class=\"wp-block-heading\">Common mistakes<\/h2>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Running orphan file removal with a short retention window on a table that has long-running writers. This is the one that can actually destroy data.<\/li>\n\n\n\n<li>Setting a bucket-wide expiration rule to &#8220;clean up the lake&#8221; and taking live data files with it.<\/li>\n\n\n\n<li>Enabling metadata deletion after the fact and expecting it to collect the already-untracked backlog.<\/li>\n\n\n\n<li>Scheduling expiry and orphan removal as two independent cron jobs that can overlap.<\/li>\n\n\n\n<li>Skipping the dry run because the last twenty were fine. The one that matters is the run after somebody changes the storage endpoint.<\/li>\n\n\n\n<li>Setting retention from a storage target rather than a recovery target. Ask how far back you need to roll back, then keep that much.<\/li>\n<\/ul>\n\n\n\n<h2 class=\"wp-block-heading\">Best practices worth adopting<\/h2>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Express retention as a table property, not a job argument. It travels with the table and survives whoever wrote the scheduler.<\/li>\n\n\n\n<li>Track referenced bytes against physical bytes per table as a metric, not an occasional query. The ratio drifting upward is the early signal; the invoice is the late one. Any time series backend you already run will do.<\/li>\n\n\n\n<li>Give maintenance its own compute and its own credentials. It is the only workload in the lake that needs delete permissions.<\/li>\n\n\n\n<li>Log the returned counters from every run. When someone asks in six months where the storage went, that log is the answer.<\/li>\n<\/ul>\n\n\n\n<h2 class=\"wp-block-heading\">Frequently asked questions<\/h2>\n\n\n\n<h3 class=\"wp-block-heading\">How often should I expire Iceberg snapshots?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">As often as your commit rate justifies. A streaming table committing every minute benefits from daily or hourly expiry; a nightly batch table is fine with weekly. Set the window from how far back you would realistically need to roll back, then keep a few extra snapshots as a floor.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Does expiring snapshots delete my data?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">It deletes data files that no retained snapshot references. Current table contents are never affected. What you lose is the ability to time travel to or roll back to the expired snapshots, so agree the window with anyone who queries the table as of a date.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Why do orphan files exist if Iceberg commits are atomic?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">The commit is atomic; the writes preceding it are not part of that atomicity. Data files are written first, then the commit makes them visible. If the job dies between those steps, the files exist and the commit never happened. Atomicity guarantees readers never see a partial table, not that failed writes clean up after themselves.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Can I use an S3 lifecycle rule instead of Iceberg maintenance?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">No, and trying is how tables get broken. Lifecycle rules act on object age. Whether Iceberg still needs a file has nothing to do with age; a data file written two years ago can be referenced by the current snapshot. Lifecycle rules are right for upload parts, noncurrent versions and delete markers, and wrong for anything the table might still point at.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Do I need Spark to run Iceberg table maintenance?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Spark has the most complete set of procedures, but it is not the only route. Flink has its own maintenance API, Trino exposes expiry and orphan removal as table procedures, and several managed catalogs run scheduled maintenance for you. The mechanics are identical regardless of engine, because the work is defined by the table format rather than the compute. Only the syntax and the exposed arguments change.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Is orphan file removal safe to automate?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Yes, once you have verified that the retention window comfortably exceeds your longest write and that the paths in your metadata match those your storage listing returns. Automate it after watching a few dry runs, not before. Erroring on a prefix mismatch is a stop signal, not something to configure away.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">The one thing to take away<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Iceberg table maintenance is not housekeeping you get to postpone. It is the step where logical deletes become physical ones, and until it runs, storage only grows. The bill is the visible symptom; the same neglect also makes query planning slower and recovery windows fuzzier.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Set retention as a table property so the policy lives with the data. Expire snapshots on a schedule matching your commit rate. Remove orphan files less often, always after expiry, always with a window longer than your slowest write. Then add the storage lifecycle rules that clean up what Iceberg cannot see. Any one alone leaves money on the table.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">Need help getting your lakehouse storage under control?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Most of the work here is not writing the procedures. It is working out which of the four categories your bytes are in, and building a maintenance path that will not break a table at three in the morning. That is the kind of thing I do:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Audit an existing lake and produce a per-table breakdown of referenced bytes against physical bytes, so you know where the money is before touching anything<\/li>\n\n\n\n<li>Build a scheduled maintenance pipeline with correct sequencing, dry-run gates and per-table retention driven by table properties<\/li>\n\n\n\n<li>Set retention against real rollback and time travel requirements, gathered from the people who depend on them<\/li>\n\n\n\n<li>Design storage lifecycle rules that clear incomplete uploads, noncurrent versions and delete markers without touching live data files<\/li>\n\n\n\n<li>Diagnose maintenance jobs that report success and free nothing, including prefix mismatches, pinned refs and disabled garbage collection<\/li>\n\n\n\n<li>Add table health metrics and alerting so drift shows up as a graph rather than a surprise on the invoice<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">If you want a second opinion, send me the output of your snapshots metadata table, your current table properties, or the counters from a run that did not do what you expected. That is usually enough to tell you where the bytes are.<\/p>\n\n\n\n<div class=\"wp-block-buttons is-layout-flex wp-block-buttons-is-layout-flex\">\n<div class=\"wp-block-button\"><a class=\"wp-block-button__link wp-element-button\" href=\"https:\/\/www.upwork.com\/freelancers\/~01f15a912ad84a6620\" target=\"_blank\" rel=\"noreferrer noopener\">Work with me on Upwork<\/a><\/div>\n<\/div>\n","protected":false},"excerpt":{"rendered":"<p>Your lake bucket keeps growing even though nobody added data. Here is where the bytes actually sit: snapshots you never expired, orphan files no metadata points at, metadata JSON accumulation, and the object store&#8217;s own hidden inventory. Plus the safe order to clean all four up.<\/p>\n","protected":false},"author":1,"featured_media":1050,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[498,674,52],"tags":[161,164,634,187,160,237,1138,959,488,1137,941,106,583,958],"class_list":["post-855","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-data-engineering","category-finops","category-technical-guides","tag-amazon-s3","tag-apache-iceberg","tag-apache-spark","tag-cost-optimization","tag-data-lake","tag-lakehouse","tag-object-storage","tag-orphan-files","tag-s3-lifecycle-rules","tag-snapshot-expiration","tag-snapshots","tag-storage","tag-table-compaction","tag-table-maintenance","entry","has-media"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v28.5 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Iceberg Table Maintenance: Stop the Storage Bill Growing<\/title>\n<meta name=\"description\" content=\"Iceberg table maintenance in practice: why expiring snapshots frees nothing, how orphan files quietly pile up, and the storage settings hiding your bill.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/john-nessime.com\/blog\/technical-guides\/iceberg-table-maintenance-storage-costs\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Iceberg Table Maintenance: Stop the Storage Bill Growing\" \/>\n<meta property=\"og:description\" content=\"Iceberg table maintenance in practice: why expiring snapshots frees nothing, how orphan files quietly pile up, and the storage settings hiding your bill.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/john-nessime.com\/blog\/technical-guides\/iceberg-table-maintenance-storage-costs\/\" \/>\n<meta property=\"og:site_name\" content=\"John Nessime\" \/>\n<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/J.Nessime\" \/>\n<meta property=\"article:author\" content=\"https:\/\/www.facebook.com\/J.Nessime\" \/>\n<meta property=\"article:published_time\" content=\"2026-09-20T06:00:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-09-20T08:56:52+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/john-nessime.com\/blog\/wp-content\/uploads\/2026\/09\/Iceberg-Table-Maintenance-Why-Your-Storage-Bill-Keeps-Growing.png\" \/>\n\t<meta property=\"og:image:width\" content=\"1536\" \/>\n\t<meta property=\"og:image:height\" content=\"1024\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"author\" content=\"John Nessime\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"John Nessime\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"13 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/technical-guides\\\/iceberg-table-maintenance-storage-costs\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/technical-guides\\\/iceberg-table-maintenance-storage-costs\\\/\"},\"author\":{\"name\":\"John Nessime\",\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/#\\\/schema\\\/person\\\/ede0b56d0c808f123f57d5d796902105\"},\"headline\":\"Iceberg Table Maintenance: Why Your Storage Bill Keeps Growing\",\"datePublished\":\"2026-09-20T06:00:00+00:00\",\"dateModified\":\"2026-09-20T08:56:52+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/technical-guides\\\/iceberg-table-maintenance-storage-costs\\\/\"},\"wordCount\":2784,\"publisher\":{\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/#\\\/schema\\\/person\\\/ede0b56d0c808f123f57d5d796902105\"},\"image\":{\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/technical-guides\\\/iceberg-table-maintenance-storage-costs\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/09\\\/Iceberg-Table-Maintenance-Why-Your-Storage-Bill-Keeps-Growing.png\",\"keywords\":[\"Amazon S3\",\"Apache Iceberg\",\"Apache Spark\",\"Cost Optimization\",\"Data Lake\",\"Lakehouse\",\"Object Storage\",\"Orphan Files\",\"S3 Lifecycle Rules\",\"Snapshot Expiration\",\"Snapshots\",\"Storage\",\"Table Compaction\",\"Table Maintenance\"],\"articleSection\":[\"Data Engineering\",\"FinOps\",\"Technical Guides\"],\"inLanguage\":\"en-US\"},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/technical-guides\\\/iceberg-table-maintenance-storage-costs\\\/\",\"url\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/technical-guides\\\/iceberg-table-maintenance-storage-costs\\\/\",\"name\":\"Iceberg Table Maintenance: Stop the Storage Bill Growing\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/technical-guides\\\/iceberg-table-maintenance-storage-costs\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/technical-guides\\\/iceberg-table-maintenance-storage-costs\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/09\\\/Iceberg-Table-Maintenance-Why-Your-Storage-Bill-Keeps-Growing.png\",\"datePublished\":\"2026-09-20T06:00:00+00:00\",\"dateModified\":\"2026-09-20T08:56:52+00:00\",\"description\":\"Iceberg table maintenance in practice: why expiring snapshots frees nothing, how orphan files quietly pile up, and the storage settings hiding your bill.\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/technical-guides\\\/iceberg-table-maintenance-storage-costs\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/john-nessime.com\\\/blog\\\/technical-guides\\\/iceberg-table-maintenance-storage-costs\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/technical-guides\\\/iceberg-table-maintenance-storage-costs\\\/#primaryimage\",\"url\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/09\\\/Iceberg-Table-Maintenance-Why-Your-Storage-Bill-Keeps-Growing.png\",\"contentUrl\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/09\\\/Iceberg-Table-Maintenance-Why-Your-Storage-Bill-Keeps-Growing.png\",\"width\":1536,\"height\":1024,\"caption\":\"Apache Iceberg table maintenance showing hidden snapshots, manifests, data files, and orphan files driving AWS storage costs higher.\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/technical-guides\\\/iceberg-table-maintenance-storage-costs\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Iceberg Table Maintenance: Why Your Storage Bill Keeps Growing\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/#website\",\"url\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/\",\"name\":\"John Nessime\",\"description\":\"Cloud, DevOps, Data &amp; AI \u2014 Built, Tested, Explained\",\"publisher\":{\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/#\\\/schema\\\/person\\\/ede0b56d0c808f123f57d5d796902105\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":[\"Person\",\"Organization\"],\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/#\\\/schema\\\/person\\\/ede0b56d0c808f123f57d5d796902105\",\"name\":\"John Nessime\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/cropped-jn.png\",\"url\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/cropped-jn.png\",\"contentUrl\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/cropped-jn.png\",\"width\":512,\"height\":512,\"caption\":\"John Nessime\"},\"logo\":{\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/cropped-jn.png\"},\"description\":\"AWS Certified Solutions Architect helping businesses build reliable cloud, data, reporting, and automation solutions. I help startups, agencies, and growing businesses replace manual processes and disconnected data with practical AWS architectures, clean data pipelines, useful dashboards, and maintainable automation.\",\"sameAs\":[\"https:\\\/\\\/john-nessime.com\\\/blog\",\"https:\\\/\\\/www.facebook.com\\\/J.Nessime\",\"https:\\\/\\\/www.linkedin.com\\\/in\\\/john-m-nessime\"],\"url\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/author\\\/johnnessime\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Iceberg Table Maintenance: Stop the Storage Bill Growing","description":"Iceberg table maintenance in practice: why expiring snapshots frees nothing, how orphan files quietly pile up, and the storage settings hiding your bill.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/john-nessime.com\/blog\/technical-guides\/iceberg-table-maintenance-storage-costs\/","og_locale":"en_US","og_type":"article","og_title":"Iceberg Table Maintenance: Stop the Storage Bill Growing","og_description":"Iceberg table maintenance in practice: why expiring snapshots frees nothing, how orphan files quietly pile up, and the storage settings hiding your bill.","og_url":"https:\/\/john-nessime.com\/blog\/technical-guides\/iceberg-table-maintenance-storage-costs\/","og_site_name":"John Nessime","article_publisher":"https:\/\/www.facebook.com\/J.Nessime","article_author":"https:\/\/www.facebook.com\/J.Nessime","article_published_time":"2026-09-20T06:00:00+00:00","article_modified_time":"2026-09-20T08:56:52+00:00","og_image":[{"width":1536,"height":1024,"url":"https:\/\/john-nessime.com\/blog\/wp-content\/uploads\/2026\/09\/Iceberg-Table-Maintenance-Why-Your-Storage-Bill-Keeps-Growing.png","type":"image\/png"}],"author":"John Nessime","twitter_card":"summary_large_image","twitter_misc":{"Written by":"John Nessime","Est. reading time":"13 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/john-nessime.com\/blog\/technical-guides\/iceberg-table-maintenance-storage-costs\/#article","isPartOf":{"@id":"https:\/\/john-nessime.com\/blog\/technical-guides\/iceberg-table-maintenance-storage-costs\/"},"author":{"name":"John Nessime","@id":"https:\/\/john-nessime.com\/blog\/#\/schema\/person\/ede0b56d0c808f123f57d5d796902105"},"headline":"Iceberg Table Maintenance: Why Your Storage Bill Keeps Growing","datePublished":"2026-09-20T06:00:00+00:00","dateModified":"2026-09-20T08:56:52+00:00","mainEntityOfPage":{"@id":"https:\/\/john-nessime.com\/blog\/technical-guides\/iceberg-table-maintenance-storage-costs\/"},"wordCount":2784,"publisher":{"@id":"https:\/\/john-nessime.com\/blog\/#\/schema\/person\/ede0b56d0c808f123f57d5d796902105"},"image":{"@id":"https:\/\/john-nessime.com\/blog\/technical-guides\/iceberg-table-maintenance-storage-costs\/#primaryimage"},"thumbnailUrl":"https:\/\/john-nessime.com\/blog\/wp-content\/uploads\/2026\/09\/Iceberg-Table-Maintenance-Why-Your-Storage-Bill-Keeps-Growing.png","keywords":["Amazon S3","Apache Iceberg","Apache Spark","Cost Optimization","Data Lake","Lakehouse","Object Storage","Orphan Files","S3 Lifecycle Rules","Snapshot Expiration","Snapshots","Storage","Table Compaction","Table Maintenance"],"articleSection":["Data Engineering","FinOps","Technical Guides"],"inLanguage":"en-US"},{"@type":"WebPage","@id":"https:\/\/john-nessime.com\/blog\/technical-guides\/iceberg-table-maintenance-storage-costs\/","url":"https:\/\/john-nessime.com\/blog\/technical-guides\/iceberg-table-maintenance-storage-costs\/","name":"Iceberg Table Maintenance: Stop the Storage Bill Growing","isPartOf":{"@id":"https:\/\/john-nessime.com\/blog\/#website"},"primaryImageOfPage":{"@id":"https:\/\/john-nessime.com\/blog\/technical-guides\/iceberg-table-maintenance-storage-costs\/#primaryimage"},"image":{"@id":"https:\/\/john-nessime.com\/blog\/technical-guides\/iceberg-table-maintenance-storage-costs\/#primaryimage"},"thumbnailUrl":"https:\/\/john-nessime.com\/blog\/wp-content\/uploads\/2026\/09\/Iceberg-Table-Maintenance-Why-Your-Storage-Bill-Keeps-Growing.png","datePublished":"2026-09-20T06:00:00+00:00","dateModified":"2026-09-20T08:56:52+00:00","description":"Iceberg table maintenance in practice: why expiring snapshots frees nothing, how orphan files quietly pile up, and the storage settings hiding your bill.","breadcrumb":{"@id":"https:\/\/john-nessime.com\/blog\/technical-guides\/iceberg-table-maintenance-storage-costs\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/john-nessime.com\/blog\/technical-guides\/iceberg-table-maintenance-storage-costs\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/john-nessime.com\/blog\/technical-guides\/iceberg-table-maintenance-storage-costs\/#primaryimage","url":"https:\/\/john-nessime.com\/blog\/wp-content\/uploads\/2026\/09\/Iceberg-Table-Maintenance-Why-Your-Storage-Bill-Keeps-Growing.png","contentUrl":"https:\/\/john-nessime.com\/blog\/wp-content\/uploads\/2026\/09\/Iceberg-Table-Maintenance-Why-Your-Storage-Bill-Keeps-Growing.png","width":1536,"height":1024,"caption":"Apache Iceberg table maintenance showing hidden snapshots, manifests, data files, and orphan files driving AWS storage costs higher."},{"@type":"BreadcrumbList","@id":"https:\/\/john-nessime.com\/blog\/technical-guides\/iceberg-table-maintenance-storage-costs\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/john-nessime.com\/blog\/"},{"@type":"ListItem","position":2,"name":"Iceberg Table Maintenance: Why Your Storage Bill Keeps Growing"}]},{"@type":"WebSite","@id":"https:\/\/john-nessime.com\/blog\/#website","url":"https:\/\/john-nessime.com\/blog\/","name":"John Nessime","description":"Cloud, DevOps, Data &amp; AI \u2014 Built, Tested, Explained","publisher":{"@id":"https:\/\/john-nessime.com\/blog\/#\/schema\/person\/ede0b56d0c808f123f57d5d796902105"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/john-nessime.com\/blog\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":["Person","Organization"],"@id":"https:\/\/john-nessime.com\/blog\/#\/schema\/person\/ede0b56d0c808f123f57d5d796902105","name":"John Nessime","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/john-nessime.com\/blog\/wp-content\/uploads\/2026\/07\/cropped-jn.png","url":"https:\/\/john-nessime.com\/blog\/wp-content\/uploads\/2026\/07\/cropped-jn.png","contentUrl":"https:\/\/john-nessime.com\/blog\/wp-content\/uploads\/2026\/07\/cropped-jn.png","width":512,"height":512,"caption":"John Nessime"},"logo":{"@id":"https:\/\/john-nessime.com\/blog\/wp-content\/uploads\/2026\/07\/cropped-jn.png"},"description":"AWS Certified Solutions Architect helping businesses build reliable cloud, data, reporting, and automation solutions. I help startups, agencies, and growing businesses replace manual processes and disconnected data with practical AWS architectures, clean data pipelines, useful dashboards, and maintainable automation.","sameAs":["https:\/\/john-nessime.com\/blog","https:\/\/www.facebook.com\/J.Nessime","https:\/\/www.linkedin.com\/in\/john-m-nessime"],"url":"https:\/\/john-nessime.com\/blog\/author\/johnnessime\/"}]}},"_links":{"self":[{"href":"https:\/\/john-nessime.com\/blog\/wp-json\/wp\/v2\/posts\/855","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/john-nessime.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/john-nessime.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/john-nessime.com\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/john-nessime.com\/blog\/wp-json\/wp\/v2\/comments?post=855"}],"version-history":[{"count":1,"href":"https:\/\/john-nessime.com\/blog\/wp-json\/wp\/v2\/posts\/855\/revisions"}],"predecessor-version":[{"id":857,"href":"https:\/\/john-nessime.com\/blog\/wp-json\/wp\/v2\/posts\/855\/revisions\/857"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/john-nessime.com\/blog\/wp-json\/wp\/v2\/media\/1050"}],"wp:attachment":[{"href":"https:\/\/john-nessime.com\/blog\/wp-json\/wp\/v2\/media?parent=855"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/john-nessime.com\/blog\/wp-json\/wp\/v2\/categories?post=855"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/john-nessime.com\/blog\/wp-json\/wp\/v2\/tags?post=855"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}