<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Late-Arriving Data | John Nessime</title>
	<atom:link href="https://john-nessime.com/blog/tag/late-arriving-data/feed/" rel="self" type="application/rss+xml" />
	<link>https://john-nessime.com/blog/tag/late-arriving-data/</link>
	<description>Cloud, DevOps, Data &#38; AI — Built, Tested, Explained</description>
	<lastBuildDate>Thu, 24 Sep 2026 06:00:06 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://john-nessime.com/blog/wp-content/uploads/2026/07/cropped-jn-32x32.png</url>
	<title>Late-Arriving Data | John Nessime</title>
	<link>https://john-nessime.com/blog/tag/late-arriving-data/</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>Building a Hotel Data Lake on AWS That Agrees With the Night Audit</title>
		<link>https://john-nessime.com/blog/cloud-computing/hotel-data-lake-aws/</link>
		
		<dc:creator><![CDATA[John Nessime]]></dc:creator>
		<pubDate>Thu, 24 Sep 2026 06:00:06 +0000</pubDate>
				<category><![CDATA[Cloud Computing]]></category>
		<category><![CDATA[Data Engineering]]></category>
		<category><![CDATA[Hospitality Technology]]></category>
		<category><![CDATA[Amazon Athena]]></category>
		<category><![CDATA[Amazon S3]]></category>
		<category><![CDATA[Apache Iceberg]]></category>
		<category><![CDATA[AWS Glue]]></category>
		<category><![CDATA[Bitemporal Modeling]]></category>
		<category><![CDATA[Booking Systems]]></category>
		<category><![CDATA[Change Data Capture]]></category>
		<category><![CDATA[Data Lake]]></category>
		<category><![CDATA[Data Modeling]]></category>
		<category><![CDATA[Data Quality]]></category>
		<category><![CDATA[Hotel PMS]]></category>
		<category><![CDATA[Lake Formation]]></category>
		<category><![CDATA[Late-Arriving Data]]></category>
		<category><![CDATA[Medallion Architecture]]></category>
		<category><![CDATA[OPERA Cloud]]></category>
		<category><![CDATA[PCI DSS]]></category>
		<category><![CDATA[PII Redaction]]></category>
		<category><![CDATA[Reconciliation]]></category>
		<category><![CDATA[RevPAR]]></category>
		<category><![CDATA[Table Compaction]]></category>
		<category><![CDATA[Travel Technology]]></category>
		<category><![CDATA[USALI]]></category>
		<guid isPermaLink="false">https://john-nessime.com/blog/?p=461</guid>

					<description><![CDATA[<p>Hotel source data is mutable in the past, so an append-only pipeline drifts away from the PMS without anyone noticing until month close. A practical guide to room-night grain, bitemporal modeling with Apache Iceberg, PMS ingestion, guest data scope, and file physics at hotel volumes.</p>
<p>The post <a href="https://john-nessime.com/blog/cloud-computing/hotel-data-lake-aws/">Building a Hotel Data Lake on AWS That Agrees With the Night Audit</a> appeared first on <a href="https://john-nessime.com/blog">John Nessime</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">Someone opens the manager&#8217;s flash report on one screen and the shiny new dashboard on the other. The RevPAR figures don&#8217;t match. Not by a rounding error either, by enough that the revenue meeting stops being about revenue and turns into an argument about which system is lying.</p>



<p class="wp-block-paragraph">The dashboard usually isn&#8217;t lying. It&#8217;s answering a slightly different question, and the cause is almost always the same one: the pipeline recorded what the property management system said on the day it read it, and the PMS has since changed its mind. A reservation got shortened. A rate was corrected. A rebate landed against a folio that closed two weeks ago. The lake never went back to look.</p>



<p class="wp-block-paragraph">This post is about building a <strong>hotel data lake on AWS</strong> that survives that. Not a service tour, because the AWS documentation already covers S3, Glue and Athena better than a blog post can. What it covers instead is the handful of modelling and ingestion decisions that determine whether your numbers reconcile with the night audit six months from now, and what each of those decisions costs you.</p>



<h2 class="wp-block-heading">The failure that shows up three weeks late</h2>



<p class="wp-block-paragraph">Most first attempts at hotel analytics are append-only. Pull yesterday&#8217;s reservations, write Parquet under a folder keyed by the day you pulled it, register the partition, done. Yesterday reconciles perfectly. So does the day after. Confidence goes up.</p>



<p class="wp-block-paragraph">The problem is that hotel source data is mutable in the past, and it stays mutable for a long time. Group blocks get released. Stays get extended at the desk. Adjustments and rebates get posted against nights that already closed. Commission corrections come back from an OTA weeks after checkout. Night audit itself is a rewrite operation: it closes the business date and posts room and tax charges for every occupied room, and if the auditor reruns it or corrects a posting, the figures for that date change after you already stored them.</p>



<p class="wp-block-paragraph">An append-only lake captures a diagonal. For each business date you keep the first thing you ever saw about it, and nothing after. Every individual day looks right on the day you check it. The month-to-date total drifts. Nobody notices until close, and by then you have no idea which nights moved or why.</p>



<p class="wp-block-paragraph">That is the whole reason to be careful about grain and table format. Everything else in a hospitality pipeline is ordinary data engineering.</p>



<h2 class="wp-block-heading">Fix the grain before you fix the pipeline</h2>



<p class="wp-block-paragraph">The core hotel metrics are per-night metrics. Occupancy is rooms sold divided by rooms available. ADR is room revenue divided by rooms sold. RevPAR is room revenue divided by rooms available, which is the same as ADR multiplied by occupancy. All three are defined against a single night.</p>



<p class="wp-block-paragraph">What the PMS hands you is reservations. A reservation spans nights, carries one confirmation number, and may have different rates on different nights of the same stay. If you store reservations and try to compute occupancy from them, you end up writing date-range arithmetic into every query, and two analysts will write it differently.</p>



<h3 class="wp-block-heading">Explode to room-nights once, in the pipeline</h3>



<p class="wp-block-paragraph">The fact table that makes everything else easy has one row per room, per night, per reservation. Stay dates, rate code, market segment, source of business, channel, room type, rate amount and the taxes as separate columns. Every downstream metric becomes a filter and a sum instead of a range join.</p>



<p class="wp-block-paragraph">Do the explosion in the transform layer, not in the BI tool. If a range join lives in a Power BI model or a QuickSight dataset, it exists once per report and drifts once per report.</p>



<h3 class="wp-block-heading">The denominator is a decision, not a fact</h3>



<p class="wp-block-paragraph">&#8220;Rooms available&#8221; sounds like a constant. It isn&#8217;t. Rooms out of order for renovation, rooms out of service for a night, house-use rooms and comp rooms all get treated differently by different properties, and the PMS report you&#8217;re reconciling against has already made a choice.</p>



<p class="wp-block-paragraph">Same question on the numerator. Whether no-show revenue counts as room revenue, whether cancellation fees count, whether package components get split out of the room rate, whether ADR is quoted net or gross of tax. Hotel accounting conventions follow the Uniform System of Accounts for the Lodging Industry, and most PMS vendors document how their own reports treat each case. Read that documentation before you write a single SQL definition.</p>



<p class="wp-block-paragraph">Write the definitions down in the repository next to the transform code. When the lake and the PMS disagree later, the first question is always whether it&#8217;s a bug or a definition, and without written definitions you cannot answer that question quickly.</p>



<h2 class="wp-block-heading">Model restatements with two dates, not one</h2>



<p class="wp-block-paragraph">Every row needs a business date, which is the night the revenue belongs to, and an as-of timestamp, which is when you learned it. Two time axes. That&#8217;s what makes the difference between &#8220;the numbers changed&#8221; and &#8220;here is exactly what changed, when, and by how much&#8221;.</p>



<p class="wp-block-paragraph">There are two ways to store that. Keep every version of every row and let the reader pick a version, or keep the current truth in a table you update in place and keep the history separately. The second is easier to query and easier to explain, and it&#8217;s the one I reach for first. It needs a table format that supports updates and deletes, which in practice on AWS means Apache Iceberg.</p>



<p class="wp-block-paragraph">Iceberg gives you three things that matter here: row-level MERGE, snapshots you can query as of a point in time, and schema evolution without rewriting the dataset. Athena, Glue ETL, EMR and Redshift all read Iceberg tables through the Glue Data Catalog, so you&#8217;re not tying the storage to one engine.</p>



<p class="wp-block-paragraph">A curated room-night table in Athena looks roughly like this. Note the partitioning, which I&#8217;ll come back to.</p>



<pre class="wp-block-code"><code>CREATE TABLE curated.room_night (
  property_code      string,
  business_date      date,
  confirmation_id    string,
  room_number        string,
  room_type_code     string,
  rate_code          string,
  market_segment     string,
  source_of_business string,
  room_revenue       decimal(12,2),
  room_tax           decimal(12,2),
  reservation_status string,
  guest_key          string,
  as_of_ts           timestamp,
  source_system      string
)
PARTITIONED BY (property_code, month(business_date))
LOCATION 's3://example-lake-curated/room_night/'
TBLPROPERTIES ('table_type' = 'ICEBERG');</code></pre>



<p class="wp-block-paragraph">Loading is a MERGE from a staging table holding the latest extract. Matched rows get overwritten with the newer version, unmatched rows get inserted, and rows that vanished from the source get marked rather than deleted.</p>



<pre class="wp-block-code"><code>MERGE INTO curated.room_night AS t
USING staging.room_night_batch AS s
  ON  t.property_code   = s.property_code
  AND t.business_date   = s.business_date
  AND t.confirmation_id = s.confirmation_id
  AND t.room_number     = s.room_number
WHEN MATCHED AND s.as_of_ts &gt; t.as_of_ts THEN
  UPDATE SET
    room_revenue       = s.room_revenue,
    room_tax           = s.room_tax,
    rate_code          = s.rate_code,
    reservation_status = s.reservation_status,
    as_of_ts           = s.as_of_ts
WHEN NOT MATCHED THEN
  INSERT (property_code, business_date, confirmation_id, room_number,
          room_revenue, room_tax, rate_code, reservation_status, as_of_ts)
  VALUES (s.property_code, s.business_date, s.confirmation_id, s.room_number,
          s.room_revenue, s.room_tax, s.rate_code, s.reservation_status, s.as_of_ts);</code></pre>



<p class="wp-block-paragraph">The <code>as_of_ts</code> comparison in the MATCHED clause is doing real work. Extracts arrive out of order more often than you&#8217;d like, particularly when a backfill runs alongside the scheduled job, and without that guard a late-arriving older snapshot overwrites newer truth. It&#8217;s a silent corruption, which is the worst kind.</p>



<h3 class="wp-block-heading">Keep a restatement log, not just the current state</h3>



<p class="wp-block-paragraph">Before the MERGE runs, write the delta out to an append-only table: key, field changed, old value, new value, as-of timestamp. It costs almost nothing and it turns &#8220;the March numbers moved&#8221; into a query that returns rows.</p>



<p class="wp-block-paragraph">You can also lean on Iceberg snapshots and time travel for this, and they&#8217;re useful for spot checks. I still prefer an explicit log, because snapshot retention will eventually expire the snapshot you want, and finance will eventually ask about a night that&#8217;s older than your retention window.</p>



<h2 class="wp-block-heading">Landing the sources without building a point-to-point mess</h2>



<p class="wp-block-paragraph">A property runs more systems than people expect. PMS, POS, channel manager, revenue management, spa and golf, parking, loyalty, and whatever the marketing team bought last year. Land each of them raw and immutable first, then transform. The raw zone is your audit trail and your ability to reprocess without going back to a vendor API.</p>



<h3 class="wp-block-heading">The PMS is the hard one</h3>



<p class="wp-block-paragraph">For OPERA Cloud, the Oracle Hospitality Integration Platform is the supported route. It exposes REST APIs over OAuth, and registration is self-service rather than going through the older partner validation process. Two things to plan for: the API surface is large and reservation, financial and configuration data live in different areas, and access depends on the property&#8217;s licensing, so confirm what the property actually has before you scope the work.</p>



<p class="wp-block-paragraph">Legacy on-premise OPERA is a different world. Integration there goes through the older interface stack and the validation programme, and scheduled report exports to a file drop are frequently the pragmatic answer. Cloud-native systems like Cloudbeds and Mews expose REST APIs directly and are far less work.</p>



<p class="wp-block-paragraph">Reading the PMS database directly, where you can even reach it, is tempting and usually a bad idea. Schema changes arrive with vendor upgrades and nothing tells you in advance, and depending on the contract it can put support out of scope. Check the agreement before you build anything on top of direct database access.</p>



<h3 class="wp-block-heading">The relay box, if you need one</h3>



<p class="wp-block-paragraph">Properties with on-premise systems usually need something that picks up scheduled exports and pushes them to S3. Keep it boring: a small always-on Linux host, a scheduled job, checksum on write, and an SNS or CloudWatch alarm when a file doesn&#8217;t turn up. A cheap VPS from somewhere like Contabo or InterServer does this fine, and it is much easier to hand over at the end of an engagement than a machine sitting under the front desk.</p>



<p class="wp-block-paragraph">For reaching that host without opening inbound ports on a hotel network, an outbound-only tunnel such as Cloudflare Tunnel is a clean fit. If staff are hitting admin interfaces from guest or public Wi-Fi, a consumer VPN client like NordVPN or Surfshark protects that transport, though it is not a substitute for a real network path into the property.</p>



<h3 class="wp-block-heading">Time zones will bite you</h3>



<p class="wp-block-paragraph">The hotel business date is not a calendar day and it is not UTC. It&#8217;s a property-local concept that rolls at night audit, which might be at 2am or 4am depending on the property, and it can be held open when the audit runs late.</p>



<p class="wp-block-paragraph">Store the source system&#8217;s own business date as the partition key. Store event timestamps in UTC alongside it. Never derive the business date by truncating a UTC timestamp, because for a group with properties in several time zones that quietly shifts revenue between nights, and the error is small enough to look like noise.</p>



<h2 class="wp-block-heading">Keep card data out and plan for erasure on day one</h2>



<p class="wp-block-paragraph">Guest records are personal data and hotels hold a lot of it: names, addresses, passport and ID details in some jurisdictions, loyalty identifiers, stay history. Payment card data sits under a separate regime again.</p>



<p class="wp-block-paragraph">The single highest-value decision is negative. Do not bring cardholder data into the analytics platform at all. Nothing in occupancy, ADR or segment analysis needs a card number, and pulling one in drags the whole lake, its query engines and everyone with access into PCI DSS scope. Take the payment token or a masked reference from the PMS and stop there.</p>



<p class="wp-block-paragraph">For guest identity, pseudonymise at ingestion. Keep a surrogate guest key in the analytics tables and hold the mapping to real identity in one small, separately governed place with its own access controls. Amazon Macie can scan the raw zone for personal data that shouldn&#8217;t be there, which is useful precisely because someone will eventually land a CSV export nobody reviewed.</p>



<p class="wp-block-paragraph">Erasure needs thinking about before you have to do it. Iceberg makes row-level deletes straightforward, but a deleted row is still present in older snapshots, and time travel will happily return it. Deletion isn&#8217;t complete until snapshot expiration has run past the snapshots containing that row and the underlying files have been removed. Set retention deliberately, know what your window is, and be able to state it. Crypto-shredding, where the identity mapping is encrypted with a per-subject key you can destroy, is the cleaner design if the requirement is strict.</p>



<p class="wp-block-paragraph">The same discipline applies to workstations. Analysts pull extracts, extracts contain guest names, and laptops get replaced. A secure erasure tool such as O&amp;O SafeErase covers that end of it, and it belongs in the same policy as the lake.</p>



<p class="wp-block-paragraph">For access control, Lake Formation is the right layer. Define column and row filters once and they&#8217;re enforced across Athena, EMR and Redshift rather than reimplemented per engine. A group operator can then give each property&#8217;s management team row access to its own data from the same tables, which is far less work than building a schema per property.</p>



<h2 class="wp-block-heading">File physics at hotel scale</h2>



<p class="wp-block-paragraph">Hotel data is small. A single property produces a few hundred room-nights a day. Even a large group is nowhere near the volumes that data lake tutorials assume, and that changes the advice in one specific direction: your problem is too many tiny files, not too much data.</p>



<p class="wp-block-paragraph">Two habits create it. Partitioning by day when a property generates kilobytes per day, and polling the PMS every fifteen minutes so each run writes another sliver. Both feel careful and both make queries slower and more expensive, because the engine spends its time opening files rather than reading them.</p>



<ul class="wp-block-list">
<li>Partition curated tables by property and month, not by day. Iceberg prunes from its own metadata, so you get file-level skipping without the partition explosion.</li>

<li>Batch the writes. Buffer frequent polls and MERGE on a schedule that matches how the numbers are actually consumed, which for most operational reporting means hourly at most.</li>

<li>Run compaction. The Glue Data Catalog offers managed table optimizers for Iceberg tables covering compaction, snapshot retention and orphan file deletion, configurable from the console, CLI or API. Check the current format and engine support against the AWS docs before you rely on it.</li>

<li>Or drive it from Athena directly with <code>OPTIMIZE</code> to rewrite data files and <code>VACUUM</code> to expire snapshots and remove orphans.</li>
</ul>



<p class="wp-block-paragraph">For the raw zone, where tables are usually plain Hive-style Parquet rather than Iceberg, Athena partition projection is worth knowing about. It computes partition values from a configured pattern instead of listing them from the catalog, which removes a metadata round trip on tables with many partitions. It does not apply to Iceberg tables, which handle pruning through their own manifests.</p>



<p class="wp-block-paragraph">Athena bills on data scanned and Glue on job runtime, so cost tracks file layout and query shape rather than how much data you own. That&#8217;s a good thing at this scale, but it does mean a single badly written dashboard query on a schedule can cost more than the entire storage bill. Tools like Vantage or CloudZero make that visible per workload if the AWS cost console isn&#8217;t giving you enough granularity.</p>



<h2 class="wp-block-heading">Troubleshooting a hotel data lake on AWS</h2>



<p class="wp-block-paragraph"><strong>The lake and the PMS report disagree by a small, consistent amount.</strong> Almost always a definition, not a bug. Compare the room count first: if rooms sold matches and revenue doesn&#8217;t, you&#8217;re including or excluding tax, packages or fees differently. If rooms sold doesn&#8217;t match, check house-use and comp rooms.</p>



<p class="wp-block-paragraph"><strong>Yesterday reconciles, last month doesn&#8217;t.</strong> The append-only failure. Query your restatement log for that period. If you don&#8217;t have one, compare the current curated table against an older Iceberg snapshot for the same business dates.</p>



<p class="wp-block-paragraph"><strong>Numbers went backwards after a backfill.</strong> An older extract overwrote a newer one. Check that the as-of guard is present in the MERGE and that the staging table is deduplicated to one row per key before the merge runs, keeping the highest as-of value.</p>



<p class="wp-block-paragraph"><strong>Occupancy exceeds one hundred percent.</strong> Usually duplicate room-nights from a reservation that was modified into a new record while the old one stayed active, or a room moved mid-stay and counted in both rooms for the same night. Add a uniqueness check on property, business date and room number to the pipeline and fail the load rather than publishing it.</p>



<p class="wp-block-paragraph"><strong>Queries got slow and nothing changed.</strong> Look at file counts before anything else. Compaction that stopped running, or a new high-frequency job appending small files, explains most of these.</p>



<p class="wp-block-paragraph"><strong>A daily load silently produced nothing.</strong> This is the one that hurts, because an empty successful run looks like a healthy pipeline. Alert on the absence of expected rows per property per business date, not just on job failure. A deadman check in CloudWatch, or in Grafana Cloud if you already run dashboards there, catches it the same morning instead of at month end.</p>



<h2 class="wp-block-heading">Common mistakes</h2>



<ul class="wp-block-list">
<li>Storing reservations instead of room-nights, then rebuilding the date logic in every report.</li>

<li>Treating the ingestion date as the business date. They diverge the first time a load fails and gets rerun the next morning.</li>

<li>Building the reporting layer before agreeing the metric definitions with the finance and revenue teams.</li>

<li>Pulling the full guest profile because the API returns it, rather than the fields the analysis needs.</li>

<li>Skipping the raw zone and transforming on ingest, which means a logic bug is unrecoverable without re-extracting from the vendor.</li>

<li>Partitioning by day out of habit, then wondering why a small dataset queries slowly.</li>

<li>Relying on time travel as the audit trail, and discovering retention expired the snapshot finance asked about.</li>
</ul>



<h2 class="wp-block-heading">Best practices worth the effort</h2>



<ul class="wp-block-list">
<li>One row per room, per night, per reservation. Everything else derives from it.</li>

<li>Business date and as-of timestamp on every fact row, always.</li>

<li>Raw zone immutable, curated zone merged, consumption zone shaped for the tool that reads it.</li>

<li>Metric definitions in version control, next to the SQL that implements them.</li>

<li>A daily reconciliation query against one PMS report, checked automatically, with a tolerance and an alert.</li>

<li>Pseudonymise guest identity at ingestion and keep the mapping somewhere separate.</li>

<li>Compaction and snapshot retention configured before the table grows, not after someone complains.</li>
</ul>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Frequently asked questions</h2>



<h3 class="wp-block-heading">Do I need Iceberg, or is plain Parquet enough?</h3>



<p class="wp-block-paragraph">Plain Parquet is fine for anything genuinely append-only, such as raw extracts and event logs. It stops being fine the moment you need to correct history, because updating a row means rewriting whole partitions and coordinating that yourself. Since correcting history is the defining characteristic of hotel data, the curated layer wants a table format with row-level updates.</p>



<h3 class="wp-block-heading">Should the data lake match the PMS exactly?</h3>



<p class="wp-block-paragraph">It should match one named PMS report, on one named metric, within a tolerance you&#8217;ve agreed. Trying to match every report is a trap, because vendor reports use different denominators and different treatment of taxes and fees, and several of them disagree with each other. Pick the report finance already trusts and reconcile to that one.</p>



<h3 class="wp-block-heading">How often should the pipeline pull from the PMS?</h3>



<p class="wp-block-paragraph">Match the decision it feeds. Financial reporting is a daily cycle anchored on night audit and needs nothing faster. Operational views like arrivals and room status benefit from intraday refresh, but that&#8217;s a different table with a different retention policy, not a reason to run the whole lake hourly.</p>



<h3 class="wp-block-heading">Can guest names go in the data lake?</h3>



<p class="wp-block-paragraph">Sometimes, with a lawful basis, retention limits and access controls. But most of the analysis people actually want, occupancy, rate, segment, channel mix, pace, works perfectly on a pseudonymous key. Start without names and add them back only for a use case that genuinely requires them. Card data is a different matter and the answer there is simply no.</p>



<h3 class="wp-block-heading">Athena or Redshift for the reporting layer?</h3>



<p class="wp-block-paragraph">At single-property and small-group volumes, Athena over Iceberg is usually enough, and there&#8217;s no cluster to size or keep warm. Redshift earns its place when you have high-concurrency BI users, heavy joins across many domains, or an existing warehouse investment. Either way, keep the storage open so the choice stays reversible.</p>



<h3 class="wp-block-heading">How do I handle a group with many properties?</h3>



<p class="wp-block-paragraph">One set of tables with property code as the leading partition key, and Lake Formation row filters for per-property access. Separate schemas per property look tidy at three properties and become unmanageable at thirty, because every schema change and every metric fix has to be applied thirty times.</p>



<h3 class="wp-block-heading">Where does benchmarking data fit?</h3>



<p class="wp-block-paragraph">Land it in the same lake as a separate domain, at whatever grain the provider supplies, usually a market aggregate rather than a room-night. Keep it in its own tables and join at report time. Mixing external benchmark figures into your own fact table is how index numbers end up in a revenue total.</p>



<h2 class="wp-block-heading">The one thing to remember</h2>



<p class="wp-block-paragraph">A hotel data lake on AWS is not hard because of the volume. It&#8217;s hard because the past keeps changing, and an architecture that assumes otherwise fails quietly and slowly rather than loudly and fast.</p>



<p class="wp-block-paragraph">Get the room-night grain right, carry a business date and an as-of timestamp on every row, merge instead of appending, and write down what each metric means before anyone builds a dashboard on it. The S3, Glue, Iceberg and Athena assembly is the easy part, and it&#8217;s well documented. Reconciling with the night audit six months later is the part that decides whether anyone trusts the thing.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Working on this with me</h2>



<p class="wp-block-paragraph">Most of the hospitality data work I get asked about is one of a small number of shapes. Here&#8217;s where I&#8217;d be useful:</p>



<ul class="wp-block-list">
<li>Designing the room-night model and the bitemporal load, including the MERGE logic and the restatement log, so your numbers reconcile with the night audit and you can show why they moved.</li>

<li>Getting data out of the PMS reliably, whether that&#8217;s OHIP against OPERA Cloud, a REST API on a cloud-native system, or a scheduled export relay for an on-premise property.</li>

<li>Auditing an existing lake that has quietly drifted, finding which business dates restated and which metric definitions diverged from the reports finance uses.</li>

<li>Cutting Athena and Glue spend on a small dataset that behaves like a big one, usually through partition layout, compaction and killing scheduled queries nobody reads.</li>

<li>Setting up Lake Formation row and column controls for a multi-property group so each property sees its own data from shared tables.</li>

<li>Building the boring safety net: freshness and row-count checks per property per business date, deadman alerts, and a daily automated reconciliation against a named PMS report.</li>
</ul>



<p class="wp-block-paragraph">If you&#8217;ve got a discrepancy you can&#8217;t explain, send me the two figures and the query behind one of them, and I&#8217;ll tell you which side is wrong.</p>



<div class="wp-block-buttons is-layout-flex wp-block-buttons-is-layout-flex">
<div class="wp-block-button"><a class="wp-block-button__link wp-element-button" href="https://www.upwork.com/freelancers/~01f15a912ad84a6620" target="_blank" rel="noreferrer noopener">Work with me on Upwork</a></div>
</div>
<p>The post <a href="https://john-nessime.com/blog/cloud-computing/hotel-data-lake-aws/">Building a Hotel Data Lake on AWS That Agrees With the Night Audit</a> appeared first on <a href="https://john-nessime.com/blog">John Nessime</a>.</p>
]]></content:encoded>
					
		
		
			</item>
	</channel>
</rss>
