<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>AWS HealthLake | John Nessime</title>
	<atom:link href="https://john-nessime.com/blog/tag/aws-healthlake/feed/" rel="self" type="application/rss+xml" />
	<link>https://john-nessime.com/blog/tag/aws-healthlake/</link>
	<description>Cloud, DevOps, Data &#38; AI — Built, Tested, Explained</description>
	<lastBuildDate>Wed, 19 Aug 2026 12:12:48 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://john-nessime.com/blog/wp-content/uploads/2026/07/cropped-jn-32x32.png</url>
	<title>AWS HealthLake | John Nessime</title>
	<link>https://john-nessime.com/blog/tag/aws-healthlake/</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>Where PHI Escapes: Building a Secure Medical Lab Data Pipeline on AWS</title>
		<link>https://john-nessime.com/blog/cloud-computing/medical-lab-data-pipeline-aws/</link>
					<comments>https://john-nessime.com/blog/cloud-computing/medical-lab-data-pipeline-aws/#respond</comments>
		
		<dc:creator><![CDATA[John Nessime]]></dc:creator>
		<pubDate>Fri, 21 Aug 2026 09:00:00 +0000</pubDate>
				<category><![CDATA[Cloud Computing]]></category>
		<category><![CDATA[Cloud Security]]></category>
		<category><![CDATA[Compliance]]></category>
		<category><![CDATA[Data Engineering]]></category>
		<category><![CDATA[Amazon Comprehend Medical]]></category>
		<category><![CDATA[Amazon Macie]]></category>
		<category><![CDATA[Amazon S3]]></category>
		<category><![CDATA[AWS]]></category>
		<category><![CDATA[AWS HealthLake]]></category>
		<category><![CDATA[AWS KMS]]></category>
		<category><![CDATA[Business Associate Agreement]]></category>
		<category><![CDATA[CLIA]]></category>
		<category><![CDATA[Clinical Data]]></category>
		<category><![CDATA[Crypto Shredding]]></category>
		<category><![CDATA[Data Classification]]></category>
		<category><![CDATA[De-identification]]></category>
		<category><![CDATA[Dead Letter Queue]]></category>
		<category><![CDATA[ePHI]]></category>
		<category><![CDATA[FHIR]]></category>
		<category><![CDATA[Healthcare Cloud]]></category>
		<category><![CDATA[HIPAA]]></category>
		<category><![CDATA[HL7 v2]]></category>
		<category><![CDATA[Log Retention]]></category>
		<category><![CDATA[Multi-Account Strategy]]></category>
		<category><![CDATA[PHI]]></category>
		<category><![CDATA[Pipeline Design]]></category>
		<category><![CDATA[Pseudonymization]]></category>
		<category><![CDATA[S3 Object Lock]]></category>
		<guid isPermaLink="false">https://john-nessime.com/blog/?p=455</guid>

					<description><![CDATA[<p>Encryption is the part everyone gets right. The PHI sitting in your logs, dead letter queues and object key names is the part that fails an audit. A boundary-by-boundary guide to building a medical lab data pipeline on AWS that survives review.</p>
<p>The post <a href="https://john-nessime.com/blog/cloud-computing/medical-lab-data-pipeline-aws/">Where PHI Escapes: Building a Secure Medical Lab Data Pipeline on AWS</a> appeared first on <a href="https://john-nessime.com/blog">John Nessime</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">The question that breaks most healthcare data designs is not technical. It is an auditor asking you to list every place protected health information is stored. Someone points at the encrypted S3 bucket, the encrypted RDS instance, the encrypted backups, and says &#8220;those three.&#8221; Then you grep CloudWatch Logs and find an accession number, a date of birth, and a patient surname sitting inside a Python traceback with no retention policy set on the log group.</p>



<p class="wp-block-paragraph">That is the failure mode that actually bites when you build a medical lab data pipeline on AWS. Not the bucket you thought about. The seven places PHI ended up because a payload got logged, an error got queued, or a filename contained a medical record number.</p>



<p class="wp-block-paragraph">This post walks through the design boundary by boundary: where lab data crosses from one trust zone into another, what leaks at each crossing, and which control closes it. It is aimed at engineers who have been handed an HL7 feed and a compliance checklist and told to make them meet. I am an engineer and not your compliance counsel, so treat the regulatory references here as pointers to read the primary text yourself.</p>



<h2 class="wp-block-heading">Why laboratory data is a harder pipeline than it looks</h2>



<p class="wp-block-paragraph">Lab data has three properties that most ETL work does not.</p>



<p class="wp-block-paragraph">First, results get corrected. A preliminary result goes out, the analyzer is recalibrated, and a corrected result follows hours or days later. Your pipeline is not append-only in the way a clickstream is. It has to carry amendment semantics all the way through, and an analytics table that silently keeps the first value is worse than no table at all.</p>



<p class="wp-block-paragraph">Second, the wire format is old and the transport is older. Most laboratory information systems still emit HL7 version 2 messages, pipe-delimited segments carried over the Minimum Lower Layer Protocol, which is framed bytes over a raw TCP socket. AWS documents MLLP as the common transport standard for HL7 v2 interoperability in its Healthcare Industry Lens. It has no built-in authentication and no built-in encryption. Whatever security it has, you bolt on around it.</p>



<p class="wp-block-paragraph">Third, the identifiers are load-bearing. Accession numbers, specimen IDs and medical record numbers are how the lab reconciles anything, so PHI is not a column you can drop. It is structural, and it will try to escape into every metadata surface you own.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Boundary one: getting messages out of the lab</h2>



<p class="wp-block-paragraph">The instrument or LIS sits on a network you probably do not control, run by a vendor who may not let you touch the configuration. That constrains the design more than anything on the AWS side.</p>



<p class="wp-block-paragraph">You get three realistic shapes, and the LIS vendor usually picks for you.</p>



<ul class="wp-block-list">
<li><strong>Live MLLP over a private link.</strong> An interface engine in the lab holds the TCP session and forwards over a Site-to-Site VPN or Direct Connect into a private subnet. Lowest latency, highest operational burden, because a dropped socket is a clinical incident and not a retry.</li>

<li><strong>Batch file drop over SFTP.</strong> The LIS writes flat files on a schedule and something picks them up. Far easier to reason about, far easier to audit, and acceptable whenever the downstream use is analytics rather than result delivery.</li>

<li><strong>An outbound API push.</strong> Newer systems will POST to an endpoint you expose. Best case, because you control authentication and you get a response code back, but rarest in practice.</li>
</ul>



<p class="wp-block-paragraph">Whichever shape you get, the same rule applies at this boundary: <strong>terminate the untrusted protocol as early as possible and convert to something you can authorize.</strong> A raw MLLP socket carries no identity. The moment those bytes land, wrap them in a message you can attribute, sign, and trace.</p>



<p class="wp-block-paragraph">Two things go wrong here. First, people treat a network tunnel as authentication. A VPN tells you the packets came from the lab&#8217;s network. It does not tell you which system sent them, or that the sender was authorized to send that patient&#8217;s results. Second, people assume consumer VPN products fill the gap. Tools like NordVPN or Surfshark are fine for your own admin laptop on untrusted wifi, but they carry no business associate agreement and they are not a substitute for AWS Site-to-Site VPN or AWS Client VPN on a clinical path.</p>



<p class="wp-block-paragraph">One piece of housekeeping that is cheap to do and expensive to skip: the business associate addendum. AWS makes this self-service through AWS Artifact, and it can be accepted across an entire organization. Only services on the current HIPAA Eligible Services Reference may hold PHI, and that list changes, so check it against the services you actually chose rather than your memory of it.</p>



<h2 class="wp-block-heading">Boundary two: the landing zone</h2>



<p class="wp-block-paragraph">Raw messages land in S3. This part is well documented, so I will focus on the parts people configure once and never verify.</p>



<h3 class="wp-block-heading">Make the bucket policy do the enforcing</h3>



<p class="wp-block-paragraph">Default bucket encryption means objects get encrypted. It does not mean a caller cannot write an object with different settings, and it does not stop a plaintext HTTP request. Two explicit denials cover the gap. The first protects the transport, the second protects the storage:</p>



<pre class="wp-block-code"><code>{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "DenyInsecureTransport",
      "Effect": "Deny",
      "Principal": "*",
      "Action": "s3:*",
      "Resource": [
        "arn:aws:s3:::lab-raw-inbound",
        "arn:aws:s3:::lab-raw-inbound/*"
      ],
      "Condition": {
        "Bool": { "aws:SecureTransport": "false" }
      }
    },
    {
      "Sid": "DenyWrongKey",
      "Effect": "Deny",
      "Principal": "*",
      "Action": "s3:PutObject",
      "Resource": "arn:aws:s3:::lab-raw-inbound/*",
      "Condition": {
        "StringNotEquals": {
          "s3:x-amz-server-side-encryption-aws-kms-key-id":
            "arn:aws:kms:REGION:ACCOUNT:key/KEY-ID"
        }
      }
    }
  ]
}</code></pre>



<p class="wp-block-paragraph">Use a customer managed KMS key, not the AWS managed one. The reason is not stronger cryptography, it is the key policy. A customer managed key gives you a second, independent authorization surface: even a principal with broad S3 permissions cannot read the object if the key policy does not grant them decrypt. That separation is what makes a blast radius argument credible to an assessor, and it is what makes crypto shredding possible later.</p>



<h3 class="wp-block-heading">Your object keys are metadata, and metadata leaks</h3>



<p class="wp-block-paragraph">This is the one I would put on a poster. Object contents are encrypted. Object <em>keys</em> are not. They show up in CloudTrail data events, S3 server access logs, S3 Inventory reports, bucket listings, Lambda event payloads, and every console screenshot anyone pastes into a ticket.</p>



<pre class="wp-block-code"><code># Leaks PHI into six places that are not the bucket
raw/2024/ORU_MRN4471902_SMITH_JOHN_A1C.hl7

# Same object, no identifiers in the key path
raw/dt=&lt;partition&gt;/src=lis01/msg=01HXYZ...ULID.hl7</code></pre>



<p class="wp-block-paragraph">Use an opaque, sortable identifier and keep the mapping from identifier to patient inside an encrypted store you actually control. Partition on ingest date and source system, never on anything derived from the patient. This one decision removes more PHI surface than any other single change in the pipeline.</p>



<h3 class="wp-block-heading">Immutability, and the trap inside it</h3>



<p class="wp-block-paragraph">Clinical laboratories in the United States operating under CLIA have record retention obligations set out in 42 CFR 493.1105. The floor for test reports is at least two years after the date of reporting, with pathology test reports at ten years, and several state requirements sit above the federal floor. Retention is therefore not something you invent; go and read the section, then check your state.</p>



<p class="wp-block-paragraph">S3 Object Lock gives you write-once-read-many enforcement, and it is the right tool. But understand the two modes before you commit, because one of them is genuinely irreversible.</p>



<ul class="wp-block-list">
<li><strong>Governance mode</strong> blocks deletion for most principals, but anyone holding <code>s3:BypassGovernanceRetention</code> can override it by sending the <code>x-amz-bypass-governance-retention:true</code> header. Note the S3 console includes that header by default.</li>

<li><strong>Compliance mode</strong> cannot be overridden by anyone, including the account root. Per AWS documentation, the only way to delete an object under compliance retention before its date expires is to delete the AWS account.</li>
</ul>



<p class="wp-block-paragraph">A typo in a retention date under compliance mode is permanent, and you pay to store the mistake for its full term. Object Lock also has to be enabled at bucket creation, alongside versioning. Start in governance mode, prove the retention values against real traffic, then move to compliance mode deliberately rather than as a default.</p>



<pre class="wp-block-code"><code># Indefinite hold for a specific object version, e.g. under litigation
aws s3api put-object-legal-hold 
  --bucket lab-raw-inbound 
  --key raw/dt=.../msg=01HXYZ.hl7 
  --legal-hold Status=ON</code></pre>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Boundary three: where PHI escapes into your observability</h2>



<p class="wp-block-paragraph">Here is the section that justifies the post. Everything above is on the compliance checklist. This is the part that is invisible until an audit or an incident makes it visible.</p>



<p class="wp-block-paragraph">Your operational plane is a PHI store you did not declare. Specifically:</p>



<ul class="wp-block-list">
<li><strong>Application logs.</strong> One <code>print(record)</code> left in a parser during debugging writes patient demographics into CloudWatch Logs, where the default is often no expiry at all.</li>

<li><strong>Exception traces.</strong> Unhandled parse errors frequently include the offending segment in the message. That segment is the PHI.</li>

<li><strong>Step Functions execution history.</strong> If you pass message bodies between states as payloads, that content is retained in the execution record and readable by anyone who can view executions.</li>

<li><strong>Dead letter queues.</strong> The whole point of a DLQ is to keep the failed message. A DLQ full of unparseable HL7 is a PHI archive with a different access policy than the bucket you designed so carefully.</li>

<li><strong>Third-party monitoring.</strong> Anything shipped to an external platform leaves your boundary. Grafana Cloud, Datadog and equivalents are perfectly reasonable choices for infrastructure telemetry, but a log line containing PHI going to a vendor without an executed BAA is a disclosure, not a metric.</li>

<li><strong>Notifications.</strong> An SNS topic that emails an on-call engineer &#8220;failed to process record for J. Smith, DOB &#8230;&#8221; has just sent PHI to a mail server nobody reviewed.</li>
</ul>



<h3 class="wp-block-heading">The pattern that fixes it: pass pointers, not payloads</h3>



<p class="wp-block-paragraph">Adopt a single rule and enforce it in review: <strong>PHI moves by reference between components, never by value.</strong></p>



<ol class="wp-block-list">
<li>The ingest function writes the raw message to the encrypted bucket and emits only an opaque object key plus a message ID.</li>

<li>Every downstream state, queue and event carries that pointer. Nothing carries the record body.</li>

<li>Error handling logs the pointer and a failure class, never the content that failed.</li>

<li>A DLQ therefore contains pointers to failures, and reprocessing means re-reading from the bucket under the same key policy as everything else.</li>
</ol>



<p class="wp-block-paragraph">The cost of this pattern is a slightly noisier debugging experience. You cannot read the failing record straight out of the queue. You have to go and fetch it, with credentials, and that fetch is logged. That inconvenience is the control working.</p>



<h3 class="wp-block-heading">Then close the log surface itself</h3>



<p class="wp-block-paragraph">Log groups created implicitly by Lambda and other services often retain data indefinitely unless you set a policy. Find the ones nobody configured:</p>



<pre class="wp-block-code"><code># List log groups with no retention configured
aws logs describe-log-groups 
  --query 'logGroups[?retentionInDays==null].logGroupName' 
  --output text

# Set an explicit retention on one of them
aws logs put-retention-policy 
  --log-group-name /aws/lambda/hl7-ingest 
  --retention-in-days 30</code></pre>



<p class="wp-block-paragraph">Run the first command as a scheduled check, not a one-off. New functions create new log groups, and the default comes back every time someone deploys.</p>



<p class="wp-block-paragraph">Amazon Macie is worth pointing at your buckets as a detective backstop. It will not stop a leak, but it will tell you when something started landing where it should not.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Boundary four: from identified data to analytics</h2>



<p class="wp-block-paragraph">At some point someone wants to query the data. This is the crossing where teams are most likely to convince themselves they have done something they have not.</p>



<p class="wp-block-paragraph">Amazon Comprehend Medical will detect PHI entities in unstructured clinical text through its <code>DetectPHI</code> operation, returning entity types such as name, address, identifier and date, each with a confidence score. It is a genuinely useful tool for narrative fields: pathology comments, specimen notes, microbiology free text.</p>



<pre class="wp-block-code"><code>aws comprehendmedical detect-phi 
  --text "Specimen received from the referring clinic on the stated date."</code></pre>



<p class="wp-block-paragraph">Now the part the marketing pages underplay, straight from the AWS developer guide: the entities Comprehend Medical detects <em>do not map one to one</em> to the identifier list specified by the Safe Harbor method, and AWS explicitly recommends additional human review or other methods to confirm accuracy for compliance use cases.</p>



<p class="wp-block-paragraph">Read that as: PHI detection is not de-identification. Under the HIPAA Privacy Rule, de-identification has two defined routes, Safe Harbor and Expert Determination. An NLP model with confidence scores is an input to either, not a substitute for either. If your analytics tier is meant to hold de-identified data, someone has to own that determination, and it will not be your pipeline.</p>



<p class="wp-block-paragraph">What works structurally:</p>



<ul class="wp-block-list">
<li>Put the analytics tier in a <strong>separate AWS account</strong> with its own KMS key. Not a separate bucket. A separate account, so the blast radius argument is enforced by IAM boundaries rather than by convention.</li>

<li>Make the flow <strong>one directional</strong>. The de-identification job reads from the identified side and writes to the analytics side. Nothing in the analytics account holds permissions pointing back.</li>

<li><strong>Replace identifiers with pseudonyms</strong> rather than deleting them, keeping the crosswalk in the identified account. This preserves your ability to join across results and to re-identify under a documented process, which the research use case usually needs.</li>

<li>Watch <strong>quasi-identifiers</strong>. Rare test panels, unusual reference ranges and precise timestamps re-identify people in small populations even after the obvious fields are gone. This is a statistics problem, not an IAM problem.</li>

<li>Use <strong>Lake Formation</strong> or equivalent for column and row filtering if analysts need partial access, so the restriction lives with the catalog rather than in each query tool.</li>
</ul>



<h2 class="wp-block-heading">Boundary five: sending results back out</h2>



<p class="wp-block-paragraph">Delivery is where clinical correctness and security pull against each other, and security usually loses quietly.</p>



<p class="wp-block-paragraph">The correction problem is the sharp edge. HL7 result messages carry a status that distinguishes preliminary, final and corrected results. If your delivery layer is idempotent on message ID alone, a corrected result with a new ID for the same observation will happily land beside the original instead of superseding it. Two conflicting values, both marked delivered, nobody alerted.</p>



<p class="wp-block-paragraph">Build the idempotency key from the identity of the <em>observation</em>, which typically means the placer or filler order identifier combined with the specific observation identifier, and carry the result status as a first-class field so a correction overwrites rather than appends. Keep every version, expose the current one.</p>



<p class="wp-block-paragraph">For the outbound path itself:</p>



<ul class="wp-block-list">
<li>Presigned URLs are convenient for report PDFs and they are also a bearer credential. Anyone holding the link has the document. Keep expiry short and treat generation as an auditable event.</li>

<li>If you expose a provider-facing portal, put a WAF in front of it. AWS WAF and Cloudflare are both credible here; the deciding factor is usually where the rest of your edge already lives, not the rule engines.</li>

<li>Never email a result. Email a notification that a result is available, behind authentication. The distinction sounds pedantic right up until someone forwards a mailbox.</li>
</ul>



<h2 class="wp-block-heading">Boundary six: deletion, and the conflict nobody plans for</h2>



<p class="wp-block-paragraph">Two obligations point in opposite directions. Retention rules say keep the record. Privacy rights and internal policy say be able to remove data. If you set compliance-mode Object Lock across everything, you have chosen one side without noticing.</p>



<p class="wp-block-paragraph">Resolve it by classifying before you lock. Records that fall under laboratory retention obligations go into the immutable tier with a retention period derived from the applicable rule. Derived artifacts, caches, intermediate parquet, analytics extracts and enrichment outputs do not belong there. They belong in buckets with lifecycle rules, and they are what you actually delete.</p>



<p class="wp-block-paragraph">Crypto shredding covers the middle ground: if a data set is encrypted under a key used for nothing else, scheduling that key for deletion renders the ciphertext unrecoverable without touching the objects. It is clean, it is verifiable, and it only works if you planned the key granularity up front. Retrofitting per-tenant or per-cohort keys onto a pipeline that used one key for everything is a full re-encryption exercise.</p>



<p class="wp-block-paragraph">Do not forget the endpoints. Instrument workstations, the interface engine box, and the laptop somebody used to test the parser all accumulate copies. Cloud controls do nothing for physical media, and secure erase tooling such as O&amp;O SafeErase exists for exactly that step in a decommissioning runbook.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Troubleshooting a medical lab data pipeline on AWS</h2>



<h3 class="wp-block-heading">Messages arrive but nothing appears downstream</h3>



<p class="wp-block-paragraph">Check the DLQ first, then check whether the DLQ itself has a consumer. A surprising number of pipelines have a correctly configured dead letter queue that nobody monitors, so failures accumulate silently and the only symptom is a gap in the data. Alarm on DLQ depth greater than zero, not on a threshold.</p>



<h3 class="wp-block-heading">Access denied from a Lambda that has the right IAM policy</h3>



<p class="wp-block-paragraph">With a customer managed KMS key, IAM permission on S3 is only half the grant. The function&#8217;s role also needs decrypt permission on the key, and the key policy has to allow it. If you are crossing accounts, both the key policy and the role policy must permit the action. This is the intended behavior of the two-lock design, and it is the single most common support question on any pipeline built this way.</p>



<h3 class="wp-block-heading">Deletion fails on an object you are certain you should be able to delete</h3>



<p class="wp-block-paragraph">Check for a legal hold before you check retention. A legal hold has no expiry and is independent of the retention period, so an object can be past its retention date and still undeletable. Also confirm which mode the retention uses, because governance and compliance produce the same error to a caller without bypass permission.</p>



<h3 class="wp-block-heading">Counts do not reconcile with the LIS</h3>



<p class="wp-block-paragraph">Compare on observations, not messages. One message can carry several results, corrections create additional messages for the same observation, and a naive message count will diverge from the LIS by exactly the amount that matters clinically. Reconcile daily and alert on drift rather than investigating at quarter end.</p>



<h3 class="wp-block-heading">Timestamps drift between systems</h3>



<p class="wp-block-paragraph">HL7 v2 timestamps do not always carry a timezone offset, and analyzers are often set to local time with no daylight saving handling. Capture the offset at ingest from the source system configuration and store everything in UTC with the original string preserved. Recovering an ambiguous timestamp after the fact is unpleasant, and around a daylight saving transition it may not be possible at all.</p>



<h2 class="wp-block-heading">Common mistakes</h2>



<ul class="wp-block-list">
<li><strong>Treating &#8220;HIPAA eligible&#8221; as &#8220;compliant.&#8221;</strong> Eligibility means AWS will cover the service under the BAA. Configuration is still entirely yours.</li>

<li><strong>Putting identifiers in object keys, queue names, or log group names.</strong> Encrypted contents, plaintext metadata.</li>

<li><strong>Enabling compliance-mode Object Lock before validating retention values.</strong> There is no undo, and no support ticket that fixes it.</li>

<li><strong>Assuming automated PHI detection equals de-identification.</strong> AWS itself says the entity list does not map one to one to Safe Harbor identifiers.</li>

<li><strong>Running the analytics tier in the same account as identified data.</strong> A bucket boundary is a policy away from collapsing. An account boundary is not.</li>

<li><strong>Shipping application logs to an external platform without checking what the log lines contain.</strong> Test it by grepping your own log group for a known test patient name.</li>

<li><strong>Ignoring corrected results until an analyst notices.</strong> Amendment handling is a day-one requirement, not a phase two feature.</li>
</ul>



<h2 class="wp-block-heading">Best practices worth the effort</h2>



<ul class="wp-block-list">
<li><strong>Write the data flow diagram first and mark every trust boundary crossing.</strong> Every arrow that crosses one needs a named control. This exercise finds more problems than any scanner.</li>

<li><strong>Encode the eligible service list as policy.</strong> Service control policies at the organization level stop someone from putting PHI into a service you never assessed.</li>

<li><strong>Keep VPC endpoints on the PHI path.</strong> Interface and gateway endpoints keep traffic to AWS services off the public internet and give you an endpoint policy as an extra choke point. Check which of your chosen services offer one.</li>

<li><strong>Build with synthetic HL7 from day one.</strong> Nobody should need production PHI to develop a parser. A cheap VPS from Contabo or InterServer is fine for a synthetic-data development box, and keeping that environment entirely outside the PHI boundary is the point.</li>

<li><strong>Alarm on absence.</strong> A feed that stops is more dangerous than a feed that errors, because errors are loud. Alert when expected message volume for a source drops below its floor.</li>

<li><strong>Test restores, not backups.</strong> Restore into an isolated account, confirm the KMS grants work there, and document how long it took.</li>

<li><strong>Review CloudTrail for data events on the PHI buckets.</strong> Object-level logging costs money and is the only record of who read what.</li>
</ul>



<h2 class="wp-block-heading">FAQ</h2>



<h3 class="wp-block-heading">Do I need a BAA with AWS before I start building?</h3>



<p class="wp-block-paragraph">Before PHI touches the environment, yes. You can architect and test with synthetic data first. AWS provides the business associate addendum through AWS Artifact as a self-service acceptance that can be applied across an organization, and only services listed on the HIPAA Eligible Services Reference may process, store or transmit PHI.</p>



<h3 class="wp-block-heading">Should I use AWS HealthLake or build on S3, Glue and Athena?</h3>



<p class="wp-block-paragraph">HealthLake is a managed FHIR datastore with FHIR APIs, medical NLP and query built in, and it is the shorter path if your consumers speak FHIR. It does not natively ingest HL7 v2; AWS points to partner tooling or your own transformation for non-FHIR input. If your consumers are analysts with SQL and the data is structured lab results rather than clinical documents, an S3 and Athena lakehouse is simpler and easier to keep inside a boundary you control. Decide on the consumer, not the source.</p>



<h3 class="wp-block-heading">How long do I have to keep laboratory records?</h3>



<p class="wp-block-paragraph">For CLIA-regulated laboratories in the United States, 42 CFR 493.1105 sets the federal floor, including at least two years for test reports after the date of reporting and at least ten years for pathology test reports. State law frequently exceeds these minimums, so the applicable period is whichever is longer for your jurisdiction. Read the regulation and confirm with your compliance lead before you encode a number into a retention policy.</p>



<h3 class="wp-block-heading">Can Amazon Comprehend Medical de-identify data for me?</h3>



<p class="wp-block-paragraph">It can detect PHI entities and give you confidence scores, which is a strong starting point for redaction workflows. It does not by itself satisfy HIPAA&#8217;s de-identification standard. AWS documents that the detected entities do not map one to one to the Safe Harbor identifier list and recommends human review for compliance use cases. Treat it as a detector feeding a documented Safe Harbor or Expert Determination process.</p>



<h3 class="wp-block-heading">Is S3 Object Lock compliance mode required for lab records?</h3>



<p class="wp-block-paragraph">Not inherently. The argument for compliance mode is a threat model in which a compromised or malicious administrator could remove retention, not a regulation that names the mode. Governance mode plus tightly controlled bypass permissions and strong audit logging is a defensible position for many laboratories. Choose deliberately and write down the reasoning, because compliance mode commits you to the storage cost for the full term with no exit.</p>



<h3 class="wp-block-heading">Where does PHI most often leak in a pipeline that looks correctly configured?</h3>



<p class="wp-block-paragraph">Logs, dead letter queues, workflow execution histories, notification messages and object key names. All five are metadata surfaces that sit outside the storage layer everyone reviews. Grep your own log groups for a known test patient identifier and you will find out quickly whether your pipeline has the problem.</p>



<h3 class="wp-block-heading">Do I need a dedicated AWS account just for the lab pipeline?</h3>



<p class="wp-block-paragraph">Separate accounts for identified and de-identified data are worth it even at small scale, because the boundary is then enforced by IAM rather than by naming conventions. Whether the PHI workload also needs isolating from your other production workloads depends on who holds admin access there. If the answer is &#8220;broadly the same people,&#8221; separate it.</p>



<h2 class="wp-block-heading">The one thing to take away</h2>



<p class="wp-block-paragraph">A secure medical lab data pipeline on AWS is not defined by the encryption on its storage layer. Encryption is the easy part, and it is the part everyone gets right. What separates a design that survives an audit from one that does not is whether PHI can travel by value through your operational plane: into logs, into queue bodies, into workflow histories, into filenames, into alert emails.</p>



<p class="wp-block-paragraph">Move PHI by reference, keep identifiers out of every metadata surface, and put an account boundary between identified and de-identified data. Do those three things and the rest of the compliance work becomes documentation rather than redesign.</p>



<h2 class="wp-block-heading">Work with me on healthcare data pipelines</h2>



<p class="wp-block-paragraph">I design and review AWS data platforms for regulated workloads, and lab data is a specific enough problem that generic cloud advice tends to miss it. Things I can help with:</p>



<ul class="wp-block-list">
<li>Mapping every trust boundary in an existing pipeline and finding the PHI surfaces nobody documented, including logs, queues and object key schemes.</li>

<li>Designing HL7 v2 or file-drop ingestion into S3 with KMS key separation, VPC endpoints and bucket policies that enforce rather than suggest.</li>

<li>Building the identified and de-identified account split, including the one-way transformation job and the pseudonym crosswalk.</li>

<li>Getting amendment and correction handling right so corrected results supersede rather than duplicate, all the way to the analytics tables.</li>

<li>Retention and deletion design: Object Lock modes, lifecycle policies, crypto shredding key granularity, and the runbook that ties them together.</li>

<li>Observability that is useful without being a disclosure risk, including CloudWatch retention hygiene and safe alerting patterns.</li>
</ul>



<p class="wp-block-paragraph">If you have an architecture diagram, a Terraform plan, or a redacted sample of the messages you are receiving, send it over and I will tell you what I would change first.</p>



<div class="wp-block-buttons is-layout-flex wp-block-buttons-is-layout-flex">
<div class="wp-block-button"><a class="wp-block-button__link wp-element-button" href="https://www.upwork.com/freelancers/~01f15a912ad84a6620" target="_blank" rel="noreferrer noopener">Work with me on Upwork</a></div>
</div>
<p>The post <a href="https://john-nessime.com/blog/cloud-computing/medical-lab-data-pipeline-aws/">Where PHI Escapes: Building a Secure Medical Lab Data Pipeline on AWS</a> appeared first on <a href="https://john-nessime.com/blog">John Nessime</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://john-nessime.com/blog/cloud-computing/medical-lab-data-pipeline-aws/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
	</channel>
</rss>
