<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Case Studies | John Nessime</title>
	<atom:link href="https://john-nessime.com/blog/case-studies/feed/" rel="self" type="application/rss+xml" />
	<link>https://john-nessime.com/blog/case-studies/</link>
	<description>Cloud, DevOps, Data &#38; AI — Built, Tested, Explained</description>
	<lastBuildDate>Sun, 20 Sep 2026 16:20:12 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://john-nessime.com/blog/wp-content/uploads/2026/07/cropped-jn-32x32.png</url>
	<title>Case Studies | John Nessime</title>
	<link>https://john-nessime.com/blog/case-studies/</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>Legal Document Intelligence on AWS: The Five Boundaries That Have to Hold</title>
		<link>https://john-nessime.com/blog/cloud-computing/legal-document-intelligence-aws/</link>
		
		<dc:creator><![CDATA[John Nessime]]></dc:creator>
		<pubDate>Tue, 25 Aug 2026 18:00:00 +0000</pubDate>
				<category><![CDATA[Case Studies]]></category>
		<category><![CDATA[Cloud Computing]]></category>
		<category><![CDATA[Web Security]]></category>
		<category><![CDATA[Amazon Bedrock]]></category>
		<category><![CDATA[Amazon Comprehend]]></category>
		<category><![CDATA[Amazon S3]]></category>
		<category><![CDATA[Amazon Textract]]></category>
		<category><![CDATA[Architecture]]></category>
		<category><![CDATA[Audit Logging]]></category>
		<category><![CDATA[AWS]]></category>
		<category><![CDATA[AWS KMS]]></category>
		<category><![CDATA[AWS Organizations]]></category>
		<category><![CDATA[Bedrock Guardrails]]></category>
		<category><![CDATA[Bedrock Knowledge Bases]]></category>
		<category><![CDATA[Cloud Security]]></category>
		<category><![CDATA[Compliance]]></category>
		<category><![CDATA[Data Residency]]></category>
		<category><![CDATA[Document Processing]]></category>
		<category><![CDATA[Encryption]]></category>
		<category><![CDATA[Human In The Loop]]></category>
		<category><![CDATA[Intelligent Document Processing]]></category>
		<category><![CDATA[Legal Tech]]></category>
		<category><![CDATA[Litigation Hold]]></category>
		<category><![CDATA[Metadata Filtering]]></category>
		<category><![CDATA[Multi-Tenant]]></category>
		<category><![CDATA[OCR]]></category>
		<category><![CDATA[OpenSearch Serverless]]></category>
		<category><![CDATA[PrivateLink]]></category>
		<category><![CDATA[RAG]]></category>
		<category><![CDATA[S3 Object Lock]]></category>
		<category><![CDATA[Step Functions]]></category>
		<category><![CDATA[Tenant Isolation]]></category>
		<category><![CDATA[Vector Database]]></category>
		<guid isPermaLink="false">https://john-nessime.com/blog/?p=272</guid>

					<description><![CDATA[<p>A confident answer with a citation from a matter the reader was walled off from. Nothing crashed, nothing alerted. Building legal document intelligence on AWS with Textract, Bedrock and OpenSearch is the easy half; the hard half is tenant isolation, verified deletion, S3 Object Lock holds, what leaves the account, and an audit trail that names the actual user.</p>
<p>The post <a href="https://john-nessime.com/blog/cloud-computing/legal-document-intelligence-aws/">Legal Document Intelligence on AWS: The Five Boundaries That Have to Hold</a> appeared first on <a href="https://john-nessime.com/blog">John Nessime</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">The demo was going well until someone asked where the third citation came from.</p>



<p class="wp-block-paragraph">The answer on screen was good. Fluent, specific, correctly hedged, three sources listed underneath it. Then a partner in the room asked a simple question: which matter is source three from? Nobody could answer it in the room, and when we went and looked, it was from a matter that half the people watching the demo were formally walled off from.</p>



<p class="wp-block-paragraph">Nothing had crashed. No alarm fired. The retrieval layer had done exactly what a similarity search does, which is return the nearest vectors, and the nearest vectors did not care about the ethical wall. That is the thing about building a legal document intelligence platform on AWS: the failures that matter almost never look like failures. They look like a confident answer with a citation attached.</p>



<p class="wp-block-paragraph">This post is about the parts of that build that are hard, organised around the five boundaries that actually have to hold. The extraction pipeline is the easy half. I will cover it, but quickly, because the AWS documentation is good and the failure modes are visible. The other four boundaries fail quietly, and that is where the engineering goes.</p>



<h2 class="wp-block-heading">The shape of the thing</h2>



<p class="wp-block-paragraph">Before the boundaries, the skeleton. A legal document intelligence platform on AWS almost always ends up looking like this, whether you plan it or arrive at it:</p>



<ul class="wp-block-list">
<li>Documents land in Amazon S3, one prefix per matter, versioning on.</li>



<li>An S3 event triggers AWS Lambda, which starts an asynchronous Amazon Textract job.</li>



<li>Textract output lands back in S3 as JSON. Something normalises it into text plus page and bounding-box coordinates.</li>



<li>Amazon Comprehend or a foundation model classifies the document and pulls entities: parties, dates, governing law, clause types.</li>



<li>Chunks get embedded and written to a vector store, usually Amazon OpenSearch Serverless behind Amazon Bedrock Knowledge Bases.</li>



<li>An application layer, typically API Gateway plus Lambda, takes a question, retrieves, and calls a model on Amazon Bedrock.</li>
</ul>



<p class="wp-block-paragraph">AWS Step Functions is worth reaching for once you have more than three stages, because Textract jobs are long-running and Lambda timeouts are not a retry strategy. That is the whole architecture. You can stand it up in a fortnight. Then you spend six months on everything below.</p>



<h2 class="wp-block-heading">Boundary one: isolation, and why metadata filtering is not optional</h2>



<p class="wp-block-paragraph">This is the one from the demo, and it is the one that ends engagements.</p>



<p class="wp-block-paragraph">A vector index has no native concept of a matter, a client, or an ethical wall. If you pool every document into one index and rely on the prompt to keep things separate, you have built a system where a well-phrased question can pull privileged material across a wall. Prompt instructions are not an access control. They are a suggestion that usually works.</p>



<p class="wp-block-paragraph">There are two shapes that do work, and the choice between them is a real trade-off rather than a best practice.</p>



<h3 class="wp-block-heading">Pooled index with server-side filtering</h3>



<p class="wp-block-paragraph">Every chunk carries a <code>matter_id</code> and a <code>client_id</code> as metadata. Every retrieval call attaches a filter. With an S3 data source in Bedrock Knowledge Bases, the metadata comes from a sidecar JSON file that sits next to the document and carries a <code>metadataAttributes</code> object.</p>



<pre class="wp-block-code"><code>{
  "metadataAttributes": {
    "matter_id": "M-4417",
    "client_id": "C-108",
    "doc_type": "engagement_letter",
    "privileged": true
  }
}</code></pre>



<p class="wp-block-paragraph">The retrieval call then narrows the search before the model ever sees a chunk:</p>



<pre class="wp-block-code"><code>"retrievalConfiguration": {
  "vectorSearchConfiguration": {
    "filter": {
      "equals": { "key": "matter_id", "value": "M-4417" }
    }
  }
}</code></pre>



<p class="wp-block-paragraph">Two rules make this safe, and both are the kind of thing that gets skipped under deadline. First, the filter value is derived on the server from the authenticated caller&#8217;s claims, never accepted from the client. If the browser can send you a <code>matter_id</code>, the browser can send you a different one. Second, the filter is applied by code that no feature request can bypass. A helper that builds every retrieval request, and a code review rule that no other path may call the retrieve API directly.</p>



<p class="wp-block-paragraph">The trap here is ordering. You cannot filter on an attribute the index does not have. Add <code>matter_id</code> after ingestion and every chunk already in the index is invisible to that filter, which means it either returns for everyone or for no one, depending on how your filter is written. Design the metadata schema before the first ingestion run, not after the first demo.</p>



<h3 class="wp-block-heading">Separate collection per client</h3>



<p class="wp-block-paragraph">The heavier option. A separate OpenSearch Serverless collection per client, which buys you a separate AWS KMS key per client, separate index settings, and a failure mode where a bug in the filtering logic cannot reach across clients at all.</p>



<p class="wp-block-paragraph">It costs you sprawl. Every collection is a resource to provision, monitor, patch policy on, and eventually delete. Ingestion jobs run per data source per knowledge base, so freshness guarantees fragment. For a firm with twelve institutional clients this is fine. For a platform onboarding a hundred small clients it becomes the main operational burden of the product.</p>



<p class="wp-block-paragraph">My default is pooled with strict server-side filtering, and per-client collections only where the client&#8217;s own contract demands a dedicated encryption key. If you cannot say out loud which of those two you are running, you are running neither properly.</p>



<h2 class="wp-block-heading">Boundary two: extraction, and the confidence score everyone ignores</h2>



<p class="wp-block-paragraph">Textract returns a confidence score between 0 and 100 with each prediction. AWS is explicit in its own best-practice guidance that applications sensitive to detection errors should enforce a minimum threshold and route anything below it for human scrutiny, and that the right threshold depends entirely on how the output gets used.</p>



<p class="wp-block-paragraph">Legal documents sit at the harsh end of that. A misread date on an archival scan is a curiosity. A misread date on a limitation period is a problem with a name on it.</p>



<p class="wp-block-paragraph">Practical notes from this layer:</p>



<ul class="wp-block-list">
<li>Use the asynchronous operations for anything multipage. <code>StartDocumentAnalysis</code> and <code>GetDocumentAnalysis</code> handle long PDFs and TIFFs; the synchronous <code>AnalyzeDocument</code> path is for single pages and will fight you on a 400-page bundle.</li>



<li>Textract&#8217;s Queries feature earns its place on structured instruments. Instead of extracting everything and grepping, you ask the document a direct question and get a scoped answer with its own confidence.</li>



<li>Keep the page number and bounding box for every chunk. When a lawyer asks where a sentence came from, &#8220;page 14, second column&#8221; is an answer. &#8220;It&#8217;s in the corpus somewhere&#8221; is not, and the platform loses credibility the first time you say it.</li>



<li>Amazon Augmented AI wires low-confidence predictions into a human review queue with a private workforce, so the reviewers are your people rather than an anonymous pool. For privileged material that distinction is the entire point.</li>
</ul>



<p class="wp-block-paragraph">What I would skip early: building a custom entity recognition model. Comprehend supports custom entity recognition and it is genuinely useful, but you need labelled data you do not have yet, and a foundation model with a decent prompt gets you far enough to find out which entities the users actually care about. Train the custom model once the requirement has stopped moving.</p>



<h2 class="wp-block-heading">Boundary three: retention, where deletion is a distributed problem</h2>



<p class="wp-block-paragraph">Here is the failure that will not show up in any test suite you write.</p>



<p class="wp-block-paragraph">A client asks for a document to be removed. Someone deletes the S3 object. The ticket closes. The document is still fully searchable, because its chunks are sitting in the vector index and will stay there until the data source is re-synced. Meanwhile the Textract JSON output is in a second bucket, the extracted text may be in a database, and if model invocation logging is on, whole passages of the document are sitting in CloudWatch Logs or another S3 bucket entirely.</p>



<p class="wp-block-paragraph">One delete, at least four places holding a copy. Write the deletion path as a real workflow with an assertion at the end, not as a single API call:</p>



<ol class="wp-block-list">
<li>Delete the source object and, if versioning is on, its versions.</li>



<li>Delete the derived artefacts: Textract JSON, normalised text, any thumbnails or page images.</li>



<li>Trigger an ingestion job so the knowledge base drops the orphaned chunks.</li>



<li>Confirm removal by querying the index for the document identifier and expecting nothing back.</li>



<li>Deal with the logs, which means either a retention policy short enough to make the problem expire or a deliberate decision, written down, that logs are out of scope.</li>
</ol>



<p class="wp-block-paragraph">Step four is the one people leave out, and it is the only step that actually proves anything.</p>



<h3 class="wp-block-heading">The opposite problem: things that must not be deleted</h3>



<p class="wp-block-paragraph">Legal work has the reverse requirement too. When a matter goes into litigation hold, the documents need to survive an administrator with delete permissions and a bad afternoon.</p>



<p class="wp-block-paragraph">S3 Object Lock is the mechanism, and it has two independent controls that people routinely conflate. A retention period protects an object version until a fixed date, in either governance mode, which privileged users can override, or compliance mode, which nobody can override and where the period cannot be shortened. A legal hold has no date at all. It stays until someone with <code>s3:PutObjectLegalHold</code> explicitly removes it. The two can be active at once, and while either is active the object version cannot be deleted or overwritten.</p>



<pre class="wp-block-code"><code># Place an indefinite hold on a specific object version
aws s3api put-object-legal-hold 
  --bucket matters-archive 
  --key M-4417/exhibit-c.pdf 
  --version-id 3sL7f2Qz9pXvB1kR 
  --legal-hold Status=ON</code></pre>



<p class="wp-block-paragraph">Two things bite here. Object Lock requires versioning and turns it on automatically, and once enabled on a bucket you cannot turn it off or suspend versioning again. And compliance mode is genuinely permanent: if you set a seven-year retention when you meant seven days, that is the answer, for everyone, including the account root. Test in governance mode first. The bypass path exists and is deliberately awkward, requiring both the <code>s3:BypassGovernanceRetention</code> permission and an explicit <code>x-amz-bypass-governance-retention:true</code> header on the request, which is exactly the level of friction you want on that operation.</p>



<p class="wp-block-paragraph">Design the hold model before you design the deletion model, because holds win. A deletion request that collides with an active hold is a legal question, not an engineering one, and the platform&#8217;s job is to surface the collision clearly rather than resolve it silently.</p>



<h2 class="wp-block-heading">Boundary four: disclosure, or what leaves the account</h2>



<p class="wp-block-paragraph">This is the boundary the client&#8217;s general counsel will ask about, usually in writing, usually before signature. Three separate controls, and they are not interchangeable.</p>



<h3 class="wp-block-heading">The AI services opt-out policy</h3>



<p class="wp-block-paragraph">AWS Organizations has a governance policy type that opts your accounts out of having content processed by certain AI services stored and used for service improvement. The list includes Amazon Textract and Amazon Comprehend, which are precisely the two doing the reading in this architecture. The policy type has to be enabled at the organisation root before you can attach anything:</p>



<pre class="wp-block-code"><code>aws organizations enable-policy-type 
  --root-id r-example 
  --policy-type AISERVICES_OPT_OUT_POLICY</code></pre>



<p class="wp-block-paragraph">Read the AWS documentation on this one carefully rather than taking my summary as gospel, because there is a caveat in it that matters: the services may still need to store your data operationally even when you have opted out of it being used for improvement. Opting out is not the same as the data never existing outside your account. Say that plainly in the client conversation. It is a much better position than being asked about it later.</p>



<h3 class="wp-block-heading">Bedrock&#8217;s data position, and the log that undoes it</h3>



<p class="wp-block-paragraph">Amazon Bedrock&#8217;s published position is that inputs and outputs are not shared with third-party model providers and are not used to train the base models, and that fine-tuning operates on a private copy. That is a strong starting point for privileged content and it is the main reason Bedrock rather than a direct provider API shows up in these builds.</p>



<p class="wp-block-paragraph">Then there is model invocation logging. It captures full prompt and response payloads to CloudWatch Logs or S3, and you want it on, because without it you cannot reconstruct what the system told someone. But understand what you have just built: a log that contains privileged document text, at a per-region setting, in a destination that probably has looser access controls than the document store it came from. Encrypt it with the same KMS key discipline, restrict it harder than you think you need to, and set a retention period on purpose.</p>



<p class="wp-block-paragraph">Worth separating in your head: CloudTrail records that an API call happened and who made it. Model invocation logging records what was in it. You need both, for different questions, and only one of them contains client confidences.</p>



<h3 class="wp-block-heading">Network path</h3>



<p class="wp-block-paragraph">Interface VPC endpoints via AWS PrivateLink keep traffic to Bedrock, Textract and the rest on the AWS network rather than out through an internet gateway or NAT. Add a gateway endpoint for S3 while you are there.</p>



<p class="wp-block-paragraph">Endpoint policies are the part that gets forgotten. An endpoint without a policy is a private path to the whole service, including into accounts that are not yours. Scope it to the operations and resources you actually use, and pair it with an S3 bucket policy that rejects requests not arriving through your endpoint. The first protects what leaves; the second protects what can be reached.</p>



<h3 class="wp-block-heading">The parts that are not on the architecture diagram</h3>



<p class="wp-block-paragraph">Two of these have caught people out badly, and neither is an AWS control.</p>



<p class="wp-block-paragraph">The first is the development environment. You do not want real client documents on a laptop or in a scratch account, so build the parsing and chunking code against synthetic documents on a small separate box. A cheap VPS from a provider like Contabo or InterServer is fine for this, and the separation is worth more than the convenience of iterating in the production account.</p>



<p class="wp-block-paragraph">The second is the reviewers. If your human review loop involves contractors working remotely, their network path and their disks are part of your boundary whether or not you drew them. A managed VPN such as NordVPN or Surfshark handles the network side, and when a matter closes and local copies have to go, a dedicated erasure tool from something like O&amp;O Software does what dragging a folder to the bin does not. Unglamorous, and it is the layer that gets audited.</p>



<h2 class="wp-block-heading">Boundary five: evidence, because someone will ask</h2>



<p class="wp-block-paragraph">At some point the question stops being &#8220;does it work&#8221; and becomes &#8220;who saw what, and when&#8221;. If you cannot answer that from stored records, the platform is not defensible regardless of how good the answers are.</p>



<ul class="wp-block-list">
<li>Turn on CloudTrail data events for the document buckets. They are off by default, they are billed separately from management events, and they are the only way to see individual object-level reads.</li>



<li>Log the retrieval, not just the generation. Store which chunks came back and which filter was applied. When somebody asks whether the wall held on a specific query, this record is the answer and there is no reconstructing it later.</li>



<li>Carry the end user&#8217;s identity through to the audit record. A Lambda execution role in the logs tells you the platform did something. It does not tell you who asked.</li>



<li>Alarm on the filter, not just on errors. A retrieval that ran without a matter filter should page someone, even though it returned HTTP 200 and a perfectly good answer.</li>
</ul>



<p class="wp-block-paragraph">That last one is the closest thing to a single takeaway in this post. The dangerous events in this architecture are successful ones.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Troubleshooting the things that will go wrong</h2>



<h3 class="wp-block-heading">Retrieval returns nothing after you add a filter</h3>



<p class="wp-block-paragraph">Nearly always a metadata problem rather than a query problem. Either the sidecar metadata file was missing at ingestion, so the chunks carry no attribute to match, or the attribute was added after the chunks were indexed. Check whether the attribute exists on a known chunk before you touch the query. If it does not, you need a re-sync, not a different filter expression.</p>



<h3 class="wp-block-heading">Textract returns text but the layout is scrambled</h3>



<p class="wp-block-paragraph">Common on two-column pleadings and on documents with headers and footers on every page. Raw text detection reads in a reading order that is not always yours. Use the layout and table features rather than plain detection, and hold onto the geometry so you can reassemble columns yourself if the default order is wrong for that document class.</p>



<h3 class="wp-block-heading">The model cites a document that does not exist</h3>



<p class="wp-block-paragraph">Usually the citation is being generated rather than passed through. If document names are being written by the model instead of read from the retrieval response, it will invent plausible ones. Build citations from the retrieval result metadata in your application code and never from the generated text. Bedrock Guardrails helps with contextual grounding, but the structural fix is not asking the model to produce the reference in the first place.</p>



<h3 class="wp-block-heading">Ingestion is slow and nobody knows where</h3>



<p class="wp-block-paragraph">Instrument per stage before you optimise anything. In most of these pipelines the time is in Textract for large scanned bundles, and the fix is parallelism at the document level rather than tuning the extraction itself. Step Functions with a distributed map over documents gets you there without a queue you have to babysit.</p>



<h3 class="wp-block-heading">Costs climb faster than volume</h3>



<p class="wp-block-paragraph">Look at re-processing first. A pipeline that re-extracts a document every time anything downstream changes will quietly multiply your Textract bill against a static corpus. Key the extraction cache on the object&#8217;s content hash, not its path, and make re-ingestion an explicit action. The billing mechanisms differ per service, so read the current pricing pages rather than trusting a number from a blog post.</p>



<h2 class="wp-block-heading">Common mistakes</h2>



<ul class="wp-block-list">
<li><strong>Treating the system prompt as an access control.</strong> It is a formatting instruction that happens to look like a rule.</li>



<li><strong>Taking the tenant identifier from the request body.</strong> If the client can name the matter, the client can name someone else&#8217;s.</li>



<li><strong>Designing the metadata schema after the first ingestion run.</strong> Retrofitting a filterable attribute means re-indexing the whole corpus.</li>



<li><strong>Assuming a deleted S3 object is gone from the platform.</strong> The chunks, the extraction output and the invocation logs all outlive it.</li>



<li><strong>Setting compliance-mode retention without testing in governance mode.</strong> There is no support ticket that fixes a seven-year mistake.</li>



<li><strong>Enabling model invocation logging without treating the log as privileged.</strong> You have just made a second copy of the documents in a less protected place.</li>



<li><strong>Building VPC endpoints and leaving the default policy on them.</strong> A private path with no policy is still a path.</li>
</ul>



<h2 class="wp-block-heading">Best practices worth the effort</h2>



<ul class="wp-block-list">
<li>One retrieval helper, server-side filter injection, and a review rule that no other code path calls the retrieve API directly.</li>



<li>Metadata schema agreed and frozen before the first production ingestion. Include the fields you might filter on later, even if unused today.</li>



<li>Page-level provenance on every chunk, surfaced in the interface. It builds trust faster than any accuracy improvement.</li>



<li>A confidence threshold chosen per document class, with a human review path for everything under it.</li>



<li>Deletion implemented as a verified workflow across every store, ending in an assertion that the content is actually unretrievable.</li>



<li>Object Lock legal holds for matters under litigation hold, tested in governance mode before anything runs in compliance mode.</li>



<li>Customer-managed KMS keys across S3, the vector store and the logs, so key access is a control you can actually revoke.</li>



<li>An alarm on unfiltered retrievals, because the failure mode is a success response.</li>
</ul>



<h2 class="wp-block-heading">Frequently asked questions</h2>



<h3 class="wp-block-heading">Is a legal document intelligence platform on AWS safe for privileged material?</h3>



<p class="wp-block-paragraph">The services support it. Bedrock&#8217;s stated position is that inputs and outputs are not shared with model providers or used to train base models, PrivateLink keeps traffic off the public internet, and KMS covers encryption at rest. What determines safety is your own configuration: tenant isolation, log handling and deletion behaviour. The platform is exactly as confidential as its weakest copy of the text.</p>



<h3 class="wp-block-heading">Should I use Bedrock Knowledge Bases or build retrieval myself?</h3>



<p class="wp-block-paragraph">Start with Knowledge Bases. It handles chunking, embedding, sync and metadata filtering, and those are weeks of work with no differentiation in them. Build your own when you need retrieval behaviour it does not expose, such as unusual re-ranking or per-tenant chunking strategies. Going custom on day one usually means reimplementing the managed service badly.</p>



<h3 class="wp-block-heading">How do I stop one client&#8217;s documents surfacing in another client&#8217;s answers?</h3>



<p class="wp-block-paragraph">Tag every chunk with a tenant identifier at ingestion, and inject the matching filter server-side on every retrieval from the authenticated caller&#8217;s identity. Never accept the identifier from the client. If the contract requires per-client encryption keys, use a separate collection per client instead, accepting the extra operational load.</p>



<h3 class="wp-block-heading">Does deleting a document from S3 remove it from the search index?</h3>



<p class="wp-block-paragraph">No. The chunks stay in the vector store until an ingestion job runs and reconciles the data source. Until then the document is deleted and still fully searchable, which is the worst combination available. Always finish a deletion by querying the index and confirming nothing comes back.</p>



<h3 class="wp-block-heading">What is the difference between an S3 legal hold and a retention period?</h3>



<p class="wp-block-paragraph">A retention period runs until a fixed date and comes in governance mode, which privileged users can override, or compliance mode, which nobody can. A legal hold has no end date and stays until someone explicitly removes it. They are independent, they can both be active on the same object version, and while either is active the object cannot be deleted or overwritten.</p>



<h3 class="wp-block-heading">Do I need human review, or is the model accurate enough?</h3>



<p class="wp-block-paragraph">Accuracy is the wrong frame. The question is what happens to a wrong answer downstream. Where output feeds a decision with consequences, you want a confidence threshold and a review queue, and Amazon Augmented AI with a private workforce gives you that without building the review tooling yourself. Where output is a search aid a person will verify anyway, review is friction you do not need.</p>



<h3 class="wp-block-heading">Can I run this in a single AWS account?</h3>



<p class="wp-block-paragraph">Technically yes, and for a pilot it is reasonable. It gets uncomfortable once you have production documents and a development environment in the same place, because an IAM mistake has nowhere to stop. Separate accounts under Organizations also gives you the policy layer, including the AI services opt-out policy, which only exists at the organisation level.</p>



<h2 class="wp-block-heading">The one thing to take away</h2>



<p class="wp-block-paragraph">A legal document intelligence platform on AWS is not hard to build. Textract reads, Bedrock reasons, OpenSearch retrieves, and the tutorial version works on the first afternoon.</p>



<p class="wp-block-paragraph">What is hard is that its failures are silent and well-formed. The cross-matter citation, the deleted document that still answers questions, the privileged passage sitting in a log bucket, the retrieval that ran without a filter and returned a beautiful paragraph. None of these throw an error. All of them are the kind of thing that ends a client relationship.</p>



<p class="wp-block-paragraph">So build the boundaries first and the features second, and make sure every one of them is something you can prove from a stored record rather than something you believe about the code. If you can only take one habit from this: alarm on the successful requests that should not have been possible.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Need a second pair of eyes on your document pipeline?</h2>



<p class="wp-block-paragraph">I work on AWS document processing and retrieval systems where the confidentiality requirements are real. Typical things I get called in for:</p>



<ul class="wp-block-list">
<li>Reviewing tenant isolation in an existing RAG setup and finding the paths where the filter can be bypassed.</li>



<li>Designing the metadata and chunking schema before ingestion, so filtering works without a re-index later.</li>



<li>Building Textract pipelines with confidence thresholds, human review routing and page-level provenance.</li>



<li>Implementing verified deletion across S3, derived artefacts, the vector index and logs, with a proof step at the end.</li>



<li>Setting up retention and legal hold with S3 Object Lock, including the governance-mode rehearsal before compliance mode.</li>



<li>Locking down the network and logging path: VPC endpoints with real policies, KMS key separation, and audit records that name the actual user.</li>
</ul>



<p class="wp-block-paragraph">If you have an architecture diagram, a Step Functions definition or a retrieval request you are unsure about, send it over and I will tell you what I would change and why.</p>



<div class="wp-block-buttons is-layout-flex wp-block-buttons-is-layout-flex">
<div class="wp-block-button"><a class="wp-block-button__link wp-element-button" href="https://www.upwork.com/freelancers/~01f15a912ad84a6620" target="_blank" rel="noreferrer noopener">Work with me on Upwork</a></div>
</div>
<p>The post <a href="https://john-nessime.com/blog/cloud-computing/legal-document-intelligence-aws/">Legal Document Intelligence on AWS: The Five Boundaries That Have to Hold</a> appeared first on <a href="https://john-nessime.com/blog">John Nessime</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>As-Planned vs As-Built Analysis: Building a Platform That Survives Cross-Examination</title>
		<link>https://john-nessime.com/blog/technical-guides/as-planned-vs-as-built-analysis-platform/</link>
		
		<dc:creator><![CDATA[John Nessime]]></dc:creator>
		<pubDate>Mon, 24 Aug 2026 13:00:00 +0000</pubDate>
				<category><![CDATA[Case Studies]]></category>
		<category><![CDATA[Technical Guides]]></category>
		<category><![CDATA[As-Built Schedule]]></category>
		<category><![CDATA[Audit Logging]]></category>
		<category><![CDATA[Capital Projects]]></category>
		<category><![CDATA[Construction Technology]]></category>
		<category><![CDATA[Data Engineering]]></category>
		<category><![CDATA[Data Lineage]]></category>
		<category><![CDATA[Delay Analysis]]></category>
		<category><![CDATA[Document Processing]]></category>
		<category><![CDATA[Extension of Time]]></category>
		<category><![CDATA[Forensic Schedule Analysis]]></category>
		<category><![CDATA[Legal Tech]]></category>
		<category><![CDATA[OCR]]></category>
		<category><![CDATA[PostgreSQL]]></category>
		<category><![CDATA[Primavera P6]]></category>
		<category><![CDATA[Project Controls]]></category>
		<category><![CDATA[Python]]></category>
		<category><![CDATA[Schema Design]]></category>
		<category><![CDATA[XER Files]]></category>
		<guid isPermaLink="false">https://john-nessime.com/blog/?p=274</guid>

					<description><![CDATA[<p>Most as-planned vs as-built analysis compares the baseline to the last P6 update and calls the result an as-built. It isn't one. This is how to build the data layer underneath a delay analysis: versioned XER ingestion, activity identity across renumbering, a first-appearance table that proves when every actual date entered the record, calendar-safe float, and record linking that proposes candidates instead of asserting cause.</p>
<p>The post <a href="https://john-nessime.com/blog/technical-guides/as-planned-vs-as-built-analysis-platform/">As-Planned vs As-Built Analysis: Building a Platform That Survives Cross-Examination</a> appeared first on <a href="https://john-nessime.com/blog">John Nessime</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">The question that kills a delay analysis is never about methodology. It is five words from the other side&#8217;s expert: &#8220;which update did that come from?&#8221;</p>



<p class="wp-block-paragraph">Picture the experts&#8217; meeting. Your exhibit shows activity C-2140, structural steel erection to level four, finishing well behind where the baseline put it. He has the same XER files you do. He opens update 19 and points out that C-2140 carries no actual finish. Same in update 20. In update 21 an actual finish appears, and the date on it sits three weeks in the past, inside the window you already closed out.</p>



<p class="wp-block-paragraph">So which file is the record? And if one activity was back-filled a month after the fact, what does that say about the other four thousand?</p>



<p class="wp-block-paragraph">That is the failure an <strong>as-planned vs as-built analysis</strong> has to survive, and no amount of care inside the Excel workbook fixes it. It is a data problem before it is a scheduling problem. This post covers the platform underneath the analysis: how to ingest a long series of Primavera P6 updates, how to prove where every actual date came from, and how to produce output a tribunal can audit instead of having to trust.</p>



<h2 class="wp-block-heading">The as-built is not a file, and treating it as one is the invisible failure</h2>



<p class="wp-block-paragraph">Almost every quick as-planned vs as-built comparison does the same thing: take the approved baseline, take the last update in the sequence, join on activity ID, subtract the dates, sort by variance. It produces a clean-looking table in an afternoon.</p>



<p class="wp-block-paragraph">It is also the weakest possible construction of an as-built, for three reasons that never show up in the output.</p>



<ul class="wp-block-list">
<li><strong>A P6 update is a forecast document.</strong> It was produced to tell the employer when the job would finish, not to record what happened. Actual dates are a by-product of that exercise, entered by a planner working to a monthly deadline.</li>

<li><strong>The last update has been edited the most.</strong> By the end of a long job it has usually been re-baselined, partially renumbered, had calendars swapped, had activities merged or deleted. Every one of those edits is invisible in a two-file comparison.</li>

<li><strong>Actual dates get back-filled.</strong> A date typed in month 21 describing month 19 is not contemporaneous evidence in the same sense as a date typed in month 19. It might still be right. It is a different quality of record, and the difference is exactly what gets probed in cross-examination.</li>
</ul>



<p class="wp-block-paragraph">AACE International&#8217;s Recommended Practice 29R-03 places the as-planned versus as-built family in its observational, static group: you compare a planned network against an as-built without inserting or removing delay events. The SCL Delay and Disruption Protocol, second edition, is blunter about what makes any of it work, pushing hard on agreeing a record-keeping regime up front and never overwriting a programme version. Both point at the same requirement. You need the whole sequence of updates, intact, and you need to know which one told you what.</p>



<p class="wp-block-paragraph">The platform&#8217;s whole job is to make that sequence queryable.</p>



<h2 class="wp-block-heading">A worked example: what thirty-four updates actually look like</h2>



<p class="wp-block-paragraph">Take a job shaped like this. Three years of monthly updates, call it thirty-four XER exports. Around four thousand activities in the current programme. Three baseline revisions, two agreed and one submitted but never accepted. Somewhere near month fourteen the contractor moved most trades from a five-day to a six-day calendar to recover. The numbers are illustrative rather than from any particular project, but the shape is ordinary.</p>



<p class="wp-block-paragraph">Run the naive comparison and you get a variance table. Run the sequence properly and four separate questions fall out, none of which the variance table can answer:</p>



<ol class="wp-block-list">
<li>Which activities existed in the baseline, disappeared mid-job, and came back under a different activity ID?</li>

<li>For each actual date in the final update, in which update did that value first appear, and how far behind the event was it?</li>

<li>When float moved, was it because work moved, or because someone changed a calendar or a duration?</li>

<li>What contemporaneous record sits underneath each actual date, and does it agree?</li>
</ol>



<p class="wp-block-paragraph">Those four questions are the platform. Everything below is how to answer each one without hand-checking four thousand rows thirty-four times.</p>



<h2 class="wp-block-heading">Failure family one: activity identity across versions</h2>



<p class="wp-block-paragraph">An XER file is tab-delimited text. It opens in a text editor. The structure is a header line beginning <code>ERMHDR</code>, then a run of tables, each introduced by <code>%T</code> with the table name, a <code>%F</code> line carrying the field names, and <code>%R</code> lines carrying rows, closed by <code>%E</code>. The tables you care about first are <code>PROJECT</code>, <code>TASK</code>, <code>TASKPRED</code> and <code>CALENDAR</code>.</p>



<pre class="wp-block-code"><code>ERMHDR	19.12	2026-01-15	Project	admin	Admin User	dbxDatabaseNoName	Project Management	USD
%T	TASK
%F	task_id	proj_id	wbs_id	clndr_id	task_code	task_name	status_code	act_start_date	act_end_date	target_start_date	target_end_date	total_float_hr_cnt
%R	1001	100	200	1	C-2140	Structural Steel Erection L4	TK_Complete	...
%E</code></pre>



<p class="wp-block-paragraph">Here is the trap. <code>task_id</code> is the primary key of the P6 database the file was exported from. It is not a stable identifier for an activity across time. If the project was ever copied between databases, restored, or exported from a different P6 instance, the same physical activity can carry a different <code>task_id</code>. Worse, two XERs from different sources can reuse the same <code>task_id</code> for completely unrelated activities.</p>



<p class="wp-block-paragraph"><code>task_code</code> is the Activity ID a human sees, and it is stable right up until someone renumbers. Neither field alone gives you identity. So do not pick one. Build an identity resolution layer and keep the evidence for every decision it makes:</p>



<ul class="wp-block-list">
<li>Use <code>(proj_id, task_code)</code> as the working join key, scoped to a single source database.</li>

<li>Store <code>task_id</code>, <code>task_name</code> and <code>wbs_id</code> on every snapshot row so a rename or a WBS move is detectable after the fact.</li>

<li>Emit an explicit <em>unresolved</em> record when an activity vanishes or appears mid-sequence. Do not guess.</li>
</ul>



<p class="wp-block-paragraph">That last point is the one people get wrong. When C-2140 disappears in update 22 and an activity with the same name but the ID C-2140A appears in update 23, there are three candidate explanations: deletion and replacement, renumbering, or a split. The platform cannot know which. A tool that silently picks one is worse than a tool that flags it, because the analyst never learns there was a question. Surface it, hand it to a human, record the decision and the reason.</p>



<h2 class="wp-block-heading">Failure family two: date provenance, and the table that makes the whole thing worth building</h2>



<p class="wp-block-paragraph">If you build one thing, build this. For every activity and every date field, record the first update in which that value appeared, and every subsequent change to it.</p>



<p class="wp-block-paragraph">Assume a snapshot table with one row per activity per update, appended and never modified. A window function does the work:</p>



<pre class="wp-block-code"><code>-- First appearance of each distinct actual finish value,
-- and how far the reported date lags the update that reported it.
WITH changes AS (
  SELECT
    task_code,
    version_no,
    data_date,
    act_end_date,
    LAG(act_end_date) OVER (
      PARTITION BY task_code ORDER BY version_no
    ) AS prev_act_end
  FROM activity_snapshot
)
SELECT
  task_code,
  version_no          AS first_reported_in_update,
  act_end_date        AS reported_actual_finish,
  data_date           AS update_data_date,
  data_date::date - act_end_date::date AS reporting_lag_days
FROM changes
WHERE act_end_date IS NOT NULL
  AND (prev_act_end IS NULL OR prev_act_end &lt;&gt; act_end_date)
ORDER BY reporting_lag_days DESC;</code></pre>



<p class="wp-block-paragraph">Three things arrive at once. You can answer the question that opened this post, for any activity, in a second: C-2140&#8217;s actual finish first appeared in update 21, carrying a date twenty-three days behind that update&#8217;s data date.</p>



<p class="wp-block-paragraph">The distribution of that lag column is then a finding in its own right. If most activities report within days and one subcontractor&#8217;s work consistently reports a month late, that is not a rounding artefact, it is a records issue you can evidence. And any actual date that changes after it was first stated deserves a hard look. Actuals are supposed to lock. A value that moves has either been corrected, which should have a paper trail, or overwritten, which should not have happened.</p>



<p class="wp-block-paragraph">P6 ships a Schedule Comparison tool, the one that used to be called Claim Digger and now lives inside Visualizer. It does a field-by-field diff and dumps a large HTML table, which is genuinely useful for checking one submitted update against the last. That is what it was built for. It was not built to reason across thirty-four files at once, and pushing its output into Excel does not change that. Deltek&#8217;s Acumen Fuse and Steelray&#8217;s Delay Analyzer go further, with what the vendors describe as half-step analysis, separating progress from scope revisions across successive versions. If your instruction allows the licence and the timetable allows the learning curve, look at them seriously before writing any code. Neither is priced publicly, both target enterprise project controls teams, and neither does the records-linking work in the last section for you.</p>



<h2 class="wp-block-heading">Failure family three: calendars, and float that quietly means nothing</h2>



<p class="wp-block-paragraph">Durations and float in an XER are stored in hours. <code>total_float_hr_cnt</code> holds hours, not days. A value of 40 is five days only if the activity runs an eight-hour calendar. Divide everything by eight and a job with mixed shift patterns produces float numbers that are confidently wrong.</p>



<p class="wp-block-paragraph">The calendar itself lives in the <code>CALENDAR</code> table, in a <code>clndr_data</code> field holding a proprietary parenthesised blob describing work weeks, exceptions and holidays. It is the hardest part of the format to parse correctly, and getting it wrong produces durations and dates that look plausible and are not.</p>



<p class="wp-block-paragraph">Two rules follow, and both are about the read path, not the maths.</p>



<ul class="wp-block-list">
<li>Store float and durations in hours exactly as exported. Convert at presentation time, using that activity&#8217;s own calendar, and print the calendar name next to the number. If you cannot resolve the calendar, print hours and say so.</li>

<li>Track calendar assignment as a versioned attribute, the same as any date. A change from a five-day to a six-day calendar changes every float figure downstream of it, with no change to logic and no change to progress. An activity can go from critical to non-critical because of an administrative edit, and a variance table will attribute that shift to the works.</li>
</ul>



<p class="wp-block-paragraph">This is also why an as-built critical path is not simply the chain with zero float in the last update. Retrospective longest path and as-planned versus as-built windows analysis exist as separate methods precisely because the float figures in a progressed programme are a function of how the programme was maintained.</p>



<h2 class="wp-block-heading">Failure family four: separating movement from editing</h2>



<p class="wp-block-paragraph">Every delta between two consecutive updates falls into one of a small number of categories, and collapsing them into a single &#8220;variance&#8221; number is where analyses lose their defensibility. Classify each change explicitly:</p>



<ul class="wp-block-list">
<li><strong>Progress:</strong> an actual start or finish appeared, or remaining duration reduced consistent with work done.</li>

<li><strong>Duration change:</strong> original or remaining duration edited without corresponding progress.</li>

<li><strong>Logic change:</strong> a relationship added, removed, retyped, or its lag altered.</li>

<li><strong>Calendar change:</strong> activity reassigned, or the calendar definition itself edited.</li>

<li><strong>Constraint change:</strong> a date constraint applied, moved or removed.</li>

<li><strong>Scope change:</strong> activity added or deleted.</li>
</ul>



<p class="wp-block-paragraph">Six buckets, one row per change, with the before value, the after value and the version pair. Once that table exists, a lot of the argument stops being an argument. &#8220;The completion date moved out eighteen days in window 12&#8221; becomes &#8220;of which fourteen days sit with progress on the critical chain and four with two logic edits applied in the same submission.&#8221; The second version is the one that gets tested and holds.</p>



<p class="wp-block-paragraph">Logic changes deserve their own view. Pull <code>TASKPRED</code> for every version, key each relationship on the predecessor and successor activity codes plus the relationship type, and diff the sets. Relationships that appear late in a job, on activities that are already in progress, are worth reading one by one.</p>



<h2 class="wp-block-heading">Failure family five: joining the schedule to the records</h2>



<p class="wp-block-paragraph">The schedule tells you when. It never tells you why. Cause comes from daily reports, site instructions, RFIs, correspondence, minutes, inspection records, weather data, and delivery notes. That material arrives as scanned PDFs and email exports, and it does not carry activity IDs. Nobody has ever written &#8220;C-2140&#8221; on a site diary.</p>



<p class="wp-block-paragraph">So the join key is not the activity ID. It is a fuzzy composite of date, location or area, and trade or discipline. The practical approach:</p>



<ol class="wp-block-list">
<li>Extract structured fields from each record: document date, report date, area, trade, headcount if present, and the narrative text. Text extraction plus a layout-aware OCR pass handles most of it. Handwriting on older site diaries will not fully automate, and you should budget for that.</li>

<li>Normalise area and trade against the schedule&#8217;s own coding. Activity codes and WBS paths are usually the best available crosswalk, since they encode area and discipline already.</li>

<li>Generate <em>candidate</em> links: for each activity&#8217;s as-built window, every record whose date falls inside it and whose area or trade matches.</li>

<li>Have a human confirm, reject or annotate each candidate. Store who decided and when.</li>
</ol>



<p class="wp-block-paragraph">Step four is not optional and it is not a nicety. A machine-proposed link is a search result. Causation is an opinion, and the expert has to own it. Build the tool so it accelerates the search and refuses to state the conclusion, and you get the productivity without handing opposing counsel an argument about black boxes.</p>



<h2 class="wp-block-heading">The data model, in about six tables</h2>



<p class="wp-block-paragraph">The grain that makes everything else easy: one immutable row per activity per schedule version. Nothing is ever updated in place.</p>



<pre class="wp-block-code"><code>CREATE TABLE schedule_version (
  version_id     bigserial PRIMARY KEY,
  version_no     integer     NOT NULL,   -- ordinal in the sequence
  source_file    text        NOT NULL,
  file_sha256    char(64)    NOT NULL,   -- proves the file was not altered
  p6_export_ts   timestamptz,            -- from the ERMHDR line
  data_date      timestamptz NOT NULL,
  ingested_at    timestamptz NOT NULL DEFAULT now(),
  notes          text
);

CREATE TABLE activity_snapshot (
  version_id           bigint  NOT NULL REFERENCES schedule_version,
  task_code            text    NOT NULL,
  task_id              bigint  NOT NULL,
  task_name            text,
  wbs_path             text,
  clndr_id             bigint,
  status_code          text,
  target_start_date    timestamptz,
  target_end_date      timestamptz,
  act_start_date       timestamptz,
  act_end_date         timestamptz,
  total_float_hr_cnt   numeric,
  PRIMARY KEY (version_id, task_code)
);</code></pre>



<p class="wp-block-paragraph">Add <code>relationship_snapshot</code> keyed on <code>(version_id, pred_task_code, succ_task_code, rel_type)</code>, a <code>calendar_snapshot</code> holding the raw <code>clndr_data</code> alongside your parsed working pattern, a <code>record_document</code> table for the contemporaneous material, and a <code>record_link</code> table carrying the human decision on each candidate link.</p>



<p class="wp-block-paragraph">Thirty-four updates at four thousand activities is under a hundred and fifty thousand rows. This is not a big-data problem and it should not be built like one. PostgreSQL on a single machine, with the parsing done in Python, will run every query in this post fast enough that you stop thinking about it.</p>



<p class="wp-block-paragraph">On the parsing side, an XER is simple enough to read with the standard library, and there are open-source parsers such as PyP6Xer if you would rather not write the tokeniser yourself. Whichever route you take, keep the raw file. Hash it on ingest, store the hash in <code>schedule_version</code>, and never write to the original. Chain of custody over the source files is the cheapest credibility you will ever buy.</p>



<h2 class="wp-block-heading">Where to run it, and the confidentiality problem nobody budgets for</h2>



<p class="wp-block-paragraph">A modest VPS is plenty. Contabo and InterServer both sell machines with more than enough memory and disk for a job this size, and the specification matters far less than the question you should ask first: does your confidentiality undertaking actually permit the disclosure material to sit on that machine, in that jurisdiction? Often the answer is no, and the analysis has to run on a workstation inside the client&#8217;s own environment. Find that out before you build anything.</p>



<p class="wp-block-paragraph">Three practical points that follow from the material being privileged rather than ordinary project data:</p>



<ul class="wp-block-list">
<li><strong>Full-disk encryption on anything that holds the files</strong>, including the laptop you take to the hearing. This is table stakes and it is still routinely skipped.</li>

<li><strong>A VPN for untrusted networks.</strong> You will open the data room from a hotel or an airport at some point. NordVPN and Surfshark both handle that case. It protects the transport; it does nothing about where the data comes to rest, so do not let it substitute for the jurisdiction question above.</li>

<li><strong>A disposal plan.</strong> Undertakings usually require destruction or return at the close of the engagement, and deleting a file does not destroy it. Overwrite-based tools such as O&amp;O SafeErase cover the workstation case. If the volume was encrypted from the start, destroying the key is faster and cleaner than overwriting the data, and easier to certify.</li>
</ul>



<p class="wp-block-paragraph">If cloud storage is permitted, write-once object storage with a retention lock suits the ingested originals well. It gives you an immutability story that does not rest on your own discipline.</p>



<h2 class="wp-block-heading">Troubleshooting the ingest</h2>



<ul class="wp-block-list">
<li><strong>The parser dies on a decode error.</strong> XER files are not reliably UTF-8, particularly when they came from a database with a non-English locale. Try a single-byte fallback encoding before assuming the file is corrupt, and log which encoding succeeded against the version record.</li>

<li><strong>Row and field counts do not match.</strong> Activity names can contain characters that confuse a naive split, and some exports wrap fields inconsistently. Parse defensively on field count per <code>%F</code> line rather than assuming.</li>

<li><strong>Duplicate <code>task_code</code> in one file.</strong> One XER can carry several projects plus the shared global data they depend on. Always scope by <code>proj_id</code> and confirm you have the right project before anything else.</li>

<li><strong>Activity counts drop between updates for no reason.</strong> Check whether the export was filtered or taken at a WBS level rather than the full project. A partial export looks exactly like a deletion event in your diff.</li>

<li><strong>Dates shift by a few hours.</strong> P6 stores times, not just dates, and calendar start hours vary. Normalise to the activity calendar&#8217;s working day before comparing, or every finish looks a shift early.</li>

<li><strong>Float figures look wrong across the board.</strong> Check the hours-versus-days conversion and the calendar assignment before checking anything else. It is almost always one of those two.</li>
</ul>



<h2 class="wp-block-heading">Common mistakes</h2>



<ul class="wp-block-list">
<li>Treating the final update as the as-built without stating that assumption anywhere in the report.</li>

<li>Joining on <code>task_id</code> because it is the primary key, and inheriting whatever database history came with the file.</li>

<li>Letting the platform resolve an ambiguous activity match silently, so the analyst never sees the ambiguity existed.</li>

<li>Converting float to days with a fixed divisor, then reporting a single variance figure that mixes progress with programme edits.</li>

<li>Automating the record-to-activity link all the way to a stated cause, which turns an evidence tool into a target.</li>

<li>Modifying source XERs in place, then having no way to demonstrate they are as received.</li>
</ul>



<h2 class="wp-block-heading">Best practices worth the effort</h2>



<ul class="wp-block-list">
<li>Hash every source file on ingest and record the hash next to the version. It takes one line and it answers a question you will otherwise answer badly.</li>

<li>Make every table append-only. Corrections become new rows with a reason, never overwrites.</li>

<li>Build the first-appearance view early. It reframes the analysis from &#8220;what do the files say&#8221; to &#8220;when did the files start saying it&#8221;, which is the more defensible question.</li>

<li>Make every exhibit reproducible from one query, so a challenged figure traces back to a file and a row rather than a spreadsheet cell.</li>

<li>Keep a written record of every judgement call the platform surfaced and a human resolved. That log is often more useful in the hearing than the analysis itself.</li>

<li>Validate against the commercial tools where you can. If Acumen Fuse or Schedule Comparison disagrees with your diff on a given version pair, one of you is wrong and you want to know which before the report goes out.</li>
</ul>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Frequently asked questions</h2>



<h3 class="wp-block-heading">Is as-planned vs as-built analysis still accepted in construction disputes?</h3>



<p class="wp-block-paragraph">Yes, as one of a recognised family of methods. AACE Recommended Practice 29R-03 includes it in its taxonomy as an observational method, and the SCL Protocol&#8217;s second edition deliberately stopped naming a single preferred technique for retrospective analysis, setting out selection factors instead. Acceptance in a given dispute turns on the records available, the contract and the forum, not on the method&#8217;s reputation in the abstract.</p>



<h3 class="wp-block-heading">How many schedule updates do I need before building a platform is worth it?</h3>



<p class="wp-block-paragraph">Roughly: below about eight updates, a careful analyst with a workbook is faster. Above twenty, manual comparison stops being reliable, not because people are careless but because the number of pairwise comparisons grows past what anyone can hold. The crossover in practice sits somewhere in the low teens, and it moves earlier if there are multiple baselines or a renumbering event.</p>



<h3 class="wp-block-heading">Can I do as-planned vs as-built analysis without a Primavera P6 licence?</h3>



<p class="wp-block-paragraph">For the data work, yes. An XER is plain text and parses without P6 installed, and browser-based and open-source viewers exist. You will still want licensed access somewhere in the engagement for anything that requires recalculating the network, because reproducing P6&#8217;s scheduling engine faithfully is not a project you want inside a disputes timetable.</p>



<h3 class="wp-block-heading">What is the difference between as-planned vs as-built and a windows analysis?</h3>



<p class="wp-block-paragraph">The simple form of as-planned versus as-built compares planned dates against actual dates observationally, without necessarily using network logic. A windows analysis divides the project into periods, usually monthly, and examines the contemporaneous critical path and the critical delay in each period before investigating the cause. Windows demands more of the records and generally carries more weight where the programme was maintained properly. This platform supports both, because both need the same version history underneath.</p>



<h3 class="wp-block-heading">What do I do when the contractor&#8217;s updates were not maintained properly?</h3>



<p class="wp-block-paragraph">State it, quantify it, and let it drive method selection. That is exactly what the first-appearance table is for: reporting lag, actuals that moved after being stated, out-of-sequence progress and unexplained logic edits are all measurable. Recommended practice treats recreating a schedule from records as a last resort, used when contemporaneous data is unavailable, and recreated schedules carry less evidential weight. Showing the deficiency with numbers is stronger than asserting it.</p>



<h3 class="wp-block-heading">Should I build this or buy commercial delay analysis software?</h3>



<p class="wp-block-paragraph">Buy for schedule diagnostics and version comparison. That problem is well solved and the established tools have years of edge cases behind them. Build for the parts specific to your instruction: the record linking, the exhibit generation in your house format, and the audit trail. The hybrid is usually right, and the deciding factor is more often the licence and the timescale than the technical fit.</p>



<h3 class="wp-block-heading">Does the platform need to reproduce P6&#8217;s critical path calculation?</h3>



<p class="wp-block-paragraph">No, and it should not try. Read the float and date values P6 already calculated and stored in each export, and treat them as evidence of what the programme said at that moment. Anywhere you genuinely need a recalculated network, do it in the scheduling tool so the result is reproducible by anyone with the same file.</p>



<h2 class="wp-block-heading">The one thing to take away</h2>



<p class="wp-block-paragraph">A credible as-planned vs as-built analysis is not a better comparison between two files. It is a complete, immutable, queryable history of what every schedule update said and when it started saying it. Build the version history and the first-appearance table first, and the exhibits fall out of it. Build the exhibits first and you will spend the hearing defending a spreadsheet whose provenance you cannot reconstruct.</p>



<p class="wp-block-paragraph">The scheduling judgement stays with the expert. The platform&#8217;s job is narrower and duller: make sure that when someone asks where a date came from, the answer takes one query and not one week.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Need the data side of this built?</h2>



<p class="wp-block-paragraph">I build the engineering layer underneath delay and quantum work, so the expert spends their time on opinion rather than on data wrangling. That usually looks like:</p>



<ul class="wp-block-list">
<li>Ingesting a full sequence of P6 XER or XML updates into a versioned, append-only PostgreSQL model with hashing and chain of custody on every source file.</li>

<li>Building the first-appearance and reporting-lag views that show when each actual date entered the record, and flagging actuals that moved after being stated.</li>

<li>Classifying every between-update change into progress, duration, logic, calendar, constraint or scope, with before and after values on each row.</li>

<li>Extracting structured fields from daily reports, RFIs and correspondence with OCR and document processing, then generating candidate activity links for an analyst to confirm or reject.</li>

<li>Producing exhibit-grade output in Power BI or as static deliverables, where every figure traces back to a query, a file and a row.</li>

<li>Standing it up on infrastructure that fits the confidentiality regime you work under, whether that is a locked-down VPS or a machine inside the client&#8217;s own environment.</li>
</ul>



<p class="wp-block-paragraph">If you have a stack of XERs and a deadline, send me two consecutive updates and whatever the record-keeping looked like, and I will tell you what the sequence can and cannot support.</p>



<div class="wp-block-buttons is-layout-flex wp-block-buttons-is-layout-flex">
<div class="wp-block-button"><a class="wp-block-button__link wp-element-button" href="https://www.upwork.com/freelancers/~01f15a912ad84a6620" target="_blank" rel="noreferrer noopener">Work with me on Upwork</a></div>
</div>
<p>The post <a href="https://john-nessime.com/blog/technical-guides/as-planned-vs-as-built-analysis-platform/">As-Planned vs As-Built Analysis: Building a Platform That Survives Cross-Examination</a> appeared first on <a href="https://john-nessime.com/blog">John Nessime</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Building an AI Construction Claims Platform on AWS That Holds Up Under Scrutiny</title>
		<link>https://john-nessime.com/blog/case-studies/ai-construction-claims-platform-aws/</link>
		
		<dc:creator><![CDATA[John Nessime]]></dc:creator>
		<pubDate>Mon, 03 Aug 2026 14:05:46 +0000</pubDate>
				<category><![CDATA[Case Studies]]></category>
		<category><![CDATA[Cloud Computing]]></category>
		<category><![CDATA[Technical Guides]]></category>
		<category><![CDATA[Amazon Athena]]></category>
		<category><![CDATA[Amazon Bedrock]]></category>
		<category><![CDATA[Amazon S3]]></category>
		<category><![CDATA[Amazon S3 Vectors]]></category>
		<category><![CDATA[Amazon Textract]]></category>
		<category><![CDATA[Architecture]]></category>
		<category><![CDATA[AWS]]></category>
		<category><![CDATA[AWS Glue]]></category>
		<category><![CDATA[Bedrock Guardrails]]></category>
		<category><![CDATA[Cloud]]></category>
		<category><![CDATA[Construction Technology]]></category>
		<category><![CDATA[Data Engineering]]></category>
		<category><![CDATA[Document Processing]]></category>
		<category><![CDATA[Embeddings]]></category>
		<category><![CDATA[Generative AI]]></category>
		<category><![CDATA[Infrastructure]]></category>
		<category><![CDATA[Legal Tech]]></category>
		<category><![CDATA[Metadata Filtering]]></category>
		<category><![CDATA[Primavera P6]]></category>
		<category><![CDATA[RAG]]></category>
		<category><![CDATA[Serverless]]></category>
		<category><![CDATA[SQL]]></category>
		<category><![CDATA[Vector Database]]></category>
		<guid isPermaLink="false">https://john-nessime.com/blog/?p=130</guid>

					<description><![CDATA[<p>Semantic search finds the most persuasive document, not the earliest one. Here is how to architect an AI construction claims and dispute intelligence platform on AWS so retrieval respects the contractual clock, schedule data stays out of the vector index, every answer resolves to a page, and privileged material never shares a retrieval path with project records.</p>
<p>The post <a href="https://john-nessime.com/blog/case-studies/ai-construction-claims-platform-aws/">Building an AI Construction Claims Platform on AWS That Holds Up Under Scrutiny</a> appeared first on <a href="https://john-nessime.com/blog">John Nessime</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">Someone hands you a shared drive and asks a question that sounds trivial: &#8220;Did we give notice of the delay event inside the contractual period, or didn&#8217;t we?&#8221;</p>



<p class="wp-block-paragraph">The answer is in there. It is one email, or one line in a site diary, sitting among forty thousand other files. Nobody can read forty thousand files, so the instinct is to point a language model at the pile and ask it. That instinct is right. The naive implementation of it is where the money goes.</p>



<p class="wp-block-paragraph">Here is the failure mode that bites hardest, and it is invisible until an expert challenges you on it. You build retrieval over the document set, ask about notice of delay, and the system confidently returns a letter that discusses the delay event in great detail. It is a good letter. It is also dated eleven months after the event, written by the claims consultant during preparation of the claim itself. It scored highest precisely because it was written to argue the point. The contemporaneous notice, the thing you actually needed, was four badly typed lines in a routine progress email that mentioned the word &#8220;delay&#8221; once.</p>



<p class="wp-block-paragraph">Semantic similarity has no concept of a deadline. That single gap is the difference between an <strong>AI construction claims platform</strong> that shortens a disclosure exercise and one that quietly manufactures a wrong answer with a citation attached to it.</p>



<p class="wp-block-paragraph">This post covers how to build that platform on AWS: how to lay out ingestion, how to make retrieval respect the contractual clock, why schedule data must never go anywhere near your vector index, how to keep privileged material out of the same retrieval path as project records, and which AWS building blocks are actually the current ones now that several of the obvious candidates have been moved to maintenance mode.</p>



<h2 class="wp-block-heading">What a claims platform actually has to answer</h2>



<p class="wp-block-paragraph">Before any architecture, be honest about the question shapes. They are not all the same problem and they do not all get solved by retrieval.</p>



<ol class="wp-block-list"><li><strong>Chronology.</strong> What happened, in what order, and on what date was it recorded? This is a retrieval and metadata problem.</li><li><strong>Entitlement.</strong> Which clause applies, and what did it require the parties to do? This is retrieval over the contract plus careful prompting.</li><li><strong>Causation.</strong> Which event moved the critical path, and by how much? This is schedule data and date arithmetic. It is not a language problem at all.</li><li><strong>Quantum.</strong> What did the disruption cost? This is cost and resource data, joined to the events above.</li></ol>



<p class="wp-block-paragraph">Treat all four as &#8220;ask the documents&#8221; and you will get fluent nonsense on two of them. The architecture below splits them deliberately.</p>



<h2 class="wp-block-heading">Failure one: retrieval that finds the best match instead of the first one</h2>



<p class="wp-block-paragraph">Two corpora live in every dispute bundle and they look identical to an embedding model.</p>



<ul class="wp-block-list"><li><strong>Contemporaneous records.</strong> Site diaries, progress emails, minutes, early warnings, RFIs, instructions. Written while the project was running, by people with no idea a dispute was coming.</li><li><strong>Claim-era material.</strong> Narratives, expert reports, without-prejudice correspondence, internal analysis. Written afterwards, specifically to be persuasive about the same events.</li></ul>



<p class="wp-block-paragraph">Claim-era material wins on cosine similarity almost every time, because it is denser in exactly the terms you searched for. If your retriever cannot distinguish them, every answer is contaminated by the argument you were trying to test.</p>



<p class="wp-block-paragraph">The fix is metadata, applied at ingestion, and it is cheap to get right and expensive to retrofit. Amazon Bedrock Knowledge Bases reads a sidecar file that sits next to each document in S3, named with the full original filename plus <code>.metadata.json</code>. So <code>letter-0421.pdf</code> gets <code>letter-0421.pdf.metadata.json</code>. The naming convention is the only link between them; there is no separate registration step.</p>



<pre class="wp-block-code"><code>{
  "metadataAttributes": {
    "doc_date": 20240314,
    "corpus": "contemporaneous",
    "doc_type": "site_correspondence",
    "matter_id": "matter-0007",
    "date_source": "email_header",
    "privileged": false
  }
}</code></pre>



<p class="wp-block-paragraph">Look closely at <code>doc_date</code>. It is an integer, not a string, and that is not a style choice. Bedrock Knowledge Bases metadata attributes support STRING, NUMBER, BOOLEAN and STRING_LIST. The range comparison operators, the ones you need to express &#8220;on or before the notice deadline&#8221;, only apply to NUMBER. Store the date as <code>"2024-03-14"</code> and your filter will not throw an error. It will just quietly match nothing, or match everything, depending on how you wrote it. You will find out weeks later when someone asks why a document they can see in the bundle never appears in results.</p>



<p class="wp-block-paragraph">With the date as a sortable integer, a query filter can express the contractual window directly.</p>



<pre class="wp-block-code"><code>{
  "andAll": [
    { "equals":              { "key": "corpus",   "value": "contemporaneous" } },
    { "equals":              { "key": "matter_id","value": "matter-0007" } },
    { "greaterThanOrEquals": { "key": "doc_date", "value": 20240301 } },
    { "lessThanOrEquals":    { "key": "doc_date", "value": 20240329 } }
  ]
}</code></pre>



<p class="wp-block-paragraph">That is the whole trick. You are no longer asking &#8220;what is the most relevant document about this delay&#8221;. You are asking &#8220;what did the parties actually write during the window in which the contract required them to write it&#8221;. Those are different questions and only one of them is worth anything in a dispute.</p>



<h3 class="wp-block-heading">Where the date comes from matters more than the date</h3>



<p class="wp-block-paragraph">Do not use the S3 object timestamp. It records when someone copied a folder, usually years after the fact and identical across ten thousand files. Derive the date from the document itself: the <code>Date:</code> header on an email, the printed date on a letter, the period covered by a diary entry.</p>



<p class="wp-block-paragraph">Sometimes you cannot, because the scanned undated fax exists in every project archive. Record that honestly with a <code>date_source</code> attribute rather than guessing, and treat unknown-date documents as a separate review pile. An extension of time argument built on an inferred date is an argument you will lose.</p>



<h2 class="wp-block-heading">Failure two: treating the programme like a document</h2>



<p class="wp-block-paragraph">This one is worse, because the output looks right.</p>



<p class="wp-block-paragraph">Oracle Primavera P6 exports XER and PMXML files. Asta Powerproject and Microsoft Project have their own formats. XER in particular is a plain text dump of relational tables, so it goes through a text pipeline without complaint. Chunk it, embed it, and you now have vectors representing fragments of a table of activity codes with no relationships attached.</p>



<p class="wp-block-paragraph">Ask that index how much float activity A1200 had at the March data date and you will get a number. It will be well formatted and it will be invented. Total float is the product of a forward and backward pass across the whole logic network under a specific calendar. It cannot be recovered from a retrieved fragment, and a language model asked to produce it will produce something plausible instead of admitting that.</p>



<p class="wp-block-paragraph">Schedule data goes into a structured store, and the model queries it rather than reasoning about it.</p>



<ol class="wp-block-list"><li>Parse each programme file into tables. <code>PyP6Xer</code> handles XER from Python; MPXJ is a Java library that reads XER, PMXML, Asta Powerproject and MSPDI among others, which matters when the bundle contains four scheduling tools.</li><li>Load activities, logic links, calendars, resource assignments and WBS into Amazon Aurora PostgreSQL for interactive work, or into S3 with AWS Glue and Amazon Athena when you have hundreds of updates and want columnar scans.</li><li>Stamp every row with the <em>data date</em> of the update it came from. This is the single most important column in the whole platform. Without it you have a pile of schedules; with it you have a time series of the project&#8217;s own view of itself.</li><li>Run windows analysis, as-planned versus as-built comparison and float erosion in SQL or Python, deterministically, so the same inputs always give the same numbers.</li><li>Expose the results to the model as a tool it can call, or as generated SQL against a defined schema. The model turns a question into a query and narrates the result. It does not do the arithmetic.</li></ol>



<p class="wp-block-paragraph">A rough shape of the query that makes float erosion visible:</p>



<pre class="wp-block-code"><code>SELECT
    a.activity_id,
    a.data_date,
    a.total_float_days,
    a.total_float_days - LAG(a.total_float_days)
        OVER (PARTITION BY a.activity_id ORDER BY a.data_date)
      AS float_change
FROM   schedule_activities a
WHERE  a.project_id = 'PRJ-01'
  AND  a.data_date BETWEEN DATE '2024-01-01' AND DATE '2024-06-30'
ORDER BY a.activity_id, a.data_date;</code></pre>



<p class="wp-block-paragraph">Nothing clever there, and that is the point. Every number is traceable to a row that came from a named XER file. When an opposing expert asks where a figure came from, the answer is a file name and a query, not &#8220;the model said so&#8221;.</p>



<p class="wp-block-paragraph">Be realistic about effort here. Programme parsing and normalisation across inconsistent updates is the hardest part of the build and the part clients always underestimate. Activity IDs get reused, calendars change mid-project, and someone will have re-baselined without telling anyone. Budget for it.</p>



<h2 class="wp-block-heading">Failure three: an answer with no paper trail</h2>



<p class="wp-block-paragraph">In most RAG applications a citation is a nice touch. In dispute work it <em>is</em> the product. An answer that cannot be traced to a page of a disclosed document is not evidence, it is a rumour with good grammar.</p>



<p class="wp-block-paragraph">Design for that from the ingestion layer, not the presentation layer.</p>



<ul class="wp-block-list"><li><strong>Keep page and position.</strong> Amazon Bedrock Data Automation returns confidence scores and bounding box data alongside extracted fields, and Amazon Textract returns geometry per block. Carry both through the pipeline so a citation resolves to a page and a region, not just a file.</li><li><strong>Route low confidence to humans.</strong> Handwritten site diaries and faxed variation orders will produce low-confidence extractions. Those should land in a review queue by default rather than silently entering the index.</li><li><strong>Reject ungrounded answers.</strong> Amazon Bedrock Guardrails includes contextual grounding checks that score whether a response is supported by the retrieved passages. It reduces confident invention. It does not eliminate it, and anyone who tells you otherwise is selling something.</li><li><strong>Keep an immutable evidential copy.</strong> S3 Versioning plus S3 Object Lock on the landing bucket means the file the platform indexed is provably the file that was disclosed.</li></ul>



<p class="wp-block-paragraph">One design rule underpins all of it: the platform shortlists evidence, it does not decide entitlement. Recognised frameworks for this work, the Society of Construction Law Delay and Disruption Protocol and AACE International&#8217;s Recommended Practice 29R-03 on forensic schedule analysis, both assume a named analyst applying a stated method and exercising judgement. A system that outputs &#8220;the contractor is entitled to 42 days&#8221; is not helping. A system that outputs &#8220;here are the eleven contemporaneous documents inside the notice window, here is the float movement across those updates, here is what is missing&#8221; is doing real work.</p>



<h2 class="wp-block-heading">Failure four: one index for privileged and non-privileged material</h2>



<p class="wp-block-paragraph">Dispute bundles contain legal advice, counsel&#8217;s opinions, without-prejudice correspondence and internal settlement analysis. Those must not be retrievable through the same path as project records.</p>



<p class="wp-block-paragraph">The tempting shortcut is a <code>privileged: false</code> metadata filter on every query. Do not rely on that as your boundary. A metadata filter is a query parameter. One missing filter in one code path, one debug endpoint, one caching layer that drops it, and privileged material surfaces in a general search. The blast radius of that mistake is not a bug report.</p>



<p class="wp-block-paragraph">Separate the indexes physically and separate the IAM roles that can reach them. Amazon S3 Vectors makes this practical: you can set a dedicated customer-managed KMS key per vector index, and you get a large number of indexes per vector bucket, so per-matter and per-sensitivity separation does not become an operational burden. Keep the metadata flag as well, because defence in depth is free, but make the identity boundary the one you actually trust.</p>



<p class="wp-block-paragraph">Amazon Macie is worth pointing at the landing bucket to find personal data you did not expect, particularly in HR records and accident reports that get swept into project archives.</p>



<h2 class="wp-block-heading">Choosing the AWS building blocks, including what not to build on</h2>



<p class="wp-block-paragraph">A lot of published architectures for this kind of platform are now pointing at services AWS has stopped developing. Two matter here, and the dates are the point.</p>



<ul class="wp-block-list"><li><strong>Amazon Kendra</strong> entered maintenance mode on 30 June 2026 and stops accepting new customers on 30 July 2026. Existing customers keep support and security fixes but no new capability. AWS directs new enterprise search and RAG work to Amazon Bedrock Knowledge Bases. If a tutorial or a proposal you are reading starts with a Kendra index, it predates that change.</li><li><strong>Amazon Bedrock Agents</strong> moved to maintenance mode in the same round of service availability changes, with Amazon Bedrock AgentCore as the successor for agentic orchestration. Check the current AWS service availability page before you commit an orchestration layer.</li></ul>



<p class="wp-block-paragraph">For the retrieval layer itself, Bedrock Knowledge Bases now comes in two shapes and the choice is a real trade-off rather than a marketing tier.</p>



<h3 class="wp-block-heading">Managed Knowledge Base</h3>



<p class="wp-block-paragraph">AWS manages the vector store, embeddings model, re-ranker and retrieval orchestration as a single primitive, with native connectors for Amazon S3, SharePoint, Confluence, Google Drive, OneDrive and a web crawler, plus automatic parsing strategy selection and a retriever that decomposes multi-step queries. The connectors pull source permissions along with content, which matters when the document set lives in the client&#8217;s SharePoint rather than a bucket you control.</p>



<p class="wp-block-paragraph">Where it wins: you get a working retrieval layer in an afternoon instead of a fortnight, and the parsing tuning that normally eats the first weeks of a build is done for you. For a first matter, or a proof of value before a client commits budget, this is the one I would reach for.</p>



<h3 class="wp-block-heading">Custom Knowledge Base</h3>



<p class="wp-block-paragraph">You bring your own vector store and control chunking, embedding model and index layout.</p>



<p class="wp-block-paragraph">Where it wins: claims work has awkward chunking requirements. A two-page letter split mid-sentence at a page boundary produces a chunk where the notice sentence has lost its date and its addressee. Controlling chunk boundaries around document structure, and controlling which index a document lands in, are both easier when you own the store. Where it doesn&#8217;t: you now own embedding model upgrades, re-indexing, sync failures and capacity, which is real ongoing work for a small team.</p>



<p class="wp-block-paragraph">Start managed, build a retrieval evaluation set of real questions with known correct documents, and only move to custom when that set demonstrates the problem is chunking. Most teams migrate on a hunch and discover the problem was metadata all along.</p>



<h3 class="wp-block-heading">Where the vector storage bill actually comes from</h3>



<p class="wp-block-paragraph">Rates change, so learn the billing mechanism rather than a number. Amazon S3 Vectors charges on three axes: upload volume by logical gigabyte, storage by logical gigabyte, and queries by data processed, where data processed scales with the size of the index being searched. Note that filtering does not reduce the data processed by a query.</p>



<p class="wp-block-paragraph">That shape suits claims work unusually well. A dispute archive is enormous and cold: millions of chunks, queried by a handful of analysts a few hundred times a day, so you pay mostly for storage, which is the cheap axis. Compare that against Amazon OpenSearch Serverless, which prices on provisioned compute units and therefore rewards high query volume against a smaller index, or Aurora PostgreSQL with pgvector when you already need Aurora for the schedule tables and would rather run one system than two.</p>



<p class="wp-block-paragraph">The practical lever is to split indexes per matter. Query cost scales with index size, so one giant index across every dispute you have ever run makes every query more expensive than it needs to be, on top of being a bad idea for confidentiality.</p>



<h2 class="wp-block-heading">A reference pipeline</h2>



<ol class="wp-block-list"><li>Everything lands in S3 under a per-matter prefix, with Versioning and Object Lock enabled on the evidential copy.</li><li>S3 event notifications trigger AWS Step Functions. Use Step Functions rather than a chain of Lambdas so that a failed extraction on page 300 of a 400-page bundle is visible and resumable.</li><li>Classify and split. Scanned bundles arrive as one PDF containing forty separate documents. Splitting them correctly is a prerequisite for dating them correctly.</li><li>Extract text with Amazon Bedrock Data Automation or Amazon Textract, keeping confidence scores and geometry.</li><li>Derive the document date and write the <code>.metadata.json</code> sidecar. Anything undated goes to the review queue.</li><li>Route by type: correspondence to the knowledge base, programme files to the XER parser and the relational store, cost data to its own tables.</li><li>Sync the knowledge base, then run your retrieval evaluation set before anyone uses it. A sync that succeeds is not the same as an index that answers correctly.</li><li>Serve through an API that refuses to return an answer without citations, and log every query with the filters that were applied.</li></ol>



<p class="wp-block-paragraph">Define the whole thing in Terraform or OpenTofu from the start. Matters are per-client and short-lived, and standing one up should be a variable file, not an afternoon in the console. Point Amazon CloudWatch, or Grafana Cloud if you already run Grafana elsewhere, at the Step Functions execution metrics so a silently failing extraction stage does not go unnoticed for a week.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Troubleshooting</h2>



<h3 class="wp-block-heading">Date filters return nothing, and no error</h3>



<p class="wp-block-paragraph">Almost always the date was stored as a string. Range operators need NUMBER. Convert to an integer in <code>YYYYMMDD</code> form and re-sync the affected documents.</p>



<h3 class="wp-block-heading">A document is in the bucket but never appears in results</h3>



<p class="wp-block-paragraph">Check the sidecar filename first. It must be the complete original filename with <code>.metadata.json</code> appended, extension included. <code>report.pdf.metadata.json</code> works; <code>report.metadata.json</code> is a file the ingestion job will happily ignore. After that, check whether a filter in the query path is excluding it.</p>



<h3 class="wp-block-heading">Answers cite the right document but the wrong passage</h3>



<p class="wp-block-paragraph">Chunking split the document somewhere structurally meaningful. Look at the raw chunks for that file. If the notice sentence and its date are in different chunks, no amount of prompt tuning fixes it. That is the signal to take control of chunking.</p>



<h3 class="wp-block-heading">Float figures do not match the client&#8217;s own analysis</h3>



<p class="wp-block-paragraph">Check calendars before you check logic. Different activity calendars, a changed default calendar, or an update where someone applied a progress override will move float without any logic change. Reconcile activity counts between your parsed tables and the source file before trusting anything downstream.</p>



<h3 class="wp-block-heading">Query costs jumped without more usage</h3>



<p class="wp-block-paragraph">An index grew. With storage-side vector search, query cost tracks the size of the index being scanned, so ingesting a large new bundle raises the price of every subsequent query against that index. Split by matter.</p>



<h2 class="wp-block-heading">Common mistakes</h2>



<ul class="wp-block-list"><li>Using the file&#8217;s storage timestamp as the document date. It records the migration, not the event.</li><li>Indexing claim narratives and contemporaneous records into the same corpus with no way to tell them apart.</li><li>Embedding programme exports because they happen to be text files.</li><li>Treating a metadata filter as a privilege boundary instead of an optimisation.</li><li>Letting the model state entitlement conclusions rather than assembling and citing evidence.</li><li>Building on services that have moved to maintenance mode because the tutorial you followed predates the change.</li><li>Shipping without a retrieval evaluation set, so you have no way to know whether a change made things better or worse.</li><li>One index for every matter, which is both a cost problem and a confidentiality problem.</li></ul>



<h2 class="wp-block-heading">Best practices</h2>



<ul class="wp-block-list"><li>Make the document date a first-class, numeric, filterable attribute, and record where it came from.</li><li>Keep an immutable evidential copy separate from the working copy the pipeline mutates.</li><li>Separate structured schedule and cost data from unstructured documents, and let the model query the former rather than reason about it.</li><li>Build a retrieval evaluation set from real questions with known correct documents before you tune anything.</li><li>Enforce citations at the API layer, so an uncited answer is impossible rather than discouraged.</li><li>Isolate privileged material by index and by IAM role, with metadata as a second layer.</li><li>Log every query with its filters, so you can reconstruct how any given answer was reached.</li><li>Define infrastructure as code so a new matter is a deployment, not a project.</li></ul>



<h2 class="wp-block-heading">FAQ</h2>



<h3 class="wp-block-heading">Can an AI construction claims platform replace a delay expert?</h3>



<p class="wp-block-paragraph">No, and building toward that goal produces something unusable. Established forensic frameworks assume a named analyst applying a stated method whose reasoning can be tested. The platform&#8217;s value is compressing weeks of document review into hours and making the schedule data queryable, so the expert spends their time on judgement rather than searching.</p>



<h3 class="wp-block-heading">Should I use Amazon Kendra for the search layer?</h3>



<p class="wp-block-paragraph">Not for a new build. Kendra entered maintenance mode on 30 June 2026 and closed to new customers on 30 July 2026, with AWS pointing to Bedrock Knowledge Bases for equivalent and more current capability. Existing Kendra deployments continue to be supported, so this is a migration assessment rather than an emergency, but starting there now means starting on a service with no roadmap.</p>



<h3 class="wp-block-heading">How do I stop the model inventing float and delay figures?</h3>



<p class="wp-block-paragraph">Do not give it the chance. Keep schedule data in a relational or columnar store and have the model generate queries against a defined schema, or call a tool that runs a fixed calculation. The arithmetic happens in SQL or Python where it is deterministic and reproducible; the model only turns questions into queries and results into sentences.</p>



<h3 class="wp-block-heading">Which vector store should I choose for a claims archive?</h3>



<p class="wp-block-paragraph">Match the store to your query pattern. Large, cold archives queried by a few analysts favour storage-priced options like Amazon S3 Vectors, where you mostly pay to keep the data. Smaller indexes hit constantly favour compute-priced options like Amazon OpenSearch Serverless. If you already run Aurora PostgreSQL for schedule data, pgvector alongside it is a legitimate way to avoid operating a second system.</p>



<h3 class="wp-block-heading">How do I handle scanned and handwritten site records?</h3>



<p class="wp-block-paragraph">Extract them with confidence scores retained, set a threshold, and route everything below it to human review before indexing. Handwritten diaries are frequently the most probative documents in a delay claim and also the least reliable to read automatically, so the review queue is not an edge case. Plan capacity for it.</p>



<h3 class="wp-block-heading">Where do documents come from if they are not already in S3?</h3>



<p class="wp-block-paragraph">Most project records live in a common data environment such as Procore, Autodesk Construction Cloud, Aconex or a client SharePoint tenancy. Bedrock Managed Knowledge Base has native connectors for SharePoint, Confluence, Google Drive and OneDrive that ingest permissions alongside content. For platforms without a native connector, export to S3 and keep the export manifest as part of the disclosure record.</p>



<h2 class="wp-block-heading">The one thing worth remembering</h2>



<p class="wp-block-paragraph">An <strong>AI construction claims platform</strong> lives or dies on whether it understands time. Every hard requirement in this build traces back to that: numeric dates so you can filter to a contractual window, a data date on every schedule row so float movement is measurable, a corpus flag so contemporaneous records are not drowned out by material written to argue about them, and citations that resolve to a page so any answer can be checked.</p>



<p class="wp-block-paragraph">Get the temporal metadata right at ingestion and the rest of the architecture is ordinary AWS work. Get it wrong and you have built a very expensive way to retrieve the most persuasive document instead of the true one.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Need help building this on AWS?</h2>



<p class="wp-block-paragraph">I design and build document and data platforms on AWS, and this kind of system sits squarely in that work. Things I can help with:</p>



<ul class="wp-block-list"><li>Designing the ingestion pipeline: S3 landing zones with Object Lock, Step Functions orchestration, splitting and classifying scanned bundles, and confidence-based routing to human review.</li><li>Getting the temporal metadata model right, including date derivation, sidecar generation and filter design against Amazon Bedrock Knowledge Bases.</li><li>Parsing Primavera P6 XER and PMXML exports into queryable tables in Aurora PostgreSQL or S3 with Glue and Athena, with a data date on every row.</li><li>Choosing and sizing the vector layer across Amazon S3 Vectors, OpenSearch Serverless and pgvector, based on your actual query pattern rather than a benchmark.</li><li>Building index and IAM separation for privileged material, plus KMS key strategy and Macie scanning of landing buckets.</li><li>Setting up retrieval evaluation, citation enforcement, query audit logging and CloudWatch or Grafana dashboards over the pipeline so failures surface early.</li></ul>



<p class="wp-block-paragraph">If you are partway into something like this already, send me a sample metadata sidecar, a Step Functions execution history, or a query that returns the wrong document, and I will tell you what I think is going on.</p>



<div class="wp-block-buttons is-layout-flex wp-block-buttons-is-layout-flex">
<div class="wp-block-button"><a class="wp-block-button__link wp-element-button" href="https://www.upwork.com/freelancers/~01f15a912ad84a6620" target="_blank" rel="noreferrer noopener">Work with me on Upwork</a></div>
</div>
<p>The post <a href="https://john-nessime.com/blog/case-studies/ai-construction-claims-platform-aws/">Building an AI Construction Claims Platform on AWS That Holds Up Under Scrutiny</a> appeared first on <a href="https://john-nessime.com/blog">John Nessime</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Build a Secure AI Medical Assistant on AWS: The Boundaries That Actually Leak</title>
		<link>https://john-nessime.com/blog/case-studies/secure-ai-medical-assistant-aws/</link>
		
		<dc:creator><![CDATA[John Nessime]]></dc:creator>
		<pubDate>Mon, 03 Aug 2026 13:45:12 +0000</pubDate>
				<category><![CDATA[Case Studies]]></category>
		<category><![CDATA[Cloud Computing]]></category>
		<category><![CDATA[DevOps]]></category>
		<category><![CDATA[Technical Guides]]></category>
		<category><![CDATA[Amazon Bedrock]]></category>
		<category><![CDATA[Amazon Comprehend Medical]]></category>
		<category><![CDATA[Amazon S3]]></category>
		<category><![CDATA[Architecture]]></category>
		<category><![CDATA[AWS]]></category>
		<category><![CDATA[Bedrock Guardrails]]></category>
		<category><![CDATA[Cloud Security]]></category>
		<category><![CDATA[CloudWatch]]></category>
		<category><![CDATA[Data Residency]]></category>
		<category><![CDATA[Generative AI]]></category>
		<category><![CDATA[Healthcare AI]]></category>
		<category><![CDATA[HIPAA]]></category>
		<category><![CDATA[IAM]]></category>
		<category><![CDATA[Logging]]></category>
		<category><![CDATA[PHI]]></category>
		<category><![CDATA[RAG]]></category>
		<category><![CDATA[Vector Database]]></category>
		<category><![CDATA[VPC]]></category>
		<guid isPermaLink="false">https://john-nessime.com/blog/?p=127</guid>

					<description><![CDATA[<p>A practical architecture for a secure AI medical assistant on AWS, organised by the boundary the data crosses: the input box, your own invocation logs, cross-Region inference routing, the retrieval index, and clinical accuracy. Includes real commands, the failure modes that stay invisible until an audit, and the trade-offs worth knowing before you build.</p>
<p>The post <a href="https://john-nessime.com/blog/case-studies/secure-ai-medical-assistant-aws/">Build a Secure AI Medical Assistant on AWS: The Boundaries That Actually Leak</a> appeared first on <a href="https://john-nessime.com/blog">John Nessime</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">The message usually arrives on a Friday afternoon: &#8220;One of the residents pasted a real discharge summary into the demo.&#8221; Nobody meant anything by it. The thing was a study aid, a chat box over a pile of reference material, running in a sandbox account with no Business Associate Addendum in place and invocation logging switched on because logging is a good habit. And now there is protected health information sitting in plaintext in a CloudWatch log group in an account that was never in scope for it.</p>



<p class="wp-block-paragraph">That is the failure mode worth planning for. Not a jailbreak, not a model saying something clinically wrong on stage. A user typing something perfectly reasonable into a box you built, and the data ending up somewhere you never drew on the diagram. This post walks through how to build a secure AI medical assistant on AWS, organised by the boundary the data actually crosses: the input box, your own logs, the Region the inference runs in, the retrieval index, and finally the answer itself. Commands are included where they explain something. Where a value depends on your account or your counsel, I say so instead of making one up.</p>



<h2 class="wp-block-heading">Two different products hiding behind one request</h2>



<p class="wp-block-paragraph">&#8220;An AI assistant for medical students and doctors&#8221; is two builds with two risk profiles, and conflating them is the root of most of the trouble.</p>



<ul class="wp-block-list">
<li><strong>The study tool.</strong> Question banks, guideline summaries, differential drills, spaced repetition. In theory it never touches patient data. Its real risk is confident wrongness and unattributed answers, not privacy.</li>

<li><strong>The clinical assistant.</strong> Note summarisation, coding support, chart question answering. It handles PHI by design, so the whole thing has to sit inside a HIPAA-designated account from day one.</li>
</ul>



<p class="wp-block-paragraph">The trap is that the study tool becomes the clinical assistant without anyone shipping a release. A student rehearses a case they saw on the ward. A doctor tries the study tool on a real chart because it is the one that is already open. The moment your input box accepts free text from someone with clinical access, you should assume PHI will arrive in it. Build accordingly, or put the study tool on infrastructure where PHI arriving is survivable.</p>



<p class="wp-block-paragraph">My default is to run both in the same HIPAA-designated account with the same controls, and keep only the marketing site, the docs and the waitlist form somewhere ordinary and cheap like InterServer or any commodity host, entirely outside the AWS organisation. Small blast radius beats clever separation you have to explain to an auditor.</p>



<h2 class="wp-block-heading">Boundary one: PHI arrives before you decide to accept it</h2>



<p class="wp-block-paragraph">Start with the thing that trips up almost every first build: HIPAA-eligible and HIPAA-compliant are not the same word. AWS designating a service as HIPAA-eligible means you are permitted to process PHI with it under an executed Business Associate Addendum. It says nothing about whether your deployment is compliant. That part is entirely yours.</p>



<p class="wp-block-paragraph">Three conditions have to hold together before any PHI touches a service: the service is on the AWS HIPAA Eligible Services Reference, you have an executed BAA, and the account is designated for HIPAA use. Amazon Bedrock and Amazon Comprehend Medical both appear on that list, and the BAA is accepted through AWS Artifact rather than negotiated by email. Check the reference page yourself before you commit to an architecture, because the list changes and every component in your diagram has to be on it, not just the model.</p>



<h3 class="wp-block-heading">Detecting PHI at the door</h3>



<p class="wp-block-paragraph">Amazon Comprehend Medical has an operation specifically for finding protected health information in unstructured clinical text. You hand it free text, it returns detected entities with a type, a confidence score and character offsets.</p>



<pre class="wp-block-code"><code>aws comprehendmedical detect-phi 
  --region us-east-1 
  --text "Patient seen for chest pain, MRN 004512, discharged Tuesday."</code></pre>



<p class="wp-block-paragraph">The offsets are the useful part. They let you redact or tokenise in place before the text goes anywhere else, rather than throwing the whole message away and telling the user to try again.</p>



<p class="wp-block-paragraph">Now the caveat that matters more than the feature. AWS states plainly that Comprehend Medical may not identify PHI in all circumstances and that it does not, on its own, meet HIPAA&#8217;s requirements for de-identification. Read that as: it is a good filter and a terrible guarantee. If your compliance story is &#8220;we strip PHI before the model sees it, so we are outside HIPAA scope,&#8221; that story does not hold. Treat detection as defence in depth inside a compliant account, not as an escape hatch from needing one.</p>



<h2 class="wp-block-heading">Boundary two: your own logs are the most likely leak</h2>



<p class="wp-block-paragraph">This is the invisible one. Bedrock model invocation logging is disabled by default and captures the full request data, response data and metadata for every call in the account, in that Region. You turn it on for a good reason, usually because someone in security asked who prompted what and when. Then it quietly becomes the largest concentration of raw clinical text you own, and it is nowhere on the architecture diagram because it is a checkbox rather than a component.</p>



<p class="wp-block-paragraph">Check whether it is on before you assume anything:</p>



<pre class="wp-block-code"><code>aws bedrock get-model-invocation-logging-configuration --region us-east-1</code></pre>



<p class="wp-block-paragraph">The configuration is per account per Region, so run it in every Region where anyone has ever opened the Bedrock console. Setting it deliberately looks like this:</p>



<pre class="wp-block-code"><code>aws bedrock put-model-invocation-logging-configuration 
  --region us-east-1 
  --logging-config '{
    "s3Config": {
      "bucketName": "med-assistant-invocation-logs",
      "keyPrefix": "bedrock/"
    },
    "textDataDeliveryEnabled": true,
    "imageDataDeliveryEnabled": false,
    "embeddingDataDeliveryEnabled": false
  }'</code></pre>



<p class="wp-block-paragraph">Each delivery flag is a separate decision about a separate category of PHI. Text is the obvious one. Image delivery matters the moment anyone uploads a photographed chart or a scan, because burned-in identifiers travel with the pixels and no text filter will ever see them. Embedding delivery is the one people leave on without thinking; vectors derived from clinical text are not a safe artefact, and they are large.</p>



<p class="wp-block-paragraph">The sharpest detail is buried in the Guardrails documentation: AWS notes you can disable invocation logs if you do not want blocked content appearing as plaintext in them. Read the implication. A guardrail can refuse a prompt, mask the identifiers, and stop the model ever seeing them, and the original text still lands in your log destination. The guardrail protects the model call. It does not protect the log.</p>



<ul class="wp-block-list">
<li>Encrypt the log destination with a customer-managed KMS key, and keep the key policy tight enough that &#8220;everyone with S3 read&#8221; is not also &#8220;everyone with chart access&#8221;.</li>

<li>Set a retention period that reflects a legal decision someone actually made, not the CloudWatch default of never expiring.</li>

<li>Send logs to a separate, tightly scoped account if your organisation is large enough to have people who need dashboards but not records.</li>

<li>Remember your application logs too. A framework that logs request bodies on error will do this to you long before Bedrock does.</li>
</ul>



<h2 class="wp-block-heading">Boundary three: where the inference actually runs</h2>



<p class="wp-block-paragraph">Cross-Region inference in Bedrock exists because on-demand capacity is uneven and bursts happen. It routes your request to another Region to get it served. There are two flavours and the difference is not cosmetic.</p>



<ul class="wp-block-list">
<li><strong>Geographic profiles</strong> keep routing inside a defined geography such as the US or the EU. A request that starts in the EU stays in EU Regions. This is the one built for residency requirements.</li>

<li><strong>Global profiles</strong> route to supported commercial Regions worldwide for maximum throughput. AWS documents that a request can be routed to a destination Region even if you never opted that Region into your account.</li>
</ul>



<p class="wp-block-paragraph">To be fair to global profiles: data is not stored in the destination Region, transfer happens encrypted across the AWS network, and your invocation logs, knowledge bases and configuration all stay in the source Region. For a workload with no geographic constraint it is a genuinely good default, and it typically carries a lower per-token rate than staying in-geography. For a clinical workload with a residency commitment in a contract, it is the wrong tool, and &#8220;the prompt left the geography but was not stored there&#8221; is a sentence you do not want to be constructing during an audit.</p>



<p class="wp-block-paragraph">Pin it in policy rather than trusting a config value in a repo. A Service Control Policy denying the Bedrock API outside your approved Regions closes both doors at once, because invoking a cross-Region profile requires model access in the destination Regions as well as the source:</p>



<pre class="wp-block-code"><code>{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "DenyBedrockOutsideApprovedRegions",
      "Effect": "Deny",
      "Action": "bedrock:*",
      "Resource": "*",
      "Condition": {
        "StringNotEquals": {
          "aws:RequestedRegion": ["us-east-1", "us-west-2"]
        }
      }
    }
  ]
}</code></pre>



<p class="wp-block-paragraph">The trade-off is real and you should know it before you apply this. If you later adopt a geographic profile whose destination list includes a Region you denied, invocations fail with an access error that looks nothing like a routing problem. Keep the approved Region list and the profile&#8217;s destination list in the same review, not in two different tickets.</p>



<h2 class="wp-block-heading">Boundary four: retrieval, and the grounding check that will not save you</h2>



<p class="wp-block-paragraph">A useful medical assistant is retrieval-augmented. The model alone is a fluent generalist; the value comes from grounding answers in a curated corpus, whether that is institutional guidelines, formulary rules or a licensed reference set. Bedrock Knowledge Bases will manage the ingestion and the vector store for you, or you can run your own index in Amazon OpenSearch Service.</p>



<p class="wp-block-paragraph">Two things bite here.</p>



<p class="wp-block-paragraph"><strong>Tenant isolation in the index.</strong> If you serve more than one hospital, department or study cohort, a shared index with a filter applied in application code is one refactor away from cross-tenant retrieval. Filters are easy to forget and impossible to notice, because a wrong answer that cites a real document looks exactly like a right one. Separate indexes per tenant, or fine-grained access control enforced below your application, cost more and fail safe.</p>



<p class="wp-block-paragraph"><strong>The grounding check has a scope limit.</strong> Guardrails contextual grounding checks detect responses that are not supported by the source material or not relevant to the question. Excellent feature. The documentation states the supported use cases are summarisation, paraphrasing and question answering, and that conversational chatbot use cases are not supported. If your product is a chat interface with turn history, do not assume this check is covering you. There is also a streaming wrinkle: relevance is assessed per chunk, so an irrelevant response can reach the user before it is marked irrelevant at the end of the stream.</p>



<p class="wp-block-paragraph">What you can rely on regardless of interface is applying a guardrail to arbitrary text directly, which is how you check an input before it enters your own pipeline:</p>



<pre class="wp-block-code"><code>import boto3

client = boto3.client("bedrock-runtime", region_name="us-east-1")

response = client.apply_guardrail(
    guardrailIdentifier=GUARDRAIL_ID,
    guardrailVersion="DRAFT",
    source="INPUT",
    content=[{"text": {"text": user_message}}],
)

if response["action"] == "GUARDRAIL_INTERVENED":
    # Stop here. Do not forward, and do not write the raw text anywhere.
    handle_blocked(response["assessments"])</code></pre>



<p class="wp-block-paragraph">Note what the comment is doing. The most common bug in this pattern is catching the intervention and then logging the offending input &#8220;for debugging&#8221;, which reintroduces exactly the leak the guardrail just prevented. Log the assessment, log a request identifier, never the text.</p>



<p class="wp-block-paragraph">Sensitive information filters give you two handling modes: block, which rejects the request outright, and mask, which replaces detected entities with placeholder tags. For a study tool, mask is usually right, because a student who typed a name by reflex gets a useful answer instead of a wall. For a clinical assistant working over charts, blocking on unexpected identifiers in an input that should have arrived structured is a better signal that something upstream is wrong.</p>



<h2 class="wp-block-heading">Boundary five: being confidently wrong</h2>



<p class="wp-block-paragraph">Everything above is about data leaving. This one is about a bad answer arriving, and for a medical audience it is the reputational failure that actually ends products.</p>



<ul class="wp-block-list">
<li><strong>Cite or refuse.</strong> Return the retrieved passages alongside the answer. If retrieval returned nothing above your relevance threshold, say so rather than letting the model answer from parametric memory. A student cannot verify what they cannot see.</li>

<li><strong>Version the corpus, not just the model.</strong> When a guideline changes, you need to know which answers were generated against the old text. Store a corpus revision identifier with every logged response.</li>

<li><strong>Use denied topics deliberately.</strong> Dosing for a named patient, and anything that reads as a treatment directive rather than reference information, are reasonable things to route to a refusal with a clear explanation.</li>

<li><strong>Get regulatory advice early.</strong> Whether clinical decision support software is regulated as a medical device depends on your jurisdiction and, critically, on the claims you make about it. This is a legal question with engineering consequences, and it is much cheaper to answer before the interface exists.</li>
</ul>



<h2 class="wp-block-heading">A build order for a secure AI medical assistant on AWS</h2>



<p class="wp-block-paragraph">Sequence matters here more than in most builds, because several of these are painful to retrofit.</p>



<ol class="wp-block-list">
<li>Accept the BAA through AWS Artifact and designate the account for HIPAA use. Do this before the first prototype, not before the first customer.</li>

<li>Pin Regions with a Service Control Policy, and decide the geographic-versus-global inference profile question in writing.</li>

<li>Configure invocation logging deliberately, with a customer-managed KMS key, an explicit retention period, and each data-type delivery flag chosen rather than defaulted.</li>

<li>Put the application in private subnets and reach Bedrock over VPC endpoints so PHI-bearing traffic does not traverse the public internet.</li>

<li>Build the guardrail before the prompt. Sensitive information filters, denied topics, and grounding checks where they apply.</li>

<li>Add Comprehend Medical detection in the ingestion path for anything you are storing, and at the input boundary for anything a user types.</li>

<li>Build retrieval with tenant isolation from the first index, not the second.</li>

<li>Only now write the assistant&#8217;s prompt and interface, and put a WAF such as Cloudflare or AWS WAF in front of the public endpoint.</li>
</ol>



<h2 class="wp-block-heading">Troubleshooting</h2>



<ul class="wp-block-list">
<li><strong>AccessDenied on a model that clearly works elsewhere.</strong> Usually one of three things: model access not requested in this Region, a cross-Region profile whose destination Regions your SCP denies, or an IAM policy that grants the model in the source Region only.</li>

<li><strong>No invocation logs appearing.</strong> The configuration is per Region and disabled by default. Confirm you queried the same Region the application calls, and that the delivery flag for the data type you expect is enabled.</li>

<li><strong>Guardrail passes text you expected it to catch.</strong> Sensitive information detection is probabilistic and context-dependent. Very short inputs give it little to work with. Test with realistic clinical phrasing, not single tokens, and add regex patterns for structured identifiers like MRNs that follow a house format.</li>

<li><strong>Grounding check appears to do nothing.</strong> Check your interface shape against the supported use cases before assuming it is misconfigured, and check whether streaming is masking the result until the response completes.</li>

<li><strong>Retrieval returns plausible but wrong documents.</strong> Look at chunking before you look at the model. Clinical guidelines chunked mid-table or mid-criteria retrieve badly no matter what embedding you use.</li>
</ul>



<h2 class="wp-block-heading">Common mistakes</h2>



<ul class="wp-block-list">
<li>Prototyping in a personal or sandbox account and promising to migrate later. The prototype is where the first real chart gets pasted.</li>

<li>Assuming the eligibility of Bedrock covers the whole stack. Every component that touches PHI needs to be eligible, including the vector store, the queue and the cache.</li>

<li>Treating PHI detection as de-identification. AWS says explicitly that it is not.</li>

<li>Leaving image and embedding log delivery enabled by copy-paste.</li>

<li>Logging blocked prompts to debug the guardrail.</li>

<li>Shipping a chat interface and citing contextual grounding as the hallucination control.</li>

<li>Filtering tenants in application code over a shared index.</li>
</ul>



<h2 class="wp-block-heading">Best practices</h2>



<ul class="wp-block-list">
<li>Write the data-flow diagram with logs, backups and the vector index drawn as first-class destinations. If PHI can land there, it is on the diagram.</li>

<li>Define everything in Terraform or OpenTofu so the guardrail, the logging configuration and the SCP are reviewable artefacts rather than console state.</li>

<li>Keep a small evaluation set of realistic clinical questions with known-good answers, and run it on every prompt or model change.</li>

<li>Alarm on guardrail intervention rate. A sudden rise usually means a change upstream, not a change in users.</li>

<li>Dashboard invocation counts, latency and intervention rates somewhere your on-call actually looks, whether that is CloudWatch, Grafana or Datadog.</li>

<li>Scope IAM to specific model ARNs and specific guardrail identifiers. A wildcard on <code>bedrock:InvokeModel</code> means any model, including ones you never evaluated.</li>

<li>Rehearse the breach path once. Knowing which bucket, which log group and which key you would need to reason about is worth an afternoon.</li>
</ul>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Frequently asked questions</h2>



<h3 class="wp-block-heading">Is Amazon Bedrock HIPAA compliant?</h3>



<p class="wp-block-paragraph">Bedrock is HIPAA-eligible, which is a different claim. Eligibility means AWS permits you to process PHI with the service under an executed BAA. Compliance is a property of your deployment: your account designation, encryption, access control, network isolation, logging and retention. A Bedrock workload with a signed BAA and public endpoints and unbounded plaintext logs is not compliant.</p>



<h3 class="wp-block-heading">Do I need a BAA if the tool is only for medical students?</h3>



<p class="wp-block-paragraph">If the tool genuinely never receives PHI, HIPAA is not engaged. The practical question is whether you can guarantee that, given that your users have clinical access and a free-text box. If you cannot, get the BAA. It costs you a self-service acceptance in AWS Artifact and removes an entire category of incident.</p>



<h3 class="wp-block-heading">Does AWS use my prompts to train the models?</h3>



<p class="wp-block-paragraph">AWS states that customer content submitted to Bedrock is not used to train the underlying foundation models or shared with model providers. That statement is the sort of thing a hospital security review will want quoted verbatim from the current AWS data protection documentation rather than from a blog, so pull the live wording when you write your assessment.</p>



<h3 class="wp-block-heading">Is Comprehend Medical enough to de-identify clinical text?</h3>



<p class="wp-block-paragraph">No. AWS documents that it may not identify PHI in all circumstances and does not by itself meet HIPAA&#8217;s de-identification requirements. Use it as a detection layer and a redaction aid inside a compliant environment. Formal de-identification, whether by the Safe Harbor method or expert determination, is a separate exercise with its own review.</p>



<h3 class="wp-block-heading">Where should the vector store live?</h3>



<p class="wp-block-paragraph">In the same account and Region as the rest of the workload, on a HIPAA-eligible service, encrypted with a customer-managed key, reachable only from private subnets. Embeddings derived from clinical text are not sanitised data and should not be treated as a lower-sensitivity artefact than the source.</p>



<h3 class="wp-block-heading">Should I use a global or geographic inference profile?</h3>



<p class="wp-block-paragraph">Geographic if you have any residency commitment, contractual or regulatory. Global if you have none and want the throughput and the lower per-token rate. Decide once, document the reasoning, and enforce it with a Service Control Policy rather than a configuration constant.</p>



<h3 class="wp-block-heading">Does an AI medical assistant count as a medical device?</h3>



<p class="wp-block-paragraph">It depends on your jurisdiction and on what you claim the software does. Software that surfaces reference information a clinician independently reviews has generally been treated differently from software that directs a clinical decision, but the boundary is fact-specific and moves. This is a question for regulatory counsel before launch, not a question for your architecture diagram.</p>



<h2 class="wp-block-heading">The one thing worth remembering</h2>



<p class="wp-block-paragraph">A secure AI medical assistant on AWS is not mainly a model problem. Bedrock, Guardrails and Comprehend Medical are the easy part, and the documentation for them is good. The hard part is that PHI leaves through the paths you did not draw: an invocation log you enabled for good reasons, a global inference profile that routes wherever capacity exists, an index shared between tenants, a debug line added at two in the morning.</p>



<p class="wp-block-paragraph">So build the boundary first and the assistant second. Get the BAA accepted, pin the Regions in policy, decide consciously what your logs are allowed to hold, and isolate retrieval per tenant before there is a second tenant. Every one of those is cheap on day one and expensive in month six.</p>



<h2 class="wp-block-heading">Need a second pair of eyes on your build?</h2>



<p class="wp-block-paragraph">I work with teams building AI on AWS where the data is sensitive and the failure modes are quiet. Things I can help with on a project like this:</p>



<ul class="wp-block-list">
<li>Reviewing a Bedrock architecture against the boundaries above and telling you where PHI can actually land</li>

<li>Setting up account separation, BAA scope, Region pinning with Service Control Policies, and VPC endpoint access to Bedrock</li>

<li>Designing and tuning Guardrails policies, including custom regex for house identifier formats, and the block-versus-mask decision per surface</li>

<li>Building the retrieval layer with per-tenant isolation, sensible clinical chunking, and citation-or-refuse behaviour</li>

<li>Auditing invocation logging, KMS key policies, retention and application-level log hygiene for accidental PHI capture</li>

<li>Putting the whole thing in Terraform or OpenTofu so your controls are reviewable instead of remembered</li>
</ul>



<p class="wp-block-paragraph">If any of that is on your plate, send me the piece you are least sure about. A redacted architecture diagram, a guardrail configuration, a logging policy, an <code>AccessDenied</code> you cannot explain. I would rather look at the real thing than talk in generalities.</p>



<div class="wp-block-buttons is-layout-flex wp-block-buttons-is-layout-flex">
<div class="wp-block-button"><a class="wp-block-button__link wp-element-button" href="https://www.upwork.com/freelancers/~01f15a912ad84a6620" target="_blank" rel="noreferrer noopener">Work with me on Upwork</a></div>
</div>
<p>The post <a href="https://john-nessime.com/blog/case-studies/secure-ai-medical-assistant-aws/">Build a Secure AI Medical Assistant on AWS: The Boundaries That Actually Leak</a> appeared first on <a href="https://john-nessime.com/blog">John Nessime</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>HIPAA Compliance on AWS: The Gaps That Pass Every Security Check</title>
		<link>https://john-nessime.com/blog/technical-guides/hipaa-compliance-aws/</link>
		
		<dc:creator><![CDATA[John Nessime]]></dc:creator>
		<pubDate>Sat, 25 Jul 2026 12:04:00 +0000</pubDate>
				<category><![CDATA[Case Studies]]></category>
		<category><![CDATA[Cloud Computing]]></category>
		<category><![CDATA[Technical Guides]]></category>
		<category><![CDATA[Web Security]]></category>
		<category><![CDATA[Amazon S3]]></category>
		<category><![CDATA[Architecture]]></category>
		<category><![CDATA[Audit Logging]]></category>
		<category><![CDATA[AWS]]></category>
		<category><![CDATA[AWS Config]]></category>
		<category><![CDATA[AWS KMS]]></category>
		<category><![CDATA[AWS Organizations]]></category>
		<category><![CDATA[Cloud]]></category>
		<category><![CDATA[Cloud Security]]></category>
		<category><![CDATA[CloudWatch]]></category>
		<category><![CDATA[Compliance]]></category>
		<category><![CDATA[Data Residency]]></category>
		<category><![CDATA[Encryption]]></category>
		<category><![CDATA[Healthcare Cloud]]></category>
		<category><![CDATA[HIPAA]]></category>
		<category><![CDATA[IAM]]></category>
		<category><![CDATA[Infrastructure]]></category>
		<category><![CDATA[Log Retention]]></category>
		<category><![CDATA[Logging]]></category>
		<category><![CDATA[PHI]]></category>
		<category><![CDATA[Restore Testing]]></category>
		<category><![CDATA[VPC]]></category>
		<guid isPermaLink="false">https://john-nessime.com/blog/?p=184</guid>

					<description><![CDATA[<p>A working engineer's guide to HIPAA compliance on AWS, organised by the gap between the control you configured and the obligation you actually carry. Covers BAA account scope, the eligible services list as a contract boundary, KMS key policy versus the encryption checkbox, what "six years" really applies to, backup and restore scope, and the subprocessor chain nobody inventories.</p>
<p>The post <a href="https://john-nessime.com/blog/technical-guides/hipaa-compliance-aws/">HIPAA Compliance on AWS: The Gaps That Pass Every Security Check</a> appeared first on <a href="https://john-nessime.com/blog">John Nessime</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">The ticket usually reads something like: &#8220;Legal wants to know if the analytics account is in scope.&#8221; So you open Security Hub. Green. You check the Config rules. Passing. Every bucket is encrypted, every volume is encrypted, MFA is on, CloudTrail is running in all Regions. You reply that the account is fine.</p>



<p class="wp-block-paragraph">Then someone points out that a nightly job has been copying a de-identified extract into that account for eight months, the de-identification script never removed the admission dates, and the account was spun up before anyone thought about the Business Associate Addendum. Nothing was misconfigured. Every control you built worked exactly as designed. And you have been out of compliance the entire time.</p>



<p class="wp-block-paragraph">That is the shape of most real failures here. Not a breach, not a misconfiguration, but a mismatch between the boundary your tooling checks and the boundary your obligation actually follows. This post covers HIPAA compliance on AWS organised by those gaps: where the contract stops, where encryption stops being a control, what &#8220;six years&#8221; genuinely applies to, and which parts of the estate people forget are in scope at all.</p>



<h2 class="wp-block-heading">Eligible is not compliant, and the difference is the whole job</h2>



<p class="wp-block-paragraph">AWS does not sell HIPAA compliance. It sells HIPAA <em>eligible</em> services, which is a genuinely different thing. Eligible means AWS has built the service so it can lawfully handle electronic protected health information and has agreed to cover it under a Business Associate Addendum. Compliant describes an entire system: your architecture, your key management, your access reviews, your policies, your staff, your vendors.</p>



<p class="wp-block-paragraph">Under the shared responsibility model, AWS secures the infrastructure. You secure everything you build on it. Nothing about signing the BAA transfers a single obligation off your side of the line. An unencrypted RDS instance, an overly broad IAM policy or an application that logs a patient identifier into stdout is your problem in exactly the same way it would be in a rack you own.</p>



<p class="wp-block-paragraph">People know this in the abstract. Where it bites is in the specifics below.</p>



<h2 class="wp-block-heading">Gap one: the BAA is a contract boundary, and nothing enforces it</h2>



<p class="wp-block-paragraph">This is the one I would fix first, because it is invisible to every security tool you own.</p>



<p class="wp-block-paragraph">The AWS BAA is self-service through AWS Artifact, at no extra cost. You can accept it for a single account, or, if you are in the management account of an AWS Organization, accept it once so that existing and future member accounts are covered. That organization-level option is the one worth using, because the per-account version quietly rots: someone creates a new account for a proof of concept, nobody repeats the Artifact step, and six months later that account is running something real.</p>



<p class="wp-block-paragraph">The second half of the boundary is the HIPAA Eligible Services Reference that AWS publishes. Only services on that list may create, receive, process, maintain or transmit ePHI under the BAA. The list is long, it changes, and some entries carry carve-outs where the service is eligible but a specific feature is not. Reading a service name on the list and assuming every feature inside it is covered is the kind of mistake that only surfaces during an audit.</p>



<p class="wp-block-paragraph">Here is the part worth internalising: <strong>there is no AWS control that stops you putting PHI into a non-eligible service.</strong> No API error, no Config rule out of the box, no GuardDuty finding. The eligible services list is a contractual construct. Your infrastructure has no idea it exists.</p>



<h3 class="wp-block-heading">Turning a contract boundary into a technical one</h3>



<p class="wp-block-paragraph">The mechanism that actually helps is Service Control Policies on the organizational unit that holds your PHI accounts. SCPs set the ceiling on what any principal in those accounts can do, including the root user, so they work as a guardrail rather than a suggestion.</p>



<p class="wp-block-paragraph">Start with the easy one. Pin the accounts to the Regions you have actually assessed, because data residency assumptions fall apart the moment someone launches something in a Region you never reviewed:</p>



<pre class="wp-block-code"><code>{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "DenyUnapprovedRegions",
      "Effect": "Deny",
      "NotAction": [
        "iam:*",
        "organizations:*",
        "route53:*",
        "cloudfront:*",
        "support:*",
        "sts:*"
      ],
      "Resource": "*",
      "Condition": {
        "StringNotEquals": {
          "aws:RequestedRegion": ["us-east-1", "us-west-2"]
        }
      }
    }
  ]
}</code></pre>



<p class="wp-block-paragraph">The <code>NotAction</code> list matters. Global services are backed by endpoints in specific Regions, so denying them wholesale by Region locks you out of IAM and breaks Route 53 and CloudFront. Those entries are exemptions, not an allow-list.</p>



<p class="wp-block-paragraph">The harder one is restricting which services can be used at all. The same <code>NotAction</code> pattern works, with the services you have approved for PHI listed as the exemptions and everything else denied. It is effective and it is blunt: every new service anyone wants becomes a change request against the policy, and if you forget a dependency you find out through a failure in production. I would only reach for it on a dedicated PHI OU where the workload is well understood, not across a general-purpose organization.</p>



<p class="wp-block-paragraph">Whichever route you take, write down the approved service list somewhere a human reviews on a schedule, and diff it against the AWS reference periodically. That review is itself a compliance artefact.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Gap two: encryption is a checkbox, the key policy is the control</h2>



<p class="wp-block-paragraph">Almost every guide to HIPAA compliance on AWS tells you to encrypt at rest and in transit. Almost none of them explain why it is worth doing properly rather than minimally, so teams enable default encryption with an AWS-managed key, watch the Config rule turn green, and move on.</p>



<p class="wp-block-paragraph">The reason to care is the Breach Notification Rule. It applies to <em>unsecured</em> PHI, meaning PHI that has not been rendered unusable, unreadable or indecipherable through a method HHS has specified. HHS guidance points at NIST-validated encryption. If PHI is encrypted to that standard and the decryption keys were not compromised alongside it, an incident involving that data generally does not trigger the notification machinery at all. No individual letters, no HHS portal submission, no press release for a large incident.</p>



<p class="wp-block-paragraph">Read that second condition again, because it is where the architecture decision lives. The safe harbour depends on the keys not being compromised with the data. If your encryption key is one an attacker inherits automatically the moment they compromise a role in the account, you have encryption but you may not have the argument.</p>



<h3 class="wp-block-heading">What that means in practice</h3>



<ul class="wp-block-list">
<li>Use customer managed KMS keys for anything holding PHI, not AWS-managed keys. Only a customer managed key gives you a key policy you can write, and only a key policy lets you deny decryption independently of the resource policy.</li>

<li>Separate the key administrators from the key users. The people who can schedule deletion of a key should not be the people whose application role uses it every second.</li>

<li>Use a distinct key per data domain rather than one key for the whole account. Blast radius and audit trail both improve, and you get the ability to revoke access to one dataset without touching another.</li>

<li>Constrain key usage with the <code>kms:ViaService</code> condition so a key that exists to encrypt RDS storage cannot be used to decrypt something a role dragged into Lambda.</li>

<li>Turn on key rotation and leave it on. It costs nothing operationally and it is the kind of thing an assessor asks about by reflex.</li>
</ul>



<p class="wp-block-paragraph">Pull the current key policy before you assume it says what you think:</p>



<pre class="wp-block-code"><code>aws kms get-key-policy 
  --key-id alias/phi-rds 
  --policy-name default 
  --output text

# Find storage that slipped through unencrypted
aws ec2 describe-volumes 
  --filters Name=encrypted,Values=false 
  --query 'Volumes[].{Id:VolumeId,AZ:AvailabilityZone}' 
  --output table

aws rds describe-db-instances 
  --query 'DBInstances[?StorageEncrypted==`false`].DBInstanceIdentifier' 
  --output text</code></pre>



<p class="wp-block-paragraph">The RDS query is the important one, because RDS encryption cannot be enabled in place. If that command returns anything, the fix is a snapshot, an encrypted copy of the snapshot, a restore, and a cutover. Plan for downtime or a replication strategy. This is the single most common &#8220;we will fix it later&#8221; item I see, and later gets expensive.</p>



<p class="wp-block-paragraph">Also switch on EBS encryption by default in every Region you use, so the next instance somebody launches from a console wizard is not a new exception:</p>



<pre class="wp-block-code"><code>aws ec2 enable-ebs-encryption-by-default --region us-east-1
aws ec2 get-ebs-encryption-by-default --region us-east-1</code></pre>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Gap three: you have logs, but you may not have evidence</h2>



<p class="wp-block-paragraph">The Security Rule requires audit controls: mechanisms that record and examine activity in systems containing ePHI. It also requires you to regularly review records of information system activity. Both of those are about having and using the records.</p>



<p class="wp-block-paragraph">Now the correction, because this one is repeated everywhere and it is wrong in a way that costs money. You will read that HIPAA requires six years of audit logs. It does not. The six-year requirement sits in the documentation standard, and it applies to the policies, procedures and records of actions, activities and assessments that the Security Rule requires you to keep, retained for six years from creation or from when the document was last in effect, whichever is later. There is no clause anywhere in the Security Rule that names a retention period for CloudTrail events.</p>



<p class="wp-block-paragraph">What this actually means is more demanding, not less. You have to <em>decide</em> your log retention period, write it into a policy, justify it against your risk analysis, and then keep that policy for six years. And an assessor will hold you to the number you wrote. Setting a CloudWatch Logs retention of thirty days while your policy claims one year is a finding. Storing seven years of everything because a blog told you to, when your policy says two, is not compliance, it is just a bill.</p>



<p class="wp-block-paragraph">So: pick a period you can defend, make the infrastructure match it exactly, and treat any gap between policy and configuration as a defect.</p>



<h3 class="wp-block-heading">Making logs into evidence</h3>



<p class="wp-block-paragraph">Retention is only half of it. The other half is being able to show that the records were not altered. CloudTrail has log file validation for exactly this, and it is off unless you turn it on:</p>



<pre class="wp-block-code"><code>aws cloudtrail update-trail 
  --name org-phi-trail 
  --enable-log-file-validation

# Later, prove a window of logs is intact
aws cloudtrail validate-logs 
  --trail-arn arn:aws:cloudtrail:us-east-1:111122223333:trail/org-phi-trail 
  --start-time "$(date -u -d '90 days ago' +%Y-%m-%dT%H:%M:%SZ)"</code></pre>



<p class="wp-block-paragraph">With validation enabled, CloudTrail writes signed digest files alongside the log files, and <code>validate-logs</code> checks them. The difference between &#8220;here are our logs&#8221; and &#8220;here are our logs, and here is a cryptographic check that nothing was modified or deleted&#8221; is the difference between an assertion and evidence.</p>



<p class="wp-block-paragraph">Put the archive bucket in a separate account that the workload accounts cannot write to or delete from, and apply S3 Object Lock in compliance mode for the retention window you committed to. Object Lock in compliance mode cannot be shortened or bypassed by anyone, including the root user, which is exactly the property you want and exactly the property that will hurt if you set the period carelessly. Test it in governance mode first.</p>



<p class="wp-block-paragraph">For the review obligation, a query interface matters more than raw storage. Athena over the CloudTrail bucket is the cheap default. If you want alerting and dashboards on top of access patterns, this is a natural place for a platform such as Grafana, Datadog or Splunk, and any of them will hold access records for you. Just remember that if those records contain PHI, that vendor needs a BAA too. See the subprocessor section below.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Gap four: backups, snapshots and the parts of scope people forget</h2>



<p class="wp-block-paragraph">The Security Rule&#8217;s contingency plan standard is not optional decoration. It requires a data backup plan, a disaster recovery plan and an emergency mode operation plan, plus testing and revision procedures. Most teams have the backups. Far fewer have the tested restore, and the tested restore is the part that gets asked about.</p>



<p class="wp-block-paragraph">Three things routinely go wrong here.</p>



<ol class="wp-block-list">
<li><strong>Copies leave the boundary.</strong> A cross-Region snapshot copy lands in a Region you did not assess. A cross-account copy for the DR account lands somewhere outside the OU your SCPs protect. The data is still PHI. The controls did not travel with it.</li>

<li><strong>Re-encryption changes the key, not just the copy.</strong> Copying an encrypted snapshot to another account requires a key the destination can use. It is easy to end up with a shared or less restrictive key protecting your backups than protects production, which inverts the risk model.</li>

<li><strong>The restore is never rehearsed.</strong> A backup you have never restored is a hypothesis. Schedule a restore into an isolated account, record the elapsed time, and file the result. That record is your evidence for the testing requirement, and it is the single easiest compliance artefact to produce for free.</li>
</ol>



<p class="wp-block-paragraph">While you are inventorying, remember the places PHI ends up without anyone deciding it should: application logs that include request bodies, database slow query logs capturing parameter values, support tickets with screenshots attached, CSV extracts in an analyst&#8217;s bucket, and non-production environments seeded from a production dump. That last one is the classic. If your staging database is a copy of production, staging is in scope, and staging is almost never built to the same standard.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Gap five: the business associate chain does not stop at AWS</h2>



<p class="wp-block-paragraph">Your BAA with AWS covers AWS. It covers nothing else in your stack.</p>



<p class="wp-block-paragraph">Every vendor that can create, receive, maintain or transmit PHI on your behalf is a business associate and needs an agreement. In a typical AWS estate that means the error tracker holding stack traces, the log aggregation platform, the APM tool, the transactional email provider, the customer support desk, the CI system if it ever touches a production dataset, and any AI or analytics service you have wired in.</p>



<p class="wp-block-paragraph">Build the inventory as a table with three columns: vendor, what PHI it can see, and whether a signed agreement exists. The third column is usually where the surprises are. Some vendors sign readily, some only on higher-priced tiers, and some decline entirely, at which point you have an architecture decision rather than a procurement one.</p>



<p class="wp-block-paragraph">One structural move that reduces this surface considerably: keep everything that does not need PHI out of the PHI accounts entirely. Your marketing site, your docs, your status page and your public API gateway for non-clinical traffic do not belong in a regulated account. Running them on ordinary infrastructure, whether that is a separate AWS account, a straightforward VPS from a host like InterServer, or a static site behind Cloudflare, shrinks the estate you have to assess, evidence and defend. Fewer things in scope is the cheapest compliance win available.</p>



<p class="wp-block-paragraph">For tracking the paperwork side, compliance automation platforms such as Vanta, Drata or Secureframe pull evidence from AWS on a schedule and keep the vendor register current. They are genuinely useful for the collection and reminder burden. They do not design your architecture, and I have seen teams treat a green dashboard in one of those tools as though it were an assessment. It is not. It is a checklist that knows what you told it.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">What is changing, and why &#8220;addressable&#8221; is a bad thing to build on</h2>



<p class="wp-block-paragraph">Since it was adopted, the Security Rule has split implementation specifications into <em>required</em> and <em>addressable</em>. Addressable never meant optional. It meant you assess whether the specification is reasonable and appropriate, and if not, you implement an equivalent alternative or document why neither is necessary. In practice, a lot of organisations turned the documented justification into the deliverable and skipped the control.</p>



<p class="wp-block-paragraph">HHS published a Notice of Proposed Rulemaking in the Federal Register in January 2025 that would remove that distinction, making implementation specifications required with limited exceptions, and would explicitly require encryption of ePHI at rest and in transit and multi-factor authentication, again with limited exceptions. The comment period closed in March 2025.</p>



<p class="wp-block-paragraph">Be precise about the status, because a lot of vendor content is not: <strong>this is a proposed rule and it is not final.</strong> The expected timeline for final action has slipped more than once, and the requirements could still change or be withdrawn. Nobody should be telling you a compliance deadline as though it were settled.</p>



<p class="wp-block-paragraph">What is worth taking from it is the direction of travel. If your current position depends on having documented that encryption or MFA was not reasonable and appropriate, that position is fragile regardless of what the final rule says. On AWS specifically, encryption at rest and MFA are both cheap and both already best practice. Building the architecture on an addressable deferral is an unforced risk.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Troubleshooting the findings you will actually hit</h2>



<h3 class="wp-block-heading">&#8220;An assessor asked which accounts are in BAA scope and nobody could answer&#8221;</h3>



<p class="wp-block-paragraph">Sign in to AWS Artifact from the management account and check the organization agreements tab to see whether the BAA was accepted at the organization level or per account. If it is per account, list your accounts, work out which hold PHI, and confirm each one individually. Then move to the organization-level agreement so this question has one answer forever.</p>



<h3 class="wp-block-heading">&#8220;Config says the bucket is encrypted but we cannot prove who read the objects&#8221;</h3>



<p class="wp-block-paragraph">Bucket encryption and object-level access logging are unrelated. CloudTrail management events do not record S3 object reads by default. You need CloudTrail data events for that bucket, or S3 server access logging, and both cost money proportional to request volume. Enable data events selectively on the buckets that hold PHI rather than account-wide.</p>



<h3 class="wp-block-heading">&#8220;We enabled an SCP and production broke&#8221;</h3>



<p class="wp-block-paragraph">Almost always a Region deny catching a global service endpoint, or a service allow-list missing a dependency the workload calls indirectly. Check CloudTrail for <code>AccessDenied</code> events with an explicit deny from an SCP, and look at the service name in the event rather than the one you expected. Attach new SCPs to a test OU with a representative workload before the PHI OU.</p>



<h3 class="wp-block-heading">&#8220;Snapshot copy to the DR account fails with a KMS error&#8221;</h3>



<p class="wp-block-paragraph">The destination account cannot use the source key. The source key policy has to grant the destination principal permission to use it, and the copy has to specify a key the destination can decrypt with. Fix it by granting explicitly on a key you control, not by falling back to an AWS-managed key, which is the tempting shortcut and gives up the key policy control you needed.</p>



<h3 class="wp-block-heading">&#8220;CloudWatch Logs retention was never set&#8221;</h3>



<p class="wp-block-paragraph">New log groups default to never expiring, which is both a cost problem and a policy mismatch. Audit them with <code>aws logs describe-log-groups</code> and look for groups with no <code>retentionInDays</code> value, then set the period your policy specifies with <code>aws logs put-retention-policy</code>.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Common mistakes</h2>



<ul class="wp-block-list">
<li>Treating the signed BAA as the finish line rather than the prerequisite. It is the thing you need before the first byte of PHI arrives, not evidence that anything is configured correctly.</li>

<li>Assuming a service is fully eligible because its name appears on the list, without reading the feature-level carve-outs next to it.</li>

<li>Quoting &#8220;six years&#8221; as a log retention requirement, then either overspending on storage or writing a policy that contradicts the actual configuration.</li>

<li>Using AWS-managed KMS keys for PHI, which leaves no key policy to write and no independent revocation path.</li>

<li>Seeding staging or test environments from production data and then holding those environments to a lower standard.</li>

<li>Forgetting that the risk analysis is a required, recurring, documented activity, not a one-off spreadsheet from the year you launched.</li>

<li>Signing a BAA with AWS and none of the ten other vendors that can see the same data.</li>

<li>Letting a compliance automation dashboard stand in for an architecture review.</li>
</ul>



<h2 class="wp-block-heading">Best practices worth the effort</h2>



<ul class="wp-block-list">
<li><strong>Isolate PHI in its own accounts and its own OU.</strong> Account boundaries are the strongest isolation AWS offers, and they make the scope question answerable in one sentence.</li>

<li><strong>Accept the BAA at the organization level.</strong> It removes an ongoing manual step that fails silently.</li>

<li><strong>Customer managed keys, one per data domain, with split admin and usage roles.</strong> This is what makes the breach safe harbour argument defensible rather than theoretical.</li>

<li><strong>Write the retention period down first, configure second.</strong> Policy and infrastructure should agree exactly, in both directions.</li>

<li><strong>Ship audit logs to a separate account with Object Lock and CloudTrail validation enabled.</strong> Immutability and integrity are what turn logs into evidence.</li>

<li><strong>Keep PHI out of everything that does not need it.</strong> De-identify early, tokenise where you can, and route non-clinical traffic through infrastructure that is not in scope.</li>

<li><strong>Define everything in Terraform or OpenTofu.</strong> A reviewable, version-controlled definition of your controls is worth more to an assessor than any screenshot, and it stops drift being invisible.</li>

<li><strong>Rehearse the restore and the breach response.</strong> Both are required, both are tested by asking for the record, and both are cheap to evidence if you actually do them.</li>
</ul>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Frequently asked questions</h2>



<h3 class="wp-block-heading">Is AWS HIPAA compliant?</h3>



<p class="wp-block-paragraph">Not on its own, and the phrasing is the problem. AWS offers HIPAA-eligible services and will sign a Business Associate Addendum, which means you can build a compliant system on it. Compliance is a property of your whole environment, including configuration, policies, vendors and staff. No provider can sell it to you as a finished product.</p>



<h3 class="wp-block-heading">How do I sign a BAA with AWS?</h3>



<p class="wp-block-paragraph">Through AWS Artifact in the console. It is self-service and there is no additional charge. Accept it for an individual account under account agreements, or from the management account of an AWS Organization under organization agreements so all current and future member accounts are covered. It should be accepted by someone with authority to bind your organisation, and it must be in place before any PHI reaches AWS.</p>



<h3 class="wp-block-heading">Does HIPAA require six years of CloudTrail logs?</h3>



<p class="wp-block-paragraph">No. The six-year requirement is a documentation retention rule covering the policies, procedures and records the Security Rule requires you to maintain, kept for six years from creation or from when they were last in effect. The audit controls standard requires the mechanism to record and examine activity but does not name a retention period for the logs themselves. You set that period in your own policy, justify it, and make the configuration match.</p>



<h3 class="wp-block-heading">Which AWS services can I use with PHI?</h3>



<p class="wp-block-paragraph">Only those on the AWS HIPAA Eligible Services Reference, and only in accounts covered by your BAA. Check the list before adopting anything new, read the feature-level exclusions noted against individual services, and re-check periodically because entries are added over time. Nothing in AWS will stop you using a non-eligible service with PHI, so this has to be an explicit process on your side.</p>



<h3 class="wp-block-heading">If encrypted PHI is exposed, do I still have to report a breach?</h3>



<p class="wp-block-paragraph">Generally no, provided the encryption meets the standard in HHS guidance and the decryption keys were not compromised along with the data. The Breach Notification Rule applies to unsecured PHI, and properly encrypted data does not meet that definition. This is why key management, not just enabling encryption, is the part that determines whether the protection is real. You still document the incident and the assessment.</p>



<h3 class="wp-block-heading">Are the new HIPAA Security Rule requirements in force?</h3>



<p class="wp-block-paragraph">Not at the time of writing. The proposals to make all implementation specifications required and to mandate encryption and multi-factor authentication came from a Notice of Proposed Rulemaking published in January 2025. The comment period has closed, but no final rule has been issued and the timeline has moved. Treat any specific compliance deadline you see quoted with suspicion and check the current status directly.</p>



<h3 class="wp-block-heading">Is a HIPAA-compliant AWS environment expensive to run?</h3>



<p class="wp-block-paragraph">The controls themselves are mostly cheap. KMS keys, CloudTrail validation, Config rules and account separation cost very little. The real costs are log storage volume, CloudTrail data events on busy buckets, running non-production environments to production standard, and staff time on risk analysis and evidence collection. Reducing what is in scope is the most effective cost lever, which is another reason to keep non-clinical workloads out of the regulated accounts.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">The one thing to take away</h2>



<p class="wp-block-paragraph">HIPAA compliance on AWS fails at the seams, not at the controls. Your encryption works. Your IAM policies are tight. What goes wrong is that the obligation follows the data into an account nobody added to the BAA, a Region nobody assessed, a staging database seeded from production, a vendor nobody signed an agreement with, or a retention period nobody wrote down.</p>



<p class="wp-block-paragraph">So build the boundary technically rather than trusting it contractually. Isolate PHI into its own accounts, wrap those accounts in guardrails that make the contract boundary enforceable, own your keys so the encryption means something legally, and make your logs provable rather than merely present. Then write the whole thing down, because in this domain the documentation genuinely is part of the control.</p>



<p class="wp-block-paragraph">None of that is exotic engineering. It is ordinary AWS work applied to a boundary that no dashboard draws for you.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Working on a healthcare workload on AWS?</h2>



<p class="wp-block-paragraph">This is the kind of work I do. If you are building or inheriting a PHI environment on AWS, I can help with:</p>



<ul class="wp-block-list">
<li><strong>Scope and boundary review:</strong> mapping which accounts, Regions, services and vendors actually touch PHI, and finding the ones nobody knew about.</li>

<li><strong>Account and OU design with enforceable guardrails:</strong> SCPs, organization-level BAA coverage, and Region and service restrictions that hold without breaking your workloads.</li>

<li><strong>KMS key architecture:</strong> customer managed keys per data domain, split administration and usage, and key policies written so the breach safe harbour argument stands up.</li>

<li><strong>Audit logging that produces evidence:</strong> centralised CloudTrail with log file validation, an isolated archive account with Object Lock, and retention that matches your written policy exactly.</li>

<li><strong>Backup, restore and contingency testing:</strong> encrypted cross-account copies that stay inside your boundary, plus rehearsed restores documented as compliance artefacts.</li>

<li><strong>Terraform or OpenTofu modules for the whole control set,</strong> so your posture is reviewable, repeatable and does not drift between audits.</li>
</ul>



<p class="wp-block-paragraph">If you would rather start with something concrete than a discovery call, send me a redacted account structure, an SCP that is causing trouble, or the output of a Config or Security Hub run, and I will tell you what I would look at first.</p>



<div class="wp-block-buttons is-layout-flex wp-block-buttons-is-layout-flex">
<div class="wp-block-button"><a class="wp-block-button__link wp-element-button" href="https://www.upwork.com/freelancers/~01f15a912ad84a6620" target="_blank" rel="noreferrer noopener">Work with me on Upwork</a></div>
</div>
<p>The post <a href="https://john-nessime.com/blog/technical-guides/hipaa-compliance-aws/">HIPAA Compliance on AWS: The Gaps That Pass Every Security Check</a> appeared first on <a href="https://john-nessime.com/blog">John Nessime</a>.</p>
]]></content:encoded>
					
		
		
			</item>
	</channel>
</rss>
