<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Amazon SNS | John Nessime</title>
	<atom:link href="https://john-nessime.com/blog/tag/amazon-sns/feed/" rel="self" type="application/rss+xml" />
	<link>https://john-nessime.com/blog/tag/amazon-sns/</link>
	<description>Cloud, DevOps, Data &#38; AI — Built, Tested, Explained</description>
	<lastBuildDate>Wed, 19 Aug 2026 12:12:34 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://john-nessime.com/blog/wp-content/uploads/2026/07/cropped-jn-32x32.png</url>
	<title>Amazon SNS | John Nessime</title>
	<link>https://john-nessime.com/blog/tag/amazon-sns/</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>Amazon Textract Data Extraction: What Breaks on Real Contracts and Reports</title>
		<link>https://john-nessime.com/blog/devops/amazon-textract-data-extraction/</link>
					<comments>https://john-nessime.com/blog/devops/amazon-textract-data-extraction/#respond</comments>
		
		<dc:creator><![CDATA[John Nessime]]></dc:creator>
		<pubDate>Thu, 20 Aug 2026 18:00:00 +0000</pubDate>
				<category><![CDATA[Cloud Computing]]></category>
		<category><![CDATA[DevOps]]></category>
		<category><![CDATA[Technical Guides]]></category>
		<category><![CDATA[Amazon Bedrock]]></category>
		<category><![CDATA[Amazon S3]]></category>
		<category><![CDATA[Amazon SNS]]></category>
		<category><![CDATA[Amazon SQS]]></category>
		<category><![CDATA[Amazon Textract]]></category>
		<category><![CDATA[AWS]]></category>
		<category><![CDATA[AWS Lambda]]></category>
		<category><![CDATA[Batch Processing]]></category>
		<category><![CDATA[Boto3]]></category>
		<category><![CDATA[Cost Optimization]]></category>
		<category><![CDATA[Data Integration]]></category>
		<category><![CDATA[Data Quality]]></category>
		<category><![CDATA[Document Processing]]></category>
		<category><![CDATA[Idempotency]]></category>
		<category><![CDATA[Intelligent Document Processing]]></category>
		<category><![CDATA[Legal Tech]]></category>
		<category><![CDATA[OCR]]></category>
		<category><![CDATA[Pipeline Design]]></category>
		<category><![CDATA[Python]]></category>
		<category><![CDATA[Serverless]]></category>
		<guid isPermaLink="false">https://john-nessime.com/blog/?p=270</guid>

					<description><![CDATA[<p>Textract rarely fails loudly. It returns a plausible result that is quietly incomplete: a truncated result set, a tick box read as an empty string, a clause split across a page break. A practitioner's guide to the failure modes that actually bite when you point Amazon Textract at contracts, technical reports and correspondence, plus how to choose between sync and async, which feature types are worth paying for, and where Textract stops being the right tool.</p>
<p>The post <a href="https://john-nessime.com/blog/devops/amazon-textract-data-extraction/">Amazon Textract Data Extraction: What Breaks on Real Contracts and Reports</a> appeared first on <a href="https://john-nessime.com/blog">John Nessime</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">The pipeline passed every test. Three sample agreements, all the right fields, clean JSON out the other side. Then it ran against the real archive and somebody in legal noticed that every contract longer than about forty pages was missing its termination clause. No errors. No failed jobs. The dashboard was green the whole time.</p>



<p class="wp-block-paragraph">The bug was a missing loop. The job had finished, the results were there, and the code had read the first chunk of them and stopped.</p>



<p class="wp-block-paragraph">That is the shape of most problems with Amazon Textract data extraction on real documents. The service rarely fails loudly. It returns a plausible-looking result that is quietly incomplete, and you find out weeks later when somebody asks a question the data cannot answer. This post covers the failures that actually bite on contracts, technical reports and correspondence: sync versus async, reading the block graph without losing fields, which feature types are worth paying for, and where Textract stops being the right tool.</p>



<h2 class="wp-block-heading">Why Amazon Textract data extraction is not just OCR</h2>



<p class="wp-block-paragraph">Plain OCR gives you words and their positions. Textract gives you a graph. Every response is a flat array of <code>Block</code> objects, and each block carries an ID plus relationships to other block IDs. A page relates to its lines, a line to its words, a key to its value, a table to its cells.</p>



<p class="wp-block-paragraph">That structure is the whole point, and it is where the pain lives. Nothing is nested for you. To read one form field you find the <code>KEY_VALUE_SET</code> block with entity type <code>KEY</code>, follow its <code>VALUE</code> relationship to another block, then follow that block&#8217;s <code>CHILD</code> relationships to the words. Three hops, and getting a hop wrong returns an empty string rather than an exception.</p>



<p class="wp-block-paragraph">Contracts and reports make this harder than invoices do. An invoice has a total. A master services agreement has a liability cap buried in a numbered sub-clause that spans a page break.</p>



<h2 class="wp-block-heading">Synchronous or asynchronous: the choice is made for you</h2>



<p class="wp-block-paragraph">People reach for <code>AnalyzeDocument</code> first because it returns results in the same call and is easy to test in a notebook. Then they hit the quotas, which AWS documents as hard limits you cannot raise. Synchronous operations cap JPEG, PNG, PDF and TIFF at 10 MB in memory, and cap PDF and TIFF at <strong>one page</strong>. Asynchronous operations keep JPEG and PNG at 10 MB but take PDF and TIFF up to 500 MB and 3,000 pages.</p>



<p class="wp-block-paragraph">One page. That single constraint decides your architecture. Any real contract goes through <code>StartDocumentAnalysis</code> and <code>GetDocumentAnalysis</code>, which means it goes through S3, which means you need somewhere to stage documents and a way to learn when the job is done.</p>



<p class="wp-block-paragraph">Three other constraints from the same quota page are worth knowing before you promise anything to a client:</p>



<ul class="wp-block-list">
<li>Text detection covers English, French, German, Italian, Portuguese and Spanish. Query detection is English only.</li>

<li>Vertical text is not supported. Rotation is fine, including odd in-plane angles, but vertically written scripts are not.</li>

<li>Password-protected PDFs are rejected and XFA-based PDFs are unsupported. Both turn up in legal archives more often than you would expect.</li>
</ul>



<p class="wp-block-paragraph">Pass a <code>ClientRequestToken</code> on submission. It is an idempotency token: the same token returns the same <code>JobId</code> instead of starting a duplicate job. If your submitter is a Lambda function behind an S3 event, and S3 events can be delivered more than once, that is the difference between paying once and paying twice.</p>



<pre class="wp-block-code"><code>import boto3

textract = boto3.client("textract")

job = textract.start_document_analysis(
    DocumentLocation={
        "S3Object": {"Bucket": "contracts-intake", "Name": "msa/acme-2.pdf"}
    },
    FeatureTypes=["FORMS", "TABLES"],
    ClientRequestToken="msa-acme-2-v1",
    NotificationChannel={
        "SNSTopicArn": "arn:aws:sns:eu-west-1:111122223333:textract-done",
        "RoleArn": "arn:aws:iam::111122223333:role/TextractSnsPublish",
    },
)</code></pre>



<p class="wp-block-paragraph">The notification channel is optional but you want it. Polling in a loop inside a Lambda function burns billed duration doing nothing, and Lambda&#8217;s execution ceiling will cut you off on long documents anyway. Publish to SNS, fan out to SQS, let a second function do the reading, and you get a natural place to hang a dead letter queue.</p>



<h2 class="wp-block-heading">The result set you never finished reading</h2>



<p class="wp-block-paragraph">This is the one that cost the termination clauses. <code>GetDocumentAnalysis</code> returns results in pages. When there are more blocks than fit in one response, the response carries a <code>NextToken</code> and you have to call again with it. If you do not, you get the beginning of the document and nothing tells you so. <code>JobStatus</code> still reports success, because the job did succeed. Your code just stopped reading.</p>



<p class="wp-block-paragraph">Short test documents fit in a single response. That is exactly why this survives testing and dies in production.</p>



<pre class="wp-block-code"><code>def fetch_all_blocks(textract, job_id):
    blocks, next_token = [], None

    while True:
        kwargs = {"JobId": job_id}
        if next_token:
            kwargs["NextToken"] = next_token

        response = textract.get_document_analysis(**kwargs)

        if response["JobStatus"] == "FAILED":
            raise RuntimeError(response.get("StatusMessage", "job failed"))

        blocks.extend(response["Blocks"])
        next_token = response.get("NextToken")

        if not next_token:
            return blocks</code></pre>



<p class="wp-block-paragraph">Then make the failure detectable. Treat the highest page number present in the blocks as a checksum against the page count of the source file. Submit a 60-page PDF, get blocks that stop at page 12, and something is wrong no matter what the job status says.</p>



<h2 class="wp-block-heading">Choosing feature types without paying for all of them</h2>



<p class="wp-block-paragraph"><code>AnalyzeDocument</code> and <code>StartDocumentAnalysis</code> take a <code>FeatureTypes</code> list: <code>FORMS</code>, <code>TABLES</code>, <code>QUERIES</code>, <code>SIGNATURES</code> and <code>LAYOUT</code>. This is not cosmetic. Textract bills per page and the per-page rate depends on which features you enabled, so turning all five on because you might need them later multiplies your bill across the whole archive. Rates change and vary by region, so check the current pricing page rather than trusting a number in a blog post. The mechanism does not change: each feature adds to the per-page cost, and you pay it on every page you submit, including blank separator sheets.</p>



<ul class="wp-block-list">
<li><strong>FORMS</strong> finds key-value pairs where the document has an explicit label. Good for cover sheets, signature blocks, report headers. Useless for prose clauses with no label.</li>

<li><strong>TABLES</strong> reconstructs rows, columns and cells, and also identifies table titles, footers and merged cells. This is what you want for rate cards and appendix tables.</li>

<li><strong>QUERIES</strong> lets you ask natural-language questions and get answers back under an alias you chose. This is the feature that makes contracts tractable.</li>

<li><strong>SIGNATURES</strong> returns the location and confidence of handwritten signatures, electronic signatures and initials. It tells you something is signed at a coordinate, not whose signature it is.</li>

<li><strong>LAYOUT</strong> groups text into titles, headers, footers, section headers, paragraphs and lists in reading order. Skippable on single-column documents. On two-column reports it is the difference between coherent text and interleaved nonsense.</li>
</ul>



<p class="wp-block-paragraph">A practical pattern: run a cheap text-detection pass to classify each document, then apply the expensive feature set only to the pages that matter. A 200-page appendix of scanned site photographs does not need FORMS and TABLES.</p>



<h2 class="wp-block-heading">Reading the block graph without losing fields</h2>



<p class="wp-block-paragraph">Here is the failure that produces silent empty strings. A key-value pair&#8217;s value is not always text. If the value is a tick box, the value block&#8217;s children include a <code>SELECTION_ELEMENT</code> block, and that block has no text at all. It has a <code>SelectionStatus</code> of <code>SELECTED</code> or <code>NOT_SELECTED</code>. Concatenate only <code>WORD</code> children and a ticked box comes back as an empty string, which looks exactly like a field nobody filled in.</p>



<pre class="wp-block-code"><code>def block_text(block, block_map):
    parts = []
    for rel in block.get("Relationships", []):
        if rel["Type"] != "CHILD":
            continue
        for child_id in rel["Ids"]:
            child = block_map[child_id]
            if child["BlockType"] == "WORD":
                parts.append(child["Text"])
            elif child["BlockType"] == "SELECTION_ELEMENT":
                parts.append(child["SelectionStatus"])
    return " ".join(parts)</code></pre>



<p class="wp-block-paragraph">Build <code>block_map</code> once as a dictionary keyed on block ID before you traverse anything. Scanning the block list linearly for each lookup turns a three-hop traversal into an accidental quadratic, which you will notice the first time a 3,000-page bundle goes through.</p>



<h2 class="wp-block-heading">Queries: asking for clauses that have no label</h2>



<p class="wp-block-paragraph">FORMS works when the document says &#8220;Effective Date:&#8221; next to the date. Contracts frequently do not. The governing law sits in a paragraph of prose halfway down page nine. Queries handles that. You pass questions in plain English, each with an <code>Alias</code> you choose, and the response pairs <code>QUERY</code> blocks with <code>QUERY_RESULT</code> blocks through an <code>ANSWER</code> relationship. The alias is what makes this usable downstream: you match on <code>governing_law</code>, not on the exact wording of the question, so you can reword questions without breaking your schema.</p>



<pre class="wp-block-code"><code>QueriesConfig = {
    "Queries": [
        {"Text": "What is the governing law?",
         "Alias": "governing_law", "Pages": ["*"]},
        {"Text": "What is the termination notice period?",
         "Alias": "termination_notice", "Pages": ["9-*"]},
    ]
}</code></pre>



<p class="wp-block-paragraph">The limits matter. AWS documents a maximum of 15 queries per page for synchronous operations and 30 per page for asynchronous ones. A fifty-field schema cannot be asked in one pass. You either split it across multiple calls, which multiplies per-page spend, or you narrow <code>Pages</code> so each query runs only where the answer plausibly lives. Signature blocks are at the end, definitions near the front. Running every query against every page is the most common way people accidentally triple their bill.</p>



<p class="wp-block-paragraph">Queries are pre-trained on a spread of business documents including paystubs, bank statements, loan applications and mortgage notes. Your niche contract template was not in that set. Expect good results on common concepts and mediocre ones on house-specific terminology.</p>



<h2 class="wp-block-heading">Documents that fight back</h2>



<h3 class="wp-block-heading">Clauses that span a page break</h3>



<p class="wp-block-paragraph">Textract analyses pages. A clause starting at the bottom of page 11 and finishing at the top of page 12 is two disconnected fragments as far as the block graph is concerned. Queries will often return the fragment on one page and miss the qualifier on the other, which is the worst outcome available because the answer looks complete. The fix is not clever: reassemble full text in reading order using LAYOUT, then run clause-level logic over that rather than over per-page answers. Use Queries to <em>locate</em> a clause and the reassembled text to <em>read</em> it.</p>



<h3 class="wp-block-heading">Two-column reports and email threads</h3>



<p class="wp-block-paragraph">Without LAYOUT, a two-column technical report reads as alternating lines from both columns. Everything downstream inherits that corruption and it passes a spot check, because individual lines look fine. LAYOUT sequences elements in reading order and returns block types for titles, headers, footers and section headers, which is also what you need to chunk a report by section.</p>



<p class="wp-block-paragraph">Printed email chains are quoted text inside quoted text, newest message at the top, the same signature block repeated five times. Textract extracts all of it faithfully, which is the problem. Splitting a thread into individual messages is a text-processing job you do afterwards, not something a feature type solves.</p>



<h3 class="wp-block-heading">PDFs that were never scanned</h3>



<p class="wp-block-paragraph">This is the cheapest win available and almost everybody misses it. A large share of contract archives are born-digital PDFs exported from a word processor, and they already contain a text layer. Running OCR over them pays a service to guess at text you could have read directly. Put a triage step in front of the pipeline; <code>pdftotext</code> from poppler-utils pulls the embedded layer if there is one:</p>



<pre class="wp-block-code"><code>pdftotext -layout contract.pdf - | head -c 2000</code></pre>



<p class="wp-block-paragraph">Substantial text back means born-digital: route it down a cheaper path and reserve Textract for genuinely scanned material. Triage across a large archive is an embarrassingly parallel batch job that sits better on a plain VPS from somewhere like Contabo or InterServer than on per-invocation serverless billing, since you are CPU-bound for minutes at a time rather than reacting to events. One caveat: born-digital does not mean clean. Some exporters produce a text layer with broken word spacing or mangled ligatures, so sample before you trust it.</p>



<h2 class="wp-block-heading">Confidence scores you can actually act on</h2>



<p class="wp-block-paragraph">Every block carries a confidence score and the instinct is to threshold on it. That is usually wrong. Confidence tells you how sure the model is about <em>which characters are on the page</em>, not whether the field is semantically right. Textract can read a date with total certainty and hand you the wrong one, because it picked up the printing date in the footer instead of the effective date in the recitals. High confidence, completely wrong answer.</p>



<p class="wp-block-paragraph">It is also per block. A key-value pair has a score on the key, another on the value and separate ones on each word, so there is no single number to gate on. Validate the extracted value instead. Does the date parse and fall in a plausible range? Does the total equal the sum of the line items? Does the party name appear in the signature block as well as the preamble? Route to human review on failed validation and treat low confidence as one input among several.</p>



<h2 class="wp-block-heading">When Custom Queries adapters are worth the effort</h2>



<p class="wp-block-paragraph">If pre-trained Queries keeps getting your house document type wrong, train an adapter. Create it with <code>CreateAdapter</code>, annotate sample documents, train a version with <code>CreateAdapterVersion</code>, then reference it at inference time.</p>



<pre class="wp-block-code"><code>AdaptersConfig = {
    "Adapters": [
        {"AdapterId": ADAPTER_ID, "Version": "1", "Pages": ["1-5"]},
        {"AdapterId": ADAPTER_ID, "Version": "1", "Pages": ["6-*"]},
    ]
}</code></pre>



<p class="wp-block-paragraph">That <code>Pages</code> string takes digits, hyphens and an asterisk with no blank spaces, a page can only have one adapter applied to it, and an asterisk meaning all pages must be the only element in the list.</p>



<p class="wp-block-paragraph">Two things to weigh. AWS sets a minimum of five samples per query level for training or testing and says plainly that more is better, so treat five as a demo floor rather than a number that survives layout variation. And successful trainings per month are capped per account, so the tight annotate-train-evaluate loop you want is not available. Adapters suit one high-volume, house-specific document type with a stable schema. They are a poor fit for a long tail of one-off templates, which is what most legal archives actually are.</p>



<h2 class="wp-block-heading">Where Textract stops and a language model starts</h2>



<p class="wp-block-paragraph">This is a real decision now, so both sides deserve a hearing. Textract&#8217;s case is strong on standardised, high-volume documents. It is deterministic in a way generative models are not, it returns bounding-box geometry so you can point at exactly where a value came from, it gives per-element confidence, and it bills predictably per page. When an auditor asks why a field has the value it does, geometry beats a model&#8217;s reasoning.</p>



<p class="wp-block-paragraph">The generative side, whether that is Amazon Bedrock Data Automation as a managed document processing service or a model called directly through Amazon Bedrock, wins where the task needs reasoning rather than pattern matching. &#8220;Does this agreement contain an assignment restriction, and if so summarise it&#8221; is a comprehension question, and Textract has no answer for it. Generative approaches also cope better with layouts that were in nobody&#8217;s training set.</p>



<p class="wp-block-paragraph">Most production pipelines use both, which is what AWS&#8217;s own reference architectures push: Textract for faithful extraction with geometry and confidence, then a model over the extracted text for classification and comprehension. Choosing from scratch, ask whether your documents are standardised and high-volume. If yes, Textract plus a thin post-processing layer is cheaper and more auditable. If they are varied and volume is modest, a managed generative service gets you working faster. Google Document AI and Azure AI Document Intelligence cover similar ground if you are not committed to AWS.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Troubleshooting</h2>



<ul class="wp-block-list">
<li><strong>Extraction stops partway through long documents.</strong> You are not following <code>NextToken</code>. Compare the highest page number in your blocks against the source page count.</li>

<li><strong>Fields come back empty on forms you know were filled in.</strong> The value is a <code>SELECTION_ELEMENT</code> and you are only reading <code>WORD</code> children.</li>

<li><strong>The document is rejected as too large.</strong> You are on the synchronous API. Multi-page PDFs must go through <code>StartDocumentAnalysis</code>, and the 10 MB cap is on in-memory size, not file size on disk.</li>

<li><strong>Textract cannot read the document at all.</strong> Check for password protection and XFA-based forms. Both are documented as unsupported.</li>

<li><strong>The service cannot access the S3 object.</strong> Usually the execution role missing <code>s3:GetObject</code>, or a bucket in a different region from the Textract endpoint you called.</li>

<li><strong>Throttling under batch load.</strong> Transactions-per-second quotas are per account per region. Smooth spiky traffic through a queue and retry with exponential backoff and jitter rather than raising concurrency.</li>

<li><strong>Garbled text on reports that look fine to you.</strong> Multi-column layout without LAYOUT enabled, or a source scan below roughly 150 DPI. Re-scanning beats any amount of post-processing.</li>
</ul>



<h2 class="wp-block-heading">Common mistakes</h2>



<ul class="wp-block-list">
<li>Testing only on short documents, which hides the pagination bug entirely.</li>

<li>Enabling every feature type on every page because it is easier than deciding.</li>

<li>Running OCR over born-digital PDFs that already carry a text layer.</li>

<li>Treating a confidence score as a correctness score.</li>

<li>Scanning the block array linearly instead of building an ID map.</li>

<li>Polling for job completion inside a Lambda function instead of using SNS.</li>

<li>Storing only extracted values and discarding the raw JSON, so reprocessing means paying again.</li>
</ul>



<h2 class="wp-block-heading">Best practices</h2>



<ul class="wp-block-list">
<li>Persist the full Textract JSON to S3 keyed on a hash of the source document. Reprocessing then costs storage, not extraction.</li>

<li>Classify first, extract second. A cheap text-detection pass tells you which expensive feature set each document needs.</li>

<li>Validate on the value, not the score. Parse dates, check ranges, cross-reference names against other parts of the document.</li>

<li>Use aliases on every query and treat them as your schema contract.</li>

<li>Set a per-job page-count assertion and alert on mismatches. That is your canary for silent truncation.</li>

<li>Keep source documents in an encrypted bucket with a lifecycle policy. Contracts and correspondence should not accumulate indefinitely by accident.</li>

<li>Track cost per document rather than per API call. Feature stacking makes the per-call figure meaningless.</li>
</ul>



<h2 class="wp-block-heading">Frequently asked questions</h2>



<h3 class="wp-block-heading">Can Amazon Textract handle multi-page PDF contracts?</h3>



<p class="wp-block-paragraph">Yes, through the asynchronous API only. Synchronous operations cap PDF and TIFF at one page. Asynchronous operations accept PDF and TIFF up to 500 MB and 3,000 pages, staged through S3.</p>



<h3 class="wp-block-heading">Why is my Textract output missing the end of the document?</h3>



<p class="wp-block-paragraph">Almost certainly because you read the first response from <code>GetDocumentAnalysis</code> and stopped. Results are paginated with a <code>NextToken</code> and the job status still reports success, so nothing signals the truncation. Loop until <code>NextToken</code> is absent.</p>



<h3 class="wp-block-heading">How many Textract Queries can I ask per page?</h3>



<p class="wp-block-paragraph">AWS documents 15 per page for synchronous operations and 30 per page for asynchronous ones. Larger schemas need multiple passes, which costs more, so narrowing each query&#8217;s page range is worth the effort.</p>



<h3 class="wp-block-heading">Does Textract work on handwriting and signatures?</h3>



<p class="wp-block-paragraph">It detects handwriting, and the SIGNATURES feature returns the location and confidence of handwritten signatures, electronic signatures and initials. It does not verify identity: it tells you something was signed and where, not by whom.</p>



<h3 class="wp-block-heading">What languages does Amazon Textract support?</h3>



<p class="wp-block-paragraph">Text detection covers English, French, German, Italian, Portuguese and Spanish. Queries detection is English only, vertically written text is unsupported, and Textract does not return the detected language in its output.</p>



<h3 class="wp-block-heading">Should I use Textract or a generative model?</h3>



<p class="wp-block-paragraph">Textract when documents are standardised, volume is high, and you need auditable geometry and per-field confidence. A generative service when layouts vary widely, volume is modest, or the task needs comprehension rather than extraction. Most production pipelines use both.</p>



<h2 class="wp-block-heading">The one thing worth remembering</h2>



<p class="wp-block-paragraph">Amazon Textract data extraction fails quietly far more often than it fails loudly. A truncated result set, a tick box read as an empty string, a clause split across a page break, a confidently extracted wrong date: none of these throw an exception and none show up on a dashboard.</p>



<p class="wp-block-paragraph">So build the assertions that make silence detectable. Compare expected page counts against extracted ones, validate values rather than trusting scores, and keep the raw JSON so you can prove what the service actually returned. Textract is genuinely good at reading documents. Your job is noticing when it did not read all of them.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Need a document extraction pipeline that does not lose fields?</h2>



<p class="wp-block-paragraph">I build and fix document processing pipelines on AWS. Typical work looks like this:</p>



<ul class="wp-block-list">
<li>Auditing an existing Textract pipeline for silent truncation, dropped selection elements and mis-scoped queries</li>

<li>Designing the async architecture end to end: S3 intake, SNS and SQS fan-out, Lambda or container workers, dead letter handling and retries</li>

<li>Cutting per-page cost by classifying first, triaging born-digital PDFs out of the OCR path and narrowing feature types per page range</li>

<li>Turning block-graph JSON into a clean schema in CSV, a relational database or a data lake, with validation rules and a review queue for exceptions</li>

<li>Building the hybrid path where Textract handles extraction and a model on Amazon Bedrock handles classification and comprehension</li>

<li>Adding the observability that makes quiet failures loud: page-count assertions, per-document cost tracking, confidence distribution monitoring</li>
</ul>



<p class="wp-block-paragraph">If you have a redacted sample document, a chunk of Textract JSON or an extraction coming back half empty, send it over and I will tell you what is going wrong with it.</p>



<div class="wp-block-buttons is-layout-flex wp-block-buttons-is-layout-flex">
<div class="wp-block-button"><a class="wp-block-button__link wp-element-button" href="https://www.upwork.com/freelancers/~01f15a912ad84a6620" target="_blank" rel="noreferrer noopener">Work with me on Upwork</a></div>
</div>
<p>The post <a href="https://john-nessime.com/blog/devops/amazon-textract-data-extraction/">Amazon Textract Data Extraction: What Breaks on Real Contracts and Reports</a> appeared first on <a href="https://john-nessime.com/blog">John Nessime</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://john-nessime.com/blog/devops/amazon-textract-data-extraction/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Building an Insurance Claims Processing Pipeline on AWS That Fails Loudly</title>
		<link>https://john-nessime.com/blog/cloud-computing/insurance-claims-processing-pipeline-aws/</link>
					<comments>https://john-nessime.com/blog/cloud-computing/insurance-claims-processing-pipeline-aws/#respond</comments>
		
		<dc:creator><![CDATA[John Nessime]]></dc:creator>
		<pubDate>Thu, 20 Aug 2026 09:00:00 +0000</pubDate>
				<category><![CDATA[Cloud Computing]]></category>
		<category><![CDATA[Insurance Technology]]></category>
		<category><![CDATA[Workflow Automation]]></category>
		<category><![CDATA[Amazon SNS]]></category>
		<category><![CDATA[Amazon Textract]]></category>
		<category><![CDATA[AWS]]></category>
		<category><![CDATA[Bedrock Data Automation]]></category>
		<category><![CDATA[Claims Automation]]></category>
		<category><![CDATA[Confidence Scoring]]></category>
		<category><![CDATA[Data Validation]]></category>
		<category><![CDATA[Dead Letter Queue]]></category>
		<category><![CDATA[Document Processing]]></category>
		<category><![CDATA[Event-Driven Architecture]]></category>
		<category><![CDATA[HIPAA]]></category>
		<category><![CDATA[Human In The Loop]]></category>
		<category><![CDATA[Idempotency]]></category>
		<category><![CDATA[Insurance Claims]]></category>
		<category><![CDATA[Intelligent Document Processing]]></category>
		<category><![CDATA[Serverless]]></category>
		<category><![CDATA[Step Functions]]></category>
		<category><![CDATA[Straight-Through Processing]]></category>
		<guid isPermaLink="false">https://john-nessime.com/blog/?p=464</guid>

					<description><![CDATA[<p>Claims pipelines rarely crash. They succeed, emit clean JSON, and hand a wrong number to a payment system. Six failure families in an insurance claims processing pipeline on AWS, with the Textract, Bedrock Data Automation and Step Functions details that decide whether a bad extraction is visible or silent.</p>
<p>The post <a href="https://john-nessime.com/blog/cloud-computing/insurance-claims-processing-pipeline-aws/">Building an Insurance Claims Processing Pipeline on AWS That Fails Loudly</a> appeared first on <a href="https://john-nessime.com/blog">John Nessime</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">The worst ticket on a claims pipeline is never the one that says the pipeline is down. A stuck queue is loud. It pages somebody, somebody restarts something, and it gets fixed before lunch. The bad ticket arrives three weeks later from finance: a run of claims was auto-approved at amounts nobody can reconcile, and every single execution in the Step Functions console is green.</p>



<p class="wp-block-paragraph">That is the failure mode that defines this problem. An insurance claims processing pipeline on AWS almost never falls over in the way you designed it to fall over. It succeeds. It emits well-formed JSON. It hands a number to a payment system, and the number is wrong, and nothing in the pipeline had any reason to think otherwise.</p>



<p class="wp-block-paragraph">This post is organized by failure family rather than by service. I&#8217;ll walk through the six ways these pipelines go quietly wrong, what each one costs, and what the fix actually looks like in Textract, Bedrock, Step Functions and S3. There&#8217;s a troubleshooting section, the mistakes I see repeated, and an FAQ at the end.</p>



<h2 class="wp-block-heading">The shape most claims pipelines end up with</h2>



<p class="wp-block-paragraph">Before the failure families make sense, the skeleton. Almost every serverless claims pipeline lands on roughly the same set of stages, whatever the vendor deck calls them:</p>



<ol class="wp-block-list">
<li><strong>Intake.</strong> A document lands in S3 from a portal upload, an SFTP drop, or a mail scanning vendor. An S3 event or EventBridge rule starts an execution.</li>

<li><strong>Classification.</strong> Work out what the packet actually contains. A first notice of loss, a CMS-style claim form, a police report, an itemized bill, forty pages of photographs.</li>

<li><strong>Extraction.</strong> Pull the fields you need. Amazon Textract for OCR, forms and tables, or Amazon Bedrock Data Automation with a blueprint that names the fields directly.</li>

<li><strong>Validation.</strong> Check the extracted values against business rules, policy data, and each other.</li>

<li><strong>Routing.</strong> Straight-through processing, human review, or rejection with a reason.</li>

<li><strong>Persistence and audit.</strong> The claim record, the extraction artifacts, and enough evidence to explain a decision months later.</li>
</ol>



<p class="wp-block-paragraph">Nothing controversial there. AWS publishes an open-source GenAI IDP Accelerator that implements exactly this shape, with a Bedrock Data Automation mode and a Textract-plus-foundation-model pipeline mode, and it&#8217;s a reasonable place to start reading. The interesting part is not the boxes. It&#8217;s what happens between them.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Failure family one: the field that was never there</h2>



<p class="wp-block-paragraph">This is the one that pays out the wrong number, and it is worth more attention than everything else in this post combined.</p>



<p class="wp-block-paragraph">Every extraction service gives you confidence scores. So the obvious design is a gate: if every field scores above some threshold, approve automatically; if anything falls below, send it to a human. That gate is sound reasoning applied to the wrong population.</p>



<p class="wp-block-paragraph">A confidence score only exists for a value that came back. When the extractor doesn&#8217;t find a field at all, there is no low score to catch, because there&#8217;s nothing to score. The gate iterates over four returned fields, finds all four above threshold, and reports a clean pass. The fifth field, the one that determines coordination of benefits or the deductible offset, is simply absent from the response. Downstream code treats absent as zero, or as null, or as &#8220;not applicable,&#8221; and the claim goes through.</p>



<p class="wp-block-paragraph">The fix is structural, not statistical. Validate <em>presence against a schema</em> before you validate confidence, and treat the two as separate gates with separate outcomes:</p>



<pre class="wp-block-code"><code># Two gates, not one. Presence first, then confidence.
# 'extracted' is the flattened field map from Textract Queries
# or a Bedrock Data Automation blueprint result.

REQUIRED = {
    "claim_number",
    "date_of_service",
    "billed_amount",
    "member_id",
    "secondary_payer_indicator",
}

def gate(extracted, scores, threshold=0.95):
    missing = REQUIRED - set(extracted)
    if missing:
        # Never silently default. This is a routing decision.
        return "HUMAN_REVIEW", {"reason": "missing_fields",
                                "fields": sorted(missing)}

    weak = [f for f in REQUIRED if scores.get(f, 0.0) &lt; threshold]
    if weak:
        return "HUMAN_REVIEW", {"reason": "low_confidence",
                                "fields": sorted(weak)}

    return "STRAIGHT_THROUGH", {}
</code></pre>



<p class="wp-block-paragraph">Two details matter here. The set difference is computed against a declared schema, not against whatever keys happen to be in the response, so an absent field becomes a first-class routing reason. And <code>scores.get(f, 0.0)</code> defaults to zero rather than to a passing value, so a field that arrives without a score fails closed.</p>



<p class="wp-block-paragraph">If you&#8217;re on Textract Queries, there&#8217;s a second reason to be explicit: Queries let you attach an alias to each question, which means your schema keys are yours rather than whatever label happened to be printed on the form. That&#8217;s the difference between &#8220;the field is missing&#8221; and &#8220;the field moved and we didn&#8217;t notice.&#8221;</p>



<h2 class="wp-block-heading">Failure family two: confidence scores that answer a different question</h2>



<p class="wp-block-paragraph">Assume you&#8217;ve fixed presence. The next trap is what the confidence number is measuring.</p>



<p class="wp-block-paragraph">OCR confidence is a statement about characters. It says the model is highly sure those pixels read <code>1,240.00</code>. It is not a statement that <code>1,240.00</code> is the billed amount rather than the allowed amount from the box directly above it, or the prior balance from a remittance summary that happened to be stapled into the same packet. Read it as a legibility score, because that&#8217;s closer to what it is.</p>



<p class="wp-block-paragraph">Bedrock Data Automation narrows this gap: blueprints define fields semantically, confidence scores come with bounding boxes, and the visual grounding lets you point at the region a value came from. Textract Queries narrow it too, by asking a question rather than harvesting a label. Neither eliminates the problem, because a high-confidence read of the wrong region still scores high.</p>



<p class="wp-block-paragraph">What actually catches this is cross-field invariants. They cost almost nothing and they fail for reasons a human can read:</p>



<ul class="wp-block-list">
<li><strong>Arithmetic.</strong> Line items sum to the claimed total. If they don&#8217;t, one of the two is wrong and you don&#8217;t yet know which.</li>

<li><strong>Temporal.</strong> Date of service falls inside the policy period and before the date of submission. A service date after the submission date is a parsing error nine times out of ten.</li>

<li><strong>Referential.</strong> The member or policy identifier resolves against your system of record. An identifier that matches the format but not a real record is a strong signal you read the wrong box.</li>

<li><strong>Range.</strong> Amounts within a plausible band for the claim type. Not a fraud model, just a tripwire for a decimal point that moved.</li>

<li><strong>Page provenance.</strong> Fields that must come from the same page or the same document within the packet. Bounding box data makes this checkable rather than assumed.</li>
</ul>



<p class="wp-block-paragraph">An invariant failure is more useful than a low score, because it names a contradiction. &#8220;Line items sum to 1,180 but the claimed total reads 1,240&#8221; is something a reviewer resolves in seconds. &#8220;Confidence 0.91&#8221; is something a reviewer stares at.</p>



<h2 class="wp-block-heading">Failure family three: the claim that stops halfway</h2>



<p class="wp-block-paragraph">Claims documents are multi-page packets, so you&#8217;ll be using Textract&#8217;s asynchronous operations. That means jobs, notifications, and a whole class of orchestration bugs that only show up under load or after a weekend.</p>



<p class="wp-block-paragraph">The asynchronous pattern is: call <code>StartDocumentAnalysis</code>, get a <code>JobId</code> back, and let Textract publish completion to an SNS topic you nominate. A request looks like this:</p>



<pre class="wp-block-code"><code>{
  "DocumentLocation": {
    "S3Object": { "Bucket": "claims-intake", "Name": "packets/abc123.pdf" }
  },
  "FeatureTypes": ["FORMS", "TABLES"],
  "ClientRequestToken": "abc123-v1",
  "JobTag": "fnol-packet",
  "NotificationChannel": {
    "SNSTopicArn": "arn:aws:sns:REGION:ACCOUNT:textract-complete",
    "RoleArn": "arn:aws:iam::ACCOUNT:role/TextractPublishRole"
  },
  "OutputConfig": {
    "S3Bucket": "claims-extraction",
    "S3Prefix": "raw/"
  },
  "KMSKeyId": "alias/claims-cmk"
}
</code></pre>



<p class="wp-block-paragraph">Three of those parameters are doing load-bearing work that is easy to skip.</p>



<p class="wp-block-paragraph"><code>ClientRequestToken</code> is the idempotency token. Reuse the same token and you get the same <code>JobId</code> back instead of a second job. Derive it from the document, not from the invocation, and a Lambda retry or a duplicated S3 event stops turning into a duplicate charge and a duplicate claim record.</p>



<p class="wp-block-paragraph"><code>OutputConfig</code> writes results into a bucket you control. Without it, results stay internal to Textract and the only way to read them is the <code>Get</code> operations, which have their own throttling limits. Under concurrency those limits become the bottleneck: you end up polling more jobs than you&#8217;re allowed to poll, backing off, and watching end-to-end latency climb for reasons that have nothing to do with the documents. Writing to S3 sidesteps the whole path.</p>



<p class="wp-block-paragraph"><code>JobTag</code> shows up in the completion notification. In a mixed pipeline where the same topic carries first notice of loss packets, itemized bills and ID documents, that tag is what lets the notification handler route without a lookup.</p>



<p class="wp-block-paragraph">One expiry to plan around: a Textract <code>JobId</code> is only valid for seven days. If your retry story is &#8220;requeue it and someone will look on Monday,&#8221; a bad weekend turns recoverable failures into full reprocessing. Persist the S3 output location, not the job identifier.</p>



<h3 class="wp-block-heading">Callbacks that never come back</h3>



<p class="wp-block-paragraph">Human review means pausing a workflow for hours or days, which in Step Functions means the callback pattern. You append <code>.waitForTaskToken</code> to the resource ARN, pass <code>$$.Task.Token</code> into the payload, and the execution parks until something calls <code>SendTaskSuccess</code> or <code>SendTaskFailure</code> with that token.</p>



<p class="wp-block-paragraph">The trap is that a callback task with no timeout waits until the execution itself hits its quota, and Standard workflow executions can run for up to a year. A reviewer who leaves, a review UI that drops the token, a queue consumer that crashes after reading the message and before writing it to the review table: all of these produce an execution that is neither failed nor finished. It just sits there. Nobody alerts on it because nothing broke.</p>



<pre class="wp-block-code"><code>"AwaitAdjusterDecision": {
  "Type": "Task",
  "Resource": "arn:aws:states:::lambda:invoke.waitForTaskToken",
  "Parameters": {
    "FunctionName": "enqueue-review-task",
    "Payload": {
      "claimId.$": "$.claimId",
      "taskToken.$": "$$.Task.Token"
    }
  },
  "TimeoutSeconds": 259200,
  "HeartbeatSeconds": 3600,
  "Catch": [{
    "ErrorEquals": ["States.Timeout"],
    "Next": "EscalateStaleReview"
  }],
  "Next": "ApplyDecision"
}
</code></pre>



<p class="wp-block-paragraph"><code>TimeoutSeconds</code> is the maximum total lifetime of the task regardless of heartbeats. <code>HeartbeatSeconds</code> is the maximum gap between <code>SendTaskHeartbeat</code> calls, so a review app that periodically confirms the item is still in someone&#8217;s queue will fail fast when that app dies, rather than at the outer limit. AWS&#8217;s own guidance is to set the heartbeat below the task timeout for exactly this reason: a heartbeat failure tells you the worker died, a timeout tells you the work took too long, and those are different incidents. Catch <code>States.Timeout</code> and route to a real state. An unhandled timeout is just a differently-shaped silence.</p>



<p class="wp-block-paragraph">One constraint worth knowing before you design around it: the callback pattern requires Standard workflows. Express workflows support request-response integrations only, so no <code>.waitForTaskToken</code> and no <code>.sync</code>. If you split your pipeline into a fast Express path and a Standard review path, the boundary between them is where the token has to live.</p>



<h2 class="wp-block-heading">Failure family four: the human review service you can no longer sign up for</h2>



<p class="wp-block-paragraph">This one catches people copying a reference architecture, and it&#8217;s the reason to read publication dates on IDP blog posts.</p>



<p class="wp-block-paragraph">Amazon Augmented AI, known as A2I, was the managed answer to human-in-the-loop review. It plugged directly into Textract&#8217;s <code>AnalyzeDocument</code>, watched confidence conditions, and spun up review tasks with a worker UI for you. It appears in a great many architecture diagrams for claims and lending workflows.</p>



<p class="wp-block-paragraph">Per the AWS documentation, SageMaker A2I is no longer open to new customers. Existing customers can keep using it, and AWS continues security and availability work, but no new features are planned. If you&#8217;re standing up a new account today, that diagram does not deploy.</p>



<p class="wp-block-paragraph">Be fair about what that costs you, because A2I genuinely removed real work: the task assignment logic, the worker UI, the private workforce plumbing through Cognito, result consolidation. Rebuilding it means owning all of that. What you get back is that the review queue becomes yours, which in practice means you can put claim-specific context on the screen instead of a generic key-value editor. For adjusters that difference is not cosmetic.</p>



<p class="wp-block-paragraph">A minimal replacement is not exotic:</p>



<ul class="wp-block-list">
<li>A DynamoDB table of review items, each holding the claim identifier, the extracted values, the bounding boxes, and the Step Functions task token.</li>

<li>A small web app for reviewers that renders the page image with the boxes overlaid, so a reviewer confirms placement rather than retyping values.</li>

<li>An API that writes the corrected values and calls <code>SendTaskSuccess</code> with the stored token.</li>

<li>Authentication in front of it. Amazon Cognito if you want to stay inside AWS, or an identity-aware proxy such as Cloudflare Access if your reviewers are external adjusters you&#8217;d rather not create AWS identities for.</li>

<li>A sweeper that finds review items older than your heartbeat window and escalates them.</li>
</ul>



<p class="wp-block-paragraph">If the review app is a small internal tool with no data residency requirement of its own, it doesn&#8217;t have to live in the same account or even the same provider. A modest VPS from a host like Contabo or InterServer running behind a Cloudflare Tunnel is a legitimate answer for a reviewer console that talks to AWS over scoped API credentials. Just be honest about what crosses that boundary, which brings us to the next family.</p>



<h2 class="wp-block-heading">Failure family five: claim data in places nobody decided to put it</h2>



<p class="wp-block-paragraph">Claims documents carry protected health information, financial identifiers, and often photographs of people and property. The pipeline you drew has three or four places that data lives. The pipeline you deployed has a dozen.</p>



<p class="wp-block-paragraph">The ones that get missed:</p>



<ul class="wp-block-list">
<li><strong>Lambda logs.</strong> One <code>print</code> of an event payload during a debugging session, and CloudWatch Logs is now a claims repository with a different retention policy and a different access model.</li>

<li><strong>Dead letter queues.</strong> A DLQ holds the full failed message. If that message carries extracted values, your DLQ is regulated data, and it is usually the least governed thing in the account.</li>

<li><strong>Step Functions execution history.</strong> State input and output are visible in the console and the history API. Passing extracted fields between states puts them there.</li>

<li><strong>Intermediate extraction output.</strong> The bucket you pointed <code>OutputConfig</code> at holds raw OCR of the whole packet, often with a lifecycle policy nobody wrote.</li>

<li><strong>Model invocation logging.</strong> Bedrock can log inputs and outputs to S3 or CloudWatch. Useful for debugging, and another copy of everything.</li>
</ul>



<p class="wp-block-paragraph">The pattern that keeps this manageable is passing pointers, not payloads. States carry an S3 key and a claim identifier; the values themselves stay in one encrypted bucket with one lifecycle policy and one access policy. It makes debugging marginally more annoying and it makes the data map fit on a page.</p>



<p class="wp-block-paragraph">On regulated workloads, check the current AWS HIPAA-eligible services list and your executed BAA for every service in the path, in the specific region you&#8217;re deploying to. Eligibility is per service and it changes. Textract has long been used for claims workflows on that basis, and Bedrock is listed as HIPAA eligible, but &#8220;I read a blog post&#8221; is not a control. Pull the list yourself before PHI touches anything.</p>



<p class="wp-block-paragraph">Two smaller things worth deciding early. Reviewers working from home should reach the console over something better than the open internet, whether that&#8217;s a corporate tunnel, a business VPN account from a provider like NordVPN or Surfshark, or an identity-aware proxy. And if reviewers ever download claim documents locally, agree what happens to those files afterward, because a deleted file is not an erased file. Tools such as O&amp;O SafeErase exist for exactly that gap on Windows endpoints.</p>



<h2 class="wp-block-heading">Failure family six: paying twice for the same page</h2>



<p class="wp-block-paragraph">Document AI services bill per page. That single fact reshapes how you think about retries, because in most pipelines a retry is free and here it isn&#8217;t.</p>



<p class="wp-block-paragraph">The expensive patterns are all shaped the same way. A poison document fails a downstream parser, gets requeued, and is re-extracted on every attempt. A batch job re-runs over an entire prefix instead of a delta. A misconfigured S3 event delivers twice. An operator reprocesses a day&#8217;s intake to fix a mapping bug in the transform stage, when the extraction stage was fine all along.</p>



<p class="wp-block-paragraph">The structural fix is separating extraction from interpretation. Extract once, write the raw result to S3 keyed by a content hash of the document, and let every downstream stage read from that. When the mapping bug shows up, you re-run interpretation over stored output and pay nothing. Combined with <code>ClientRequestToken</code>, most accidental double-charges disappear.</p>



<p class="wp-block-paragraph">Also route the packet before you extract it. Forty pages of accident photographs do not need forms and tables analysis. Classification is cheaper than extraction, and page-level routing is often the single largest lever on the bill.</p>



<p class="wp-block-paragraph">For attributing that spend, cost allocation tags on the buckets and functions give you the AWS-native view, and platforms like Vantage or CloudZero are worth a look if you need per-claim or per-client unit costs rather than per-service totals. Whatever you use, the metric that matters is cost per claim processed, split by straight-through versus reviewed. Those two numbers tell you whether the automation is earning its keep.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Troubleshooting an insurance claims processing pipeline on AWS</h2>



<p class="wp-block-paragraph">Symptoms you&#8217;ll actually see, and where to look first.</p>



<ul class="wp-block-list">
<li><strong>Executions succeed but downstream amounts are wrong.</strong> Check whether required fields are present, not just confident. Diff the schema against the response keys for a sample of recent claims. This is failure family one until proven otherwise.</li>

<li><strong>Executions stuck in Running for days.</strong> A callback task with no timeout. List running executions ordered by start time and look for the state name of your review task.</li>

<li><strong>Throttling on the extraction stage under load.</strong> If you&#8217;re polling <code>Get</code> operations, move to <code>OutputConfig</code> and SNS notification and stop polling. If you&#8217;re already there, check the start-operation limits rather than assuming the whole service is slow.</li>

<li><strong>The same claim appearing twice.</strong> Look for a missing or per-invocation <code>ClientRequestToken</code>, and check whether your S3 event handler is idempotent. Delivery is at-least-once.</li>

<li><strong>Extraction quality dropped for one document type.</strong> Usually the form changed, not the model. Compare bounding boxes for the affected field against an older sample. If the box moved, that&#8217;s a layout change, and query aliases or a blueprint update is the fix.</li>

<li><strong>Review queue growing faster than reviewers clear it.</strong> Break the routing reasons apart. If most items are low confidence on one field, that&#8217;s an extraction problem wearing a staffing problem&#8217;s clothes.</li>

<li><strong>SNS notification arrives, handler can&#8217;t find the results.</strong> Confirm the notification role has permission to publish and the handler is reading the S3 prefix rather than calling <code>Get</code> with an expired job identifier.</li>
</ul>



<h2 class="wp-block-heading">Common mistakes</h2>



<ul class="wp-block-list">
<li>Gating only on confidence, so a missing field is indistinguishable from a clean extraction.</li>

<li>Defaulting absent values to zero or null in the transform layer instead of raising a routing decision.</li>

<li>Copying an architecture diagram that includes A2I into a new AWS account.</li>

<li>Callback tasks with no <code>TimeoutSeconds</code> and no heartbeat.</li>

<li>Passing extracted claim values through Step Functions state rather than passing an S3 pointer.</li>

<li>Treating a single global confidence threshold as adequate for every field. A name and a dollar amount do not carry the same downstream risk.</li>

<li>Running forms and tables analysis over every page of a packet including the photographs.</li>

<li>No metric for straight-through rate, so nobody notices when it quietly drops.</li>
</ul>



<h2 class="wp-block-heading">Best practices worth the effort</h2>



<ul class="wp-block-list">
<li><strong>Declare the schema, then validate presence, then confidence, then invariants.</strong> Four gates, four distinct rejection reasons, four things a reviewer can act on.</li>

<li><strong>Set per-field thresholds by consequence.</strong> Get the payable amount wrong and money moves. Get a street suffix wrong and a letter is slightly odd.</li>

<li><strong>Keep a labeled regression set.</strong> A few dozen real packets with known-correct values, run on every blueprint or query change. Bedrock Data Automation can use ground-truth examples to refine blueprint instructions, which only works if you maintain the ground truth.</li>

<li><strong>Store bounding boxes alongside values.</strong> They make review faster and they turn &#8220;quality dropped&#8221; from a guess into a comparison.</li>

<li><strong>Alarm on rates, not just errors.</strong> Straight-through rate, review-queue age, and cost per claim. A pipeline that stops approving anything is broken even though nothing threw.</li>

<li><strong>Make every stage idempotent on a content hash.</strong> Reprocessing is normal. It should be safe and cheap.</li>

<li><strong>Instrument the pipeline like a pipeline.</strong> CloudWatch covers the AWS surface; if you&#8217;re consolidating with on-premises claims systems, a platform like Grafana Cloud gives you one place to correlate both sides.</li>
</ul>



<h2 class="wp-block-heading">Frequently asked questions</h2>



<h3 class="wp-block-heading">Should I use Amazon Textract or Bedrock Data Automation for claims extraction?</h3>



<p class="wp-block-paragraph">Textract is the sharper tool when your documents are standardized forms and you want deterministic OCR with forms, tables and targeted queries. Bedrock Data Automation is stronger on mixed packets, because it splits along logical document boundaries, classifies each part, and applies a blueprint per document type, with confidence scores and visual grounding on the output. Claims intake is usually mixed packets, which tilts toward Data Automation, but the honest answer is to run both against a sample of your real documents. The evaluation costs a day and it decides your architecture.</p>



<h3 class="wp-block-heading">What replaces Amazon A2I for human review?</h3>



<p class="wp-block-paragraph">For new AWS accounts, a custom review path: a queue or table of review items, a reviewer UI, and the Step Functions callback pattern to resume the workflow. It&#8217;s more code than A2I but not a large amount, and it gives you a review screen designed around claims rather than around generic key-value pairs. Existing A2I customers can continue as they are, though building on a service with no planned features is a decision to make deliberately rather than by default.</p>



<h3 class="wp-block-heading">What straight-through processing rate should I expect?</h3>



<p class="wp-block-paragraph">Anyone quoting you a number without seeing your documents is guessing. It depends almost entirely on document quality and how many fields you require. What&#8217;s reliable is the method: measure your current rate, split failures by reason, and fix the largest reason. Requiring one rarely-present field can dominate everything else, and that&#8217;s a policy decision as much as an engineering one.</p>



<h3 class="wp-block-heading">How do I keep an insurance claims processing pipeline on AWS HIPAA-aligned?</h3>



<p class="wp-block-paragraph">Start from the AWS HIPAA-eligible services list and an executed BAA, and confirm eligibility for each service in your specific region. Then do the unglamorous work: customer-managed KMS keys, no PHI in logs or state payloads, scoped IAM roles per stage, VPC endpoints where the service supports them, retention policies on every bucket and queue including dead letter queues, and CloudTrail configured so you can answer who accessed which claim. Eligibility is permission to build; the controls are yours.</p>



<h3 class="wp-block-heading">Can I run the whole pipeline with Step Functions Express workflows?</h3>



<p class="wp-block-paragraph">Not the part that waits for a human. Express workflows support request-response integrations only, so the callback pattern requires Standard. A common split is Express for the high-volume deterministic stages and Standard for anything holding a task token, with the two connected by an event or a queue.</p>



<h3 class="wp-block-heading">How do I stop duplicate claims from duplicate events?</h3>



<p class="wp-block-paragraph">Treat every trigger as at-least-once. Derive an idempotency key from the document itself, usually a content hash plus a version marker, pass it as <code>ClientRequestToken</code> to the extraction call, and use it as the conditional write key when you create the claim record. Then a duplicate event is a no-op rather than a second claim.</p>



<h3 class="wp-block-heading">Is it worth starting from the AWS GenAI IDP Accelerator?</h3>



<p class="wp-block-paragraph">As a reference for structure and as a way to get a working pipeline in front of stakeholders quickly, yes. As a production system you inherit wholesale, be careful: you&#8217;re adopting someone else&#8217;s opinions about classification, review and storage, and you&#8217;ll be reading that code anyway the first time something behaves oddly. Read it, borrow the patterns, own what you deploy.</p>



<h2 class="wp-block-heading">The one thing to take away</h2>



<p class="wp-block-paragraph">An insurance claims processing pipeline on AWS is not hard to build. Textract, Bedrock Data Automation, Step Functions and S3 will get you a working pipeline in a couple of weeks. What&#8217;s hard is making it fail in ways you can see.</p>



<p class="wp-block-paragraph">Every expensive failure in this space shares one shape: the pipeline had no opinion about what it did not receive. A field that didn&#8217;t come back scored nothing, a callback that never fired errored nothing, a duplicate event failed nothing. If you take one design rule from this, take that one. Declare what a complete claim looks like, check for its absence explicitly, and route anything incomplete to a human with a reason attached.</p>



<p class="wp-block-paragraph">Green executions are not evidence. Reconciled numbers are.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Need help with your claims pipeline?</h2>



<p class="wp-block-paragraph">I work with teams building document-heavy workflows on AWS, and claims pipelines are one of the places where a small amount of design care prevents a large amount of reconciliation work. Things I can help with:</p>



<ul class="wp-block-list">
<li>Reviewing an existing extraction pipeline for silent-failure paths, particularly missing-field handling and default values in the transform layer</li>

<li>Designing the routing logic: schema gates, per-field thresholds, cross-field invariants, and the escalation rules that sit behind them</li>

<li>Building a human review path with the Step Functions callback pattern, including timeouts, heartbeats and a sweeper for stale tasks</li>

<li>Running a structured evaluation of Amazon Textract against Bedrock Data Automation on your actual documents, with a labeled regression set you keep afterward</li>

<li>Tracing where claim data actually lands across logs, queues, execution history and intermediate buckets, then shrinking that footprint</li>

<li>Instrumenting straight-through rate, review-queue age and cost per claim so regressions surface before finance finds them</li>
</ul>



<p class="wp-block-paragraph">If you&#8217;d like a second opinion, send me a state machine definition, a sample extraction response with the values redacted, or the routing code that decides what goes to review. That&#8217;s usually enough to spot the gap.</p>



<div class="wp-block-buttons is-layout-flex wp-block-buttons-is-layout-flex">
<div class="wp-block-button"><a class="wp-block-button__link wp-element-button" href="https://www.upwork.com/freelancers/~01f15a912ad84a6620" target="_blank" rel="noreferrer noopener">Work with me on Upwork</a></div>
</div>
<p>The post <a href="https://john-nessime.com/blog/cloud-computing/insurance-claims-processing-pipeline-aws/">Building an Insurance Claims Processing Pipeline on AWS That Fails Loudly</a> appeared first on <a href="https://john-nessime.com/blog">John Nessime</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://john-nessime.com/blog/cloud-computing/insurance-claims-processing-pipeline-aws/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
	</channel>
</rss>
