<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Boto3 | John Nessime</title>
	<atom:link href="https://john-nessime.com/blog/tag/boto3/feed/" rel="self" type="application/rss+xml" />
	<link>https://john-nessime.com/blog/tag/boto3/</link>
	<description>Cloud, DevOps, Data &#38; AI — Built, Tested, Explained</description>
	<lastBuildDate>Sat, 08 Aug 2026 09:50:59 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://john-nessime.com/blog/wp-content/uploads/2026/07/cropped-jn-32x32.png</url>
	<title>Boto3 | John Nessime</title>
	<link>https://john-nessime.com/blog/tag/boto3/</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>Amazon Textract Data Extraction: What Breaks on Real Contracts and Reports</title>
		<link>https://john-nessime.com/blog/devops/amazon-textract-data-extraction/</link>
					<comments>https://john-nessime.com/blog/devops/amazon-textract-data-extraction/#respond</comments>
		
		<dc:creator><![CDATA[John Nessime]]></dc:creator>
		<pubDate>Thu, 20 Aug 2026 18:00:00 +0000</pubDate>
				<category><![CDATA[Cloud Computing]]></category>
		<category><![CDATA[DevOps]]></category>
		<category><![CDATA[Technical Guides]]></category>
		<category><![CDATA[Amazon Bedrock]]></category>
		<category><![CDATA[Amazon S3]]></category>
		<category><![CDATA[Amazon SNS]]></category>
		<category><![CDATA[Amazon SQS]]></category>
		<category><![CDATA[Amazon Textract]]></category>
		<category><![CDATA[AWS]]></category>
		<category><![CDATA[AWS Lambda]]></category>
		<category><![CDATA[Batch Processing]]></category>
		<category><![CDATA[Boto3]]></category>
		<category><![CDATA[Cost Optimization]]></category>
		<category><![CDATA[Data Integration]]></category>
		<category><![CDATA[Data Quality]]></category>
		<category><![CDATA[Document Processing]]></category>
		<category><![CDATA[Idempotency]]></category>
		<category><![CDATA[Intelligent Document Processing]]></category>
		<category><![CDATA[Legal Tech]]></category>
		<category><![CDATA[OCR]]></category>
		<category><![CDATA[Pipeline Design]]></category>
		<category><![CDATA[Python]]></category>
		<category><![CDATA[Serverless]]></category>
		<guid isPermaLink="false">https://john-nessime.com/blog/?p=270</guid>

					<description><![CDATA[<p>Textract rarely fails loudly. It returns a plausible result that is quietly incomplete: a truncated result set, a tick box read as an empty string, a clause split across a page break. A practitioner's guide to the failure modes that actually bite when you point Amazon Textract at contracts, technical reports and correspondence, plus how to choose between sync and async, which feature types are worth paying for, and where Textract stops being the right tool.</p>
<p>The post <a href="https://john-nessime.com/blog/devops/amazon-textract-data-extraction/">Amazon Textract Data Extraction: What Breaks on Real Contracts and Reports</a> appeared first on <a href="https://john-nessime.com/blog">John Nessime</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">The pipeline passed every test. Three sample agreements, all the right fields, clean JSON out the other side. Then it ran against the real archive and somebody in legal noticed that every contract longer than about forty pages was missing its termination clause. No errors. No failed jobs. The dashboard was green the whole time.</p>



<p class="wp-block-paragraph">The bug was a missing loop. The job had finished, the results were there, and the code had read the first chunk of them and stopped.</p>



<p class="wp-block-paragraph">That is the shape of most problems with Amazon Textract data extraction on real documents. The service rarely fails loudly. It returns a plausible-looking result that is quietly incomplete, and you find out weeks later when somebody asks a question the data cannot answer. This post covers the failures that actually bite on contracts, technical reports and correspondence: sync versus async, reading the block graph without losing fields, which feature types are worth paying for, and where Textract stops being the right tool.</p>



<h2 class="wp-block-heading">Why Amazon Textract data extraction is not just OCR</h2>



<p class="wp-block-paragraph">Plain OCR gives you words and their positions. Textract gives you a graph. Every response is a flat array of <code>Block</code> objects, and each block carries an ID plus relationships to other block IDs. A page relates to its lines, a line to its words, a key to its value, a table to its cells.</p>



<p class="wp-block-paragraph">That structure is the whole point, and it is where the pain lives. Nothing is nested for you. To read one form field you find the <code>KEY_VALUE_SET</code> block with entity type <code>KEY</code>, follow its <code>VALUE</code> relationship to another block, then follow that block&#8217;s <code>CHILD</code> relationships to the words. Three hops, and getting a hop wrong returns an empty string rather than an exception.</p>



<p class="wp-block-paragraph">Contracts and reports make this harder than invoices do. An invoice has a total. A master services agreement has a liability cap buried in a numbered sub-clause that spans a page break.</p>



<h2 class="wp-block-heading">Synchronous or asynchronous: the choice is made for you</h2>



<p class="wp-block-paragraph">People reach for <code>AnalyzeDocument</code> first because it returns results in the same call and is easy to test in a notebook. Then they hit the quotas, which AWS documents as hard limits you cannot raise. Synchronous operations cap JPEG, PNG, PDF and TIFF at 10 MB in memory, and cap PDF and TIFF at <strong>one page</strong>. Asynchronous operations keep JPEG and PNG at 10 MB but take PDF and TIFF up to 500 MB and 3,000 pages.</p>



<p class="wp-block-paragraph">One page. That single constraint decides your architecture. Any real contract goes through <code>StartDocumentAnalysis</code> and <code>GetDocumentAnalysis</code>, which means it goes through S3, which means you need somewhere to stage documents and a way to learn when the job is done.</p>



<p class="wp-block-paragraph">Three other constraints from the same quota page are worth knowing before you promise anything to a client:</p>



<ul class="wp-block-list">
<li>Text detection covers English, French, German, Italian, Portuguese and Spanish. Query detection is English only.</li>

<li>Vertical text is not supported. Rotation is fine, including odd in-plane angles, but vertically written scripts are not.</li>

<li>Password-protected PDFs are rejected and XFA-based PDFs are unsupported. Both turn up in legal archives more often than you would expect.</li>
</ul>



<p class="wp-block-paragraph">Pass a <code>ClientRequestToken</code> on submission. It is an idempotency token: the same token returns the same <code>JobId</code> instead of starting a duplicate job. If your submitter is a Lambda function behind an S3 event, and S3 events can be delivered more than once, that is the difference between paying once and paying twice.</p>



<pre class="wp-block-code"><code>import boto3

textract = boto3.client("textract")

job = textract.start_document_analysis(
    DocumentLocation={
        "S3Object": {"Bucket": "contracts-intake", "Name": "msa/acme-2.pdf"}
    },
    FeatureTypes=["FORMS", "TABLES"],
    ClientRequestToken="msa-acme-2-v1",
    NotificationChannel={
        "SNSTopicArn": "arn:aws:sns:eu-west-1:111122223333:textract-done",
        "RoleArn": "arn:aws:iam::111122223333:role/TextractSnsPublish",
    },
)</code></pre>



<p class="wp-block-paragraph">The notification channel is optional but you want it. Polling in a loop inside a Lambda function burns billed duration doing nothing, and Lambda&#8217;s execution ceiling will cut you off on long documents anyway. Publish to SNS, fan out to SQS, let a second function do the reading, and you get a natural place to hang a dead letter queue.</p>



<h2 class="wp-block-heading">The result set you never finished reading</h2>



<p class="wp-block-paragraph">This is the one that cost the termination clauses. <code>GetDocumentAnalysis</code> returns results in pages. When there are more blocks than fit in one response, the response carries a <code>NextToken</code> and you have to call again with it. If you do not, you get the beginning of the document and nothing tells you so. <code>JobStatus</code> still reports success, because the job did succeed. Your code just stopped reading.</p>



<p class="wp-block-paragraph">Short test documents fit in a single response. That is exactly why this survives testing and dies in production.</p>



<pre class="wp-block-code"><code>def fetch_all_blocks(textract, job_id):
    blocks, next_token = [], None

    while True:
        kwargs = {"JobId": job_id}
        if next_token:
            kwargs["NextToken"] = next_token

        response = textract.get_document_analysis(**kwargs)

        if response["JobStatus"] == "FAILED":
            raise RuntimeError(response.get("StatusMessage", "job failed"))

        blocks.extend(response["Blocks"])
        next_token = response.get("NextToken")

        if not next_token:
            return blocks</code></pre>



<p class="wp-block-paragraph">Then make the failure detectable. Treat the highest page number present in the blocks as a checksum against the page count of the source file. Submit a 60-page PDF, get blocks that stop at page 12, and something is wrong no matter what the job status says.</p>



<h2 class="wp-block-heading">Choosing feature types without paying for all of them</h2>



<p class="wp-block-paragraph"><code>AnalyzeDocument</code> and <code>StartDocumentAnalysis</code> take a <code>FeatureTypes</code> list: <code>FORMS</code>, <code>TABLES</code>, <code>QUERIES</code>, <code>SIGNATURES</code> and <code>LAYOUT</code>. This is not cosmetic. Textract bills per page and the per-page rate depends on which features you enabled, so turning all five on because you might need them later multiplies your bill across the whole archive. Rates change and vary by region, so check the current pricing page rather than trusting a number in a blog post. The mechanism does not change: each feature adds to the per-page cost, and you pay it on every page you submit, including blank separator sheets.</p>



<ul class="wp-block-list">
<li><strong>FORMS</strong> finds key-value pairs where the document has an explicit label. Good for cover sheets, signature blocks, report headers. Useless for prose clauses with no label.</li>

<li><strong>TABLES</strong> reconstructs rows, columns and cells, and also identifies table titles, footers and merged cells. This is what you want for rate cards and appendix tables.</li>

<li><strong>QUERIES</strong> lets you ask natural-language questions and get answers back under an alias you chose. This is the feature that makes contracts tractable.</li>

<li><strong>SIGNATURES</strong> returns the location and confidence of handwritten signatures, electronic signatures and initials. It tells you something is signed at a coordinate, not whose signature it is.</li>

<li><strong>LAYOUT</strong> groups text into titles, headers, footers, section headers, paragraphs and lists in reading order. Skippable on single-column documents. On two-column reports it is the difference between coherent text and interleaved nonsense.</li>
</ul>



<p class="wp-block-paragraph">A practical pattern: run a cheap text-detection pass to classify each document, then apply the expensive feature set only to the pages that matter. A 200-page appendix of scanned site photographs does not need FORMS and TABLES.</p>



<h2 class="wp-block-heading">Reading the block graph without losing fields</h2>



<p class="wp-block-paragraph">Here is the failure that produces silent empty strings. A key-value pair&#8217;s value is not always text. If the value is a tick box, the value block&#8217;s children include a <code>SELECTION_ELEMENT</code> block, and that block has no text at all. It has a <code>SelectionStatus</code> of <code>SELECTED</code> or <code>NOT_SELECTED</code>. Concatenate only <code>WORD</code> children and a ticked box comes back as an empty string, which looks exactly like a field nobody filled in.</p>



<pre class="wp-block-code"><code>def block_text(block, block_map):
    parts = []
    for rel in block.get("Relationships", []):
        if rel["Type"] != "CHILD":
            continue
        for child_id in rel["Ids"]:
            child = block_map[child_id]
            if child["BlockType"] == "WORD":
                parts.append(child["Text"])
            elif child["BlockType"] == "SELECTION_ELEMENT":
                parts.append(child["SelectionStatus"])
    return " ".join(parts)</code></pre>



<p class="wp-block-paragraph">Build <code>block_map</code> once as a dictionary keyed on block ID before you traverse anything. Scanning the block list linearly for each lookup turns a three-hop traversal into an accidental quadratic, which you will notice the first time a 3,000-page bundle goes through.</p>



<h2 class="wp-block-heading">Queries: asking for clauses that have no label</h2>



<p class="wp-block-paragraph">FORMS works when the document says &#8220;Effective Date:&#8221; next to the date. Contracts frequently do not. The governing law sits in a paragraph of prose halfway down page nine. Queries handles that. You pass questions in plain English, each with an <code>Alias</code> you choose, and the response pairs <code>QUERY</code> blocks with <code>QUERY_RESULT</code> blocks through an <code>ANSWER</code> relationship. The alias is what makes this usable downstream: you match on <code>governing_law</code>, not on the exact wording of the question, so you can reword questions without breaking your schema.</p>



<pre class="wp-block-code"><code>QueriesConfig = {
    "Queries": [
        {"Text": "What is the governing law?",
         "Alias": "governing_law", "Pages": ["*"]},
        {"Text": "What is the termination notice period?",
         "Alias": "termination_notice", "Pages": ["9-*"]},
    ]
}</code></pre>



<p class="wp-block-paragraph">The limits matter. AWS documents a maximum of 15 queries per page for synchronous operations and 30 per page for asynchronous ones. A fifty-field schema cannot be asked in one pass. You either split it across multiple calls, which multiplies per-page spend, or you narrow <code>Pages</code> so each query runs only where the answer plausibly lives. Signature blocks are at the end, definitions near the front. Running every query against every page is the most common way people accidentally triple their bill.</p>



<p class="wp-block-paragraph">Queries are pre-trained on a spread of business documents including paystubs, bank statements, loan applications and mortgage notes. Your niche contract template was not in that set. Expect good results on common concepts and mediocre ones on house-specific terminology.</p>



<h2 class="wp-block-heading">Documents that fight back</h2>



<h3 class="wp-block-heading">Clauses that span a page break</h3>



<p class="wp-block-paragraph">Textract analyses pages. A clause starting at the bottom of page 11 and finishing at the top of page 12 is two disconnected fragments as far as the block graph is concerned. Queries will often return the fragment on one page and miss the qualifier on the other, which is the worst outcome available because the answer looks complete. The fix is not clever: reassemble full text in reading order using LAYOUT, then run clause-level logic over that rather than over per-page answers. Use Queries to <em>locate</em> a clause and the reassembled text to <em>read</em> it.</p>



<h3 class="wp-block-heading">Two-column reports and email threads</h3>



<p class="wp-block-paragraph">Without LAYOUT, a two-column technical report reads as alternating lines from both columns. Everything downstream inherits that corruption and it passes a spot check, because individual lines look fine. LAYOUT sequences elements in reading order and returns block types for titles, headers, footers and section headers, which is also what you need to chunk a report by section.</p>



<p class="wp-block-paragraph">Printed email chains are quoted text inside quoted text, newest message at the top, the same signature block repeated five times. Textract extracts all of it faithfully, which is the problem. Splitting a thread into individual messages is a text-processing job you do afterwards, not something a feature type solves.</p>



<h3 class="wp-block-heading">PDFs that were never scanned</h3>



<p class="wp-block-paragraph">This is the cheapest win available and almost everybody misses it. A large share of contract archives are born-digital PDFs exported from a word processor, and they already contain a text layer. Running OCR over them pays a service to guess at text you could have read directly. Put a triage step in front of the pipeline; <code>pdftotext</code> from poppler-utils pulls the embedded layer if there is one:</p>



<pre class="wp-block-code"><code>pdftotext -layout contract.pdf - | head -c 2000</code></pre>



<p class="wp-block-paragraph">Substantial text back means born-digital: route it down a cheaper path and reserve Textract for genuinely scanned material. Triage across a large archive is an embarrassingly parallel batch job that sits better on a plain VPS from somewhere like Contabo or InterServer than on per-invocation serverless billing, since you are CPU-bound for minutes at a time rather than reacting to events. One caveat: born-digital does not mean clean. Some exporters produce a text layer with broken word spacing or mangled ligatures, so sample before you trust it.</p>



<h2 class="wp-block-heading">Confidence scores you can actually act on</h2>



<p class="wp-block-paragraph">Every block carries a confidence score and the instinct is to threshold on it. That is usually wrong. Confidence tells you how sure the model is about <em>which characters are on the page</em>, not whether the field is semantically right. Textract can read a date with total certainty and hand you the wrong one, because it picked up the printing date in the footer instead of the effective date in the recitals. High confidence, completely wrong answer.</p>



<p class="wp-block-paragraph">It is also per block. A key-value pair has a score on the key, another on the value and separate ones on each word, so there is no single number to gate on. Validate the extracted value instead. Does the date parse and fall in a plausible range? Does the total equal the sum of the line items? Does the party name appear in the signature block as well as the preamble? Route to human review on failed validation and treat low confidence as one input among several.</p>



<h2 class="wp-block-heading">When Custom Queries adapters are worth the effort</h2>



<p class="wp-block-paragraph">If pre-trained Queries keeps getting your house document type wrong, train an adapter. Create it with <code>CreateAdapter</code>, annotate sample documents, train a version with <code>CreateAdapterVersion</code>, then reference it at inference time.</p>



<pre class="wp-block-code"><code>AdaptersConfig = {
    "Adapters": [
        {"AdapterId": ADAPTER_ID, "Version": "1", "Pages": ["1-5"]},
        {"AdapterId": ADAPTER_ID, "Version": "1", "Pages": ["6-*"]},
    ]
}</code></pre>



<p class="wp-block-paragraph">That <code>Pages</code> string takes digits, hyphens and an asterisk with no blank spaces, a page can only have one adapter applied to it, and an asterisk meaning all pages must be the only element in the list.</p>



<p class="wp-block-paragraph">Two things to weigh. AWS sets a minimum of five samples per query level for training or testing and says plainly that more is better, so treat five as a demo floor rather than a number that survives layout variation. And successful trainings per month are capped per account, so the tight annotate-train-evaluate loop you want is not available. Adapters suit one high-volume, house-specific document type with a stable schema. They are a poor fit for a long tail of one-off templates, which is what most legal archives actually are.</p>



<h2 class="wp-block-heading">Where Textract stops and a language model starts</h2>



<p class="wp-block-paragraph">This is a real decision now, so both sides deserve a hearing. Textract&#8217;s case is strong on standardised, high-volume documents. It is deterministic in a way generative models are not, it returns bounding-box geometry so you can point at exactly where a value came from, it gives per-element confidence, and it bills predictably per page. When an auditor asks why a field has the value it does, geometry beats a model&#8217;s reasoning.</p>



<p class="wp-block-paragraph">The generative side, whether that is Amazon Bedrock Data Automation as a managed document processing service or a model called directly through Amazon Bedrock, wins where the task needs reasoning rather than pattern matching. &#8220;Does this agreement contain an assignment restriction, and if so summarise it&#8221; is a comprehension question, and Textract has no answer for it. Generative approaches also cope better with layouts that were in nobody&#8217;s training set.</p>



<p class="wp-block-paragraph">Most production pipelines use both, which is what AWS&#8217;s own reference architectures push: Textract for faithful extraction with geometry and confidence, then a model over the extracted text for classification and comprehension. Choosing from scratch, ask whether your documents are standardised and high-volume. If yes, Textract plus a thin post-processing layer is cheaper and more auditable. If they are varied and volume is modest, a managed generative service gets you working faster. Google Document AI and Azure AI Document Intelligence cover similar ground if you are not committed to AWS.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Troubleshooting</h2>



<ul class="wp-block-list">
<li><strong>Extraction stops partway through long documents.</strong> You are not following <code>NextToken</code>. Compare the highest page number in your blocks against the source page count.</li>

<li><strong>Fields come back empty on forms you know were filled in.</strong> The value is a <code>SELECTION_ELEMENT</code> and you are only reading <code>WORD</code> children.</li>

<li><strong>The document is rejected as too large.</strong> You are on the synchronous API. Multi-page PDFs must go through <code>StartDocumentAnalysis</code>, and the 10 MB cap is on in-memory size, not file size on disk.</li>

<li><strong>Textract cannot read the document at all.</strong> Check for password protection and XFA-based forms. Both are documented as unsupported.</li>

<li><strong>The service cannot access the S3 object.</strong> Usually the execution role missing <code>s3:GetObject</code>, or a bucket in a different region from the Textract endpoint you called.</li>

<li><strong>Throttling under batch load.</strong> Transactions-per-second quotas are per account per region. Smooth spiky traffic through a queue and retry with exponential backoff and jitter rather than raising concurrency.</li>

<li><strong>Garbled text on reports that look fine to you.</strong> Multi-column layout without LAYOUT enabled, or a source scan below roughly 150 DPI. Re-scanning beats any amount of post-processing.</li>
</ul>



<h2 class="wp-block-heading">Common mistakes</h2>



<ul class="wp-block-list">
<li>Testing only on short documents, which hides the pagination bug entirely.</li>

<li>Enabling every feature type on every page because it is easier than deciding.</li>

<li>Running OCR over born-digital PDFs that already carry a text layer.</li>

<li>Treating a confidence score as a correctness score.</li>

<li>Scanning the block array linearly instead of building an ID map.</li>

<li>Polling for job completion inside a Lambda function instead of using SNS.</li>

<li>Storing only extracted values and discarding the raw JSON, so reprocessing means paying again.</li>
</ul>



<h2 class="wp-block-heading">Best practices</h2>



<ul class="wp-block-list">
<li>Persist the full Textract JSON to S3 keyed on a hash of the source document. Reprocessing then costs storage, not extraction.</li>

<li>Classify first, extract second. A cheap text-detection pass tells you which expensive feature set each document needs.</li>

<li>Validate on the value, not the score. Parse dates, check ranges, cross-reference names against other parts of the document.</li>

<li>Use aliases on every query and treat them as your schema contract.</li>

<li>Set a per-job page-count assertion and alert on mismatches. That is your canary for silent truncation.</li>

<li>Keep source documents in an encrypted bucket with a lifecycle policy. Contracts and correspondence should not accumulate indefinitely by accident.</li>

<li>Track cost per document rather than per API call. Feature stacking makes the per-call figure meaningless.</li>
</ul>



<h2 class="wp-block-heading">Frequently asked questions</h2>



<h3 class="wp-block-heading">Can Amazon Textract handle multi-page PDF contracts?</h3>



<p class="wp-block-paragraph">Yes, through the asynchronous API only. Synchronous operations cap PDF and TIFF at one page. Asynchronous operations accept PDF and TIFF up to 500 MB and 3,000 pages, staged through S3.</p>



<h3 class="wp-block-heading">Why is my Textract output missing the end of the document?</h3>



<p class="wp-block-paragraph">Almost certainly because you read the first response from <code>GetDocumentAnalysis</code> and stopped. Results are paginated with a <code>NextToken</code> and the job status still reports success, so nothing signals the truncation. Loop until <code>NextToken</code> is absent.</p>



<h3 class="wp-block-heading">How many Textract Queries can I ask per page?</h3>



<p class="wp-block-paragraph">AWS documents 15 per page for synchronous operations and 30 per page for asynchronous ones. Larger schemas need multiple passes, which costs more, so narrowing each query&#8217;s page range is worth the effort.</p>



<h3 class="wp-block-heading">Does Textract work on handwriting and signatures?</h3>



<p class="wp-block-paragraph">It detects handwriting, and the SIGNATURES feature returns the location and confidence of handwritten signatures, electronic signatures and initials. It does not verify identity: it tells you something was signed and where, not by whom.</p>



<h3 class="wp-block-heading">What languages does Amazon Textract support?</h3>



<p class="wp-block-paragraph">Text detection covers English, French, German, Italian, Portuguese and Spanish. Queries detection is English only, vertically written text is unsupported, and Textract does not return the detected language in its output.</p>



<h3 class="wp-block-heading">Should I use Textract or a generative model?</h3>



<p class="wp-block-paragraph">Textract when documents are standardised, volume is high, and you need auditable geometry and per-field confidence. A generative service when layouts vary widely, volume is modest, or the task needs comprehension rather than extraction. Most production pipelines use both.</p>



<h2 class="wp-block-heading">The one thing worth remembering</h2>



<p class="wp-block-paragraph">Amazon Textract data extraction fails quietly far more often than it fails loudly. A truncated result set, a tick box read as an empty string, a clause split across a page break, a confidently extracted wrong date: none of these throw an exception and none show up on a dashboard.</p>



<p class="wp-block-paragraph">So build the assertions that make silence detectable. Compare expected page counts against extracted ones, validate values rather than trusting scores, and keep the raw JSON so you can prove what the service actually returned. Textract is genuinely good at reading documents. Your job is noticing when it did not read all of them.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Need a document extraction pipeline that does not lose fields?</h2>



<p class="wp-block-paragraph">I build and fix document processing pipelines on AWS. Typical work looks like this:</p>



<ul class="wp-block-list">
<li>Auditing an existing Textract pipeline for silent truncation, dropped selection elements and mis-scoped queries</li>

<li>Designing the async architecture end to end: S3 intake, SNS and SQS fan-out, Lambda or container workers, dead letter handling and retries</li>

<li>Cutting per-page cost by classifying first, triaging born-digital PDFs out of the OCR path and narrowing feature types per page range</li>

<li>Turning block-graph JSON into a clean schema in CSV, a relational database or a data lake, with validation rules and a review queue for exceptions</li>

<li>Building the hybrid path where Textract handles extraction and a model on Amazon Bedrock handles classification and comprehension</li>

<li>Adding the observability that makes quiet failures loud: page-count assertions, per-document cost tracking, confidence distribution monitoring</li>
</ul>



<p class="wp-block-paragraph">If you have a redacted sample document, a chunk of Textract JSON or an extraction coming back half empty, send it over and I will tell you what is going wrong with it.</p>



<div class="wp-block-buttons is-layout-flex wp-block-buttons-is-layout-flex">
<div class="wp-block-button"><a class="wp-block-button__link wp-element-button" href="https://www.upwork.com/freelancers/~01f15a912ad84a6620" target="_blank" rel="noreferrer noopener">Work with me on Upwork</a></div>
</div>
<p>The post <a href="https://john-nessime.com/blog/devops/amazon-textract-data-extraction/">Amazon Textract Data Extraction: What Breaks on Real Contracts and Reports</a> appeared first on <a href="https://john-nessime.com/blog">John Nessime</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://john-nessime.com/blog/devops/amazon-textract-data-extraction/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
	</channel>
</rss>
