{"id":390,"date":"2026-09-22T09:00:00","date_gmt":"2026-09-22T06:00:00","guid":{"rendered":"https:\/\/john-nessime.com\/blog\/?p=390"},"modified":"2026-09-14T16:12:56","modified_gmt":"2026-09-14T13:12:56","slug":"amazon-textract-structured-data","status":"publish","type":"post","link":"https:\/\/john-nessime.com\/blog\/technical-guides\/amazon-textract-structured-data\/","title":{"rendered":"Turning External Documents Into Structured Data With Amazon Textract"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">A vendor sends over a year of delivery notes. It arrives as one PDF, several hundred pages, scanned on a machine that was clearly overdue for service. Someone asks how long it will take to get the reference numbers into the warehouse table, and the honest answer is that the OCR call is the easy part.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Document pipelines are not hard because optical character recognition is inaccurate. Amazon Textract reads text well. They are hard because everything after the API call belongs to you, and because Textract almost never fails loudly. You get a 200 response, a list of blocks, a confidence score, and a record with an empty field where the purchase order number should have been. Nothing alerts. The row lands in the warehouse with a null, and someone in finance notices three weeks later while reconciling.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">That silent gap is the whole problem. Getting reliable Amazon Textract structured data out of files you did not design has less to do with OCR accuracy than with everything around it: which operation to call for which document family, why the response is a graph rather than a record, how to route on confidence instead of trusting it, and the limits that quietly truncate your data. All of it assumes documents you cannot ask anyone to change.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Why external documents break pipelines that passed testing<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Internal documents are generated by a template you control. External documents vary along three axes at once, and each one has its own failure mode:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Container.<\/strong> One logical document per file, or forty stapled into one PDF, or a single page split across three JPEGs.<\/li>\n\n<li><strong>Text origin.<\/strong> Born-digital PDFs with a real text layer, flatbed scans, and photographs taken at an angle in bad light. Same file extension, wildly different accuracy.<\/li>\n\n<li><strong>Layout stability.<\/strong> The vendor redesigns their invoice template and moves the PO number from the header to a footer block. Nobody tells you.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">You cannot fix any of that upstream. You can only detect it at ingest and route accordingly. The sections below are organized by the failure family each one produces.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Failure family one: the file is out of spec before Textract sees it<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Textract accepts JPEG, PNG, PDF, and TIFF. It does not support XFA-based PDFs, and it will not open a password-protected PDF. Images are accepted up to 10,000 pixels on any side. For PDFs, the maximum page dimensions are 40 inches and 9,000 points.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Two more constraints catch people out. Text detection covers English, French, German, Italian, Portuguese, and Spanish, and several capabilities are narrower still: handwriting, invoices and receipts, identity documents, and Queries are English only. Vertical text, common in Japanese and Chinese documents, is not supported at all. Rotation is fine, including heavy in-plane rotation. Vertical writing direction is not.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Build an ingest gate that checks encryption, page count, byte size, and dimensions before it spends an API call, and give each rejection a reason code a human can act on. Then catch exceptions properly instead of wrapping the call in a bare retry. <code>BadDocumentException<\/code> means Textract could not read the file, and retrying fails identically. <code>DocumentTooLargeException<\/code> means route to the asynchronous path or split the file. <code>InvalidS3ObjectException<\/code> is usually IAM or a bucket-region mismatch, not a document problem. Only <code>InternalServerErrorException<\/code> deserves a retry with backoff.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Failure family two: synchronous versus asynchronous is a correctness gate<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">This is the one that silently eats data, so it goes early.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The synchronous operations, <code>DetectDocumentText<\/code> and <code>AnalyzeDocument<\/code>, cap at 10 MB in memory. For PDF and TIFF input they also cap at <strong>one page<\/strong>. Not one page per call with automatic paging. One page, full stop. Hand a synchronous call a 40-page PDF and you get page one, and a pipeline that reports success.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The asynchronous operations are a different shape. PDF and TIFF go up to 500 MB and 3,000 pages, while JPEG and PNG stay at 10 MB. You start a job, Textract works in the background, and you collect results by job ID.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A synchronous call looks like this. The document lives in S3, and you name the features you want:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>aws textract analyze-document \n  --document '{\"S3Object\":{\"Bucket\":\"my-doc-bucket\",\"Name\":\"inbound\/note-0042.png\"}}' \n  --feature-types '[\"FORMS\",\"TABLES\"]' \n  --region us-east-1<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">The asynchronous version returns a job identifier immediately and publishes completion to an SNS topic. That is what you want in production, because polling a job inside a Lambda function means paying for a function that sleeps:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>import boto3\n\ntextract = boto3.client(\"textract\")\n\nstart = textract.start_document_analysis(\n    DocumentLocation={\n        \"S3Object\": {\"Bucket\": \"my-doc-bucket\", \"Name\": \"inbound\/batch-0042.pdf\"}\n    },\n    FeatureTypes=[\"FORMS\", \"TABLES\"],\n    NotificationChannel={\n        \"SNSTopicArn\": \"arn:aws:sns:us-east-1:111122223333:textract-jobs\",\n        \"RoleArn\": \"arn:aws:iam::111122223333:role\/TextractSnsPublishRole\",\n    },\n)\n\njob_id = start[\"JobId\"]<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Results come back paginated, and forgetting the continuation token is a second, quieter way to lose most of a file:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>blocks = []\nnext_token = None\n\nwhile True:\n    kwargs = {\"JobId\": job_id}\n    if next_token:\n        kwargs[\"NextToken\"] = next_token\n\n    page = textract.get_document_analysis(**kwargs)\n    blocks.extend(page[\"Blocks\"])\n\n    next_token = page.get(\"NextToken\")\n    if not next_token:\n        break<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Check <code>JobStatus<\/code> before you trust <code>Blocks<\/code>. A job can come back partially successful, which means some pages processed and some did not. If you treat that as success you will publish a document that is missing its middle.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Failure family three: the response is a graph, not a record<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Textract does not return fields. It returns a flat list of <code>Block<\/code> objects, each with an ID and relationships pointing at other IDs. Reassembling those into a row is your code, and it is where most of the real work lives.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For form data, the shape is specific. A key-value pair is two blocks of type <code>KEY_VALUE_SET<\/code>, distinguished by their <code>EntityTypes<\/code> array containing either <code>KEY<\/code> or <code>VALUE<\/code>. The key block carries a relationship of type <code>VALUE<\/code> holding the ID of its value block, and a relationship of type <code>CHILD<\/code> holding the IDs of the <code>WORD<\/code> blocks that spell out the label. The value block carries only <code>CHILD<\/code> relationships pointing at its own words. So reading one field means three hops.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>blocks_by_id = {b[\"Id\"]: b for b in blocks}\n\ndef child_text(block):\n    parts = []\n    for rel in block.get(\"Relationships\", []):\n        if rel[\"Type\"] != \"CHILD\":\n            continue\n        for cid in rel[\"Ids\"]:\n            child = blocks_by_id[cid]\n            if child[\"BlockType\"] == \"WORD\":\n                parts.append(child[\"Text\"])\n            elif child[\"BlockType\"] == \"SELECTION_ELEMENT\":\n                parts.append(child[\"SelectionStatus\"])\n    return \" \".join(parts)\n\nfields = {}\nfor b in blocks:\n    if b[\"BlockType\"] != \"KEY_VALUE_SET\":\n        continue\n    if \"KEY\" not in b.get(\"EntityTypes\", []):\n        continue\n\n    value_text = \"\"\n    for rel in b.get(\"Relationships\", []):\n        if rel[\"Type\"] == \"VALUE\":\n            for vid in rel[\"Ids\"]:\n                value_text = child_text(blocks_by_id[vid])\n\n    fields[child_text(b)] = (value_text, b[\"Confidence\"])<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Two things about that snippet. The <code>SELECTION_ELEMENT<\/code> branch matters because checkboxes and radio buttons are values too, and their content is a selection status rather than text. And the confidence on the key block is not a separate key confidence: Textract returns the same value for both halves of a pair. If you planned to compare label confidence against value confidence to catch mismatches, that signal does not exist.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Tables follow the same pattern with different block types, and cells carry row and column indices plus span values for merged cells. Layout adds another set, including title, header, footer, section header, page number, list, and figure blocks, and it sequences text into human reading order. That last property is why Layout is worth requesting when the output is headed for a language model rather than a database: chunking on reading order beats chunking on raw line order.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Before hand-rolling any of this, look at the AWS-maintained parsers. The <code>amazon-textract-response-parser<\/code> package and the higher-level <code>amazon-textract-textractor<\/code> library both turn the block graph into objects with sensible properties. Writing your own traversal is a good way to understand the model and a poor way to spend two sprints.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Getting Amazon Textract structured data out of the right API<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Textract is a family of operations, not one endpoint, and the biggest single accuracy win available is usually picking a narrower one. Features are billed per page and they stack, so requesting <code>TABLES<\/code> on a corpus with no tables costs money and returns noise.<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>DetectDocumentText.<\/strong> Raw text and geometry. Use it when you only need a search index or a text layer.<\/li>\n\n<li><strong>AnalyzeDocument with FORMS.<\/strong> Key-value pairs discovered from labels on the page. Works when the label is physically present and stable.<\/li>\n\n<li><strong>AnalyzeDocument with TABLES.<\/strong> Grid reconstruction with row and column indices. The right call for line-item grids and statements.<\/li>\n\n<li><strong>AnalyzeDocument with QUERIES.<\/strong> You ask natural language questions and get answers back. This is the fix for layout drift.<\/li>\n\n<li><strong>AnalyzeDocument with LAYOUT and SIGNATURES.<\/strong> Reading order and structural elements; signature locations for compliance checks.<\/li>\n\n<li><strong>AnalyzeExpense.<\/strong> Purpose-built for invoices and receipts, with normalized field types.<\/li>\n\n<li><strong>AnalyzeID.<\/strong> Purpose-built for identity documents.<\/li>\n<\/ul>\n\n\n\n<h3 class=\"wp-block-heading\">Why Queries beats FORMS on documents you don&#8217;t control<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">FORMS gives you whatever labels the document happens to contain. That means your downstream schema is hostage to a vendor&#8217;s wording. When they change &#8220;P.O. No.&#8221; to &#8220;Order Reference&#8221;, your parser looks for a key that no longer exists and writes a null.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Queries inverts this. You supply the question and an alias, and the alias comes back attached to the answer, so your output key is fixed by you rather than by the document:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>aws textract analyze-document \n  --document '{\"S3Object\":{\"Bucket\":\"my-doc-bucket\",\"Name\":\"inbound\/note-0042.png\"}}' \n  --feature-types '[\"QUERIES\"]' \n  --queries-config '{\"Queries\":[\n      {\"Text\":\"What is the purchase order number?\",\"Alias\":\"PO_NUMBER\"},\n      {\"Text\":\"What is the delivery date?\",\"Alias\":\"DELIVERY_DATE\"}\n  ]}' \n  --region us-east-1<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">The alias is the entire point. It is a contract between your schema and the extraction layer that survives a vendor redesign. There are ceilings: 15 queries per page synchronously, 30 per page asynchronously, English only. Phrase questions using words that actually appear on the page, and add positional hints when you know the layout, because the model uses them. If Queries underperforms on your corpus, adapters let you tune the feature against your own examples and pass an adapter ID and version at call time.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Invoices are their own problem<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">For invoices and receipts, reach for AnalyzeExpense first. Its response is shaped around the domain rather than around the page: each <code>ExpenseDocument<\/code> holds <code>SummaryFields<\/code> for header-level data and <code>LineItemGroups<\/code> containing the line items. Fields carry a normalized <code>Type<\/code> such as <code>VENDOR_NAME<\/code>, plus a <code>ValueDetection<\/code> and an optional <code>LabelDetection<\/code>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">That optional label is the thing to notice. AnalyzeExpense can identify a vendor name printed only inside a logo, and pull line items out of a grid with no column headers. FORMS structurally cannot do either, because FORMS needs a key to exist. On receipts that is not a small accuracy delta, it is the difference between a usable field and an empty one.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Failure family four: confidence is per field and it is not your risk model<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Every extracted element carries a confidence score. Two mistakes follow from that.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The first is averaging into a document-level score. A document where every field is strong except the invoice total is not a mostly-good document. It is a document with one field that must not be trusted, surrounded by noise. Score fields, not documents.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The second is treating one threshold as universal. A misread address on a delivery note is an annoyance. A misread bank account number is an incident. Set thresholds per field from the cost of being wrong, then route into three lanes:<\/p>\n\n\n\n<ol class=\"wp-block-list\">\n<li><strong>Auto-accept.<\/strong> Above threshold and passing a format check. Straight into the target table.<\/li>\n\n<li><strong>Review.<\/strong> Below threshold, or above it but failing a validation rule. A human sees the value and the cropped region it came from.<\/li>\n\n<li><strong>Reject.<\/strong> Field absent entirely, or the document failed the ingest gate. Back to the sender with a reason.<\/li>\n<\/ol>\n\n\n\n<p class=\"wp-block-paragraph\">Validation rules do more work than thresholds here. A total that does not equal the sum of its line items is wrong at any confidence. Check-digit rules on account and tax identifiers catch character-level errors the model was perfectly confident about. Confidence tells you how sure the model is. Validation tells you whether the answer is even possible.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Textract can hand low-confidence results to Amazon Augmented AI through a human loop configuration on the API call, which saves building a review queue. Whether that beats a small internal tool depends on how much domain context your reviewers need on screen. Either way, capture the corrections. Reviewer edits are the only labeled data you get for free, and they are how you calibrate thresholds against your own corpus instead of guessing.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Failure family five: reprocessing you didn&#8217;t budget for<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Textract bills per page, per feature. Your parser, meanwhile, will change repeatedly in the first months of a pipeline&#8217;s life, because that is what parsers do.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">So separate extraction from parsing, permanently. Write the raw response JSON to S3 keyed by a hash of the document bytes plus the operation and feature set requested, then parse from that stored artifact. Reprocessing becomes free, deterministic, and diffable: you can replay a parser change across the whole back catalogue and compare outputs before promoting it. Without that cache, every parser bug fix is a bill.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>import hashlib\n\ndigest = hashlib.sha256(doc_bytes).hexdigest()\nfeatures = \"-\".join(sorted([\"FORMS\", \"TABLES\"]))\nkey = f\"textract-raw\/{digest}\/analyze-document\/{features}.json\"<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Hashing content rather than filename also deduplicates the same scan arriving twice under different names, which happens constantly with email-driven intake. For the asynchronous operations, a client request token gives idempotency at the job level; reusing a token with different parameters raises a mismatch error rather than quietly starting a second job.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">One caveat: source documents often contain personal data, so a cache is a retention decision. Set S3 lifecycle rules on both the originals and the raw JSON, and if working copies land on a workstation or a build box, delete them properly. Tools like O&amp;O SafeErase exist for exactly that gap between &#8220;moved to trash&#8221; and &#8220;actually gone&#8221;.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Where Textract stops being the right tool<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Being fair about this matters, because AWS now offers overlapping paths to the same outcome. Amazon Bedrock Data Automation is a managed document processing service that handles classification and extraction behind one API, priced per document rather than per page and per feature. AWS positions it as the default starting point for new builds, and for mixed document types with no dominant format that recommendation is sound. You write far less orchestration.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Textract keeps the advantage where the corpus is high volume and standardized, where geometry and per-field confidence are needed for auditability, where regional availability constrains you, or where you want extraction to be deterministic and separately testable. Bounding boxes and confidence scores are not a nice-to-have in a regulated workflow, they are the evidence trail.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The hybrid is common and sensible: Textract for extraction, a Bedrock model for reasoning over the extracted text. Just do not send a multimodal model a full-page image because OCR felt inconvenient. It costs considerably more per page and gives you no geometry.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Troubleshooting<\/h2>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Only the first page came through.<\/strong> A synchronous call on a multi-page PDF or TIFF. Move to the asynchronous operations, and assert the page count Textract reports against the count you expected.<\/li>\n\n<li><strong>Output truncates partway through a long document.<\/strong> Missing continuation token handling on the get operation. Loop until the token is absent.<\/li>\n\n<li><strong>Fields visible on the page never appear in output.<\/strong> The label is implied rather than printed, or sits inside a logo. Switch that field to Queries, or to AnalyzeExpense if it is an invoice.<\/li>\n\n<li><strong>Quality collapsed for one sender only.<\/strong> Template change. Diff the current layout against a stored sample, which is the argument for keeping raw responses.<\/li>\n\n<li><strong>Throttling during a backfill.<\/strong> Transactions per second and concurrent job counts are per-account, per-region quotas. Queue submissions rather than fanning out, and request an increase before the migration.<\/li>\n\n<li><strong>Table cells land in the wrong columns.<\/strong> Usually a skewed scan. Deskew and raise contrast before submitting, because image preparation is cheaper than post-processing repairs.<\/li>\n<\/ul>\n\n\n\n<h2 class=\"wp-block-heading\">Common mistakes<\/h2>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Testing on three clean sample documents supplied by the vendor rather than a random sample of real intake.<\/li>\n\n<li>Requesting every feature type on every page because it is simpler than routing, then being surprised by the bill.<\/li>\n\n<li>Using FORMS for invoices when AnalyzeExpense exists and understands unlabeled fields.<\/li>\n\n<li>Treating a confidence score as an accuracy guarantee instead of pairing it with validation rules.<\/li>\n\n<li>Discarding the raw JSON after parsing, which turns every parser fix into a re-extraction charge.<\/li>\n<\/ul>\n\n\n\n<h2 class=\"wp-block-heading\">Best practices<\/h2>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Classify before you extract. Route each document to the narrowest API that covers it.<\/li>\n\n<li>Use Queries with explicit aliases for fields your schema depends on, so vendor wording changes cannot break your column names.<\/li>\n\n<li>Store raw responses immutably, keyed by content hash, and parse from storage.<\/li>\n\n<li>Set per-field confidence thresholds from the cost of being wrong, and pair every threshold with a format or arithmetic validation.<\/li>\n\n<li>Instrument it like any other pipeline: extraction latency, straight-through rate, review queue depth, per-field null rate. A slow rise in nulls on one field is the earliest signal of a template change. Prometheus with Grafana, or Grafana Cloud if you would rather not run it yourself, covers this well.<\/li>\n\n<li>Keep documents and pipeline in the same region, use VPC endpoints for sensitive data, and apply lifecycle rules to originals and extracted output alike.<\/li>\n\n<li>Keep a golden set of real documents with known correct answers, and run it on every parser change.<\/li>\n<\/ul>\n\n\n\n<h2 class=\"wp-block-heading\">Frequently asked questions<\/h2>\n\n\n\n<h3 class=\"wp-block-heading\">Can Amazon Textract process multi-page PDFs?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Yes, but only through the asynchronous operations. Synchronous calls accept a single page for PDF and TIFF input. Asynchronous processing handles PDF and TIFF up to 500 MB and 3,000 pages.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">What is the difference between FORMS and Queries?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">FORMS returns whatever key-value pairs it finds from labels printed on the page, so your output keys follow the document. Queries lets you ask for values by natural language question and attach your own alias to each answer, so output keys stay fixed when the layout changes. For a schema you must keep stable, Queries is the safer contract.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">How much does Amazon Textract cost?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Billing is per page and varies by operation and feature, and requesting several features on one page adds their costs together. Check the current rate card for your region and multiply by pages, not documents. The practical lever is routing: send each document to the narrowest operation that answers your question, and cache raw responses so parser iteration is free.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Does Textract handle handwriting?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">It detects handwritten text, in English only. Treat handwritten fields as a separate confidence tier with lower auto-accept thresholds, because the variance across writers is much wider than for printed text.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Which languages does Amazon Textract support?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Text detection covers English, French, German, Italian, Portuguese, and Spanish. Handwriting, invoices and receipts, identity documents, and Queries are English only. Vertically written text is not supported.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Should I use Textract or Bedrock Data Automation?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Bedrock Data Automation is the lower-effort default for mixed document types, with classification and extraction behind a single managed API. Textract fits better at high volume on standardized documents, where per-field confidence and geometry matter for audit, or where extraction needs to be deterministic and independently testable. Combining them is normal.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">Conclusion<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">If one thing survives this post, make it this: Amazon Textract structured data extraction fails quietly, not loudly. The API returns success while your record is missing its most important field, because a synchronous call read one page of forty, or because a vendor renamed a label, or because a value came back at low confidence and nothing was watching.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">So build for the silence. Gate at ingest, use asynchronous operations by default, pick the narrowest API per document family, pin your schema with query aliases, store the raw response, and route on per-field confidence backed by validation rules that know what a possible answer looks like. Do that and the OCR really does become the easy part.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Need help turning documents into a pipeline you can trust?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Most of the difficulty in document extraction shows up after the proof of concept, once real intake replaces the sample files. That is the part I work on:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Designing the ingest gate and classification step so each document family reaches the right Textract operation instead of one call doing everything.<\/li>\n\n<li>Building the asynchronous path properly: S3 events, SNS completion, paginated retrieval, dead-letter handling, idempotent reprocessing.<\/li>\n\n<li>Replacing brittle FORMS parsing with query aliases or AnalyzeExpense, and validating the change against a golden set of your own documents.<\/li>\n\n<li>Setting per-field confidence thresholds and validation rules, then wiring a review queue whose corrections feed back into tuning.<\/li>\n\n<li>Adding observability that catches template drift early: per-field null rates, straight-through rate, and review queue depth on a dashboard.<\/li>\n\n<li>Getting the cost curve under control through routing, raw-response caching, and lifecycle rules on stored documents and JSON.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">If you have a pipeline misbehaving, send me a sample document with the fields you need and the raw Textract JSON it produced. That is usually enough to see where it is going wrong.<\/p>\n\n\n\n<div class=\"wp-block-buttons is-layout-flex wp-block-buttons-is-layout-flex\">\n<div class=\"wp-block-button\"><a class=\"wp-block-button__link wp-element-button\" href=\"https:\/\/www.upwork.com\/freelancers\/~01f15a912ad84a6620\" target=\"_blank\" rel=\"noreferrer noopener\">Work with me on Upwork<\/a><\/div>\n<\/div>\n","protected":false},"excerpt":{"rendered":"<p>Amazon Textract almost never fails loudly. A working guide to turning external PDFs and scans into structured data: picking the right operation per document family, parsing the block graph, routing on per-field confidence, and catching the limits that silently truncate your records.<\/p>\n","protected":false},"author":1,"featured_media":391,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[25,498,52],"tags":[161,429,216,93,186,568,431,569,231,217,440,313,436,435,124,172,233,207],"class_list":["post-390","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-cloud-computing","category-data-engineering","category-technical-guides","tag-amazon-s3","tag-amazon-sns","tag-amazon-textract","tag-aws","tag-aws-lambda","tag-bedrock-data-automation","tag-boto3","tag-confidence-scoring","tag-data-quality","tag-document-processing","tag-human-in-the-loop","tag-idempotency","tag-intelligent-document-processing","tag-ocr","tag-pipeline-design","tag-python","tag-schema-drift","tag-serverless","entry","has-media"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v28.5 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Amazon Textract: Structured Data From Documents<\/title>\n<meta name=\"description\" content=\"Turn external PDFs and scans into structured data with Amazon Textract: API choice, block parsing, confidence routing, and the failures that stay silent.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/john-nessime.com\/blog\/technical-guides\/amazon-textract-structured-data\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Amazon Textract: Structured Data From Documents\" \/>\n<meta property=\"og:description\" content=\"Turn external PDFs and scans into structured data with Amazon Textract: API choice, block parsing, confidence routing, and the failures that stay silent.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/john-nessime.com\/blog\/technical-guides\/amazon-textract-structured-data\/\" \/>\n<meta property=\"og:site_name\" content=\"John Nessime\" \/>\n<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/J.Nessime\" \/>\n<meta property=\"article:author\" content=\"https:\/\/www.facebook.com\/J.Nessime\" \/>\n<meta property=\"article:published_time\" content=\"2026-09-22T06:00:00+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/john-nessime.com\/blog\/wp-content\/uploads\/2026\/08\/amazon-textract-structured-data.png\" \/>\n\t<meta property=\"og:image:width\" content=\"1200\" \/>\n\t<meta property=\"og:image:height\" content=\"627\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"author\" content=\"John Nessime\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"John Nessime\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"15 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/technical-guides\\\/amazon-textract-structured-data\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/technical-guides\\\/amazon-textract-structured-data\\\/\"},\"author\":{\"name\":\"John Nessime\",\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/#\\\/schema\\\/person\\\/ede0b56d0c808f123f57d5d796902105\"},\"headline\":\"Turning External Documents Into Structured Data With Amazon Textract\",\"datePublished\":\"2026-09-22T06:00:00+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/technical-guides\\\/amazon-textract-structured-data\\\/\"},\"wordCount\":3221,\"publisher\":{\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/#\\\/schema\\\/person\\\/ede0b56d0c808f123f57d5d796902105\"},\"image\":{\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/technical-guides\\\/amazon-textract-structured-data\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/amazon-textract-structured-data.png\",\"keywords\":[\"Amazon S3\",\"Amazon SNS\",\"Amazon Textract\",\"AWS\",\"AWS Lambda\",\"Bedrock Data Automation\",\"Boto3\",\"Confidence Scoring\",\"Data Quality\",\"Document Processing\",\"Human In The Loop\",\"Idempotency\",\"Intelligent Document Processing\",\"OCR\",\"Pipeline Design\",\"Python\",\"Schema Drift\",\"Serverless\"],\"articleSection\":[\"Cloud Computing\",\"Data Engineering\",\"Technical Guides\"],\"inLanguage\":\"en-US\"},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/technical-guides\\\/amazon-textract-structured-data\\\/\",\"url\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/technical-guides\\\/amazon-textract-structured-data\\\/\",\"name\":\"Amazon Textract: Structured Data From Documents\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/technical-guides\\\/amazon-textract-structured-data\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/technical-guides\\\/amazon-textract-structured-data\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/amazon-textract-structured-data.png\",\"datePublished\":\"2026-09-22T06:00:00+00:00\",\"description\":\"Turn external PDFs and scans into structured data with Amazon Textract: API choice, block parsing, confidence routing, and the failures that stay silent.\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/technical-guides\\\/amazon-textract-structured-data\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/john-nessime.com\\\/blog\\\/technical-guides\\\/amazon-textract-structured-data\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/technical-guides\\\/amazon-textract-structured-data\\\/#primaryimage\",\"url\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/amazon-textract-structured-data.png\",\"contentUrl\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/amazon-textract-structured-data.png\",\"width\":1200,\"height\":627,\"caption\":\"Diagram showing an Amazon Textract response as a block graph of KEY_VALUE_SET and WORD nodes linked by VALUE and CHILD relationships, transformed by a parser into a single table row whose fields are routed to auto-accept, review, or reject lanes by confidence.\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/technical-guides\\\/amazon-textract-structured-data\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Turning External Documents Into Structured Data With Amazon Textract\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/#website\",\"url\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/\",\"name\":\"John Nessime\",\"description\":\"Cloud, DevOps, Data &amp; AI \u2014 Built, Tested, Explained\",\"publisher\":{\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/#\\\/schema\\\/person\\\/ede0b56d0c808f123f57d5d796902105\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":[\"Person\",\"Organization\"],\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/#\\\/schema\\\/person\\\/ede0b56d0c808f123f57d5d796902105\",\"name\":\"John Nessime\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/cropped-jn.png\",\"url\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/cropped-jn.png\",\"contentUrl\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/cropped-jn.png\",\"width\":512,\"height\":512,\"caption\":\"John Nessime\"},\"logo\":{\"@id\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/cropped-jn.png\"},\"description\":\"AWS Certified Solutions Architect helping businesses build reliable cloud, data, reporting, and automation solutions. I help startups, agencies, and growing businesses replace manual processes and disconnected data with practical AWS architectures, clean data pipelines, useful dashboards, and maintainable automation.\",\"sameAs\":[\"https:\\\/\\\/john-nessime.com\\\/blog\",\"https:\\\/\\\/www.facebook.com\\\/J.Nessime\",\"https:\\\/\\\/www.linkedin.com\\\/in\\\/john-m-nessime\"],\"url\":\"https:\\\/\\\/john-nessime.com\\\/blog\\\/author\\\/johnnessime\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Amazon Textract: Structured Data From Documents","description":"Turn external PDFs and scans into structured data with Amazon Textract: API choice, block parsing, confidence routing, and the failures that stay silent.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/john-nessime.com\/blog\/technical-guides\/amazon-textract-structured-data\/","og_locale":"en_US","og_type":"article","og_title":"Amazon Textract: Structured Data From Documents","og_description":"Turn external PDFs and scans into structured data with Amazon Textract: API choice, block parsing, confidence routing, and the failures that stay silent.","og_url":"https:\/\/john-nessime.com\/blog\/technical-guides\/amazon-textract-structured-data\/","og_site_name":"John Nessime","article_publisher":"https:\/\/www.facebook.com\/J.Nessime","article_author":"https:\/\/www.facebook.com\/J.Nessime","article_published_time":"2026-09-22T06:00:00+00:00","og_image":[{"width":1200,"height":627,"url":"https:\/\/john-nessime.com\/blog\/wp-content\/uploads\/2026\/08\/amazon-textract-structured-data.png","type":"image\/png"}],"author":"John Nessime","twitter_card":"summary_large_image","twitter_misc":{"Written by":"John Nessime","Est. reading time":"15 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/john-nessime.com\/blog\/technical-guides\/amazon-textract-structured-data\/#article","isPartOf":{"@id":"https:\/\/john-nessime.com\/blog\/technical-guides\/amazon-textract-structured-data\/"},"author":{"name":"John Nessime","@id":"https:\/\/john-nessime.com\/blog\/#\/schema\/person\/ede0b56d0c808f123f57d5d796902105"},"headline":"Turning External Documents Into Structured Data With Amazon Textract","datePublished":"2026-09-22T06:00:00+00:00","mainEntityOfPage":{"@id":"https:\/\/john-nessime.com\/blog\/technical-guides\/amazon-textract-structured-data\/"},"wordCount":3221,"publisher":{"@id":"https:\/\/john-nessime.com\/blog\/#\/schema\/person\/ede0b56d0c808f123f57d5d796902105"},"image":{"@id":"https:\/\/john-nessime.com\/blog\/technical-guides\/amazon-textract-structured-data\/#primaryimage"},"thumbnailUrl":"https:\/\/john-nessime.com\/blog\/wp-content\/uploads\/2026\/08\/amazon-textract-structured-data.png","keywords":["Amazon S3","Amazon SNS","Amazon Textract","AWS","AWS Lambda","Bedrock Data Automation","Boto3","Confidence Scoring","Data Quality","Document Processing","Human In The Loop","Idempotency","Intelligent Document Processing","OCR","Pipeline Design","Python","Schema Drift","Serverless"],"articleSection":["Cloud Computing","Data Engineering","Technical Guides"],"inLanguage":"en-US"},{"@type":"WebPage","@id":"https:\/\/john-nessime.com\/blog\/technical-guides\/amazon-textract-structured-data\/","url":"https:\/\/john-nessime.com\/blog\/technical-guides\/amazon-textract-structured-data\/","name":"Amazon Textract: Structured Data From Documents","isPartOf":{"@id":"https:\/\/john-nessime.com\/blog\/#website"},"primaryImageOfPage":{"@id":"https:\/\/john-nessime.com\/blog\/technical-guides\/amazon-textract-structured-data\/#primaryimage"},"image":{"@id":"https:\/\/john-nessime.com\/blog\/technical-guides\/amazon-textract-structured-data\/#primaryimage"},"thumbnailUrl":"https:\/\/john-nessime.com\/blog\/wp-content\/uploads\/2026\/08\/amazon-textract-structured-data.png","datePublished":"2026-09-22T06:00:00+00:00","description":"Turn external PDFs and scans into structured data with Amazon Textract: API choice, block parsing, confidence routing, and the failures that stay silent.","breadcrumb":{"@id":"https:\/\/john-nessime.com\/blog\/technical-guides\/amazon-textract-structured-data\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/john-nessime.com\/blog\/technical-guides\/amazon-textract-structured-data\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/john-nessime.com\/blog\/technical-guides\/amazon-textract-structured-data\/#primaryimage","url":"https:\/\/john-nessime.com\/blog\/wp-content\/uploads\/2026\/08\/amazon-textract-structured-data.png","contentUrl":"https:\/\/john-nessime.com\/blog\/wp-content\/uploads\/2026\/08\/amazon-textract-structured-data.png","width":1200,"height":627,"caption":"Diagram showing an Amazon Textract response as a block graph of KEY_VALUE_SET and WORD nodes linked by VALUE and CHILD relationships, transformed by a parser into a single table row whose fields are routed to auto-accept, review, or reject lanes by confidence."},{"@type":"BreadcrumbList","@id":"https:\/\/john-nessime.com\/blog\/technical-guides\/amazon-textract-structured-data\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/john-nessime.com\/blog\/"},{"@type":"ListItem","position":2,"name":"Turning External Documents Into Structured Data With Amazon Textract"}]},{"@type":"WebSite","@id":"https:\/\/john-nessime.com\/blog\/#website","url":"https:\/\/john-nessime.com\/blog\/","name":"John Nessime","description":"Cloud, DevOps, Data &amp; AI \u2014 Built, Tested, Explained","publisher":{"@id":"https:\/\/john-nessime.com\/blog\/#\/schema\/person\/ede0b56d0c808f123f57d5d796902105"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/john-nessime.com\/blog\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":["Person","Organization"],"@id":"https:\/\/john-nessime.com\/blog\/#\/schema\/person\/ede0b56d0c808f123f57d5d796902105","name":"John Nessime","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/john-nessime.com\/blog\/wp-content\/uploads\/2026\/07\/cropped-jn.png","url":"https:\/\/john-nessime.com\/blog\/wp-content\/uploads\/2026\/07\/cropped-jn.png","contentUrl":"https:\/\/john-nessime.com\/blog\/wp-content\/uploads\/2026\/07\/cropped-jn.png","width":512,"height":512,"caption":"John Nessime"},"logo":{"@id":"https:\/\/john-nessime.com\/blog\/wp-content\/uploads\/2026\/07\/cropped-jn.png"},"description":"AWS Certified Solutions Architect helping businesses build reliable cloud, data, reporting, and automation solutions. I help startups, agencies, and growing businesses replace manual processes and disconnected data with practical AWS architectures, clean data pipelines, useful dashboards, and maintainable automation.","sameAs":["https:\/\/john-nessime.com\/blog","https:\/\/www.facebook.com\/J.Nessime","https:\/\/www.linkedin.com\/in\/john-m-nessime"],"url":"https:\/\/john-nessime.com\/blog\/author\/johnnessime\/"}]}},"_links":{"self":[{"href":"https:\/\/john-nessime.com\/blog\/wp-json\/wp\/v2\/posts\/390","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/john-nessime.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/john-nessime.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/john-nessime.com\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/john-nessime.com\/blog\/wp-json\/wp\/v2\/comments?post=390"}],"version-history":[{"count":1,"href":"https:\/\/john-nessime.com\/blog\/wp-json\/wp\/v2\/posts\/390\/revisions"}],"predecessor-version":[{"id":416,"href":"https:\/\/john-nessime.com\/blog\/wp-json\/wp\/v2\/posts\/390\/revisions\/416"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/john-nessime.com\/blog\/wp-json\/wp\/v2\/media\/391"}],"wp:attachment":[{"href":"https:\/\/john-nessime.com\/blog\/wp-json\/wp\/v2\/media?parent=390"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/john-nessime.com\/blog\/wp-json\/wp\/v2\/categories?post=390"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/john-nessime.com\/blog\/wp-json\/wp\/v2\/tags?post=390"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}