<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Multi-Tenant | John Nessime</title>
	<atom:link href="https://john-nessime.com/blog/tag/multi-tenant/feed/" rel="self" type="application/rss+xml" />
	<link>https://john-nessime.com/blog/tag/multi-tenant/</link>
	<description>Cloud, DevOps, Data &#38; AI — Built, Tested, Explained</description>
	<lastBuildDate>Sun, 20 Sep 2026 16:20:12 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://john-nessime.com/blog/wp-content/uploads/2026/07/cropped-jn-32x32.png</url>
	<title>Multi-Tenant | John Nessime</title>
	<link>https://john-nessime.com/blog/tag/multi-tenant/</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>Tenant Isolation on AWS: Building a Multi-Tenant Workshop Platform That Doesn&#8217;t Leak</title>
		<link>https://john-nessime.com/blog/cloud-computing/tenant-isolation-aws-multi-tenant-saas/</link>
		
		<dc:creator><![CDATA[John Nessime]]></dc:creator>
		<pubDate>Mon, 31 Aug 2026 09:00:00 +0000</pubDate>
				<category><![CDATA[Cloud Computing]]></category>
		<category><![CDATA[Cloud Security]]></category>
		<category><![CDATA[SaaS Engineering]]></category>
		<category><![CDATA[ABAC]]></category>
		<category><![CDATA[Amazon Aurora]]></category>
		<category><![CDATA[Amazon S3]]></category>
		<category><![CDATA[AWS STS]]></category>
		<category><![CDATA[Cognito]]></category>
		<category><![CDATA[Cost Allocation Tags]]></category>
		<category><![CDATA[DynamoDB]]></category>
		<category><![CDATA[IAM]]></category>
		<category><![CDATA[Least Privilege]]></category>
		<category><![CDATA[Multi-Tenant]]></category>
		<category><![CDATA[Noisy Neighbor]]></category>
		<category><![CDATA[Pool Model]]></category>
		<category><![CDATA[Row-Level Security]]></category>
		<category><![CDATA[SaaS Architecture]]></category>
		<category><![CDATA[Session Policies]]></category>
		<category><![CDATA[Silo Model]]></category>
		<category><![CDATA[Tenant Isolation]]></category>
		<category><![CDATA[Tenant Onboarding]]></category>
		<category><![CDATA[Token Vending Machine]]></category>
		<category><![CDATA[Workshop Management]]></category>
		<guid isPermaLink="false">https://john-nessime.com/blog/?p=538</guid>

					<description><![CDATA[<p>A missing tenant filter doesn't throw an error, it returns a 200 with too many rows. This is how to build a multi-tenant workshop management platform on AWS where the isolation boundary sits below your application code: STS session tags feeding IAM conditions, DynamoDB leading keys, scoped S3 prefixes, forced PostgreSQL row-level security, and a control plane that verifies each new tenant is fenced before anyone logs in.</p>
<p>The post <a href="https://john-nessime.com/blog/cloud-computing/tenant-isolation-aws-multi-tenant-saas/">Tenant Isolation on AWS: Building a Multi-Tenant Workshop Platform That Doesn&#8217;t Leak</a> appeared first on <a href="https://john-nessime.com/blog">John Nessime</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">A ticket comes in from one of the workshops on the platform. A service advisor pulled the weekly job report and there&#8217;s a vehicle on it that never came through their door. Wrong registration, wrong customer, wrong shop.</p>



<p class="wp-block-paragraph">By the time that ticket lands, the leak already happened. You can patch the query in an hour. What you cannot do is tell the affected garage how many other reports were wrong, how long it had been wrong, or whether anyone downloaded a CSV. There&#8217;s no log that answers those questions, because nothing ever denied anything. The database happily returned the rows. The API happily serialised them.</p>



<p class="wp-block-paragraph">That&#8217;s the shape of the problem. Tenant isolation on AWS is not really about writing careful code. It&#8217;s about arranging things so that careless code fails loudly instead of quietly returning somebody else&#8217;s data.</p>



<p class="wp-block-paragraph">This post walks through building a multi-tenant workshop management platform on AWS with isolation baked in from the first commit: where tenant context comes from, how the boundary changes shape per storage service, what the control plane has to own, and how you prove any of it works. The example is a shop management system, with job cards, vehicle histories, parts inventory and technician timesheets, but the patterns apply to any vertical SaaS product where one customer&#8217;s records must never touch another&#8217;s.</p>



<h2 class="wp-block-heading">The failure mode that stays invisible</h2>



<p class="wp-block-paragraph">Most multi-tenant systems start with a <code>tenant_id</code> column and a convention: every query filters on it. That works right up until it doesn&#8217;t.</p>



<p class="wp-block-paragraph">The convention breaks in ordinary ways. Someone adds a reporting endpoint and copies a query from a script that ran as an admin. A new join pulls in a table nobody remembered to filter. An ORM lazy-loads a relationship and the filter lives on the parent, not the child. A background job that recalculates parts margins runs without any tenant in scope at all, because it processes everything.</p>



<p class="wp-block-paragraph">None of these throw. That&#8217;s the whole issue. A missing authorisation check produces a 403 you&#8217;ll notice in staging. A missing tenant filter produces a 200 with too many rows, and 200s don&#8217;t page anyone.</p>



<p class="wp-block-paragraph">The fix is not more discipline. It&#8217;s moving the filter somewhere the application cannot forget it: into IAM, into the database engine, or both. AWS makes this point directly in its own SaaS guidance, and it&#8217;s the right one. If your only defence against cross-tenant reads is that developers remember, you don&#8217;t have a defence, you have a habit.</p>



<h2 class="wp-block-heading">Pick the isolation model before you write code</h2>



<p class="wp-block-paragraph">There are three shapes, and they&#8217;re usually described as pool, silo and bridge.</p>



<ul class="wp-block-list">
<li><strong>Pool.</strong> Every tenant shares the same tables, buckets and compute. Cheapest to run, cheapest to deploy, and the model where isolation has to be enforced explicitly because nothing physical separates anyone.</li>



<li><strong>Silo.</strong> Each tenant gets dedicated resources: its own database, its own bucket, sometimes its own account. Isolation is close to free, operations are not. Migrations, deploys and monitoring all multiply by tenant count.</li>



<li><strong>Bridge.</strong> Mixed. Shared compute, dedicated storage, or pooled for the standard tier and siloed for the enterprise tier that asked hard questions in procurement.</li>
</ul>



<p class="wp-block-paragraph">For a workshop platform, most independent garages will be small, and pooling is the only sane starting point. The trap is treating that as permanent. Sooner or later a dealer group with forty sites will ask for a dedicated database, and if you have not left room for a per-tenant routing decision, retrofitting it means rewriting your data access layer.</p>



<p class="wp-block-paragraph">What I&#8217;d actually do: build pooled, but put the storage target behind a resolver from day one. A function that takes a tenant ID and returns a connection, a table name or a bucket prefix. When the first silo tenant arrives, you change the resolver, not four hundred call sites.</p>



<h2 class="wp-block-heading">Where tenant context comes from</h2>



<p class="wp-block-paragraph">This is the part people get subtly wrong, and it undermines everything downstream. The tenant identifier must come from the authenticated identity, never from the request. Not a header, not a query parameter, not a field in the JSON body. If a client can influence the tenant ID, your isolation model is decoration.</p>



<p class="wp-block-paragraph">In practice that means a custom claim in the token your identity provider issues. Amazon Cognito can carry a custom attribute for this, and so can any external IdP you federate with. The API layer reads the claim, and from that point the tenant is a fact about the caller rather than an input to the call.</p>



<p class="wp-block-paragraph">Then you push that fact down into AWS itself using a session tag. When your service assumes a role, it attaches the tenant as a tag on the session. Every subsequent AWS API call made with those credentials carries the tag in the request context, where IAM policies can reference it as <code>aws:PrincipalTag</code>.</p>



<pre class="wp-block-code"><code>import boto3

def scoped_session(tenant_id, role_arn):
    sts = boto3.client("sts")
    resp = sts.assume_role(
        RoleArn=role_arn,
        RoleSessionName=f"workshop-{tenant_id}",
        Tags=[{"Key": "TenantID", "Value": tenant_id}],
    )
    c = resp["Credentials"]
    return boto3.Session(
        aws_access_key_id=c["AccessKeyId"],
        aws_secret_access_key=c["SecretAccessKey"],
        aws_session_token=c["SessionToken"],
    )</code></pre>



<p class="wp-block-paragraph">Two things make this work, and both are easy to miss. The role&#8217;s trust policy has to allow <code>sts:TagSession</code> alongside the assume-role action, or the call fails. And the calling principal needs permission to pass that tag. Get either wrong and you&#8217;ll spend an afternoon reading an error that sounds like it&#8217;s about the role rather than the tag.</p>



<p class="wp-block-paragraph">One session tag, one role, one policy. That&#8217;s attribute-based access control, and it&#8217;s the reason ABAC scales where a role per tenant does not. Roles are a finite resource in an AWS account. Tags are not.</p>



<h2 class="wp-block-heading">Tenant isolation on AWS changes shape per service</h2>



<p class="wp-block-paragraph">There is no single isolation control. Each service exposes a different lever, and you have to learn every one your architecture touches.</p>



<h3 class="wp-block-heading">DynamoDB: the partition key is the boundary</h3>



<p class="wp-block-paragraph">If job cards live in DynamoDB with the tenant ID as the partition key, IAM can pin every read and write to that key. The condition key is <code>dynamodb:LeadingKeys</code>.</p>



<pre class="wp-block-code"><code>{
  "Effect": "Allow",
  "Action": [
    "dynamodb:GetItem",
    "dynamodb:PutItem",
    "dynamodb:UpdateItem",
    "dynamodb:Query"
  ],
  "Resource": "arn:aws:dynamodb:REGION:ACCOUNT:table/JobCards",
  "Condition": {
    "ForAllValues:StringEquals": {
      "dynamodb:LeadingKeys": ["${aws:PrincipalTag/TenantID}"]
    }
  }
}</code></pre>



<p class="wp-block-paragraph">Now a Query that omits the tenant partition key doesn&#8217;t return other tenants&#8217; job cards. It gets denied. That&#8217;s the behaviour you want: loud, logged, and impossible to miss in CloudTrail.</p>



<p class="wp-block-paragraph">Two caveats worth knowing before you commit. Scan operations don&#8217;t have a leading key to constrain, so granting <code>dynamodb:Scan</code> alongside this condition undermines the whole arrangement. And global secondary indexes have their own key structure, so an index whose partition key isn&#8217;t the tenant needs separate thought. Design the access patterns so that no query ever needs to look across tenants, and this stops being a problem.</p>



<h3 class="wp-block-heading">S3: two surfaces, not one</h3>



<p class="wp-block-paragraph">Vehicle photos, inspection PDFs and signed job sheets go to S3 under a per-tenant prefix. The mistake is scoping only the object actions and leaving <code>ListBucket</code> open, which lets a tenant enumerate every other garage&#8217;s filenames even without reading them. Filenames leak plenty: customer names, registration plates, invoice numbers.</p>



<p class="wp-block-paragraph">Object actions are scoped through the resource ARN. Listing is scoped through the <code>s3:prefix</code> request condition. You need both statements, because they protect different operations.</p>



<pre class="wp-block-code"><code>[
  {
    "Effect": "Allow",
    "Action": "s3:ListBucket",
    "Resource": "arn:aws:s3:::workshop-tenant-files",
    "Condition": {
      "StringLike": {
        "s3:prefix": ["${aws:PrincipalTag/TenantID}/*"]
      }
    }
  },
  {
    "Effect": "Allow",
    "Action": ["s3:GetObject", "s3:PutObject"],
    "Resource": "arn:aws:s3:::workshop-tenant-files/${aws:PrincipalTag/TenantID}/*"
  }
]</code></pre>



<p class="wp-block-paragraph">The policy variable in the resource ARN is doing real work there. One policy, every tenant, no template rendering at request time.</p>



<h3 class="wp-block-heading">Aurora PostgreSQL: row-level security, and the trap in it</h3>



<p class="wp-block-paragraph">Relational data is where most workshop platforms actually live, because job cards, parts lines and labour rates are relational. IAM cannot see inside a table, so the boundary moves into PostgreSQL itself via row-level security.</p>



<pre class="wp-block-code"><code>ALTER TABLE job_cards ENABLE ROW LEVEL SECURITY;
ALTER TABLE job_cards FORCE ROW LEVEL SECURITY;

CREATE POLICY tenant_isolation ON job_cards
  USING (tenant_id = current_setting('app.tenant_id', true))
  WITH CHECK (tenant_id = current_setting('app.tenant_id', true));</code></pre>



<p class="wp-block-paragraph"><code>USING</code> controls which rows are visible to reads, updates and deletes. <code>WITH CHECK</code> controls what can be written. Without the second clause, a tenant can read only its own rows but insert a row stamped with someone else&#8217;s tenant ID. The first protects the read path, the second protects the write path, and you want both.</p>



<p class="wp-block-paragraph">Now the trap, and it&#8217;s the one that gives teams false confidence. PostgreSQL superusers and roles carrying <code>BYPASSRLS</code> ignore row security entirely, and by default so does the table owner. If your application connects as the same role that ran the migrations, your policies are not in effect and your tests pass anyway. That&#8217;s why <code>FORCE ROW LEVEL SECURITY</code> is in the snippet above, and why the application should connect as a dedicated non-owner, non-superuser role.</p>



<p class="wp-block-paragraph">The second trap is connection reuse. Set the tenant with <code>SET LOCAL</code> inside a transaction, or with <code>set_config</code> using the transaction-local flag. Plain <code>SET</code> persists for the life of the connection, and a pooled connection outlives the request. That&#8217;s how one garage&#8217;s context ends up serving the next garage&#8217;s query.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Test row-level security as the application role, against a pooled connection, with at least two tenants in the table. Testing as the owner tells you nothing.</p>
</blockquote>



<h2 class="wp-block-heading">Compute isolation is a different question</h2>



<p class="wp-block-paragraph">Data isolation stops one tenant reading another&#8217;s records. It does nothing about one tenant consuming everyone&#8217;s capacity.</p>



<p class="wp-block-paragraph">Workshop platforms have a specific version of this. End of month, every garage runs its invoicing and MOT reminder batch at roughly the same time. A dealer group importing three years of service history will happily saturate a shared worker pool while forty independents wait for their job cards to load.</p>



<p class="wp-block-paragraph">Levers worth knowing, roughly in order of how much they cost you:</p>



<ul class="wp-block-list">
<li>API Gateway usage plans, keyed per tenant, to cap request rates at the edge before anything expensive runs.</li>



<li>Separate queues, or at minimum separate consumer concurrency, for bulk imports versus interactive requests. Bulk work should never share a lane with a screen someone is waiting on.</li>



<li>Reserved or provisioned concurrency on the Lambda functions serving interactive paths, so a batch surge cannot starve them.</li>



<li>Dedicated compute for premium tenants. This is silo by another name, and it&#8217;s the honest answer when a customer&#8217;s load profile genuinely doesn&#8217;t fit the pool.</li>
</ul>



<p class="wp-block-paragraph">Be honest with yourself about which problem you&#8217;re solving. Throttling is not isolation. It limits blast radius, it doesn&#8217;t create a boundary.</p>



<h2 class="wp-block-heading">What the control plane has to own</h2>



<p class="wp-block-paragraph">Separate the control plane from the application plane early. The control plane manages tenants; the application plane serves them. Mixing the two is how a bug in the onboarding flow ends up with production credentials.</p>



<p class="wp-block-paragraph">Onboarding a new workshop is a sequence, and it should be a single idempotent workflow rather than a checklist someone follows:</p>



<ol class="wp-block-list">
<li>Generate a non-guessable tenant identifier. Lowercase alphanumeric, no customer name in it, because it will end up inside resource ARNs and key prefixes.</li>



<li>Write the tenant record: tier, isolation model, storage target, status.</li>



<li>Provision identity. User pool group or IdP mapping, with the tenant claim wired in.</li>



<li>Provision storage. For a pooled tenant that&#8217;s a prefix and a seeded row. For a siloed one it&#8217;s real infrastructure, which is why this step must be asynchronous.</li>



<li>Apply tags used for cost allocation and reporting.</li>



<li>Run a verification step that proves the new tenant can reach its own data and cannot reach a canary tenant&#8217;s data.</li>
</ol>



<p class="wp-block-paragraph">Step six is the one teams skip. It&#8217;s also the only step that tells you the previous five worked.</p>



<p class="wp-block-paragraph">On per-tenant cost: activated cost allocation tags attribute anything that is a distinct tagged resource, which covers siloed tenants nicely. Pooled resources will not split by themselves, because a shared table doesn&#8217;t know which garage caused which read. If you need per-tenant margin, emit consumption as a metric dimension from the application: request counts, storage bytes, document pages processed. Tools like Vantage or CloudZero can allocate shared spend afterwards, but only from the signal you produce. Nothing recovers attribution you never recorded.</p>



<p class="wp-block-paragraph">The AWS SaaS Builder Toolkit is worth a look here. It codifies control plane concepts as CDK constructs and will save you real time on onboarding plumbing. Read its own guidance first: the project describes itself as sample code, and expects you to review security fit before production. That&#8217;s a fair description rather than a warning label, but treat it as a starting point, not a finished platform.</p>



<h2 class="wp-block-heading">Proving isolation actually holds</h2>



<p class="wp-block-paragraph">An isolation model you haven&#8217;t tried to break is a design document, not a control.</p>



<p class="wp-block-paragraph">The tests that earn their keep are negative ones. Seed two tenants. Authenticate as the first. Then deliberately do the wrong thing: request the second tenant&#8217;s job card by ID, list the second tenant&#8217;s S3 prefix, run a query with the session variable set to the wrong value. Every one of those should fail, and the test should assert on the failure.</p>



<ul class="wp-block-list">
<li>Run the cross-tenant suite on every pull request, not nightly. It&#8217;s the regression that matters most and the one most likely to be introduced by an innocent refactor.</li>



<li>Use the IAM policy simulator to check a policy change before it ships, particularly when someone widens a resource ARN.</li>



<li>Turn on IAM Access Analyzer so external access grants surface without anyone having to notice them.</li>



<li>Audit the database on a schedule: tables with RLS enabled but no policy attached, tables missing FORCE, application roles that have quietly acquired ownership or BYPASSRLS during an incident.</li>



<li>Alert on AccessDenied volume per tenant. A spike is either a bug you introduced or someone probing, and both are worth knowing about.</li>
</ul>



<p class="wp-block-paragraph">Keep a permanent canary tenant in every environment, including production, holding nothing but synthetic data. Every negative test targets it. It costs almost nothing and it means your isolation tests never need real customer records.</p>



<h2 class="wp-block-heading">Troubleshooting the failures you&#8217;ll actually hit</h2>



<h3 class="wp-block-heading">AccessDenied on a policy that looks correct</h3>



<p class="wp-block-paragraph">Almost always the session tag isn&#8217;t present. If the tag is missing from the request context, the condition can&#8217;t match and the statement doesn&#8217;t apply. Call <code>sts:GetCallerIdentity</code> with the scoped credentials and confirm you&#8217;re on the assumed role you think you are, then check CloudTrail for the AssumeRole event and look at whether the tag was actually passed. Tag keys are case sensitive, and <code>TenantId</code> is not <code>TenantID</code>.</p>



<h3 class="wp-block-heading">Queries return zero rows instead of the right rows</h3>



<p class="wp-block-paragraph">Classic RLS symptom. The session variable is unset, so the policy predicate compares against null and nothing matches. Check <code>current_setting</code> inside the same transaction as the query, not in a separate connection from your SQL client. Empty results are the safe failure here, which is exactly why they&#8217;re easy to misread as a data problem.</p>



<h3 class="wp-block-heading">One tenant intermittently sees another&#8217;s data</h3>



<p class="wp-block-paragraph">Intermittent means state reuse. Look at connection pooling first, then at any per-request context stored in thread-local or async-local storage that isn&#8217;t reset when the request finishes. If you&#8217;re using a proxy in front of the database, understand how it handles session state, because some proxies pin a connection to a client once session-level settings are detected, which changes the behaviour you tested against.</p>



<h3 class="wp-block-heading">Isolation works in the API but not in background jobs</h3>



<p class="wp-block-paragraph">Because the job has no request, so it has no token, so it has no tenant. Whatever runs asynchronously needs the tenant carried on the message and a session assumed per tenant when the work is processed. A worker that loops over all tenants with admin credentials is the single most common place isolation quietly stops applying.</p>



<h3 class="wp-block-heading">Session tags don&#8217;t survive a second AssumeRole</h3>



<p class="wp-block-paragraph">Session tags are not automatically carried forward when you chain roles unless they were marked transitive. If your architecture hops through more than one role, this is where the tenant context evaporates. Flatten the chain if you can; if you can&#8217;t, mark the tag transitive deliberately and document why.</p>



<h2 class="wp-block-heading">Common mistakes</h2>



<ul class="wp-block-list">
<li>Taking the tenant ID from a request header or body instead of the authenticated token.</li>



<li>Using the customer&#8217;s name or a sequential integer as the tenant identifier, then putting it in bucket prefixes where it becomes both guessable and enumerable.</li>



<li>Scoping S3 object actions but leaving bucket listing wide open.</li>



<li>Running the application as the PostgreSQL table owner, so RLS is enabled and silently inert.</li>



<li>Writing a USING clause with no WITH CHECK, leaving the write path open.</li>



<li>Creating one IAM role per tenant and discovering the account ceiling somewhere around the point the business gets interesting.</li>



<li>Assuming cost allocation tags will attribute pooled spend. They won&#8217;t, and by the time you need the numbers the history is gone.</li>



<li>Testing isolation only through the UI, where the frontend is already sending the right tenant every time.</li>
</ul>



<h2 class="wp-block-heading">Best practices worth the effort</h2>



<ul class="wp-block-list">
<li>Derive tenant context from the token, propagate it as a session tag, and never let application code choose it.</li>



<li>Enforce the boundary at the layer below your code: IAM conditions for AWS resources, RLS for relational rows.</li>



<li>Put storage targets behind a resolver so moving a tenant from pool to silo is a configuration change.</li>



<li>Make onboarding one idempotent workflow that ends in a verification step.</li>



<li>Emit tenant as a dimension on logs and metrics from the start, so cost and performance questions stay answerable.</li>



<li>Keep a canary tenant and run cross-tenant negative tests in CI on every change.</li>



<li>Put a WAF or edge layer such as Cloudflare in front of tenant subdomains, and keep tenant routing decisions out of the origin application where you can.</li>
</ul>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Frequently asked questions</h2>



<h3 class="wp-block-heading">Is a shared database with a tenant_id column ever acceptable?</h3>



<p class="wp-block-paragraph">Yes, provided the column is enforced by the engine rather than by convention. A shared table with row-level security, forced, queried by a non-owner role, is a real boundary. A shared table where every query is expected to include the filter is not. The schema is the same; the guarantee is completely different.</p>



<h3 class="wp-block-heading">ABAC or dynamically generated IAM policies?</h3>



<p class="wp-block-paragraph">ABAC for most cases. One role, one policy, tenant supplied per session, and nothing grows as you add customers. Dynamic policy generation, sometimes called a token vending machine, earns its place when a single policy genuinely cannot express the rule, such as when per-tenant resource names have to be injected rather than a key prefix. The cost is that you now own the correctness of a policy generator, plus per-request latency, and session policies have a size ceiling you can hit.</p>



<h3 class="wp-block-heading">Should each tenant get its own AWS account?</h3>



<p class="wp-block-paragraph">It&#8217;s the strongest boundary available and the most expensive to operate. For a workshop platform serving independent garages it&#8217;s overkill. It becomes reasonable when a customer&#8217;s contract, regulator or data residency requirement makes shared infrastructure a non-starter, and at that point you&#8217;re pricing it as a premium tier rather than absorbing it.</p>



<h3 class="wp-block-heading">How do I handle a user who works at two workshops?</h3>



<p class="wp-block-paragraph">Model it as one identity with multiple tenant memberships and an explicit active tenant per session, rather than a token carrying a list. The active tenant becomes the session tag. Switching workshops means a new session, which is exactly the behaviour you want because it makes the switch visible in your audit trail.</p>



<h3 class="wp-block-heading">Does row-level security hurt query performance?</h3>



<p class="wp-block-paragraph">It adds a predicate the planner has to satisfy, so the answer depends on your indexes. Index the tenant column, and index it as the leading column of composite indexes that support your common filters. Compare plans as the application role before and after enabling policies, because a plan captured as the owner may not reflect what the application actually runs.</p>



<h3 class="wp-block-heading">Where does application-level authorisation fit?</h3>



<p class="wp-block-paragraph">Alongside, not instead. Tenant isolation answers &#8220;which organisation&#8217;s data is this&#8221;. Authorisation answers &#8220;may this technician void an invoice&#8221;. Different questions, different layers. A policy engine such as Amazon Verified Permissions handles the second cleanly, and keeping them separate stops role logic creeping into your isolation boundary.</p>



<h3 class="wp-block-heading">Can I retrofit isolation onto a platform that already has tenants?</h3>



<p class="wp-block-paragraph">You can, and it&#8217;s tedious rather than impossible. Enforce at the database first, because that&#8217;s where the leak actually happens: table by table, FORCE enabled, application moved to a non-owner role. Then move the tenant ID out of request payloads and into the token. Add IAM conditions last, since they&#8217;re the least likely source of a live leak. Expect to find at least one background job with no tenant scope at all.</p>



<h2 class="wp-block-heading">The one thing to remember</h2>



<p class="wp-block-paragraph">Tenant isolation on AWS works when the boundary sits below your application code, in a layer that denies rather than trusts. IAM conditions on session tags for AWS resources, forced row-level security for relational data, and a control plane that verifies a new tenant is properly fenced before anyone logs in.</p>



<p class="wp-block-paragraph">Retrofitting that onto a running multi-tenant platform is possible but grim, because you&#8217;re doing it while real workshops have real data in the system. Doing it on day one costs a week. That&#8217;s the entire trade, and it&#8217;s not a close call.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Need a second pair of eyes on your multi-tenant architecture?</h2>



<p class="wp-block-paragraph">Most of the isolation problems I see are not exotic. They&#8217;re a missing WITH CHECK clause, an application connecting as the table owner, or a background job nobody scoped. Things I can help with:</p>



<ul class="wp-block-list">
<li>Reviewing an existing multi-tenant design and finding where the boundary is enforced by convention rather than by the platform.</li>



<li>Implementing ABAC with STS session tags across DynamoDB, S3 and Aurora, including the trust policy wiring that trips people up.</li>



<li>Setting up PostgreSQL row-level security correctly, with forced policies, a dedicated application role and pooling that doesn&#8217;t leak session state.</li>



<li>Building a tenant onboarding workflow that provisions identity, storage and tagging idempotently and verifies itself.</li>



<li>Writing the cross-tenant negative test suite and wiring it into CI so isolation regressions fail the build.</li>



<li>Adding per-tenant usage metering so cost, performance and tier decisions rest on data instead of guesses.</li>
</ul>



<p class="wp-block-paragraph">Send me a policy document, an RLS definition or a CloudTrail AccessDenied event and I&#8217;ll tell you what it&#8217;s actually enforcing.</p>



<div class="wp-block-buttons is-layout-flex wp-block-buttons-is-layout-flex">
<div class="wp-block-button"><a class="wp-block-button__link wp-element-button" href="https://www.upwork.com/freelancers/~01f15a912ad84a6620" target="_blank" rel="noreferrer noopener">Work with me on Upwork</a></div>
</div>
<p>The post <a href="https://john-nessime.com/blog/cloud-computing/tenant-isolation-aws-multi-tenant-saas/">Tenant Isolation on AWS: Building a Multi-Tenant Workshop Platform That Doesn&#8217;t Leak</a> appeared first on <a href="https://john-nessime.com/blog">John Nessime</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Legal Document Intelligence on AWS: The Five Boundaries That Have to Hold</title>
		<link>https://john-nessime.com/blog/cloud-computing/legal-document-intelligence-aws/</link>
		
		<dc:creator><![CDATA[John Nessime]]></dc:creator>
		<pubDate>Tue, 25 Aug 2026 18:00:00 +0000</pubDate>
				<category><![CDATA[Case Studies]]></category>
		<category><![CDATA[Cloud Computing]]></category>
		<category><![CDATA[Web Security]]></category>
		<category><![CDATA[Amazon Bedrock]]></category>
		<category><![CDATA[Amazon Comprehend]]></category>
		<category><![CDATA[Amazon S3]]></category>
		<category><![CDATA[Amazon Textract]]></category>
		<category><![CDATA[Architecture]]></category>
		<category><![CDATA[Audit Logging]]></category>
		<category><![CDATA[AWS]]></category>
		<category><![CDATA[AWS KMS]]></category>
		<category><![CDATA[AWS Organizations]]></category>
		<category><![CDATA[Bedrock Guardrails]]></category>
		<category><![CDATA[Bedrock Knowledge Bases]]></category>
		<category><![CDATA[Cloud Security]]></category>
		<category><![CDATA[Compliance]]></category>
		<category><![CDATA[Data Residency]]></category>
		<category><![CDATA[Document Processing]]></category>
		<category><![CDATA[Encryption]]></category>
		<category><![CDATA[Human In The Loop]]></category>
		<category><![CDATA[Intelligent Document Processing]]></category>
		<category><![CDATA[Legal Tech]]></category>
		<category><![CDATA[Litigation Hold]]></category>
		<category><![CDATA[Metadata Filtering]]></category>
		<category><![CDATA[Multi-Tenant]]></category>
		<category><![CDATA[OCR]]></category>
		<category><![CDATA[OpenSearch Serverless]]></category>
		<category><![CDATA[PrivateLink]]></category>
		<category><![CDATA[RAG]]></category>
		<category><![CDATA[S3 Object Lock]]></category>
		<category><![CDATA[Step Functions]]></category>
		<category><![CDATA[Tenant Isolation]]></category>
		<category><![CDATA[Vector Database]]></category>
		<guid isPermaLink="false">https://john-nessime.com/blog/?p=272</guid>

					<description><![CDATA[<p>A confident answer with a citation from a matter the reader was walled off from. Nothing crashed, nothing alerted. Building legal document intelligence on AWS with Textract, Bedrock and OpenSearch is the easy half; the hard half is tenant isolation, verified deletion, S3 Object Lock holds, what leaves the account, and an audit trail that names the actual user.</p>
<p>The post <a href="https://john-nessime.com/blog/cloud-computing/legal-document-intelligence-aws/">Legal Document Intelligence on AWS: The Five Boundaries That Have to Hold</a> appeared first on <a href="https://john-nessime.com/blog">John Nessime</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">The demo was going well until someone asked where the third citation came from.</p>



<p class="wp-block-paragraph">The answer on screen was good. Fluent, specific, correctly hedged, three sources listed underneath it. Then a partner in the room asked a simple question: which matter is source three from? Nobody could answer it in the room, and when we went and looked, it was from a matter that half the people watching the demo were formally walled off from.</p>



<p class="wp-block-paragraph">Nothing had crashed. No alarm fired. The retrieval layer had done exactly what a similarity search does, which is return the nearest vectors, and the nearest vectors did not care about the ethical wall. That is the thing about building a legal document intelligence platform on AWS: the failures that matter almost never look like failures. They look like a confident answer with a citation attached.</p>



<p class="wp-block-paragraph">This post is about the parts of that build that are hard, organised around the five boundaries that actually have to hold. The extraction pipeline is the easy half. I will cover it, but quickly, because the AWS documentation is good and the failure modes are visible. The other four boundaries fail quietly, and that is where the engineering goes.</p>



<h2 class="wp-block-heading">The shape of the thing</h2>



<p class="wp-block-paragraph">Before the boundaries, the skeleton. A legal document intelligence platform on AWS almost always ends up looking like this, whether you plan it or arrive at it:</p>



<ul class="wp-block-list">
<li>Documents land in Amazon S3, one prefix per matter, versioning on.</li>



<li>An S3 event triggers AWS Lambda, which starts an asynchronous Amazon Textract job.</li>



<li>Textract output lands back in S3 as JSON. Something normalises it into text plus page and bounding-box coordinates.</li>



<li>Amazon Comprehend or a foundation model classifies the document and pulls entities: parties, dates, governing law, clause types.</li>



<li>Chunks get embedded and written to a vector store, usually Amazon OpenSearch Serverless behind Amazon Bedrock Knowledge Bases.</li>



<li>An application layer, typically API Gateway plus Lambda, takes a question, retrieves, and calls a model on Amazon Bedrock.</li>
</ul>



<p class="wp-block-paragraph">AWS Step Functions is worth reaching for once you have more than three stages, because Textract jobs are long-running and Lambda timeouts are not a retry strategy. That is the whole architecture. You can stand it up in a fortnight. Then you spend six months on everything below.</p>



<h2 class="wp-block-heading">Boundary one: isolation, and why metadata filtering is not optional</h2>



<p class="wp-block-paragraph">This is the one from the demo, and it is the one that ends engagements.</p>



<p class="wp-block-paragraph">A vector index has no native concept of a matter, a client, or an ethical wall. If you pool every document into one index and rely on the prompt to keep things separate, you have built a system where a well-phrased question can pull privileged material across a wall. Prompt instructions are not an access control. They are a suggestion that usually works.</p>



<p class="wp-block-paragraph">There are two shapes that do work, and the choice between them is a real trade-off rather than a best practice.</p>



<h3 class="wp-block-heading">Pooled index with server-side filtering</h3>



<p class="wp-block-paragraph">Every chunk carries a <code>matter_id</code> and a <code>client_id</code> as metadata. Every retrieval call attaches a filter. With an S3 data source in Bedrock Knowledge Bases, the metadata comes from a sidecar JSON file that sits next to the document and carries a <code>metadataAttributes</code> object.</p>



<pre class="wp-block-code"><code>{
  "metadataAttributes": {
    "matter_id": "M-4417",
    "client_id": "C-108",
    "doc_type": "engagement_letter",
    "privileged": true
  }
}</code></pre>



<p class="wp-block-paragraph">The retrieval call then narrows the search before the model ever sees a chunk:</p>



<pre class="wp-block-code"><code>"retrievalConfiguration": {
  "vectorSearchConfiguration": {
    "filter": {
      "equals": { "key": "matter_id", "value": "M-4417" }
    }
  }
}</code></pre>



<p class="wp-block-paragraph">Two rules make this safe, and both are the kind of thing that gets skipped under deadline. First, the filter value is derived on the server from the authenticated caller&#8217;s claims, never accepted from the client. If the browser can send you a <code>matter_id</code>, the browser can send you a different one. Second, the filter is applied by code that no feature request can bypass. A helper that builds every retrieval request, and a code review rule that no other path may call the retrieve API directly.</p>



<p class="wp-block-paragraph">The trap here is ordering. You cannot filter on an attribute the index does not have. Add <code>matter_id</code> after ingestion and every chunk already in the index is invisible to that filter, which means it either returns for everyone or for no one, depending on how your filter is written. Design the metadata schema before the first ingestion run, not after the first demo.</p>



<h3 class="wp-block-heading">Separate collection per client</h3>



<p class="wp-block-paragraph">The heavier option. A separate OpenSearch Serverless collection per client, which buys you a separate AWS KMS key per client, separate index settings, and a failure mode where a bug in the filtering logic cannot reach across clients at all.</p>



<p class="wp-block-paragraph">It costs you sprawl. Every collection is a resource to provision, monitor, patch policy on, and eventually delete. Ingestion jobs run per data source per knowledge base, so freshness guarantees fragment. For a firm with twelve institutional clients this is fine. For a platform onboarding a hundred small clients it becomes the main operational burden of the product.</p>



<p class="wp-block-paragraph">My default is pooled with strict server-side filtering, and per-client collections only where the client&#8217;s own contract demands a dedicated encryption key. If you cannot say out loud which of those two you are running, you are running neither properly.</p>



<h2 class="wp-block-heading">Boundary two: extraction, and the confidence score everyone ignores</h2>



<p class="wp-block-paragraph">Textract returns a confidence score between 0 and 100 with each prediction. AWS is explicit in its own best-practice guidance that applications sensitive to detection errors should enforce a minimum threshold and route anything below it for human scrutiny, and that the right threshold depends entirely on how the output gets used.</p>



<p class="wp-block-paragraph">Legal documents sit at the harsh end of that. A misread date on an archival scan is a curiosity. A misread date on a limitation period is a problem with a name on it.</p>



<p class="wp-block-paragraph">Practical notes from this layer:</p>



<ul class="wp-block-list">
<li>Use the asynchronous operations for anything multipage. <code>StartDocumentAnalysis</code> and <code>GetDocumentAnalysis</code> handle long PDFs and TIFFs; the synchronous <code>AnalyzeDocument</code> path is for single pages and will fight you on a 400-page bundle.</li>



<li>Textract&#8217;s Queries feature earns its place on structured instruments. Instead of extracting everything and grepping, you ask the document a direct question and get a scoped answer with its own confidence.</li>



<li>Keep the page number and bounding box for every chunk. When a lawyer asks where a sentence came from, &#8220;page 14, second column&#8221; is an answer. &#8220;It&#8217;s in the corpus somewhere&#8221; is not, and the platform loses credibility the first time you say it.</li>



<li>Amazon Augmented AI wires low-confidence predictions into a human review queue with a private workforce, so the reviewers are your people rather than an anonymous pool. For privileged material that distinction is the entire point.</li>
</ul>



<p class="wp-block-paragraph">What I would skip early: building a custom entity recognition model. Comprehend supports custom entity recognition and it is genuinely useful, but you need labelled data you do not have yet, and a foundation model with a decent prompt gets you far enough to find out which entities the users actually care about. Train the custom model once the requirement has stopped moving.</p>



<h2 class="wp-block-heading">Boundary three: retention, where deletion is a distributed problem</h2>



<p class="wp-block-paragraph">Here is the failure that will not show up in any test suite you write.</p>



<p class="wp-block-paragraph">A client asks for a document to be removed. Someone deletes the S3 object. The ticket closes. The document is still fully searchable, because its chunks are sitting in the vector index and will stay there until the data source is re-synced. Meanwhile the Textract JSON output is in a second bucket, the extracted text may be in a database, and if model invocation logging is on, whole passages of the document are sitting in CloudWatch Logs or another S3 bucket entirely.</p>



<p class="wp-block-paragraph">One delete, at least four places holding a copy. Write the deletion path as a real workflow with an assertion at the end, not as a single API call:</p>



<ol class="wp-block-list">
<li>Delete the source object and, if versioning is on, its versions.</li>



<li>Delete the derived artefacts: Textract JSON, normalised text, any thumbnails or page images.</li>



<li>Trigger an ingestion job so the knowledge base drops the orphaned chunks.</li>



<li>Confirm removal by querying the index for the document identifier and expecting nothing back.</li>



<li>Deal with the logs, which means either a retention policy short enough to make the problem expire or a deliberate decision, written down, that logs are out of scope.</li>
</ol>



<p class="wp-block-paragraph">Step four is the one people leave out, and it is the only step that actually proves anything.</p>



<h3 class="wp-block-heading">The opposite problem: things that must not be deleted</h3>



<p class="wp-block-paragraph">Legal work has the reverse requirement too. When a matter goes into litigation hold, the documents need to survive an administrator with delete permissions and a bad afternoon.</p>



<p class="wp-block-paragraph">S3 Object Lock is the mechanism, and it has two independent controls that people routinely conflate. A retention period protects an object version until a fixed date, in either governance mode, which privileged users can override, or compliance mode, which nobody can override and where the period cannot be shortened. A legal hold has no date at all. It stays until someone with <code>s3:PutObjectLegalHold</code> explicitly removes it. The two can be active at once, and while either is active the object version cannot be deleted or overwritten.</p>



<pre class="wp-block-code"><code># Place an indefinite hold on a specific object version
aws s3api put-object-legal-hold 
  --bucket matters-archive 
  --key M-4417/exhibit-c.pdf 
  --version-id 3sL7f2Qz9pXvB1kR 
  --legal-hold Status=ON</code></pre>



<p class="wp-block-paragraph">Two things bite here. Object Lock requires versioning and turns it on automatically, and once enabled on a bucket you cannot turn it off or suspend versioning again. And compliance mode is genuinely permanent: if you set a seven-year retention when you meant seven days, that is the answer, for everyone, including the account root. Test in governance mode first. The bypass path exists and is deliberately awkward, requiring both the <code>s3:BypassGovernanceRetention</code> permission and an explicit <code>x-amz-bypass-governance-retention:true</code> header on the request, which is exactly the level of friction you want on that operation.</p>



<p class="wp-block-paragraph">Design the hold model before you design the deletion model, because holds win. A deletion request that collides with an active hold is a legal question, not an engineering one, and the platform&#8217;s job is to surface the collision clearly rather than resolve it silently.</p>



<h2 class="wp-block-heading">Boundary four: disclosure, or what leaves the account</h2>



<p class="wp-block-paragraph">This is the boundary the client&#8217;s general counsel will ask about, usually in writing, usually before signature. Three separate controls, and they are not interchangeable.</p>



<h3 class="wp-block-heading">The AI services opt-out policy</h3>



<p class="wp-block-paragraph">AWS Organizations has a governance policy type that opts your accounts out of having content processed by certain AI services stored and used for service improvement. The list includes Amazon Textract and Amazon Comprehend, which are precisely the two doing the reading in this architecture. The policy type has to be enabled at the organisation root before you can attach anything:</p>



<pre class="wp-block-code"><code>aws organizations enable-policy-type 
  --root-id r-example 
  --policy-type AISERVICES_OPT_OUT_POLICY</code></pre>



<p class="wp-block-paragraph">Read the AWS documentation on this one carefully rather than taking my summary as gospel, because there is a caveat in it that matters: the services may still need to store your data operationally even when you have opted out of it being used for improvement. Opting out is not the same as the data never existing outside your account. Say that plainly in the client conversation. It is a much better position than being asked about it later.</p>



<h3 class="wp-block-heading">Bedrock&#8217;s data position, and the log that undoes it</h3>



<p class="wp-block-paragraph">Amazon Bedrock&#8217;s published position is that inputs and outputs are not shared with third-party model providers and are not used to train the base models, and that fine-tuning operates on a private copy. That is a strong starting point for privileged content and it is the main reason Bedrock rather than a direct provider API shows up in these builds.</p>



<p class="wp-block-paragraph">Then there is model invocation logging. It captures full prompt and response payloads to CloudWatch Logs or S3, and you want it on, because without it you cannot reconstruct what the system told someone. But understand what you have just built: a log that contains privileged document text, at a per-region setting, in a destination that probably has looser access controls than the document store it came from. Encrypt it with the same KMS key discipline, restrict it harder than you think you need to, and set a retention period on purpose.</p>



<p class="wp-block-paragraph">Worth separating in your head: CloudTrail records that an API call happened and who made it. Model invocation logging records what was in it. You need both, for different questions, and only one of them contains client confidences.</p>



<h3 class="wp-block-heading">Network path</h3>



<p class="wp-block-paragraph">Interface VPC endpoints via AWS PrivateLink keep traffic to Bedrock, Textract and the rest on the AWS network rather than out through an internet gateway or NAT. Add a gateway endpoint for S3 while you are there.</p>



<p class="wp-block-paragraph">Endpoint policies are the part that gets forgotten. An endpoint without a policy is a private path to the whole service, including into accounts that are not yours. Scope it to the operations and resources you actually use, and pair it with an S3 bucket policy that rejects requests not arriving through your endpoint. The first protects what leaves; the second protects what can be reached.</p>



<h3 class="wp-block-heading">The parts that are not on the architecture diagram</h3>



<p class="wp-block-paragraph">Two of these have caught people out badly, and neither is an AWS control.</p>



<p class="wp-block-paragraph">The first is the development environment. You do not want real client documents on a laptop or in a scratch account, so build the parsing and chunking code against synthetic documents on a small separate box. A cheap VPS from a provider like Contabo or InterServer is fine for this, and the separation is worth more than the convenience of iterating in the production account.</p>



<p class="wp-block-paragraph">The second is the reviewers. If your human review loop involves contractors working remotely, their network path and their disks are part of your boundary whether or not you drew them. A managed VPN such as NordVPN or Surfshark handles the network side, and when a matter closes and local copies have to go, a dedicated erasure tool from something like O&amp;O Software does what dragging a folder to the bin does not. Unglamorous, and it is the layer that gets audited.</p>



<h2 class="wp-block-heading">Boundary five: evidence, because someone will ask</h2>



<p class="wp-block-paragraph">At some point the question stops being &#8220;does it work&#8221; and becomes &#8220;who saw what, and when&#8221;. If you cannot answer that from stored records, the platform is not defensible regardless of how good the answers are.</p>



<ul class="wp-block-list">
<li>Turn on CloudTrail data events for the document buckets. They are off by default, they are billed separately from management events, and they are the only way to see individual object-level reads.</li>



<li>Log the retrieval, not just the generation. Store which chunks came back and which filter was applied. When somebody asks whether the wall held on a specific query, this record is the answer and there is no reconstructing it later.</li>



<li>Carry the end user&#8217;s identity through to the audit record. A Lambda execution role in the logs tells you the platform did something. It does not tell you who asked.</li>



<li>Alarm on the filter, not just on errors. A retrieval that ran without a matter filter should page someone, even though it returned HTTP 200 and a perfectly good answer.</li>
</ul>



<p class="wp-block-paragraph">That last one is the closest thing to a single takeaway in this post. The dangerous events in this architecture are successful ones.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Troubleshooting the things that will go wrong</h2>



<h3 class="wp-block-heading">Retrieval returns nothing after you add a filter</h3>



<p class="wp-block-paragraph">Nearly always a metadata problem rather than a query problem. Either the sidecar metadata file was missing at ingestion, so the chunks carry no attribute to match, or the attribute was added after the chunks were indexed. Check whether the attribute exists on a known chunk before you touch the query. If it does not, you need a re-sync, not a different filter expression.</p>



<h3 class="wp-block-heading">Textract returns text but the layout is scrambled</h3>



<p class="wp-block-paragraph">Common on two-column pleadings and on documents with headers and footers on every page. Raw text detection reads in a reading order that is not always yours. Use the layout and table features rather than plain detection, and hold onto the geometry so you can reassemble columns yourself if the default order is wrong for that document class.</p>



<h3 class="wp-block-heading">The model cites a document that does not exist</h3>



<p class="wp-block-paragraph">Usually the citation is being generated rather than passed through. If document names are being written by the model instead of read from the retrieval response, it will invent plausible ones. Build citations from the retrieval result metadata in your application code and never from the generated text. Bedrock Guardrails helps with contextual grounding, but the structural fix is not asking the model to produce the reference in the first place.</p>



<h3 class="wp-block-heading">Ingestion is slow and nobody knows where</h3>



<p class="wp-block-paragraph">Instrument per stage before you optimise anything. In most of these pipelines the time is in Textract for large scanned bundles, and the fix is parallelism at the document level rather than tuning the extraction itself. Step Functions with a distributed map over documents gets you there without a queue you have to babysit.</p>



<h3 class="wp-block-heading">Costs climb faster than volume</h3>



<p class="wp-block-paragraph">Look at re-processing first. A pipeline that re-extracts a document every time anything downstream changes will quietly multiply your Textract bill against a static corpus. Key the extraction cache on the object&#8217;s content hash, not its path, and make re-ingestion an explicit action. The billing mechanisms differ per service, so read the current pricing pages rather than trusting a number from a blog post.</p>



<h2 class="wp-block-heading">Common mistakes</h2>



<ul class="wp-block-list">
<li><strong>Treating the system prompt as an access control.</strong> It is a formatting instruction that happens to look like a rule.</li>



<li><strong>Taking the tenant identifier from the request body.</strong> If the client can name the matter, the client can name someone else&#8217;s.</li>



<li><strong>Designing the metadata schema after the first ingestion run.</strong> Retrofitting a filterable attribute means re-indexing the whole corpus.</li>



<li><strong>Assuming a deleted S3 object is gone from the platform.</strong> The chunks, the extraction output and the invocation logs all outlive it.</li>



<li><strong>Setting compliance-mode retention without testing in governance mode.</strong> There is no support ticket that fixes a seven-year mistake.</li>



<li><strong>Enabling model invocation logging without treating the log as privileged.</strong> You have just made a second copy of the documents in a less protected place.</li>



<li><strong>Building VPC endpoints and leaving the default policy on them.</strong> A private path with no policy is still a path.</li>
</ul>



<h2 class="wp-block-heading">Best practices worth the effort</h2>



<ul class="wp-block-list">
<li>One retrieval helper, server-side filter injection, and a review rule that no other code path calls the retrieve API directly.</li>



<li>Metadata schema agreed and frozen before the first production ingestion. Include the fields you might filter on later, even if unused today.</li>



<li>Page-level provenance on every chunk, surfaced in the interface. It builds trust faster than any accuracy improvement.</li>



<li>A confidence threshold chosen per document class, with a human review path for everything under it.</li>



<li>Deletion implemented as a verified workflow across every store, ending in an assertion that the content is actually unretrievable.</li>



<li>Object Lock legal holds for matters under litigation hold, tested in governance mode before anything runs in compliance mode.</li>



<li>Customer-managed KMS keys across S3, the vector store and the logs, so key access is a control you can actually revoke.</li>



<li>An alarm on unfiltered retrievals, because the failure mode is a success response.</li>
</ul>



<h2 class="wp-block-heading">Frequently asked questions</h2>



<h3 class="wp-block-heading">Is a legal document intelligence platform on AWS safe for privileged material?</h3>



<p class="wp-block-paragraph">The services support it. Bedrock&#8217;s stated position is that inputs and outputs are not shared with model providers or used to train base models, PrivateLink keeps traffic off the public internet, and KMS covers encryption at rest. What determines safety is your own configuration: tenant isolation, log handling and deletion behaviour. The platform is exactly as confidential as its weakest copy of the text.</p>



<h3 class="wp-block-heading">Should I use Bedrock Knowledge Bases or build retrieval myself?</h3>



<p class="wp-block-paragraph">Start with Knowledge Bases. It handles chunking, embedding, sync and metadata filtering, and those are weeks of work with no differentiation in them. Build your own when you need retrieval behaviour it does not expose, such as unusual re-ranking or per-tenant chunking strategies. Going custom on day one usually means reimplementing the managed service badly.</p>



<h3 class="wp-block-heading">How do I stop one client&#8217;s documents surfacing in another client&#8217;s answers?</h3>



<p class="wp-block-paragraph">Tag every chunk with a tenant identifier at ingestion, and inject the matching filter server-side on every retrieval from the authenticated caller&#8217;s identity. Never accept the identifier from the client. If the contract requires per-client encryption keys, use a separate collection per client instead, accepting the extra operational load.</p>



<h3 class="wp-block-heading">Does deleting a document from S3 remove it from the search index?</h3>



<p class="wp-block-paragraph">No. The chunks stay in the vector store until an ingestion job runs and reconciles the data source. Until then the document is deleted and still fully searchable, which is the worst combination available. Always finish a deletion by querying the index and confirming nothing comes back.</p>



<h3 class="wp-block-heading">What is the difference between an S3 legal hold and a retention period?</h3>



<p class="wp-block-paragraph">A retention period runs until a fixed date and comes in governance mode, which privileged users can override, or compliance mode, which nobody can. A legal hold has no end date and stays until someone explicitly removes it. They are independent, they can both be active on the same object version, and while either is active the object cannot be deleted or overwritten.</p>



<h3 class="wp-block-heading">Do I need human review, or is the model accurate enough?</h3>



<p class="wp-block-paragraph">Accuracy is the wrong frame. The question is what happens to a wrong answer downstream. Where output feeds a decision with consequences, you want a confidence threshold and a review queue, and Amazon Augmented AI with a private workforce gives you that without building the review tooling yourself. Where output is a search aid a person will verify anyway, review is friction you do not need.</p>



<h3 class="wp-block-heading">Can I run this in a single AWS account?</h3>



<p class="wp-block-paragraph">Technically yes, and for a pilot it is reasonable. It gets uncomfortable once you have production documents and a development environment in the same place, because an IAM mistake has nowhere to stop. Separate accounts under Organizations also gives you the policy layer, including the AI services opt-out policy, which only exists at the organisation level.</p>



<h2 class="wp-block-heading">The one thing to take away</h2>



<p class="wp-block-paragraph">A legal document intelligence platform on AWS is not hard to build. Textract reads, Bedrock reasons, OpenSearch retrieves, and the tutorial version works on the first afternoon.</p>



<p class="wp-block-paragraph">What is hard is that its failures are silent and well-formed. The cross-matter citation, the deleted document that still answers questions, the privileged passage sitting in a log bucket, the retrieval that ran without a filter and returned a beautiful paragraph. None of these throw an error. All of them are the kind of thing that ends a client relationship.</p>



<p class="wp-block-paragraph">So build the boundaries first and the features second, and make sure every one of them is something you can prove from a stored record rather than something you believe about the code. If you can only take one habit from this: alarm on the successful requests that should not have been possible.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Need a second pair of eyes on your document pipeline?</h2>



<p class="wp-block-paragraph">I work on AWS document processing and retrieval systems where the confidentiality requirements are real. Typical things I get called in for:</p>



<ul class="wp-block-list">
<li>Reviewing tenant isolation in an existing RAG setup and finding the paths where the filter can be bypassed.</li>



<li>Designing the metadata and chunking schema before ingestion, so filtering works without a re-index later.</li>



<li>Building Textract pipelines with confidence thresholds, human review routing and page-level provenance.</li>



<li>Implementing verified deletion across S3, derived artefacts, the vector index and logs, with a proof step at the end.</li>



<li>Setting up retention and legal hold with S3 Object Lock, including the governance-mode rehearsal before compliance mode.</li>



<li>Locking down the network and logging path: VPC endpoints with real policies, KMS key separation, and audit records that name the actual user.</li>
</ul>



<p class="wp-block-paragraph">If you have an architecture diagram, a Step Functions definition or a retrieval request you are unsure about, send it over and I will tell you what I would change and why.</p>



<div class="wp-block-buttons is-layout-flex wp-block-buttons-is-layout-flex">
<div class="wp-block-button"><a class="wp-block-button__link wp-element-button" href="https://www.upwork.com/freelancers/~01f15a912ad84a6620" target="_blank" rel="noreferrer noopener">Work with me on Upwork</a></div>
</div>
<p>The post <a href="https://john-nessime.com/blog/cloud-computing/legal-document-intelligence-aws/">Legal Document Intelligence on AWS: The Five Boundaries That Have to Hold</a> appeared first on <a href="https://john-nessime.com/blog">John Nessime</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Redshift Cost Optimization for SaaS Analytics: The Levers That Actually Move the Bill</title>
		<link>https://john-nessime.com/blog/devops/redshift-cost-optimization-saas-analytics/</link>
		
		<dc:creator><![CDATA[John Nessime]]></dc:creator>
		<pubDate>Tue, 04 Aug 2026 09:18:00 +0000</pubDate>
				<category><![CDATA[Cloud Computing]]></category>
		<category><![CDATA[DevOps]]></category>
		<category><![CDATA[Technical Guides]]></category>
		<category><![CDATA[Amazon Redshift]]></category>
		<category><![CDATA[Architecture]]></category>
		<category><![CDATA[AWS]]></category>
		<category><![CDATA[Business Intelligence]]></category>
		<category><![CDATA[Cloud]]></category>
		<category><![CDATA[Cost Optimization]]></category>
		<category><![CDATA[Data Engineering]]></category>
		<category><![CDATA[Data Warehouse]]></category>
		<category><![CDATA[Embedded Analytics]]></category>
		<category><![CDATA[FinOps]]></category>
		<category><![CDATA[Monitoring]]></category>
		<category><![CDATA[Multi-Tenant]]></category>
		<category><![CDATA[Row-Level Security]]></category>
		<category><![CDATA[SQL]]></category>
		<guid isPermaLink="false">https://john-nessime.com/blog/?p=123</guid>

					<description><![CDATA[<p>In a SaaS analytics product, the Redshift bill tracks how often queries arrive, not how much data they touch. Here is how the meter actually works, why connection pools bill you while nobody is using the product, how to attribute spend to a tenant, and which isolation choices quietly cost more than they save.</p>
<p>The post <a href="https://john-nessime.com/blog/devops/redshift-cost-optimization-saas-analytics/">Redshift Cost Optimization for SaaS Analytics: The Levers That Actually Move the Bill</a> appeared first on <a href="https://john-nessime.com/blog">John Nessime</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">The message usually comes from whoever owns the AWS bill, and it is never dramatic. &#8220;Redshift is up again this month. Did we onboard someone big?&#8221; Nobody onboarded anyone. Nobody shipped a new dashboard. Query volume looks flat on the Grafana board. The bill moved anyway.</p>



<p class="wp-block-paragraph">That gap between what you think you are paying for and what you are actually paying for is what makes Redshift cost optimization awkward in a SaaS analytics product. You are not running one nightly batch against a warehouse that sleeps the rest of the day. You are serving hundreds of small, latency-sensitive queries that fire whenever a customer opens a dashboard, plus ingestion, plus whatever your BI layer and your connection pool are doing when nobody is watching.</p>



<p class="wp-block-paragraph">This post covers the levers that genuinely move that number: how the meter works, why idle-looking connections still bill, how to work out which tenant is expensive, and which isolation choices cost more than they save. Where the popular advice is wrong for SaaS specifically, I will say so.</p>



<h2 class="wp-block-heading">How Amazon Redshift actually charges you</h2>



<p class="wp-block-paragraph">Three buckets, and they behave very differently.</p>



<ul class="wp-block-list"><li><strong>Compute.</strong> On Redshift Serverless this is RPU-hours, metered per second. On provisioned clusters it is node-hours, plus separate line items for concurrency scaling and Spectrum.</li><li><strong>Storage.</strong> Redshift Managed Storage, billed by GB per month, independent of compute. Snapshots are storage too.</li><li><strong>Everything else.</strong> Cross-region data sharing and snapshot replication, machine learning, data transfer outside the usual in-region S3 paths.</li></ul>



<p class="wp-block-paragraph">In a SaaS analytics workload compute dominates, often overwhelmingly. And the important part: compute is a function of how long the warehouse is awake and at what capacity, not how many rows you touched. Two teams can scan identical data volumes and get bills that differ by a factor of five, purely because of how their queries arrive.</p>



<h2 class="wp-block-heading">The billing mechanic that catches SaaS teams out</h2>



<p class="wp-block-paragraph">Read the serverless billing notes properly once and a lot of mysterious spend stops being mysterious. The parts that matter:</p>



<ul class="wp-block-list"><li>The minimum charge is 60 seconds of resource usage, metered per second beyond that. This is a minimum for the warehouse, not for each individual query.</li><li>Usage is recorded when a transaction <em>completes</em>, rolls back, or is stopped. A transaction that runs for hours shows up in your usage view only at the end.</li><li>Cancel a query before it finishes and you still pay for the time it ran.</li><li>Querying system tables is billed like any other query. Your monitoring loop is a workload.</li><li>After a burst, capacity can stay elevated for a period after the load drops. Scale-down is not instant.</li></ul>



<p class="wp-block-paragraph">Put those together and you reach a conclusion that irritates most engineers: on serverless, ten small queries crammed into one minute are cheaper than the same ten queries spread across ten minutes. Every wake-up costs you a minimum billing window multiplied by your base capacity. That is the opposite of the instinct you have from tuning an OLTP service, where you smooth load out to protect tail latency.</p>



<p class="wp-block-paragraph">Before you change anything, get the real numbers out of the warehouse rather than out of Cost Explorer, which lags and aggregates.</p>



<pre class="wp-block-code"><code>-- Daily billed RPU-seconds converted to RPU-hours.
-- Multiply by your region's on-demand RPU-hour rate for dollars.
SELECT trunc(start_time) AS day,
       sum(charged_seconds) / 3600::double precision AS rpu_hours
FROM   sys_serverless_usage
GROUP  BY 1
ORDER  BY 1 DESC;</code></pre>



<p class="wp-block-paragraph"><code>charged_seconds</code> is the column to build cost reporting on. <code>compute_seconds</code> is informative but it is not what the invoice is derived from, and the two can disagree within a given interval. Two constraints worth knowing before you wire this into a dashboard: the view holds roughly a week of history, and it is visible only to superusers. If you want month-over-month trends, UNLOAD it to S3 on a schedule and query the archive with Amazon Athena instead.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Lever one: connections that look idle and are not</h2>



<p class="wp-block-paragraph">This is the one that bites hardest and shows up last, because there is nothing to see. AWS documents it plainly: Redshift Serverless treats all incoming queries as billable user activity, including lightweight health-check queries sent by connection pools. It does not matter whether the statement came from your application, a JDBC driver, or a pooling framework doing its job.</p>



<p class="wp-block-paragraph">So a pool that fires <code>SELECT 1</code> every thirty seconds to validate connections is a warehouse that never gets to sleep. Your product has no users at 3am and you are still paying the minimum window, over and over, multiplied by base capacity. HikariCP, Apache Commons DBCP and PgBouncer all have some form of this behaviour, and the defaults are tuned for OLTP databases where a validation query costs nothing.</p>



<p class="wp-block-paragraph">Open transactions are the same problem wearing a different hat. A <code>BEGIN</code> without a matching <code>COMMIT</code> or <code>ROLLBACK</code> keeps consuming RPUs until the session ends. Session timeouts exist precisely because this happens.</p>



<p class="wp-block-paragraph">What I would check, in this order:</p>



<ol class="wp-block-list"><li>Disable the pool&#8217;s validation or heartbeat query entirely if the driver allows it. If it does not, stretch the interval as far as your failure tolerance permits.</li><li>Drop idle pool size to something honest. A pool sized for peak that stays warm overnight is pure waste on this pricing model.</li><li>Fix any code path that opens a transaction and returns early on error without ending it.</li><li>Set a session timeout per application role so a leaked connection cannot bill indefinitely.</li></ol>



<pre class="wp-block-code"><code>-- Cap idle sessions for the application role.
-- Value is in seconds; the documented range is 60 to 1,728,000.
ALTER USER analytics_app SESSION TIMEOUT 1800;

-- Cap how many connections a single role can hold open at once.
ALTER USER analytics_app CONNECTION LIMIT 40;

-- What is connected right now, and with what timeout.
SELECT * FROM stv_sessions;</code></pre>



<p class="wp-block-paragraph">Session timeout changes apply to new sessions only, so recycle the pool afterwards or you will conclude the setting does nothing.</p>



<h2 class="wp-block-heading">Lever two: base capacity, max capacity and usage limits are three different things</h2>



<p class="wp-block-paragraph">These get conflated constantly, and two of them will not save you a cent on their own.</p>



<ul class="wp-block-list"><li><strong>Base capacity (base RPU).</strong> The floor. It multiplies every billed second, including that 60-second minimum. Halving base capacity roughly halves the cost of a warehouse dominated by short queries. It also halves the compute those queries get, so watch p95 latency alongside the bill.</li><li><strong>Max capacity (MaxRPU).</strong> A ceiling on how far automatic scaling can go. It caps compute available to the workgroup, it does not stop queries and it does not interrupt anything running. Useful as a guard rail against a runaway scan, useless as a budget.</li><li><strong>Usage limits.</strong> An actual budget, expressed in RPU-hours over a daily, weekly or monthly period. The breach actions are: log to a system table, raise an SNS alert, or turn off user queries.</li></ul>



<p class="wp-block-paragraph">Only the third one can stop you spending money, and only the third one can take your product down at 2pm on a Tuesday. Set it to alert first, live with it for a full billing cycle so you learn the shape of a normal week, then decide whether you are genuinely willing to have queries turned off. In a customer-facing SaaS product the answer is usually no, and the limit stays as an alarm feeding PagerDuty or whatever you already page from.</p>



<p class="wp-block-paragraph">There is also the price-performance target, the slider that hands scaling decisions to AWS in exchange for a stated cost or speed preference. AWS recommends it for mid-range base capacities and advises against it at the very bottom and very top of the RPU scale, so check the current guidance against your base setting before enabling it. It is worth trying on a staging workgroup with a replayed query mix; it is not worth switching on blind in production.</p>



<p class="wp-block-paragraph">On provisioned clusters the equivalent controls are per-feature usage limits: concurrency scaling measured in time, Spectrum measured in data scanned, cross-region data sharing, and extra compute for automatic optimization. Each takes a breach action of log, emit a metric, or disable the feature. Concurrency scaling also earns free credits as the main cluster runs, which is why a moderately bursty provisioned cluster often shows no concurrency scaling charge at all until it suddenly does.</p>



<h2 class="wp-block-heading">Lever three: Redshift cost optimization starts with knowing which tenant is expensive</h2>



<p class="wp-block-paragraph">Be clear-eyed about what is possible here. On serverless you cannot get an exact dollar figure per query, because billing happens at the warehouse level and the minimum charge is shared across whatever else was running in that window. What you can build is a defensible apportionment, and that is enough to find the customer whose scheduled export is quietly eating your margin.</p>



<p class="wp-block-paragraph">Start by labelling every statement your API issues on a tenant&#8217;s behalf.</p>



<pre class="wp-block-code"><code>-- Set in the pool's per-checkout init SQL, or per request.
SET query_group TO 'tenant_4417';

SELECT metric_date, sum(events)
FROM   fact_events
WHERE  tenant_id = 4417
  AND  metric_date &gt;= dateadd(day, -30, current_date)
GROUP  BY 1;

RESET query_group;</code></pre>



<p class="wp-block-paragraph">The label lands in the query log and surfaces as <code>query_label</code> in the SYS monitoring views. Keep it short: the older query log views truncate the label to 30 characters, so a tenant slug beats a UUID with prefixes bolted on.</p>



<pre class="wp-block-code"><code>-- Seven days of activity grouped by tenant label.
-- Note: time columns in the SYS views are microseconds;
-- confirm units before converting anything to money.
SELECT trim(query_label)   AS tenant,
       count(*)            AS queries,
       sum(execution_time) AS exec_time,
       sum(queue_time)     AS queue_time
FROM   sys_query_history
WHERE  start_time &gt; dateadd(day, -7, sysdate)
  AND  query_label LIKE 'tenant_%'
GROUP  BY 1
ORDER  BY exec_time DESC;</code></pre>



<p class="wp-block-paragraph">Three columns in that view earn their keep beyond the obvious ones. <code>result_cache_hit</code> tells you which dashboard queries are already free, which is often a bigger share than people expect. The split between <code>queue_time</code> and <code>execution_time</code> tells you whether you have a tuning problem or a capacity problem, and those have opposite fixes. And <code>user_query_hash</code> groups repeated queries with different literals, which is exactly what an embedded dashboard produces, so it is the fastest way to find the one panel that fifty tenants are running badly.</p>



<p class="wp-block-paragraph">From there, apportion the day&#8217;s <code>charged_seconds</code> by each tenant&#8217;s share of execution time. It is an approximation and you should label it as one when you show it to finance. It is still the difference between &#8220;Redshift costs us a lot&#8221; and &#8220;eleven percent of our warehouse spend is one customer pulling an unbounded date range every fifteen minutes.&#8221;</p>



<h2 class="wp-block-heading">Lever four: the isolation model you picked is a cost decision</h2>



<p class="wp-block-paragraph">AWS&#8217;s SaaS guidance describes three partitioning models, and each one has a distinct cost signature on Redshift.</p>



<ul class="wp-block-list"><li><strong>Pool.</strong> All tenants share tables with a tenant identifier column. Cheapest by a wide margin, one warehouse to keep warm, one set of statistics. You pay for it in noisy-neighbour risk and in the access-control work you now have to do yourself.</li><li><strong>Bridge.</strong> Separate schemas or databases inside one cluster. Sounds like a compromise, behaves like neither. AWS&#8217;s own whitepaper is fairly blunt that the isolation profile does not usually justify it, since cluster-level access grants reach across the databases anyway.</li><li><strong>Silo.</strong> A warehouse per tenant. Clean boundaries and per-tenant cost visibility for free. On serverless it is also the most expensive thing you can do, because every workgroup carries its own base capacity floor and its own 60-second minimums. Twenty small tenants means twenty warehouses waking up independently.</li></ul>



<p class="wp-block-paragraph">Data sharing sits between these and is the pattern I reach for when workload interference is the real problem. One producer handles ingestion and transformation; consumers read the shared data without copying it, and a consumer&#8217;s load does not touch the producer. Genuinely useful for separating a heavy ETL window from customer-facing reads. But be honest about the arithmetic: every consumer is its own billable warehouse. Data sharing buys you performance isolation, not cheaper compute.</p>



<p class="wp-block-paragraph">In a pooled model, the thing I set up first is a sort key that leads with the tenant identifier followed by the time column everyone filters on. That lets Redshift prune blocks before it reads them instead of scanning broadly and filtering afterwards. Combine it with row-level security so the tenant predicate cannot be forgotten by an application bug, and you have removed both the largest cost driver and the scariest failure mode in one change.</p>



<h2 class="wp-block-heading">Lever five: scan less, refresh less</h2>



<p class="wp-block-paragraph">Classic warehouse hygiene still applies, it just pays differently here. Shorter queries mean fewer billed seconds at your base capacity.</p>



<ul class="wp-block-list"><li><strong>Sort keys that match your real predicates.</strong> Not the ones from the design doc. Pull the top twenty query hashes and read their WHERE clauses.</li><li><strong>Materialized views for the panels every tenant loads.</strong> Real savings on the read path, but refresh is compute you pay for. A view refreshed every five minutes and read twice an hour is a net loss.</li><li><strong>Let the result cache work.</strong> Identical query text against unchanged data is free. Anything your BI layer does that injects a timestamp or a random parameter into otherwise identical SQL is throwing that away. Worth checking in Amazon QuickSight, Metabase or whatever sits in front.</li><li><strong>Tune zero-ETL refresh intervals.</strong> The refresh interval on the target database is adjustable via <code>ALTER DATABASE</code>. Shorter is fresher and more expensive. For reporting and historical analysis, a longer interval is usually the right call and nobody notices.</li><li><strong>Keep cold history out of managed storage.</strong> Partitioned Parquet or Apache Iceberg tables in S3, catalogued in AWS Glue, queried through the lake. On serverless those queries bill at the same RPU rate rather than as a separate Spectrum line, so the win is in scan efficiency and storage cost, not in dodging a charge.</li></ul>



<p class="wp-block-paragraph">One reassuring detail: the automatic optimization work Redshift does in the background is not billed by default. It becomes billable only if you explicitly enable extra compute resources so those operations can run during busy periods. That is a deliberate trade, not an accident, and it is worth knowing before you turn it on.</p>



<h2 class="wp-block-heading">Provisioned or serverless: how I would decide</h2>



<p class="wp-block-paragraph">Both have a genuine case and the honest answer depends on the shape of your load, not on which is newer.</p>



<p class="wp-block-paragraph">Serverless wins when demand is spiky or concentrated in business hours, when you cannot forecast capacity, and for dev and test environments that sit idle most of the week. It also folds concurrency scaling and data-lake queries into a single rate, which removes two line items people routinely forget to model.</p>



<p class="wp-block-paragraph">Provisioned RA3 wins when load is steady around the clock, because a reserved commitment on nodes can beat accumulated on-demand RPU-hours, and because you get the full workload management surface: queues, query priority, query monitoring rules with the complete set of controls. If you need to guarantee that a tenant&#8217;s export can never starve the interactive path, that machinery is more expressive than a price-performance slider.</p>



<p class="wp-block-paragraph">Commitment discounts now exist on both sides, including reservations for serverless managed at the payer account level. Rates and terms change, so price it against your own numbers rather than a blog post.</p>



<p class="wp-block-paragraph">The tell is simple. Pull a week of <code>charged_seconds</code> bucketed by hour and plot it. A flat line means you are paying serverless rates for provisioned behaviour. A sawtooth with long dead zones means the opposite.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Troubleshooting: the bill moved and nothing shipped</h2>



<ol class="wp-block-list"><li><strong>Get hourly billed seconds first.</strong> Aggregate <code>charged_seconds</code> by hour from the usage view. If the increase is spread evenly across all 24 hours, it is background activity: a pool, a monitor, a health check. If it is concentrated, it is a workload.</li><li><strong>Check for anything running or queued right now.</strong> A single stuck statement explains a lot of otherwise inexplicable spend.</li><li><strong>Look for transactions that never ended.</strong> A deploy that changed error handling can leave transactions open without a single failed request in your logs.</li><li><strong>Compare query counts against query cost.</strong> Flat count with rising cost points at base capacity changes, scale-down lag, or data growth making the same queries slower.</li><li><strong>Group by <code>user_query_hash</code> and diff against last week.</strong> New shapes appearing means a shipped change. Old shapes getting slower means data or statistics.</li><li><strong>Only then look at storage.</strong> Managed storage grows quietly and it is rarely the cause of a sudden jump, but it is often the cause of a slow one.</li></ol>



<pre class="wp-block-code"><code>-- Anything currently running or waiting.
SELECT user_id, query_id, transaction_id, session_id, status,
       trim(database_name) AS database_name,
       start_time, queue_time, execution_time
FROM   sys_query_history
WHERE  status IN ('running','queued')
ORDER  BY start_time;</code></pre>



<h2 class="wp-block-heading">Common mistakes</h2>



<ul class="wp-block-list"><li>Treating max capacity as a spending cap. It caps compute, not cost, and it will not stop a workload that simply runs for a long time.</li><li>Optimising individual slow queries while ignoring a connection pool that wakes the warehouse every thirty seconds all night.</li><li>Smoothing scheduled jobs out across the hour to be gentle on the warehouse. On serverless this is backwards; batching into fewer windows costs less.</li><li>Building cost dashboards on <code>compute_seconds</code> instead of <code>charged_seconds</code>, then wondering why the totals never reconcile with the invoice.</li><li>Giving every tenant their own workgroup for isolation, then discovering that base capacity floors and minimum charges multiply by tenant count.</li><li>Setting a usage limit to &#8220;turn off user queries&#8221; on the first day, before anyone knows what a normal week looks like.</li><li>Leaving the monitoring loop itself unbounded. Polling system views every few seconds is a workload that bills like any other.</li></ul>



<h2 class="wp-block-heading">Best practices worth the effort</h2>



<ul class="wp-block-list"><li>Label every tenant-originated query with <code>query_group</code> from day one. Retrofitting attribution is far more painful than adding a SET statement to your pool&#8217;s init SQL.</li><li>UNLOAD the serverless usage view to S3 on a schedule. Seven days of retention is not enough to argue about a monthly invoice.</li><li>Keep at least one usage limit configured as an alert, permanently, even if you never set a hard cap.</li><li>Review base capacity quarterly against p95 latency, not just against cost. The right number moves as your workload changes.</li><li>Put a hard date bound on every customer-facing query in the application layer. Unbounded ranges are the single most common source of surprise spend in embedded analytics.</li><li>Model concurrency scaling and data-lake charges explicitly if you are on provisioned. They are the line items people forget until they appear.</li><li>Tag workgroups and clusters consistently so cost tooling, whether that is AWS Cost Explorer or something like CloudZero or Vantage, can split spend by environment without guesswork.</li></ul>



<h2 class="wp-block-heading">FAQ</h2>



<h3 class="wp-block-heading">Does Redshift Serverless really charge me when nobody is using the product?</h3>



<p class="wp-block-paragraph">Idle time itself is not billed, but anything that sends a query is. AWS states explicitly that health-check queries from connection pools count as billable user activity. If your pool validates connections on a timer overnight, you are paying minimum billing windows all night. Check the pool before you conclude the pricing model is broken.</p>



<h3 class="wp-block-heading">How do I calculate the cost of a single query?</h3>



<p class="wp-block-paragraph">You cannot, exactly. Serverless bills the warehouse, and the 60-second minimum is shared with whatever else ran in that window. The workable approach is apportionment: take <code>charged_seconds</code> for a period and divide it by each labelled tenant&#8217;s share of execution time from the query history view. Useful for finding outliers, not precise enough for per-customer invoicing.</p>



<h3 class="wp-block-heading">Should I lower base capacity to save money?</h3>



<p class="wp-block-paragraph">Often yes, and it is the single highest-leverage change for a workload made of many short queries, because base capacity multiplies every billed second including the minimum. The catch is that it also reduces the compute each query gets. Change it in one step, watch p95 latency and queue time together for a full week, then decide whether to go further.</p>



<h3 class="wp-block-heading">Is a warehouse per tenant a good idea?</h3>



<p class="wp-block-paragraph">Only when tenants are large enough to keep a warehouse genuinely busy, or when a contract requires that level of separation. For a long tail of small tenants it is the most expensive option available, since each warehouse carries its own capacity floor and its own minimum charges. Pooled tables with row-level security and a tenant-leading sort key gets you most of the isolation for a fraction of the compute.</p>



<h3 class="wp-block-heading">Does concurrency scaling cost extra?</h3>



<p class="wp-block-paragraph">On Redshift Serverless, no, scaling is included in the RPU rate. On provisioned clusters it is a separate charge, offset by credits that accrue while the main cluster runs. That difference catches out teams migrating between the two, in both directions.</p>



<h3 class="wp-block-heading">Will a usage limit take my product down?</h3>



<p class="wp-block-paragraph">It will if you configure the breach action to turn off user queries. The logging and alerting actions are safe and are what you want in a customer-facing system. Treat the hard stop as a deliberate business decision about which is worse, an unexpected invoice or an outage, rather than as a default setting.</p>



<h3 class="wp-block-heading">Why does my cost report never match the AWS invoice?</h3>



<p class="wp-block-paragraph">Usually one of three things: using <code>compute_seconds</code> rather than <code>charged_seconds</code>, forgetting that usage is recorded only when a transaction completes so long transactions land in a later interval, or leaving storage and cross-region transfer out of the model entirely.</p>



<h2 class="wp-block-heading">The one thing worth remembering</h2>



<p class="wp-block-paragraph">Redshift cost optimization for a SaaS analytics product is mostly not a query tuning exercise. It is a question of how often something wakes the warehouse up and at what capacity. Query tuning matters, sort keys matter, materialized views matter, but a connection pool with default settings will quietly outspend all of them combined.</p>



<p class="wp-block-paragraph">So start at the meter. Pull hourly billed seconds, look at the overnight hours when your product has no users, and see whether the line goes to zero. If it does not, you have found your first and cheapest win before touching a single line of SQL.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Need help getting your Redshift bill under control?</h2>



<p class="wp-block-paragraph">I work with SaaS and data teams on exactly this problem, usually somewhere between the warehouse and the application that is hammering it. Things I can help with:</p>



<ul class="wp-block-list"><li>Auditing an existing Redshift Serverless or RA3 workload and producing a ranked list of what is actually driving spend, with the numbers pulled from your own system views.</li><li>Building per-tenant cost attribution: query labelling, a usage archive in S3, and a dashboard your product and finance teams can both read.</li><li>Fixing the connection and session layer, including pool configuration, validation queries, session timeouts and transaction hygiene.</li><li>Right-sizing base and max capacity against measured latency, and setting usage limits and alerts that warn without risking an outage.</li><li>Reviewing multi-tenant data models: sort and distribution keys, row-level security, and whether data sharing or a pooled model fits your tenant mix.</li><li>Deciding between provisioned and serverless with a workload profile behind the recommendation rather than a rule of thumb.</li></ul>



<p class="wp-block-paragraph">If you have a week of usage data, a suspicious hourly cost chart, or a pool configuration you are not sure about, send it over and I will tell you what I see in it.</p>



<div class="wp-block-buttons is-layout-flex wp-block-buttons-is-layout-flex">
<div class="wp-block-button"><a class="wp-block-button__link wp-element-button" href="https://www.upwork.com/freelancers/~01f15a912ad84a6620" target="_blank" rel="noreferrer noopener">Work with me on Upwork</a></div>
</div>
<p>The post <a href="https://john-nessime.com/blog/devops/redshift-cost-optimization-saas-analytics/">Redshift Cost Optimization for SaaS Analytics: The Levers That Actually Move the Bill</a> appeared first on <a href="https://john-nessime.com/blog">John Nessime</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Embedded Analytics on AWS: The Four Decisions That Bite Later</title>
		<link>https://john-nessime.com/blog/devops/embedded-analytics-on-aws/</link>
		
		<dc:creator><![CDATA[John Nessime]]></dc:creator>
		<pubDate>Tue, 04 Aug 2026 06:07:00 +0000</pubDate>
				<category><![CDATA[Cloud Computing]]></category>
		<category><![CDATA[DevOps]]></category>
		<category><![CDATA[Technical Guides]]></category>
		<category><![CDATA[Amazon Athena]]></category>
		<category><![CDATA[Amazon Redshift]]></category>
		<category><![CDATA[Amazon S3]]></category>
		<category><![CDATA[Architecture]]></category>
		<category><![CDATA[AWS]]></category>
		<category><![CDATA[Business Intelligence]]></category>
		<category><![CDATA[Cost Optimization]]></category>
		<category><![CDATA[Data Lake]]></category>
		<category><![CDATA[Data Warehouse]]></category>
		<category><![CDATA[Embedded Analytics]]></category>
		<category><![CDATA[IAM]]></category>
		<category><![CDATA[Multi-Tenant]]></category>
		<category><![CDATA[QuickSight]]></category>
		<category><![CDATA[Row-Level Security]]></category>
		<category><![CDATA[SPICE]]></category>
		<category><![CDATA[SQL]]></category>
		<guid isPermaLink="false">https://john-nessime.com/blog/?p=120</guid>

					<description><![CDATA[<p>Rendering a dashboard inside your app is the easy part. Tenant isolation, session cost and query mode are what break. A practical walkthrough of the four decisions behind embedded analytics on AWS, the API constraints that lock you in, and the errors you will actually see.</p>
<p>The post <a href="https://john-nessime.com/blog/devops/embedded-analytics-on-aws/">Embedded Analytics on AWS: The Four Decisions That Bite Later</a> appeared first on <a href="https://john-nessime.com/blog">John Nessime</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">Getting a dashboard to render inside your own application is the easy part. You publish it, call the embed API, drop the iframe in, and it shows up. The hard question arrives about a day later, usually from someone in security or from the first customer who logs in: how exactly does tenant B not see tenant A&#8217;s rows?</p>



<p class="wp-block-paragraph">That is where embedded analytics on AWS stops being a front-end task and becomes an architecture decision. The awkward part is that the choice you make first, how the viewer is identified, quietly decides which isolation mechanisms remain available to you afterwards. Get that order backwards and you rebuild the data layer, not the iframe.</p>



<p class="wp-block-paragraph">This post covers the four decisions that determine whether the build holds: identity model, tenant isolation, query mode, and session economics. Then the embed handshake itself, the errors you will actually see in the browser console, and what to check first when it fails.</p>



<h2 class="wp-block-heading">Before anything else: the product got renamed</h2>



<p class="wp-block-paragraph">Amazon QuickSight was folded into a broader platform called Amazon Quick Suite, and the BI product inside it is now called Amazon Quick Sight. AWS documentation has since moved again under an &#8220;Amazon Quick&#8221; umbrella. You will land on all three naming conventions depending on which search result you click, which makes finding the right doc page genuinely annoying.</p>



<p class="wp-block-paragraph">The practical upshot: the APIs, SDKs and IAM action names did not change. You are still calling <code>quicksight:GenerateEmbedUrlForRegisteredUser</code> against ARNs in the <code>quicksight</code> namespace, and the JavaScript SDK is still published as <code>amazon-quicksight-embedding-sdk</code>. Nothing in your code breaks. Only your bookmarks do. I mention it because half the confusion in a first embedded build comes from following a doc page that describes a UI menu that has since been reorganised.</p>



<h2 class="wp-block-heading">Decision one: registered users or anonymous sessions</h2>



<p class="wp-block-paragraph">Two API operations generate embed URLs. <code>GenerateEmbedUrlForRegisteredUser</code> issues a session for a user who exists inside the BI account. <code>GenerateEmbedUrlForAnonymousUser</code> issues a session for someone who does not, and never will.</p>



<p class="wp-block-paragraph">This reads like a convenience choice. It is not. Row-level security using session tags, the mechanism most SaaS products want, is supported <em>only</em> for anonymous embedding. It does not work with <code>GenerateEmbedUrlForRegisteredUser</code>, it does not work with the older <code>GetDashboardEmbedUrl</code> operation, and it is not supported with the IAM identity type. That constraint is documented, easy to miss, and it is the single most expensive thing to discover late.</p>



<p class="wp-block-paragraph">So the fork is really this. If you register every viewer, you get per-user features (bookmarks, threshold alerts, scheduled snapshots) and you enforce isolation with username or group rules on the dataset. You also inherit the job of provisioning, deprovisioning and reconciling a user directory that mirrors your own. If you go anonymous, you skip all of that and filter with session tags at embed time, but per-user features are off the table because there is no persistent user to hang them on.</p>



<p class="wp-block-paragraph">The registered-user request body is small. Everything interesting is in <code>ExperienceConfiguration</code>:</p>



<pre class="wp-block-code"><code>POST /accounts/&lt;aws-account-id&gt;/embed-url/registered-user

{
  "UserArn": "arn:aws:quicksight:&lt;region&gt;:&lt;account&gt;:user/default/&lt;user&gt;",
  "SessionLifetimeInMinutes": 60,
  "AllowedDomains": ["https://app.example.com"],
  "ExperienceConfiguration": {
    "Dashboard": {
      "InitialDashboardId": "&lt;dashboard-id&gt;",
      "FeatureConfigurations": {
        "Bookmarks": { "Enabled": true }
      }
    }
  }
}</code></pre>



<p class="wp-block-paragraph">One trap on the anonymous path that deserves its own sentence. Anonymous sessions belong to a namespace, and any dashboard shared with that namespace is reachable by a session in it, whether or not you listed the dashboard in <code>AuthorizedResourceArns</code>. If you were treating that parameter as your allowlist, it is not. Namespace membership is the real boundary.</p>



<h2 class="wp-block-heading">Decision two: where tenant isolation actually lives</h2>



<p class="wp-block-paragraph">There are three places you can put the filter, and only one of them scales.</p>



<ul class="wp-block-list">
<li><strong>A dashboard per tenant.</strong> Works for five customers. Becomes a deployment problem at fifty and a change-management disaster at five hundred, because every visual fix is now a fan-out.</li>



<li><strong>A dataset per tenant, filtered in SQL.</strong> Better isolation guarantees, genuinely defensible in a compliance review, but you multiply refresh jobs and in-memory footprint by tenant count.</li>



<li><strong>One dashboard, one dataset, row-level security.</strong> The standard answer. One artifact to maintain, filtering applied per session.</li>
</ul>



<p class="wp-block-paragraph">With anonymous embedding, RLS is driven by tags. You declare tag keys against columns on the dataset, then supply values at embed time. The filter is evaluated server-side against the session, so a viewer poking at the iframe cannot lift it.</p>



<pre class="wp-block-code"><code>POST /accounts/&lt;aws-account-id&gt;/embed-url/anonymous-user

{
  "Namespace": "default",
  "SessionLifetimeInMinutes": 60,
  "AuthorizedResourceArns": [
    "arn:aws:quicksight:&lt;region&gt;:&lt;account&gt;:dashboard/&lt;dashboard-id&gt;"
  ],
  "SessionTags": [
    { "Key": "tenant_id", "Value": "acme-corp" },
    { "Key": "region",    "Value": "emea" }
  ],
  "AllowedDomains": ["https://app.example.com"],
  "ExperienceConfiguration": {
    "Dashboard": { "InitialDashboardId": "&lt;dashboard-id&gt;" }
  }
}</code></pre>



<p class="wp-block-paragraph">The value in <code>SessionTags</code> must come from your server-side session, never from a request parameter, a cookie your client can write, or a JWT claim you have not verified. This is the whole security boundary. Tag rules support combining conditions, so a manager who should see several sites is expressible without a second dashboard.</p>



<p class="wp-block-paragraph">One quiet limit worth knowing before it bites: when RLS is applied to in-memory datasets, each field has a maximum length in Unicode characters, and fields exceeding it are truncated during ingestion rather than rejected. If your tenant identifiers are long opaque strings, test that a truncated value cannot collide with another tenant&#8217;s. Silent truncation plus a prefix collision is exactly the kind of bug that produces a cross-tenant data leak with no error anywhere in the logs.</p>



<h2 class="wp-block-heading">Decision three: SPICE or direct query against your AWS data</h2>



<p class="wp-block-paragraph">Every dataset runs in one of two modes, and the difference shows up on a bill somewhere else in your account.</p>



<p class="wp-block-paragraph"><strong>Direct query</strong> sends a live query to the source each time a visual renders. Against Amazon Athena that means an S3 scan per dashboard open, billed by bytes scanned. Against Amazon Redshift it means a concurrent query slot per viewer. Freshness is perfect. The failure mode is that dashboard load is now coupled to warehouse load, and your analytics traffic competes with everything else running there. Two hundred people opening a dashboard at 9am is two hundred queries, and Redshift concurrency is finite.</p>



<p class="wp-block-paragraph"><strong>SPICE</strong> imports a snapshot into an in-memory engine and serves every viewer from it. One scan on refresh, then arbitrarily many reads. For an embedded product where the same aggregate is served to thousands of sessions, this is usually the right call, and the Athena cost difference between &#8220;scan once per refresh&#8221; and &#8220;scan once per pageview&#8221; is not subtle. What you give up is freshness, bounded by your refresh schedule, plus a capacity dimension to manage and incremental refresh to configure if the dataset is large.</p>



<p class="wp-block-paragraph">The pattern I reach for first on a data-lake backend is a hybrid: recent partitions in SPICE with an incremental refresh on a look-back window, historical data left on direct query for the rare deep query. It costs more design effort up front and it is the thing most teams skip, but it is the only shape that keeps both the bill and the load time flat as history grows.</p>



<p class="wp-block-paragraph">Whichever you pick, note that visual generation has a timeout, and data-source-specific timeouts apply on top of it. A query that is merely slow in a console tab renders as a broken visual in a customer&#8217;s browser. Model your worst partition, not your average one.</p>



<h2 class="wp-block-heading">Decision four: what a session actually costs</h2>



<p class="wp-block-paragraph">I am not going to quote figures, because AWS changes them and you should read the current pricing page. The mechanism is what matters, and it is genuinely different from seat-based BI licensing.</p>



<ul class="wp-block-list">
<li>A reader session is a fixed 30-minute window. Not a pageview, not a query. Reopening the dashboard twenty minutes later is still the same session.</li>



<li><strong>Per-user pricing</strong> charges per session with a monthly cap per reader. Predictable when the same people return daily.</li>



<li><strong>Capacity pricing</strong> buys sessions in bulk with no user provisioning at all. This is the model built for embedding, and it is the one that pairs with anonymous sessions.</li>



<li>Capacity pricing is also the prerequisite for programmatic dashboard refresh, so if near-real-time rendering is a product requirement, that decision is already made for you.</li>



<li>Annual commitments to capacity unlock removing the &#8220;Powered by&#8221; attribution footer. If white-labelling is a contractual requirement, factor that in early rather than discovering it during a customer demo.</li>



<li>Enabling certain Pro-tier and generative Q&amp;A capabilities triggers an account-level monthly infrastructure fee that exists whether or not anyone uses the feature.</li>
</ul>



<p class="wp-block-paragraph">The cost failure mode nobody plans for is architectural rather than commercial. If you embed the dashboard on a tab that loads by default, you bill a session for every user who lands on that page and looks at something else. Lazy-load the iframe on explicit interaction. That one change is often the largest single lever on the bill, and it costs an afternoon.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">The embed handshake, and the three things that break it</h2>



<p class="wp-block-paragraph">The flow is short. Your backend authenticates the user with your own identity system, calls the embed URL API with the right tags or user ARN, returns the URL to the browser, and the SDK mounts an iframe against it.</p>



<ol class="wp-block-list">
<li>The generated URL carries a temporary bearer token valid for five minutes, and it is single use once redeemed. Generate it per page load from your backend. Never cache it, never put it in a build artifact, never log it.</li>



<li>Session lifetime is separate from URL validity, set with <code>SessionLifetimeInMinutes</code>, and ranges from fifteen minutes to ten hours with ten hours as the default. Ten hours is almost never what you want for a customer-facing product. Match it to your own session, or shorter.</li>



<li>Domains must be allowed explicitly. An administrator configures static domains in the admin menu, and <code>AllowedDomains</code> on the API call can override that with up to three domains or subdomains per request. Add an <code>AllowedEmbeddingDomains</code> condition to the IAM policy of the calling role, or any developer with that permission can list any domain on the internet.</li>
</ol>



<p class="wp-block-paragraph">On the browser side, the v2 SDK creates an embedding context (which appends its own zero-pixel iframe to <code>body</code> for message passing) and then mounts the experience:</p>



<pre class="wp-block-code"><code>import { createEmbeddingContext } from 'amazon-quicksight-embedding-sdk';

const context = await createEmbeddingContext();

await context.embedDashboard(
  {
    url: embedUrl,                        // fetched from your backend, just now
    container: '#analytics',
    height: '600px',                      // acts as loading height below
    resizeHeightOnSizeChangedEvent: true,
  },
  {
    toolbarOptions: { export: false, undoRedo: false, reset: false },
    attributionOptions: { overlayContent: true },
    onMessage: async (event) =&gt; {
      if (event.eventName === 'ERROR_OCCURRED') {
        console.error(event.message.errorCode);
      }
    },
  }
);</code></pre>



<p class="wp-block-paragraph">Two details in there earn their place. <code>resizeHeightOnSizeChangedEvent</code> turns the <code>height</code> value into a loading placeholder and lets the frame grow to fit content, which is what stops the dashboard rendering into a 600px letterbox with an inner scrollbar. And <code>overlayContent</code> tells the layout to overlay the attribution footer rather than reserve extra height at the bottom for it.</p>



<h2 class="wp-block-heading">Troubleshooting embedded analytics on AWS</h2>



<p class="wp-block-paragraph">Almost every failure lands in one of these. Read the error code out of the <code>ERROR_OCCURRED</code> message before doing anything else.</p>



<ul class="wp-block-list">
<li><strong><code>Forbidden</code></strong> means the URL&#8217;s authentication code expired. You held the URL longer than five minutes, or you served it from a cache, or a retry redeemed it twice. Fix the generation path, not the permissions.</li>



<li><strong><code>Unauthorized</code></strong> means the session obtained from that code expired. Different problem, different fix: your <code>SessionLifetimeInMinutes</code> is shorter than how long people keep the tab open. Handle it by re-fetching a fresh URL and re-mounting rather than letting the frame sit there dead.</li>



<li><strong>Frame never appears at all.</strong> Check the <code>onChange</code> handler for <code>NO_CONTAINER</code> or <code>INVALID_CONTAINER</code>, which usually means you mounted before your target element existed, and for <code>INVALID_URL</code>, which means the URL shape does not match the experience method you called.</li>



<li><strong>Frame appears, dashboard does not.</strong> Nine times out of ten this is the domain allowlist. The request domain has to match what was allowed, including scheme and any subdomain, and a staging hostname that nobody added is the usual culprit.</li>



<li><strong>Modals render off-screen.</strong> A known consequence of auto-resizing height: an export dialog can open above the visible viewport. Listen for <code>MODAL_OPENED</code> and scroll the parent page to the frame position.</li>



<li><strong>Toolbar features silently missing.</strong> Bookmarks, threshold alerts and scheduling require both the SDK toolbar flag and the matching entry under <code>FeatureConfigurations</code> in the embed URL request, and they only exist on the registered-user path. Setting the client flag alone does nothing.</li>



<li><strong>First render is slow, later ones are fine.</strong> Direct query against a cold warehouse. Compare the same query in Athena or Redshift directly to confirm before blaming the BI layer.</li>
</ul>



<h2 class="wp-block-heading">Common mistakes</h2>



<ul class="wp-block-list">
<li>Choosing registered-user embedding for the identity story, then discovering session-tag RLS is unavailable on that path.</li>



<li>Treating <code>AuthorizedResourceArns</code> as the security boundary instead of namespace membership.</li>



<li>Deriving a session tag value from anything the client can influence.</li>



<li>Generating the embed URL at build time, or caching it in a CDN, and then not understanding the <code>Forbidden</code> errors.</li>



<li>Leaving session lifetime at the ten-hour default in a customer-facing app.</li>



<li>Putting the dashboard on a default-loaded tab and paying for sessions nobody asked for.</li>



<li>Building the first version on direct query against Athena because it is quicker to wire up, then meeting the scan bill.</li>



<li>Forgetting that embedding and row-level security sit in the Enterprise tier, so a Standard-tier proof of concept proves nothing.</li>
</ul>



<h2 class="wp-block-heading">Best practices</h2>



<ul class="wp-block-list">
<li>Decide the identity model before you build a single dataset. Everything downstream inherits it.</li>



<li>Put the embed URL call behind one server-side endpoint that reads tenant scope from your own session and nowhere else. One function, one place to audit.</li>



<li>Constrain the calling IAM role with an <code>AllowedEmbeddingDomains</code> condition and scope resources to specific namespaces rather than a wildcard.</li>



<li>Write an automated test that requests tenant A&#8217;s embed URL and asserts tenant B&#8217;s rows are absent. Run it on every dataset change, because RLS breaks silently.</li>



<li>Default to SPICE with a refresh schedule matched to a stated freshness SLA, and only reach for direct query where the SLA genuinely demands it.</li>



<li>Track refresh failures as a first-class alert in CloudWatch or whatever you already run, whether that is Grafana, Datadog or something in-house. A stale dashboard that still renders is worse than one that errors, because nobody notices.</li>



<li>Lazy-load the iframe on user intent, not on page mount.</li>



<li>Keep the embedded surface read-only unless authoring is a real product requirement. Console embedding is a much larger permissions surface than dashboard embedding.</li>
</ul>



<h2 class="wp-block-heading">Is managed BI even the right call?</h2>



<p class="wp-block-paragraph">Worth asking honestly, because the answer is not always yes. The case for the AWS-native route is real: no connector layer to maintain against Athena, Redshift, S3 and Aurora, IAM you already understand, and a usage-based cost model that beats per-seat licensing when your viewers are bursty. If most of your data already sits in AWS, that adds up.</p>



<p class="wp-block-paragraph">The case against is equally real. Visual customisation is limited compared to charting directly against your own API, the attribution footer needs a commitment to remove, and if you want full control of the front end you may be better served by Apache Superset or Metabase self-hosted, or by Grafana where the workload is closer to operational metrics than customer-facing BI. Those come with an operational burden you now own. That is the trade: you either run the BI layer or you rent it, and renting it means living inside its constraints.</p>



<h2 class="wp-block-heading">Frequently asked questions</h2>



<h3 class="wp-block-heading">Do my users need AWS accounts to view an embedded dashboard?</h3>



<p class="wp-block-paragraph">No. With anonymous embedding they need no AWS account and no BI user record at all. Your application authenticates them however you already do, and your backend maps that identity to session tags when it requests the embed URL.</p>



<h3 class="wp-block-heading">Can I use row-level security with registered-user embedding?</h3>



<p class="wp-block-paragraph">Yes, but only with username or group based rules, not with session tags. Tag-based RLS is restricted to the anonymous embedding path. If you need tags, you need anonymous sessions.</p>



<h3 class="wp-block-heading">How long does an embed URL stay valid?</h3>



<p class="wp-block-paragraph">The URL itself carries a bearer token valid for five minutes and usable once. The session it opens is separate and lasts between fifteen minutes and ten hours depending on <code>SessionLifetimeInMinutes</code>, defaulting to ten hours.</p>



<h3 class="wp-block-heading">Should I use SPICE or direct query for embedded analytics on AWS?</h3>



<p class="wp-block-paragraph">SPICE for anything with many viewers per refresh, which describes most embedded products. Direct query where the data must be current to the second, or where the dataset exceeds what you want to hold in memory. A hybrid split by data age is often the right answer and is under-used.</p>



<h3 class="wp-block-heading">Why do I get a Forbidden error when the dashboard worked yesterday?</h3>



<p class="wp-block-paragraph"><code>Forbidden</code> points at the URL, not at permissions. The most common causes are caching the URL, generating it more than five minutes before use, or a client retry redeeming the same single-use token twice. If it is <code>Unauthorized</code> instead, the session expired and you need a fresh URL.</p>



<h3 class="wp-block-heading">Can I white-label the embedded dashboard completely?</h3>



<p class="wp-block-paragraph">Largely. Themes control colours and typography, the SDK hides toolbar controls, and parameters let your own UI drive the dashboard. Removing the attribution footer entirely is tied to an annual capacity commitment, so confirm that against current terms before you promise it to a customer.</p>



<h3 class="wp-block-heading">Does natural-language querying work in an embedded context?</h3>



<p class="wp-block-paragraph">Yes. The SDK exposes a generative Q&amp;A experience alongside dashboards and visuals, driven by curated topics rather than raw tables. It is billed on its own capacity dimension and gates behind the Pro tiers, so treat it as a separate cost decision rather than a free addition.</p>



<h2 class="wp-block-heading">The one thing worth remembering</h2>



<p class="wp-block-paragraph">Embedded analytics on AWS is not a rendering problem. The iframe is the last five percent. The part that decides whether the build survives contact with a second customer is the identity model, because it silently determines which isolation mechanism you are allowed to use, and that in turn shapes your dataset design, your refresh strategy and your bill.</p>



<p class="wp-block-paragraph">Pick that first. Write the cross-tenant test before you write the dashboard. Everything else is recoverable in an afternoon.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Need help with an embedded analytics build?</h2>



<p class="wp-block-paragraph">This is the kind of work I do. Things I can help with directly:</p>



<ul class="wp-block-list">
<li>Reviewing an existing embed integration for cross-tenant leakage, including the session-tag path and the namespace boundary.</li>



<li>Designing the identity and row-level security model before you commit to a dataset layout.</li>



<li>Cutting Athena scan and Redshift concurrency cost by moving the right datasets into SPICE with incremental refresh.</li>



<li>Building the backend embed-URL service with scoped IAM roles, domain conditions and sane session lifetimes.</li>



<li>Setting up refresh failure alerting so a stale dashboard does not quietly serve last week&#8217;s numbers.</li>



<li>Automated cross-tenant isolation tests wired into CI, so an RLS regression fails the build instead of the customer.</li>
</ul>



<p class="wp-block-paragraph">Send me the actual thing: your embed URL request payload with secrets stripped, the browser console error, or the dataset RLS rules. It is much faster to reason about a real payload than a description of one.</p>



<div class="wp-block-buttons is-layout-flex wp-block-buttons-is-layout-flex">
<div class="wp-block-button"><a class="wp-block-button__link wp-element-button" href="https://www.upwork.com/freelancers/~01f15a912ad84a6620" target="_blank" rel="noreferrer noopener">Work with me on Upwork</a></div>
</div>
<p>The post <a href="https://john-nessime.com/blog/devops/embedded-analytics-on-aws/">Embedded Analytics on AWS: The Four Decisions That Bite Later</a> appeared first on <a href="https://john-nessime.com/blog">John Nessime</a>.</p>
]]></content:encoded>
					
		
		
			</item>
	</channel>
</rss>
