Your API Has Roles. That Does Not Mean Access Control Works

A practical test plan for object-level authorization, tenant isolation, and API access control bugs that survive happy-path role checks.

11 min read
ibrahimsql
2,060 words

Your API Has Roles. That Does Not Mean Access Control Works#

Most API access control bugs are not caused by missing login. They happen after login, when the backend forgets to ask a more specific question: should this user perform this action on this object right now?

Roles help, but they are not the whole model. A user can be admin in one workspace, a viewer in another, removed from a project yesterday, and still hold a token issued last week. That is where neat RBAC diagrams start leaking data.

Test the object, not just the role#

The weak version of authorization checks only the role:

if (session.user.role !== 'admin') { throw new Error('Forbidden') }

The stronger version checks role, object ownership, tenant membership, and action:

const report = await db.report.findFirst({ where: { id: reportId, tenantId: session.tenantId, project: { memberships: { some: { userId: session.user.id, role: { in: ['owner', 'analyst'] }, }, }, }, }, }) if (!report) { throw new Error('Not found') }

This is not just cleaner code. It changes the failure mode. Unauthorized objects and nonexistent objects now look the same from the outside, which reduces enumeration and makes the contract easier to test.

Build a matrix before touching Burp#

Random request tampering finds bugs, but a matrix finds classes of bugs. Start with the relationships your product actually has.

CaseUser AUser BExpected result
Same tenant, lower roleOwnerViewerViewer cannot write
Different tenantOwner in tenant 1Owner in tenant 2No cross-tenant read
Removed memberFormer memberActive memberOld token loses access
Shared objectCreatorInvited viewerViewer gets only allowed fields
Soft-deleted objectOwnerOwnerDetail, list, export agree
Internal endpointSupport userCustomer userCustomer cannot call support action

Run this matrix against list, detail, update, delete, export, and search endpoints. The boring endpoints matter. Exports and search are where polished products often leak old assumptions.

Watch for authorization hidden inside the UI#

Frontend checks are useful for user experience. They are useless as an API boundary.

If the frontend hides the "delete project" button for viewers, still send the request as a viewer. If the UI only shows the first 20 invoices, still request invoice 21 directly. If the admin panel uses a separate route, check whether the API route is actually separate or just hidden behind a different navigation item.

The backend should survive a client that ignores every UI decision.

Tenant filters belong in the query#

Pulling an object and checking ownership afterwards is easy to get wrong:

const file = await db.file.findUnique({ where: { id: fileId } }) if (file.tenantId !== session.tenantId) { throw new Error('Forbidden') }

This can still leak timing differences, object existence, or fields returned by logging and error handling. Prefer making authorization part of the data access:

const file = await db.file.findFirst({ where: { id: fileId, tenantId: session.tenantId, }, })

When the query returns nothing, the rest of the app has less sensitive information to accidentally expose.

What to log when a check fails#

Do not log raw secrets or full request bodies. Do log enough to debug the authorization decision.

Good fields:

  • userId
  • tenantId
  • resourceType
  • resourceId
  • action
  • policyName
  • decision
  • requestId

This makes access control testable in production without turning logs into a second database of sensitive data.

Closing checklist#

Before calling an API authorization fix done, answer these:

  • Does the query include tenant or ownership constraints?
  • Does the policy distinguish read, write, export, and delete?
  • Do removed users lose access before token expiry?
  • Do list, detail, search, and export use the same policy?
  • Do unauthorized and nonexistent resources produce the same external behavior?
  • Is there a regression test for the relationship that failed?

Access control is not a middleware checkbox. It is a product model encoded in code. If the model is fuzzy, the API will be fuzzy too.

Where the Trust Boundary Actually Sits#

client --> edge middleware --> route handler --> policy check --> query --> DB | | | rate limit session parse object scope

Each hop re-derives the decision. Middleware can drop obviously anonymous traffic, but the object-level check belongs next to the query, not in a helper that runs after the row is already loaded.

// Wrong layer: row is fetched, then checked const doc = await db.doc.findUnique({ where: { id } }) if (doc.tenantId !== session.tenantId) throw new Error('Forbidden') // Right layer: the check is the query const doc = await db.doc.findFirst({ where: { id, tenantId: session.tenantId }, }) if (!doc) throw new Error('Not found')

In the first shape the row crosses every logging, tracing, and error-handling path before the check runs. In the second it never leaves the database unless it is allowed to.

Commands That Catch the Boring Cases#

# Replay a second user's session against every known endpoint while read ep; do curl -s -o /dev/null -w "%{http_code} $ep\n" \ -H "Cookie: session=$USER_B" "$BASE$ep" done < endpoints.txt

Expected: 403 or 404 for anything owned by user A. A 200 carrying another tenant's payload is a finding, not a warning.

# Compare list, detail, and export for the same probe object curl -s "$BASE/reports?tenantId=$TENANT" | jq '.items[].id' curl -s "$BASE/reports/$RID" | jq '.id' curl -s "$BASE/reports/export?tenantId=$TENANT" | head -5

If an object shows up in the export but not in the list, the export path skipped the policy. That asymmetry is the most common soft bug in mature APIs.

# Check that removed members actually lose access curl -s -o /dev/null -w "%{http_code}\n" -H "Cookie: session=$OLD_MEMBER" \ "$BASE/projects/$PID"

Do this right after membership revocation, not after token expiry. If the answer is 200, the token outlived the relationship.

Detection in Production#

SignalWhat it means
decision=deny spikes per accountEnumeration attempt
Viewer touching hundreds of distinct IDs per hourID walking, not reading
Same object allowed via export, denied via detailPolicy drift between endpoints
Old tokens passing after membership removalCache keyed without role version

Log userId, tenantId, resourceType, resourceId, action, policyName, and decision on every denial. Never log the token or the body. Denial logs are the cheapest tripwire you will ever have.

Limitations#

  • A matrix covers relationships you know about. New features need a policy review step, or they re-enter the matrix as untested rows.
  • Returning 404 for everything breaks legitimate deep links and confuses cache debugging. Pair it with request IDs so support can separate "no access" from "does not exist".
  • CDN and token caches can re-serve a response after revocation. Personalized responses need Cache-Control: private, no-store.
  • This plan finds application-layer bugs. It will not catch a database user with excessive grants, which is why the same review list belongs in infrastructure review too.

Cikarilar#

  • Role checks guard the door; object checks guard the rooms.
  • Put tenant and ownership filters inside the query, not after it.
  • 403 and 404 should look the same from the outside for authorization failures.
  • Every new endpoint joins the test matrix the day it ships.
  • Denial logs with policy names turn silence into a signal.

Wiring the Matrix Into CI#

The matrix is only useful if it runs. A minimal version is a table of (user, endpoint, payload, expected status) executed against a seeded staging database:

FixtureEndpointMethodBody field probedExpected
owner-a/invoices/inv_bGET-404
viewer-a/invoices/inv_aPATCHamount403
ex-member/projects/p1GET-404
support/admin/refundPOSTuserId403
stranger/search?q=invoiceGETresults200, own tenant only

Run it on every pull request that touches routes, policies, or the ORM schema. A new route without a new row in the table should fail review. The failure mode you are buying down is not a clever hacker; it is an intern adding a copied service method six months from now.

Fields That Belong in Every Denial Log#

  • requestId so a user report maps to one server-side decision
  • policyName so you can see which rule fired
  • resourceType and resourceId hashed if IDs are sensitive
  • actorTenantId vs resourceTenantId on cross-tenant denials

Keep the denial log small enough that it can be written synchronously in the request path. If logging itself needs a queue, the queue will drop entries exactly when the system is under attack.

A Note on Token Lifetime#

Short-lived access tokens plus a revocation check on the policy layer is the realistic middle ground. Stateless JWTs with long expiry make "removed user keeps access" the default instead of the edge case. Whatever you choose, the matrix row for removed members is where the truth shows up.

The Two Places Teams Forget#

GraphQL field resolvers#

A resolver that returns an object from cache without re-checking membership leaks data through batching: one query can fan out to fifty objects from fifty tenants. Every resolver that touches a stored object needs the same object-scope filter the REST path uses. Sharing the policy function is the only maintainable option.

// Each field that resolves an object re-applies the scope Report: { export: async (parent, _args, ctx) => { const allowed = await policy.can(ctx.user, 'export', parent.id) if (!allowed) throw new ForbiddenError() return renderExport(parent.id) }, }

Webhooks and background jobs#

An outbound webhook signed with your secret is trusted by the receiver. If the job that builds the payload read the object without tenant scoping, the receiver gets the wrong tenant's data and your signature makes it look official. The fix is the same discipline as the request path: the job's query carries the tenant filter from the originating event, not from a global default.

await db.invoice.update({ where: { id: event.invoiceId, tenantId: event.tenantId }, data: { status: 'paid' }, })

Performance Note#

Object-scoped queries add a join or an extra filter. That is cheaper than the alternative: post-load checks push rows into memory, middleware-only checks duplicate into every handler, and both force you to log from two places. A composite index on (tenantId, id) makes the scoped query roughly the same cost as the unscoped one.

Final Gate#

Before shipping, answer the original six questions again against the new code, then answer three more: which resolver leaks if this policy drifts, which export skips the check, and which stale token still works. If those answers live in someone's head instead of a test, the fix is a test away.

A Worked Example#

Say a support engineer needs read access to a customer's invoice but never write access. Model it explicitly, not as "admin can do anything":

const allowed = await policy.check({ actor: session.user, action: 'invoice:read', resource: { type: 'invoice', id: invoiceId, tenantId }, })
if (!allowed) { audit.deny({ userId: session.user.id, tenantId, resourceId: invoiceId, action: 'invoice:read' }) throw new NotFoundError() }

Now the same check wraps the PATCH handler for invoice:write, the export handler for invoice:export, and the support panel action for invoice:admin-view. One policy module, one log format, one test fixture per role. When a new role appears, you add a row to the policy and a row to the matrix, and the CI gate tells you whether the rest of the surface behaves.

That is what "access control works" means in practice: one decision function, scoped queries, identical external behavior for denied and missing, and a matrix that runs on every deploy. Everything else is decoration.

Run that gate on every merge, and cross-tenant reads stay impossible instead of merely unlikely.

The habit matters more than any single check.

---
Share this post:

What do you think?

React to show your appreciation

Related Posts