Every permission system starts as if (user.isAdmin).
It works for months. Then a customer asks for someone who can view billing but not change it. Then for a manager who can approve expenses, but only for their own team, and only up to a limit. Each request is reasonable. Each one adds another if.
A year later nobody can say with confidence who is able to do what, and the honest answer to “can this user delete that?” is “let me read the code.”
Identity and access management has a reputation for being complicated. Most of the complication comes from making decisions in the wrong order: choosing a model before knowing the questions, and adding flexibility before knowing where it’s needed.
This is how I’d approach designing it now. What to settle first, what to leave rigid on purpose, where flexibility pays for itself, and what to refuse to build until someone needs it. The examples use a made-up product: a multi-tenant invoicing tool where companies have teams, and teams raise and approve invoices.
01IAM is four jobs, not one
“IAM” gets used as if it were a single feature. It’s four separate jobs that happen to share a login page.
Identity
Who is this?
users, service accounts, tenants
Authentication
Can they prove it?
passwords, SSO, passkeys, API keys
Authorization
What may they do?
roles, permissions, policies
Audit
What did they do?
who granted what, who did what
They fail differently, and they deserve different amounts of your own engineering.
Authentication is the one to build the least of. Password storage, session handling, multi-factor, SSO: these are solved problems with sharp edges, and nothing about them is specific to your product. Use an identity provider or a well-maintained library, and speak a standard protocol such as OpenID Connect so you can change your mind later.
Authorization is the opposite. No vendor knows that your approvers shouldn’t approve their own invoices. The rules are your product’s rules, so the design work is yours even if you run someone else’s engine underneath.
That’s where the rest of this article spends its time. Broken access control sits at the top of the OWASP Top 10 for a reason: it’s the part everyone writes themselves.
02Start with sentences, not a model
The usual first question is “should we use RBAC or ABAC?” It’s the wrong first question, because you can’t evaluate a model without knowing what it has to express.
Write the access rules down as plain sentences instead. For the invoicing tool:
- Anyone in a company can see that company’s invoices.
- Sam can create invoices for the Design team, and only that team.
- Priya can approve invoices up to $5,000 for the teams she manages.
- Nobody can approve an invoice they created.
- An owner can invite people and change their roles.
- The accounting integration can read invoices and nothing else.
- No one, ever, sees another customer’s data.
Seven sentences is a realistic starting list. If you can’t write yours, you’re not ready to design anything, and no framework will rescue you.
Now look at the shape they share:
A principal wants to perform an action on a resource, in some context. Every access-control model ever published is a different way of storing the answer to that one question.
The context is the part to read carefully. “For the teams she manages” is a scope. “Up to $5,000” is a condition on the resource’s data. “Created by someone else” is a relationship between the principal and the resource. Which of those appear in your sentences decides everything that follows.
03The models are a ladder
The models aren’t competing schools of thought. Each one adds a single capability to the one before it.
Admin flag
“An owner can do everything.”
Costs: Nothing, until the second kind of user shows up
Roles
“Approvers can approve invoices.”
Costs: A roles table, and discipline about what a role means
Scoped roles
“Sam can create invoices for the Design team only.”
Costs: Every check needs to know where the resource lives
Conditions
“Priya can approve up to $5,000.”
Costs: Rules that depend on data, so they are harder to list and explain
Relationships
“Anyone this folder is shared with can see what’s inside it.”
Costs: A graph of who relates to what, and usually a dedicated service
Go back to the seven sentences and find the highest step any of them needs. Most are plain roles. Two need a scope. One needs a condition, the $5,000 limit. One, “nobody can approve an invoice they created,” looks like a relationship, but it’s a single fixed rule and not a sharing system.
So this product wants scoped roles, one kind of condition, and one hard-coded rule. Not a policy engine. Not a relationship graph.
The rule I’d put above all the others:
choose the lowest step that can say every sentence you actually have.
Each step up is harder to explain to a customer, harder to list (“show me everyone who can approve this”), and harder to debug at 2 a.m. You can climb later. Climbing down is a rewrite.
04Needs and wants
IAM attracts speculative features more than almost any other part of a system, because every one of them sounds like security. It helps to split the list in two.
Needs: true for every product, from the first user
Tenant isolation that can’t be configured away
A leak between customers is the one mistake you can’t apologise your way out of.
Deny by default
A missing rule should mean no, not yes.
One function that makes every decision
You can’t fix or audit logic that is spread across two hundred handlers.
Code that checks actions, not role names
It’s what lets roles change later without touching code.
A record of who granted what
The first security question you’ll be asked is who gave this person access.
Revocation that works
When someone leaves, their access has to end within minutes, not when a token expires next week.
Wants: real, but only once something triggers them
| Feature | Build it when |
|---|---|
| Custom roles | When a customer’s org chart doesn’t fit your built-in roles |
| Conditions on grants | When a sentence contains a limit: an amount, a region, a time |
| SSO and automatic provisioning | When you sell to a company with an IT department |
| Temporary and break-glass access | When support staff or on-call engineers need to get in |
| Field-level permissions | When one column is more sensitive than the rest of the row |
| A policy language | When rules change faster than you can deploy, or non-engineers must write them |
The test for moving something from the second list to the first is simple: can you name the customer and write the sentence? “Enterprises will want custom roles” is a guess. “Acme’s finance lead needs to approve but must not see salaries” is a requirement.
The needs are cheap if you do them at the start and painful to retrofit. The wants are the reverse: expensive to build early, and usually straightforward to add later if the needs were done properly.
05Where to put the flexibility
This is the decision that separates a system that ages well from one that doesn’t. Flexibility isn’t good or bad. It’s good in specific places.
Fixed in code
changes with a deploy
The list of actions
Resource types
The tenant wall
Invariants no role can override
Data you control
changes with a migration or admin tool
Built-in roles
Which actions each role bundles
Plan and feature limits
Data customers control
changes in their settings page
Who has which role, and where
Limits on a grant
Custom roles, once someone needs them
The line to hold is between actions and roles.
An action is something your code does: approve an invoice, invite a member. It exists because a developer wrote the feature, so the list of actions belongs in code and changes with a deploy.
A role is a name for a bundle of actions. Which bundle counts as an “Approver” is a business opinion, and business opinions differ between customers and change over time. That belongs in data.
Code should only ever ask about actions:
// Brittle: the code knows what a "manager" is.
if (user.role === "manager" || user.role === "admin") {
await approve(invoice);
}
// Durable: the code knows what it is about to do.
if ((await can(user, "invoice.approve", invoice)).allowed) {
await approve(invoice);
}With the first version, adding a “Finance lead” role means finding every place that mentions a manager. With the second, it’s a row in a table.
The action list can be an ordinary typed constant. Keeping it in code means a typo is a compile error instead of a permission that silently never matches:
export const PERMISSIONS = [
"invoice.read",
"invoice.create",
"invoice.approve",
"invoice.void",
"member.invite",
"role.assign",
] as const;
export type Permission = (typeof PERMISSIONS)[number];Two things should never become settings.
The first is the tenant boundary. No role, no flag, and no support override should be able to make one customer’s data visible to another. If staff need access, that’s a separate, explicit, logged mechanism, not a permission.
The second is any rule that exists to stop a mistake or a fraud. “You can’t approve your own invoice” isn’t a preference a customer admin should be able to untick. Rules like that are invariants, and they live in code next to the tenant check.
06A shape that fits most products
Here’s a concrete design for the invoicing tool. It isn’t the only workable one, but it covers a surprising range of products, and every piece of it traces back to one of the seven sentences.
CREATE TABLE roles (
id uuid PRIMARY KEY,
tenant_id uuid REFERENCES tenants, -- NULL: a built-in role every tenant gets
name text NOT NULL
);
CREATE TABLE role_permissions (
role_id uuid NOT NULL REFERENCES roles,
permission text NOT NULL, -- 'invoice.approve'
PRIMARY KEY (role_id, permission)
);
CREATE TABLE role_assignments (
tenant_id uuid NOT NULL REFERENCES tenants,
user_id uuid NOT NULL REFERENCES users,
role_id uuid NOT NULL REFERENCES roles,
scope_type text NOT NULL, -- 'tenant' or 'team'
scope_id uuid NOT NULL,
conditions jsonb NOT NULL DEFAULT '{}', -- {"max_amount": 5000}
granted_by uuid NOT NULL REFERENCES users,
expires_at timestamptz,
PRIMARY KEY (user_id, role_id, scope_type, scope_id)
);A few of these columns are doing more than they appear to.
roles.tenant_id being nullable gives you built-in roles and custom roles in one table. You ship Owner, Approver, and Member with a null tenant. The day a customer needs their own role, it’s a row with their tenant id, and no code changes.
scope_type and scope_id turn a role into a role somewhere. Sam is a Member on the Design team. Priya is an Approver on two teams. An owner is an Owner on the whole tenant. That’s one extra pair of columns, and it removes the need for roles like “design-team-member.”
conditions holds the limit for this particular grant. Priya’s Approver assignment carries a maximum amount of 5,000, and someone else’s can carry a different one. The role stays the same.
granted_by and expires_at cost nothing to add now and are miserable to backfill. One answers “who gave this person access?” The other makes temporary access possible later without a migration.
07One function, and the query too
Every decision goes through one function. It takes the four parts of the sentence and returns an answer with a reason.
1Is the resource in the caller’s tenant?
never configurable
no → deny
2Does a hard rule forbid it?
for example, approving your own invoice
yes → deny
3Does a role they hold here include the action?
team first, then tenant-wide
no → deny
4Are that grant’s conditions met?
for example, the amount limit
no → deny
export async function can(
who: Principal,
action: Permission,
resource: Resource,
): Promise<Decision> {
// 1. The wall. Not configurable, checked first.
if (resource.tenantId !== who.tenantId) return deny("other tenant");
// 2. Invariants that no role can override.
if (action === "invoice.approve" && resource.createdBy === who.id) {
return deny("cannot approve your own invoice");
}
// 3. Roles that apply here: on the resource's team, or tenant-wide.
const grants = await grantsFor(who, [
{ type: "team", id: resource.teamId },
{ type: "tenant", id: resource.tenantId },
]);
// 4. Does one of them include the action, with its conditions met?
for (const grant of grants) {
if (!grant.permissions.has(action)) continue;
const limit = grant.conditions.max_amount;
if (limit != null && resource.amount > limit) continue;
return allow(grant.role + " on " + grant.scope.type);
}
// 5. Nothing said yes.
return deny("no role grants " + action);
}This is thirty lines, and it’s deliberately boring. It never throws to signal a denial, it returns a reason every time, and the default at the bottom is no.
The reason matters more than it looks. It’s what you log, what support reads when a customer says “I can’t approve this,” and what lets you build a “why can’t I?” screen later without touching the logic.
The list problem
A function like can answers “may Priya see this invoice?” Most screens ask a different question: “which invoices may Priya see?”
Loading every row and filtering through can one at a time doesn’t survive real data, and it breaks pagination. The same rules have to be expressible as a query:
SELECT i.*
FROM invoices i
WHERE i.tenant_id = $1 -- the wall, on every query, always
AND (
$2 -- true if they hold invoice.read tenant-wide
OR i.team_id = ANY($3) -- teams where they hold invoice.read
)
ORDER BY i.created_at DESC
LIMIT 50;This is the strongest practical argument for staying low on the ladder. “Teams where the user holds a role” translates into a WHERE clause. An arbitrary policy written in a general language often doesn’t, and then every list endpoint becomes its own research project.
Keep the two paths next to each other in the code, and test them against each other: anything the list returns must pass can, and the reverse.
08What goes in the token
It’s tempting to put the user’s permissions in their access token so that services don’t need a lookup. Compare two payloads:
{
"sub": "u_123",
"tenant": "t_42",
"roles": ["approver"],
"permissions": [
"invoice.read",
"invoice.approve"
],
"exp": 1792144800
}{
"sub": "u_123",
"tenant": "t_42",
"sid": "s_9f2c",
"iat": 1791540000,
"exp": 1791540900
}The first token is valid for a week and says what the user could do at the moment it was issued. Remove their Approver role on Tuesday and they keep approving until the following Monday. It also can’t express scopes or limits without growing with every team they join.
The second says who they are, which tenant they’re acting in, and which session this is. It expires in fifteen minutes. What they’re allowed to do is looked up on the server, where it can be cached for seconds and invalidated when an assignment changes.
A useful way to keep the two apart:
a token proves identity. Authorization is a question you ask fresh, every time.
There’s a real trade-off here. A lookup per request costs something, and a system with many services may accept coarse claims in a very short-lived token to avoid it. That can be a sound choice. Just make it knowing that the token’s lifetime is now your revocation delay.
09What not to do
Most of what goes wrong is one of a small number of mistakes. In rough order of how often they cause trouble:
Checking role names in code
Every new role becomes a code change, and the meaning of “manager” drifts between files.
Minting a role per exception
“Manager, north region, read-only” is not a role. It’s a role, a scope, and a limit that were never separated.
Taking the tenant from the request
The tenant comes from the session, on the server. A tenant id in a URL or body is a suggestion from the client.
Deciding in the frontend
Hiding a button is a courtesy. The API still has to say no.
A silent super-admin
Staff access that bypasses every check will be used, and one day questioned. Make it explicit, time-limited, and logged.
Deny rules on day one
As soon as rules can both allow and forbid, order matters, and nobody can predict the result. Start with allow-only and default deny.
Building an engine before you have rules
A policy language with three policies is a compiler you maintain for no reason.
Forgetting the non-humans
API keys, integrations, and background jobs are principals too, and usually the over-privileged ones.
The second one deserves an example, because it’s the most common way a clean role model decays. A customer asks for a regional manager who can only view. The quick fix is a new role. Twenty requests later there are forty roles, and each is a slightly different combination of the same three ideas: what they can do, where, and within what limit.
If your assignments already carry a scope and conditions, that request is a new assignment of an existing role, and the roles table stays at five rows.
10Fit it to your use case
The right design depends less on what’s fashionable than on what your sentences look like. Four common situations, and where I’d start with each:
Internal tool, a few dozen users
SSO and three hard-coded roles
You know every user by name. Spend the effort on audit logs, not on a permissions UI nobody will open.
B2B SaaS with customers who have teams
Scoped roles, a handful of conditions
This is the shape in this article. Tenant wall, built-in roles, assignments per team, limits where a sentence demands one.
Anything with a Share button
Relationships
When access follows folders, documents, and who shared what with whom, roles stop fitting. Look at Zanzibar-style systems before inventing one.
Regulated: money, health, infrastructure
Scoped roles plus process
The model is ordinary. The extras matter: separation of duties, time-limited grants, access reviews, and a log an auditor can read.
A few questions will usually place you. Do customers have their own internal structure, such as teams, regions, or projects? Then you need scopes. Do rules mention amounts, dates, or statuses? Then you need conditions. Can a user grant access to a single object, the way you share a document? Then you need relationships. Will an auditor read your logs? Then the audit trail is a feature, not plumbing.
Build or adopt?
For authentication, adopt. For authorization, it depends on where you are on the ladder.
Up to scoped roles with a few conditions, a few tables and one function is less to operate than any external system, and you can read all of it in an afternoon. Past that, you’re writing an engine, and good ones already exist: Open Policy Agent and Cedar for policies written as code, OpenFGA and SpiceDB for relationship graphs in the style of Google’s Zanzibar.
Whichever you choose, keep the can function as the only thing your application calls. Then moving from tables to an engine later is a change behind one interface instead of a change in every handler.
11Your sentences are your tests
The sentences from the start of this article aren’t only a design tool. They’re the test suite.
const cases = [
["Priya approves a $3,000 invoice from her team", priya, "invoice.approve", invoice({ amount: 3000 }), true],
["Priya approves a $9,000 invoice", priya, "invoice.approve", invoice({ amount: 9000 }), false],
["Priya approves her own invoice", priya, "invoice.approve", invoice({ createdBy: priya.id }), false],
["Sam creates an invoice for another team", sam, "invoice.create", invoice({ teamId: marketing.id }), false],
["Sam reads another tenant's invoice", sam, "invoice.read", invoice({ tenantId: other.id }), false],
] as const;
test.each(cases)("%s", async (_name, who, action, resource, expected) => {
const decision = await can(who, action, resource);
expect(decision.allowed).toBe(expected);
});Notice that most of the cases expect a denial. That’s the right proportion. Features get tested by people using them. Nobody tries the thing they aren’t supposed to be able to do, except an attacker.
When a customer reports that someone could see something they shouldn’t, the fix starts with a new row in this table.
12The short version
- 01Write the access questions as sentences before choosing a model.
- 02Pick the lowest step of the ladder that can say all of them.
- 03Make the tenant boundary a rule in code, not a setting.
- 04Check actions in code. Let roles be data.
- 05Put every decision behind one function that returns a reason.
- 06Apply the same rules to list queries as to single lookups.
- 07Keep tokens short-lived and small. Look permissions up on the server.
- 08Log grants and denials from the first day.
- 09Turn the sentences into tests.
- 10Add flexibility when a real customer’s sentence needs it, not before.
None of this is exotic. A good access system is mostly a small number of boring decisions made early and then defended: one wall, one function, a short list of actions, and roles that are only ever data.
The interesting part is restraint. The system that holds up isn’t the one that can express anything. It’s the one where anyone on the team can answer “who can do this, and why?” without opening the code.
Further reading
- Authorization Cheat Sheet, OWASP
- Role Based Access Control, NIST
- Zanzibar: Google’s Consistent, Global Authorization System
- Open Policy Agent documentation
- Cedar policy language
- OpenFGA
The invoicing product, its people, and its schema are invented for this article. Treat the code as a sketch of the shape, not a library to copy.