← Writing

Designing IAMWhat to Decide Before You Write a Permission Check

Start from the questions your product has to answer, pick the simplest model that can say them, and put flexibility only where customers actually differ.

Oct 09, 2026 · 17 min read

Every permission system starts as if (user.isAdmin).

It works for months. Then a customer asks for someone who can view billing but not change it. Then for a manager who can approve expenses, but only for their own team, and only up to a limit. Each request is reasonable. Each one adds another if.

A year later nobody can say with confidence who is able to do what, and the honest answer to “can this user delete that?” is “let me read the code.”

Identity and access management has a reputation for being complicated. Most of the complication comes from making decisions in the wrong order: choosing a model before knowing the questions, and adding flexibility before knowing where it’s needed.

This is how I’d approach designing it now. What to settle first, what to leave rigid on purpose, where flexibility pays for itself, and what to refuse to build until someone needs it. The examples use a made-up product: a multi-tenant invoicing tool where companies have teams, and teams raise and approve invoices.

01IAM is four jobs, not one

“IAM” gets used as if it were a single feature. It’s four separate jobs that happen to share a login page.

They fail differently, and they deserve different amounts of your own engineering.

Authentication is the one to build the least of. Password storage, session handling, multi-factor, SSO: these are solved problems with sharp edges, and nothing about them is specific to your product. Use an identity provider or a well-maintained library, and speak a standard protocol such as OpenID Connect so you can change your mind later.

Authorization is the opposite. No vendor knows that your approvers shouldn’t approve their own invoices. The rules are your product’s rules, so the design work is yours even if you run someone else’s engine underneath.

That’s where the rest of this article spends its time. Broken access control sits at the top of the OWASP Top 10 for a reason: it’s the part everyone writes themselves.

02Start with sentences, not a model

The usual first question is “should we use RBAC or ABAC?” It’s the wrong first question, because you can’t evaluate a model without knowing what it has to express.

Write the access rules down as plain sentences instead. For the invoicing tool:

  • Anyone in a company can see that company’s invoices.
  • Sam can create invoices for the Design team, and only that team.
  • Priya can approve invoices up to $5,000 for the teams she manages.
  • Nobody can approve an invoice they created.
  • An owner can invite people and change their roles.
  • The accounting integration can read invoices and nothing else.
  • No one, ever, sees another customer’s data.

Seven sentences is a realistic starting list. If you can’t write yours, you’re not ready to design anything, and no framework will rescue you.

Now look at the shape they share:

Every access question has the same four parts. The last one is where the models differ.

A principal wants to perform an action on a resource, in some context. Every access-control model ever published is a different way of storing the answer to that one question.

The context is the part to read carefully. “For the teams she manages” is a scope. “Up to $5,000” is a condition on the resource’s data. “Created by someone else” is a relationship between the principal and the resource. Which of those appear in your sentences decides everything that follows.

03The models are a ladder

The models aren’t competing schools of thought. Each one adds a single capability to the one before it.

Climb a step only when a real sentence from your list can’t be said on the one below.

Admin flag

“An owner can do everything.”

Costs: Nothing, until the second kind of user shows up

Roles

“Approvers can approve invoices.”

Costs: A roles table, and discipline about what a role means

Scoped roles

“Sam can create invoices for the Design team only.”

Costs: Every check needs to know where the resource lives

Conditions

“Priya can approve up to $5,000.”

Costs: Rules that depend on data, so they are harder to list and explain

Relationships

“Anyone this folder is shared with can see what’s inside it.”

Costs: A graph of who relates to what, and usually a dedicated service

Go back to the seven sentences and find the highest step any of them needs. Most are plain roles. Two need a scope. One needs a condition, the $5,000 limit. One, “nobody can approve an invoice they created,” looks like a relationship, but it’s a single fixed rule and not a sharing system.

So this product wants scoped roles, one kind of condition, and one hard-coded rule. Not a policy engine. Not a relationship graph.

The rule I’d put above all the others:

choose the lowest step that can say every sentence you actually have.

Each step up is harder to explain to a customer, harder to list (“show me everyone who can approve this”), and harder to debug at 2 a.m. You can climb later. Climbing down is a rewrite.

04Needs and wants

IAM attracts speculative features more than almost any other part of a system, because every one of them sounds like security. It helps to split the list in two.

Needs: true for every product, from the first user

Tenant isolation that can’t be configured away

A leak between customers is the one mistake you can’t apologise your way out of.

Deny by default

A missing rule should mean no, not yes.

One function that makes every decision

You can’t fix or audit logic that is spread across two hundred handlers.

Code that checks actions, not role names

It’s what lets roles change later without touching code.

A record of who granted what

The first security question you’ll be asked is who gave this person access.

Revocation that works

When someone leaves, their access has to end within minutes, not when a token expires next week.

Wants: real, but only once something triggers them

FeatureBuild it when
Custom rolesWhen a customer’s org chart doesn’t fit your built-in roles
Conditions on grantsWhen a sentence contains a limit: an amount, a region, a time
SSO and automatic provisioningWhen you sell to a company with an IT department
Temporary and break-glass accessWhen support staff or on-call engineers need to get in
Field-level permissionsWhen one column is more sensitive than the rest of the row
A policy languageWhen rules change faster than you can deploy, or non-engineers must write them

The test for moving something from the second list to the first is simple: can you name the customer and write the sentence? “Enterprises will want custom roles” is a guess. “Acme’s finance lead needs to approve but must not see salaries” is a requirement.

The needs are cheap if you do them at the start and painful to retrofit. The wants are the reverse: expensive to build early, and usually straightforward to add later if the needs were done properly.

05Where to put the flexibility

This is the decision that separates a system that ages well from one that doesn’t. Flexibility isn’t good or bad. It’s good in specific places.

Flexibility belongs on the right. Meaning belongs on the left.

The line to hold is between actions and roles.

An action is something your code does: approve an invoice, invite a member. It exists because a developer wrote the feature, so the list of actions belongs in code and changes with a deploy.

A role is a name for a bundle of actions. Which bundle counts as an “Approver” is a business opinion, and business opinions differ between customers and change over time. That belongs in data.

Code should only ever ask about actions:

// Brittle: the code knows what a "manager" is.
if (user.role === "manager" || user.role === "admin") {
  await approve(invoice);
}

// Durable: the code knows what it is about to do.
if ((await can(user, "invoice.approve", invoice)).allowed) {
  await approve(invoice);
}

With the first version, adding a “Finance lead” role means finding every place that mentions a manager. With the second, it’s a row in a table.

The action list can be an ordinary typed constant. Keeping it in code means a typo is a compile error instead of a permission that silently never matches:

export const PERMISSIONS = [
  "invoice.read",
  "invoice.create",
  "invoice.approve",
  "invoice.void",
  "member.invite",
  "role.assign",
] as const;

export type Permission = (typeof PERMISSIONS)[number];

Two things should never become settings.

The first is the tenant boundary. No role, no flag, and no support override should be able to make one customer’s data visible to another. If staff need access, that’s a separate, explicit, logged mechanism, not a permission.

The second is any rule that exists to stop a mistake or a fraud. “You can’t approve your own invoice” isn’t a preference a customer admin should be able to untick. Rules like that are invariants, and they live in code next to the tenant check.

06A shape that fits most products

Here’s a concrete design for the invoicing tool. It isn’t the only workable one, but it covers a surprising range of products, and every piece of it traces back to one of the seven sentences.

The assignment is the interesting table: it says who, which role, and where.
CREATE TABLE roles (
  id         uuid PRIMARY KEY,
  tenant_id  uuid REFERENCES tenants,     -- NULL: a built-in role every tenant gets
  name       text NOT NULL
);

CREATE TABLE role_permissions (
  role_id    uuid NOT NULL REFERENCES roles,
  permission text NOT NULL,               -- 'invoice.approve'
  PRIMARY KEY (role_id, permission)
);

CREATE TABLE role_assignments (
  tenant_id  uuid NOT NULL REFERENCES tenants,
  user_id    uuid NOT NULL REFERENCES users,
  role_id    uuid NOT NULL REFERENCES roles,
  scope_type text NOT NULL,               -- 'tenant' or 'team'
  scope_id   uuid NOT NULL,
  conditions jsonb NOT NULL DEFAULT '{}', -- {"max_amount": 5000}
  granted_by uuid NOT NULL REFERENCES users,
  expires_at timestamptz,
  PRIMARY KEY (user_id, role_id, scope_type, scope_id)
);

A few of these columns are doing more than they appear to.

roles.tenant_id being nullable gives you built-in roles and custom roles in one table. You ship Owner, Approver, and Member with a null tenant. The day a customer needs their own role, it’s a row with their tenant id, and no code changes.

scope_type and scope_id turn a role into a role somewhere. Sam is a Member on the Design team. Priya is an Approver on two teams. An owner is an Owner on the whole tenant. That’s one extra pair of columns, and it removes the need for roles like “design-team-member.”

conditions holds the limit for this particular grant. Priya’s Approver assignment carries a maximum amount of 5,000, and someone else’s can carry a different one. The role stays the same.

granted_by and expires_at cost nothing to add now and are miserable to backfill. One answers “who gave this person access?” The other makes temporary access possible later without a migration.

07One function, and the query too

Every decision goes through one function. It takes the four parts of the sentence and returns an answer with a reason.

Four questions, in this order. Only a full run of right answers ends in allow.
export async function can(
  who: Principal,
  action: Permission,
  resource: Resource,
): Promise<Decision> {
  // 1. The wall. Not configurable, checked first.
  if (resource.tenantId !== who.tenantId) return deny("other tenant");

  // 2. Invariants that no role can override.
  if (action === "invoice.approve" && resource.createdBy === who.id) {
    return deny("cannot approve your own invoice");
  }

  // 3. Roles that apply here: on the resource's team, or tenant-wide.
  const grants = await grantsFor(who, [
    { type: "team", id: resource.teamId },
    { type: "tenant", id: resource.tenantId },
  ]);

  // 4. Does one of them include the action, with its conditions met?
  for (const grant of grants) {
    if (!grant.permissions.has(action)) continue;
    const limit = grant.conditions.max_amount;
    if (limit != null && resource.amount > limit) continue;
    return allow(grant.role + " on " + grant.scope.type);
  }

  // 5. Nothing said yes.
  return deny("no role grants " + action);
}

This is thirty lines, and it’s deliberately boring. It never throws to signal a denial, it returns a reason every time, and the default at the bottom is no.

The reason matters more than it looks. It’s what you log, what support reads when a customer says “I can’t approve this,” and what lets you build a “why can’t I?” screen later without touching the logic.

The list problem

A function like can answers “may Priya see this invoice?” Most screens ask a different question: “which invoices may Priya see?”

Loading every row and filtering through can one at a time doesn’t survive real data, and it breaks pagination. The same rules have to be expressible as a query:

SELECT i.*
FROM invoices i
WHERE i.tenant_id = $1               -- the wall, on every query, always
  AND (
    $2                               -- true if they hold invoice.read tenant-wide
    OR i.team_id = ANY($3)           -- teams where they hold invoice.read
  )
ORDER BY i.created_at DESC
LIMIT 50;

This is the strongest practical argument for staying low on the ladder. “Teams where the user holds a role” translates into a WHERE clause. An arbitrary policy written in a general language often doesn’t, and then every list endpoint becomes its own research project.

Keep the two paths next to each other in the code, and test them against each other: anything the list returns must pass can, and the reverse.

08What goes in the token

It’s tempting to put the user’s permissions in their access token so that services don’t need a lookup. Compare two payloads:

Tempting
{
  "sub": "u_123",
  "tenant": "t_42",
  "roles": ["approver"],
  "permissions": [
    "invoice.read",
    "invoice.approve"
  ],
  "exp": 1792144800
}
Better
{
  "sub": "u_123",
  "tenant": "t_42",
  "sid": "s_9f2c",
  "iat": 1791540000,
  "exp": 1791540900
}

The first token is valid for a week and says what the user could do at the moment it was issued. Remove their Approver role on Tuesday and they keep approving until the following Monday. It also can’t express scopes or limits without growing with every team they join.

The second says who they are, which tenant they’re acting in, and which session this is. It expires in fifteen minutes. What they’re allowed to do is looked up on the server, where it can be cached for seconds and invalidated when an assignment changes.

A useful way to keep the two apart:

a token proves identity. Authorization is a question you ask fresh, every time.

There’s a real trade-off here. A lookup per request costs something, and a system with many services may accept coarse claims in a very short-lived token to avoid it. That can be a sound choice. Just make it knowing that the token’s lifetime is now your revocation delay.

09What not to do

Most of what goes wrong is one of a small number of mistakes. In rough order of how often they cause trouble:

Checking role names in code

Every new role becomes a code change, and the meaning of “manager” drifts between files.

Minting a role per exception

“Manager, north region, read-only” is not a role. It’s a role, a scope, and a limit that were never separated.

Taking the tenant from the request

The tenant comes from the session, on the server. A tenant id in a URL or body is a suggestion from the client.

Deciding in the frontend

Hiding a button is a courtesy. The API still has to say no.

A silent super-admin

Staff access that bypasses every check will be used, and one day questioned. Make it explicit, time-limited, and logged.

Deny rules on day one

As soon as rules can both allow and forbid, order matters, and nobody can predict the result. Start with allow-only and default deny.

Building an engine before you have rules

A policy language with three policies is a compiler you maintain for no reason.

Forgetting the non-humans

API keys, integrations, and background jobs are principals too, and usually the over-privileged ones.

The second one deserves an example, because it’s the most common way a clean role model decays. A customer asks for a regional manager who can only view. The quick fix is a new role. Twenty requests later there are forty roles, and each is a slightly different combination of the same three ideas: what they can do, where, and within what limit.

If your assignments already carry a scope and conditions, that request is a new assignment of an existing role, and the roles table stays at five rows.

10Fit it to your use case

The right design depends less on what’s fashionable than on what your sentences look like. Four common situations, and where I’d start with each:

Internal tool, a few dozen users

SSO and three hard-coded roles

You know every user by name. Spend the effort on audit logs, not on a permissions UI nobody will open.

B2B SaaS with customers who have teams

Scoped roles, a handful of conditions

This is the shape in this article. Tenant wall, built-in roles, assignments per team, limits where a sentence demands one.

Anything with a Share button

Relationships

When access follows folders, documents, and who shared what with whom, roles stop fitting. Look at Zanzibar-style systems before inventing one.

Regulated: money, health, infrastructure

Scoped roles plus process

The model is ordinary. The extras matter: separation of duties, time-limited grants, access reviews, and a log an auditor can read.

A few questions will usually place you. Do customers have their own internal structure, such as teams, regions, or projects? Then you need scopes. Do rules mention amounts, dates, or statuses? Then you need conditions. Can a user grant access to a single object, the way you share a document? Then you need relationships. Will an auditor read your logs? Then the audit trail is a feature, not plumbing.

Build or adopt?

For authentication, adopt. For authorization, it depends on where you are on the ladder.

Up to scoped roles with a few conditions, a few tables and one function is less to operate than any external system, and you can read all of it in an afternoon. Past that, you’re writing an engine, and good ones already exist: Open Policy Agent and Cedar for policies written as code, OpenFGA and SpiceDB for relationship graphs in the style of Google’s Zanzibar.

Whichever you choose, keep the can function as the only thing your application calls. Then moving from tables to an engine later is a change behind one interface instead of a change in every handler.

11Your sentences are your tests

The sentences from the start of this article aren’t only a design tool. They’re the test suite.

const cases = [
  ["Priya approves a $3,000 invoice from her team", priya, "invoice.approve", invoice({ amount: 3000 }),          true],
  ["Priya approves a $9,000 invoice",               priya, "invoice.approve", invoice({ amount: 9000 }),          false],
  ["Priya approves her own invoice",                priya, "invoice.approve", invoice({ createdBy: priya.id }),   false],
  ["Sam creates an invoice for another team",       sam,   "invoice.create",  invoice({ teamId: marketing.id }),  false],
  ["Sam reads another tenant's invoice",            sam,   "invoice.read",    invoice({ tenantId: other.id }),    false],
] as const;

test.each(cases)("%s", async (_name, who, action, resource, expected) => {
  const decision = await can(who, action, resource);
  expect(decision.allowed).toBe(expected);
});

Notice that most of the cases expect a denial. That’s the right proportion. Features get tested by people using them. Nobody tries the thing they aren’t supposed to be able to do, except an attacker.

When a customer reports that someone could see something they shouldn’t, the fix starts with a new row in this table.

12The short version

  1. 01Write the access questions as sentences before choosing a model.
  2. 02Pick the lowest step of the ladder that can say all of them.
  3. 03Make the tenant boundary a rule in code, not a setting.
  4. 04Check actions in code. Let roles be data.
  5. 05Put every decision behind one function that returns a reason.
  6. 06Apply the same rules to list queries as to single lookups.
  7. 07Keep tokens short-lived and small. Look permissions up on the server.
  8. 08Log grants and denials from the first day.
  9. 09Turn the sentences into tests.
  10. 10Add flexibility when a real customer’s sentence needs it, not before.

None of this is exotic. A good access system is mostly a small number of boring decisions made early and then defended: one wall, one function, a short list of actions, and roles that are only ever data.

The interesting part is restraint. The system that holds up isn’t the one that can express anything. It’s the one where anyone on the team can answer “who can do this, and why?” without opening the code.

Further reading

The invoicing product, its people, and its schema are invented for this article. Treat the code as a sketch of the shape, not a library to copy.