Security
Part 5 of 6 · AuthorizationPolicy Engines — OPA/Rego vs Cedar vs Custom
Comparative OPA/Rego vs Cedar vs custom if/else; treat policies as audited code.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Pick the engine for the job you actually have
Prefer
A constrained AuthZ language when the problem is allow or deny
Cedar-style permit and forbid stays analyzable. You can ask what a change denies before you ship it.
- Application authorization, not a general programming language.
- Explicit forbids beat a comment that says default deny.
- Still ship tests, bundles, and a canary.
Alternative
Rego everywhere, or a fork of if statements
OPA is the right hammer for admission control and data-heavy rules. It is a heavy hammer for a five-line document check. Custom forks are how two services disagree.
- General-purpose policy can be slow or surprising.
- Custom code is fastest until the third service copies it wrong.
- Either way, default allow left in from a prototype is the outage.
Policy as code, not a wiki page
The decision path and the release path are both part of the design.
- 1
Shared logic or not
One service and a stable rule can stay a tested module. Many services need one PDP. - 2
Choose the language
Broad data-driven or infra policy leans OPA. Analyzable application AuthZ leans Cedar. Tiny and stable can stay custom, with an exit plan. - 3
Bundle and pin
Policies ship like artifacts. Environments pin a version. Rollback is a version change, not a hotfix in a pod. - 4
Test, shadow, canary
Table of inputs to allow or deny in CI. Log the new decision before you enforce it. Watch deny spikes on a slice of traffic.
Overview
Once conditionals scatter, you need a policy engine: a PDP that evaluates versioned policies against structured input. The shortlist in this cluster is OPA/Rego, Cedar, and custom code. Trade expressiveness, how analyzable the language is, latency, testing, and audit.
Interviews ask where policies live. Production asks who broke the deploy with a policy push. Engines give you a place to test and a decision log. They do not replace judgment about default deny.
This page does not re-teach token validation. That is JWT vs Opaque Tokens — Validation, JWKS & Revocation. Input to the engine is a subject, an action, and a resource.
Decision path
Flow
- 1
1. Shared AuthZ across services
- next2. Broad rules point at OPA
- 2
2. Broad rules point at OPA
- next3. Analyzable AuthZ points at Cedar
- 3
3. Analyzable AuthZ points at Cedar
- next4. Tiny stable rules stay custom
- 4
4. Tiny stable rules stay custom
- next5. Tests, bundle, canary
- 5
5. Tests, bundle, canary
Lesson map
Policy Engines — OPA/Rego vs Cedar vs Custom
Comparative OPA/Rego vs Cedar vs custom if/else; treat policies as audited code.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Shared AuthZ across services"] b["2. Broad rules point at OPA"] c["3. Analyzable AuthZ points at Cedar"] d["4. Tiny stable rules stay custom"] a -->|1. Shared AuthZ across services| b b -->|2. Broad rules point at OPA| c c -->|3. Analyzable AuthZ points at Cedar| d
Rule of thumb: prefer Cedar, or a similarly constrained language, when the problem is application allow/deny and you want analysis. Prefer OPA when you already use it for admission or you need broader data-driven rules. Prefer custom only for tiny, stable checks with ruthless tests, and write down the exit.
Matrix
| Dimension | OPA/Rego | Cedar | Custom conditionals |
|---|---|---|---|
| Expressiveness | Very high, general purpose | Constrained AuthZ language | Unlimited and unstructured |
| Analysis | Powerful; easy to write a slow surprise | Built for decidable permit and forbid | Whatever the tests cover |
| Latency | Sidecar or library; watch input size | Fast in-process evaluation | Fastest while it stays tiny |
| Ecosystem | Kubernetes admission, Envoy, CI | Verified Permissions; portable Cedar | None |
| Audit | Decision logs and bundles | Explicit permit and forbid | Ad-hoc logs |
| Testing | opa test | Policy tests and analysis | Ordinary unit tests |
| Learning curve | Rego is a different paradigm | Smaller AuthZ surface | Low at first, high later |
Practices that survive a bad push
- Bundles and versions. Ship policies as artifacts. Pin them per environment.
- Unit tests. Every important allow and deny is a row in CI.
- Dry-run or shadow. Log what the new policy would have decided before it enforces.
- Canary. Enforce on a slice. Watch deny rate and latency.
- Decision logs. Subject, action, resource, policy id, effect. Leave raw secrets and full bodies out.
- Break-glass. An out-of-band emergency allow with a ticket and an expiry. The enforcement lesson owns the runbook.
Conceptual shapes, not trivia to memorize:
- Rego: default deny, then
allowwhen the input's role and action match. - Cedar:
permitwhen the resource owner is the principal, plus an explicitforbidyou can analyze.
Memorize default deny, tests, and who is allowed to push policy.
Sandbox: custom fork vs a data table
The two functions agree on the cases below. They will not agree after the next service copies only the first function and edits it.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
The log has ids and a policy version. It does not have a bearer token, a request body, or an email you do not need for the forensics question you actually ask.
Interview Q&A
Why not keep AuthZ as conditionals in each service?
Answer
They drift. Denies disagree. Audits become a search across repos. There is no single test suite. An engine centralizes evaluation and logging. A tiny custom module is acceptable only while one team owns it and the tests are ruthless.
OPA vs Cedar in one breath?
Answer
OPA/Rego is general-purpose policy for infra and application data. Cedar is an authorization language aimed at analyzable permit and forbid. Pick from the problem, not from the logo on the last conference talk.
What is a policy bundle?
Answer
A versioned artifact of policies, sometimes with data, shipped to PDP instances. Pin it per environment so rollback is a version, not an edit on one replica.
How do you test policies?
Answer
Table-driven cases: fixed input, expected allow or deny. Run them in CI on every change. Shadow in production before enforce, then canary.
What latency should you expect?
Answer
Often under a millisecond to a few milliseconds in-process. A network hop to a PDP needs a timeout, a cache with a policy version in the key, and fail-closed behavior. Caching details are the next lesson.
Can an engine replace a ReBAC store?
Answer
Partially. Engines evaluate rules. Relationship graphs still need tuples or equivalent data. OPA can ingest or call relation data. Zanzibar-style systems specialize in check and list. Do not pretend a Rego file is the share graph.
What is the biggest operational risk?
Answer
A hot reload that denies everything, or a typo that fails open because default allow survived the prototype. Default deny, canary, and a metric on decision outcomes.
What is shadow mode?
Answer
The new policy computes a decision and you log it next to the decision you actually enforce. You compare them before you flip enforcement. It is how you see a deny spike without causing one.
Why is HTTP inside a policy dangerous?
Answer
Evaluation latency becomes the callee's latency. A policy that fetches arbitrary URLs is also an injection path. Load data before evaluation, or from a bounded PIP. Do not let the policy browse.
Pitfalls
You have twelve services, a document share graph, and a region constraint. Say which system owns tuples, which system owns the region rule, and whether the first release is Cedar, OPA, or a custom module. Name the CI test and the canary metric you would watch on day one.