← Blog
Agents

Is your AI governed, or just stuck?

Too little governance creates surprises. Too much creates a waiting room. Here is how to give AI enough authority to finish the job, and clear limits on what it can do.

Bench6 min read
A brushed steel arrow pointing right. Its shaft is wrapped in a tangled knot of red ribbon on the left, while the arrowhead and two thin rails are clear on the right, against a dark background.
Conceptual illustration generated with AI.

You ask your AI to fix something. It understands the problem. It can describe the solution. Then it asks for permission to write the plan, permission to hand off the plan, and permission to start the work you already asked it to do.

By the end, you have become the project manager for a tool that was supposed to give you time back.

Now picture the opposite. An AI changes customer records, sends a message, or promises a discount without knowing whether it has the authority. Work moves quickly. You discover the consequences later.

Both situations expose a governance problem. One leaves too little room for judgment. The other leaves too much room for assumptions.

The useful question is: What decisions should this AI be able to make, under what conditions, and who owns the result?

What governance actually means

Governance is how an organization assigns authority, sets boundaries, and holds someone accountable for outcomes.

You already use it. A crew lead can adjust the order of work. An office manager can correct a scheduling mistake. A salesperson may be able to offer a discount up to a defined limit. Each person knows where their judgment applies and when they need to involve someone else.

AI needs the same clarity, especially when it can take actions through software.

A useful starting point is to answer four questions:

  • Authority: What may it decide and do on its own?
  • Boundaries: What data, systems, money, and commitments are in scope?
  • Evidence: How will we know the work was done correctly?
  • Accountability: Who owns the outcome and handles an exception?

This emphasis on clear responsibilities and ongoing review is consistent with the NIST AI Risk Management Framework. The practical examples here are our way of applying that principle to everyday work. They are not a formal compliance checklist.

When there is too little governance

Imagine an AI helping a contractor follow up on estimates. Drafting a follow-up from an approved template is one kind of action. Sending it to every contact in the database is another. Offering a price reduction to close the sale is another still.

If the instruction is simply “handle follow-up,” the system may have to guess where its authority ends.

Warning signs include changes nobody can explain, customer promises nobody authorized, access to records unrelated to the task, and no clear way to stop or correct a mistake. Another is a confident “done” with no evidence that the intended result actually happened.

The remedy is specific: define the audience, permitted actions, limits, and conditions that require help. Use permissions and software controls to enforce consequential boundaries. A sentence in a prompt is useful guidance; it cannot substitute for access controls.

When there is too much governance

Now imagine that same AI must get a fresh approval to look up the estimate, draft the follow-up, save the draft, and put it in the review queue, even though you already asked it to prepare the follow-up.

Each step might sound reasonable in isolation. Together, they make you supervise the mechanics instead of making the decisions that matter.

Look for repeated requests to approve the same intent, reviews that add no new information, and routine fixes waiting on several people with overlapping authority. Watch for agents producing more handoff documents than finished work.

A particularly revealing test: Can anyone explain what failure a required step prevents?

If the answer is “that’s the process,” inspect the step. It may be a useful control whose purpose needs to be made clear. It may also be a rule that outlived the problem it was created to solve.

Match the control to the consequence

The right amount of governance depends on the action. Consider its impact, the people affected, the sensitivity of the information, and how easily a mistake can be detected and reversed.

These are example defaults to adapt to your business:

WorkUseful defaultEvidence or boundary
Draft an internal summaryLet it proceedLink the source material; identify uncertainty
Prepare a customer follow-upLet it draftUse the approved context; review before sending unless sending authority is already defined
Fix a bug on an isolated branchLet it build and testOpen a pull request with relevant checks and a clear explanation
Change prices, payment details, or access permissionsRequire specific authorityVerify scope and record the decision before applying it
Delete production recordsRequire an explicit decision and recovery planConfirm the affected records and recovery method

A reversible action can still expose sensitive information or affect many people. “We can undo it” is one consideration, not blanket permission.

The aim is to place a decision at the point where the consequences change. Preparing a change and applying it to a live business are different moments. An AI should be able to get the work ready while preserving the decision that belongs to a person.

What we changed in our own workflow

We ran into this ourselves. Our local Aurelius instructions had accumulated mandatory specialist handoffs, a six-section commission, and another approval before filing work with Bench’s Forge.

The rules were intended to improve quality. They also made routine work harder to move forward.

We changed the local default: describe the issue, state the expected result, send it through Forge, and own the follow-through to a pull request and a verified outcome. A pull request is the proposed code change that can be reviewed before it is merged.

Specialists remain available when they improve the work. Detailed specifications remain useful when the problem needs them. They are no longer prerequisites for every actionable issue.

We kept the distinction between progress and completion. Filed does not mean fixed. A pull request does not mean live. And a “live” status does not prove that the original problem is gone.

This is a change to our local operating instructions, not a claim that every part of Bench’s delivery system has changed, and not a claim that we have already measured a performance improvement. The test is whether work reaches a verified result with fewer unnecessary interruptions and without more harmful surprises.

A practical reset for your team

Choose one recurring workflow that keeps getting stuck. Follow a real request from start to finish and look at every place it waits.

For each stop, ask: what decision is being made, what consequence does it control, and does the reviewer have information the AI lacks? Keep the stops that protect meaningful decisions. Combine duplicate reviews. Remove approvals for work that is already authorized.

Then write one plain-language agreement:

You may do these actions within these limits. Show this evidence when you finish. Stop and involve this person when these conditions occur.

Review what happens next. Track avoidable approval loops alongside rework, unauthorized actions, and failures that reach customers. Speed alone cannot tell you whether the balance is right.

Start with one workflow. Make the authority clear. Keep the evidence. Give the work a path to done.

New posts

We'll email you when a new post lands.

Email signup opens soon. Until then, new posts show up right here.

More from the blog

Blog

Back to all posts

How Bench works. Notes on the models, the software, and the jobs we run.

Next step

See Bench on your pipeline.

You approve everything. Your data is yours — even if you leave. $249/mo, unlimited users.