Insights
TrustOctenta Team · · 6 min read

Shadow Mode: Why We Refuse to Let AI Execute on Day One

Every Octenta employee starts by doing the whole job and executing none of it. Here is how the deployment model works, and what the audit trail gives you.

A humanoid figure observing silently behind glass in a darkened office — shadow mode

There is a specific failure mode in enterprise automation that we designed the entire platform around avoiding. It goes like this: a system is procured on the strength of a demo, configured over a few weeks against assumptions nobody tested, switched on across a live process, and then quietly switched off three months later after it made a category of mistake that nobody had thought to check for. The organisation concludes that the technology was not ready. Usually the technology was fine and the deployment was reckless.

Shadow Mode is our answer. Every Octenta AI Employee begins in it, without exception and without a paid upgrade to turn it on. In Shadow Mode the employee is connected to real systems and real data. It reads live records, performs its full workload, and reaches a decision on every item exactly as it would in production. Then it stops. It writes the decision down, marks it awaiting approval, and executes nothing.

Shadow is not a sandbox

This is the distinction people miss. A sandbox uses copied or synthetic data, so it tells you how the system behaves on data you already understand. Shadow Mode uses production data, so it tells you how the system behaves on the messy, exception-heavy reality you actually have — the supplier who invoices in a different currency than the PO, the employee record with two conflicting start dates, the shipment that was split across three consignments.

The only thing missing in Shadow Mode is the write. Nothing is posted, sent, paid, booked, approved or filed. The consequence is that the cost of a wrong decision during evaluation is exactly zero, which is what makes it safe to run the evaluation on the real thing.

What you are actually measuring

Running in shadow generates a decision ledger you can grade. Over a typical four-to-eight week shadow period, a finance or operations lead reviews that ledger and answers four questions:

  • Agreement rate — on what proportion of decisions did the AI Employee reach the same conclusion a competent human did?
  • Disagreement shape — when it differed, was it wrong, or was it right in a way the current process is not? Both happen, and the second one is often the more valuable finding.
  • Coverage — what proportion of the workload did it handle at all, versus escalate? A high agreement rate on 20% of volume is not a deployment.
  • Confidence calibration — when the employee said it was sure, was it? A model that is confidently wrong is far more dangerous than one that is uncertain and says so.

That last one is why confidence is displayed on every decision rather than kept internal. You should be able to see the distribution and decide your own threshold, rather than accept ours.

The audit trail

Every decision, in shadow and in live, is written to an append-only log. An entry records the timestamp, the employee, the source records it read, the decision it reached, the confidence attached to it, the rule or precedent it relied on, whether it executed or was held, and the identity of the human who approved, rejected or amended it.

Nothing in that log can be edited or deleted, including by an administrator. Amendments are appended as new entries that reference the original. This is not a compliance nicety. It is the mechanism that lets you answer the only question that ever really matters after an incident: what did the system do, when, on what basis, and who signed it off.

An automated decision you cannot reconstruct after the fact is not automation. It is an unrecorded liability.

Going live is granular, and reversible

Live is not a single switch across the whole role. Permissions are granted category by category. A finance team commonly promotes reconciliation of exact matches to live first, because it is high-volume, low-judgement and easy to verify. Invoice coding follows once agreement rates hold. Payment release, in most deployments, never goes live at all — and we regard that as a healthy outcome rather than an incomplete one.

Every category can be returned to shadow instantly, by one person, without a support ticket or a deployment window. If something in the business changes — a new entity, a restructured chart of accounts, a policy shift — the correct response is to put the affected category back into shadow, watch it for a fortnight, and promote it again. We would rather a customer used that lever often than hesitated to use it once.

Why we make this the default

It would be commercially easier to let a keen customer switch everything on in week one. It would also produce exactly the failure described at the top of this piece, and the cost of that failure is not just a churned account — it is an operations leader who now believes, with some justification, that this category of software cannot be trusted.

So the rule is fixed. Your AI Employee earns each permission with evidence you gathered yourself, on your own data, over a period you chose. Toggle to live only when you are at one hundred percent confidence — and keep the ability to toggle back.

Related reading

Schedule Your Workforce Audit

We reply within one business day with a specific plan and a specific number.

Schedule Workforce Audit