AI & AGENTS · HOW WE DEPLOY AI
AI you can sign off on.
Data boundaries, permissions, staged rollout and a clear human-machine division of labour: the page for the people who have to approve this.
We do not start with a model. We start with an accountable workflow: what the agent may read, what it may produce, who confirms it, what gets logged, and where exceptions go.
PRINCIPLES
Four rules we deploy by.
| Action | Who holds it |
|---|---|
| Draft a quote, schedule or report | The agent |
| Propose a re-sequence or correction | The agent |
| Flag a risk, conflict or low-confidence item | The agent |
| Approve commercial, legal or site-critical decisions | Always a person |
- Agents run on systems, not vibes: Every agent sits on an operational system with real data and real rules, not a chatbot guessing from the internet.
- Your data stays yours: Explicit data boundaries per engagement: what the agent can read, where it runs, what never leaves.
- Humans keep the decisions: Agents draft, propose and flag. People confirm. The confirmation points are designed in, not bolted on.
- Prove it small, then scale: Deployment is staged: one workflow, measured, then wider. No big-bang AI programmes.
THE RESEARCH POSITION
The model fills a schema.
It never writes the conclusion.
That rule is not a preference - it is the finding our WSBE26 compliance research is built on. Where a decision carries regulatory weight, the language model is constrained to populating fixed fields from retrieved source text, and to saying so when a field is absent rather than inferring a plausible value.
Report the value, the threshold and the margin, then stop. The architecture matches the claim: where there is no automated verdict, there is no automated verdict to be wrong. Any system operating on regulation, law or safety should be able to say what it does not decide.
| Rule | What it prevents |
|---|---|
| Use only the retrieved text | The model reaching for another jurisdiction's rules from training data |
| Do not infer or invent | A plausible threshold that no clause actually states |
| Say "not specified" when a field is missing | A silent omission instead of a reviewable gap |
| Keep the source wording verbatim | A reviewer having to leave the report to verify an extraction |
| Return multiple controls as separate records | Two different thresholds merged into one averaged value |
| Emit structured output only | Unverifiable narrative commentary in a high-stakes decision |
The boundary comes first.
- Define the boundary: Which data the agent may access, and where processing happens, is agreed before any build.
- Set the permissions: Agent access mirrors your role-based permissions: it sees what the role it serves would see.
- Log everything: Agent actions are auditable: what it read, what it produced, who confirmed it.
The industry adopts slowly
for rational reasons.
The Productivity Commission attributes construction's slow digital uptake partly to structure: very small firms, project-by-project margins, and innovation risk that a single job cannot absorb. A staged deployment - shadow first, nothing used until the comparison holds up - is adoption built for exactly that risk profile.
Figures as published in the Commission's February 2025 research paper, analysing data mostly to 2023-24.
Staged deployment, human division of labour.
Adoption follows a sequence, not a switch. Nothing is used in the live workflow until the previous stage holds.
- Discovery & boundary: The workflow is mapped and the data boundary agreed before anything is built.
- Shadow: The agent drafts alongside the current process; outputs are compared, not used.
- Assisted: The agent's drafts enter the workflow behind a human confirmation step.
- Routine: Routine cases flow through with spot-check review; exceptions always route to people.
- Governance review: Logs, exceptions and outcomes are reviewed on a set cadence, and the boundary is re-confirmed.
What the audit trail holds.
Accountable AI is checkable AI. Every agent action leaves a record your reviewers, auditors and directors can follow.
From 10 December 2026, APP 1.7 of the Privacy Act asks covered organisations to state in their privacy policy what their programs decide and what they substantially contribute to. The "who confirmed it" record is what makes those sentences writable.
- What the agent read: The sources and data it accessed, inside the agreed boundary.
- What it produced: Every draft, proposal and flag, kept with its inputs.
- Who confirmed it: The named reviewer and the decision they made.
- What was corrected: Edits are recorded and feed back into the workflow.
- Where exceptions went: Routed to a person, owned, and closed out visibly.
Questions the approvers ask.
How do you keep client data safe?
Every engagement starts with an explicit data boundary: what the agent may read, where processing happens, and what never leaves. Agent access mirrors your role-based permissions, and all access is logged.
Does Cyberate train public models on our data?
No. Your data is used only inside the agreed boundary, for your workflow. It is not used to train public models.
Who approves AI outputs?
The named owner in your team. Commercial, legal and site-critical outputs always carry a human confirmation step before they take effect.
What is shadow mode?
The agent runs alongside your current process and drafts in parallel. Its outputs are compared against what your team actually did, and nothing it produces is used until you decide the comparison holds up.
How do we start safely?
Pick one repetitive workflow. We map it, agree the data boundary, run the agent in shadow, and only move to assisted operation when the results earn it.