CONSTRUCTION AI PILOT

A construction AI pilot starts
with one workflow.

An agent has to start somewhere, and the choice is usually made on whichever process annoys people most that month. This page sets out the four tests that actually decide it, who on your side ends up holding the output, and what the first run has to produce before anything it drafts is allowed into the workflow.

This is the choosing step. Data boundaries, permissions, logging and the approval record - the page for whoever has to sign the deployment off - are set out on How we deploy AI.

  • Workflows to choose from: 8, each with an agent already built
  • First run: Drafts compared, nothing used
  • Release: A named person on your side

How we deploy AI

Four tests decide it.
Not which process annoys you most.

Annoyance is a good way to find candidates and a poor way to rank them. A workflow that passes all four is one an agent can be measured on within a single cycle; a workflow that fails the last two is usually not the wrong workflow, only the wrong order.

Nothing on this page needs a model chosen, a budget set or a vendor picked. It needs a workflow, a stack of decided cases and a name.

  • It comes around often enough to measure: Weekly beats quarterly. Frequency is what lets you compare the agent against your own people inside one cycle instead of waiting a year for a verdict.
  • The evidence is already in something you hold: A drawing, a specification, an order record, an approval letter. If the answer lives in a conversation nobody wrote down, the agent has nothing to read and the pilot measures your memory instead.
  • You have already decided a stack of cases: The first run is a comparison, so it needs a known answer to compare with. Recent decided work is the asset that makes the trial honest.
  • One person can be named to release the output: Not a committee and not a job title: a person, with the time to read drafts during the trial. That person's availability is the real constraint on which workflow can go first.

Every workflow has
a name against it.

The workflow and the agent below are read from the same list the agent finder renders, so the two pages cannot drift apart. What this page adds is the last column: before a trial starts the output already has an owner, and that is the commitment worth testing you can actually make.

None of these people are replaced by the agent during a trial, and none of them are asked to do extra work they were not already doing: the release step is the one they perform today.

WorkflowBest-fit agentWho releases the output
Quotes take daysQuoting AgentYour estimator, who releases every quoteAgent →
Schedule changes rippleScheduling AgentYour scheduler, who accepts or rejects each proposalAgent →
Reports assembled by handReporting AgentThe person who signs the cycle's report todayAgent →
Documents re-typed into systemsDocument IntelligenceWhoever owns those documents, against answers they already knowAgent →
Approval conditions slipCompliance Review AgentYour compliance lead, who owns the register either wayAgent →
Long-lead items surprise the siteProcurement AgentYour procurement lead, who decides which flags are realAgent →
Group numbers arrive lateFinance & Cashflow AgentYour finance lead, against the close they run todayAgent →
The board pack eats the weekExecutive Briefing AgentThe owner, who takes the pack to the boardAgent →

Match a workflow to its agent · The agents we already run

Not every workflow
starts the same way.

How the first run is arranged is decided by what already exists underneath, not by preference. The agent is the thin part: what makes a draft worth reading is the system holding your structure, your rules and your records. So the workflow you pick also picks how much configuration happens before the agent does anything at all.

How the accounting and site tools you intend to keep connect to that system, and how records move across in verified stages, is set out once on Integrations & Data rather than repeated here.

  • Shadow: The agent drafts beside a process that carries on exactly as it does now. Its output goes into the comparison and nowhere else. (Suits: quoting, programme changes, reporting cycles, approval conditions, orders and board packs; You supply: work you have already decided, plus the rules you applied to it; It ends when: the comparison holds up, or the log explains why it does not): The default shape, and the only one most people mean by a pilot.
  • Parsing trial: Reading documents has no decision to sit beside. You hand over a document set and read the structured output and its flags before anything is wired to a system. (Suits: getting drawings, specifications and contracts out of PDFs and into fields; You supply: documents whose answers you already know; It ends when: you have read the misses, not only the hits): What matters here is what it got wrong, and whether it admitted it.
  • Platform first: Where the numbers do not yet sit in one place, there is nothing for an agent to draft from. The platform holds a consolidated position first, and the agent shadows a cycle after that. (Suits: group consolidation, and anything reported across several entities; You supply: the entity structure, the books each company keeps, and the rules you consolidate by; It ends when: the platform's position agrees with the close you run by hand): A system engagement with an agent at the end of it, and better said out loud early.

Integrations & data

Grade the decisions first.
Then rank the workflows.

A workflow is a bundle of decisions. Write each one down and ask what evidence would prove it, and the bundle sorts itself into three kinds. The grading came out of our WSBE26 compliance research, where it was used to scope a drawing-review method.

We use it to order candidates rather than to judge one on its own: a workflow that sorts mostly into A has a finite first scope, while one that sorts mostly into C is asking the agent for a call no document can settle. Grade B is the interesting middle: nothing gets quietly dropped, but the human input each of those checks still waits on has to be written down, and that is what gives a second stage its order.

GradeWhat it meansHow the agent handles it
ADeterministic geometry the drawing already containsAutomated, with evidence attached to every number
BProvable once one named human input is suppliedKept in scope with the missing input named - each becomes Class A the day that input can be automated
CContext or judgement a drawing cannot settleRe-specified as an advisory flag with an explicit threshold, never as a verdict

The research behind it

What the first run
has to produce.

Running an agent beside your process only means something if both sides wrote down beforehand what would count as a result. Four things are agreed before it drafts anything, and all four belong to you rather than to us.

None of this puts the output to work. Drafts enter the workflow only at the next stage, behind the confirmation step, and the staged rollout that governs it is on How we deploy AI.

  1. The comparison set: Which decided cases the agent will draft against, chosen by you and fixed before it sees them.
  2. The measure: What is being counted: where the draft and the decision agree, how long each path took, and how often the agent declined to answer.
  3. The disagreement log: Every difference between draft and decision, sorted into missing data, a rule we got wrong, and genuine judgement.
  4. The stop condition: The result that would end the trial instead of extending it, written while it is still easy to say.

When we would say
start somewhere else.

A trial that cannot produce a readable answer is worse than not running one, because it makes the next one harder to get approved. Four cases where we would tell you so before anything is scoped.

  • Nobody can be named to release the output: Without an owner the trial produces drafts that nobody is accountable for reading, and the comparison never happens. Find the person first; the workflow is usually fine.
  • There is nothing decided to compare against: A workflow with no record of what was decided, or why, gives the first run nothing to be right or wrong about. Recording decisions for a few cycles is the real first project.
  • The system underneath does not exist yet: Group consolidation is the clearest case: until a platform holds the position, no draft is worth checking. That is a configuration or build engagement, and the agent follows it.
  • The decisions are judgement, not evidence: When the grading comes back mostly C, an agent can still gather, flag and prompt, but the part you wanted taken off someone's desk is the part that stays there.

The group finance platform · How we develop construction AI

Questions from the people
who have to choose.

Which construction workflow should an AI agent start with?

Start where the work repeats often enough to measure inside one cycle, the evidence is already in a document you hold, you have a stack of cases you have already decided, and one person can be named to release the output. Quoting, programme changes, reporting cycles, approval conditions, purchase orders and board packs usually pass that test. Which of them goes first comes down to which named person has time to read drafts during the trial.

Should a first deployment cover more than one workflow?

One. Two at once doubles the number of people reviewing drafts during the trial, and if the result is disappointing nobody can tell which half caused it. The second workflow is easier to argue for once the first has an answer either way.

What do we have to supply before the first run?

The records the workflow already produces, the rules your people apply today, and a set of cases you have already decided. What each workflow needs specifically is listed against its agent, and nothing on the list is new work: it is the material the workflow generates anyway.

What if there is no system under the workflow yet?

Then the system is the first piece of work and the agent arrives at the end of it. Group consolidation is the usual example: the platform has to hold a consolidated position before a finance lead has anything to check a draft against. We say so at the start rather than calling it a pilot.

How do we know whether the first run worked?

Because the measure was agreed before it started: the comparison set, what is counted, the disagreement log and the stop condition. A result that does not clear the bar is still a result, and reading it honestly is worth more than extending the trial until it starts to look better.

Who on our side has to be involved?

The person who releases the output today, and whoever owns the data boundary, usually IT, operations or the director who signs vendor arrangements. Everyone else keeps working exactly as they do now, which is what makes the comparison possible.

Trust & data

Bring us the workflow
that keeps breaking.

Whether the problem is quoting, scheduling, reporting, procurement, document review, group finance or site coordination, Cyberate starts with the way your operation actually works.

Start with one workflow. If the system proves value, go deeper.

Talk to us · See live deployments