AI for Construction & AEC
Strategic · M18 · lesson 18 of 23 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Proof-of-Concept Design
📖
now learning

Proof-of-Concept Design

15 min

A national GC ran a field-vision pilot for a full year. The cameras went up, the captures uploaded, the dashboards lit, and at twelve months the question nobody had answered going in was still open: did it work? No one could say. There was no number it was supposed to move, no decision it was supposed to inform, no date it was supposed to die on if it did not deliver. The pilot had become a permanent line item that proved nothing, defended by the people who liked the dashboards and resented by the people paying the subscription, and the firm spent the next quarter arguing about a tool instead of deciding about it. As the strategy lead running vendor pilots, your job is not to run a demo that looks good; it is to design a proof of concept that returns a clear yes or no on a named question, on a named date, against criteria you wrote down before the first camera went up. A POC is a decision instrument, not a science fair, and a decision instrument that cannot produce a decision is just spend. By the end of this lesson you will be able to design a POC with defined scope, named success criteria and named sunset criteria, explicit data-handling and contract terms, and you will produce the named artifact: the statement of work for a 60-day POC of two competing field-vision platforms on the same job.

The POC Is a Decision Instrument, Not a Demo

The first discipline is to fix what a POC is for. A demo is designed to impress: the vendor controls the project, data, and narrative, and the question it answers is "can this tool look good under favorable conditions," which is always yes. A proof of concept is designed to decide: you control the project, question, and criteria, and it answers "does this tool move a metric we care about on a job that looks like our jobs, enough to justify the cost and the change." The most common failure in vendor pilots is running a demo while calling it a POC, then being surprised when twelve months of good-looking captures produce no decision.

The reframe that holds the lesson together is borrowed from the program's verification discipline: you define what success is measured against before you start, exactly as a licensed professional defines design intent before checking a deliverable against it. The cardinal rule is verify before you commit, and the gate is not a mood, it is a written standard set in advance. A POC is the same shape: the success criteria are the standard, and the pilot is the verification against it. Write the criteria after seeing the results and you have not run a POC, you have run a demo and rationalized it, because nothing the result could have been would have failed.

So treat the POC as a measuring instrument you calibrate before you take a reading. The controlling analogy for this lesson is the medical trial: a trial that decides whether a drug works states its primary endpoint, enrollment, duration, and stopping rules before the first patient is dosed, precisely so no one can move the goalposts when the data comes in. Your POC needs the same parts: the endpoint (success criteria), the enrollment (scope), the duration (60 days), the stopping rule (sunset criteria), and the protocol for what happens to the data and who owns it (data-handling and contract terms). A pilot missing any of these cannot return a clean decision, the only thing a POC exists to do.

Scope: Run Competing Tools on the Same Job

Scope is the enrollment of your trial, and the single most important scoping decision in a competitive POC is that the competing tools run on the same job. If you pilot OpenSpace on a hospital tower and Buildots on a warehouse fit-out, you have not compared the tools, you have compared the projects, and you will never untangle which differences came from the platform and which from the job. The same job is the control. Same building, trades, schedule, superintendent, capture cadence: only the platform varies, so any difference in the reading is attributable to the platform. This is the field equivalent of the A/B test the program teaches.

Defining scope tightly also keeps the POC honest about cost and effort. Scope answers what part of the project is in the pilot (which floors, phases, trades), how often capture happens (weekly walks, or every trade handoff), who captures (the super, a dedicated engineer, the vendor's field team), and what the tool is asked to do (progress tracking against schedule, percent-complete by trade, rework detection, safety observation). A POC with unbounded scope cannot finish in 60 days and cannot be costed, so you bound it deliberately: a representative slice large enough to exercise the tool against real conditions and small enough to run and judge inside the window.

The "representative slice" judgment is where your knowledge of your own portfolio earns its keep. The slice has to look like the jobs you would deploy the tool on if you bought it, because a POC that succeeds on an easy slice and a job you never build tells you nothing about your real work. Pick a phase with the messiness your projects actually have: congested MEP coordination, overlapping trades, a schedule with real pressure. If the tool can track progress and surface rework on that, the result transfers; pilot it on a clean, slow, single-trade phase and the success does not generalize. Scope is not just "how much," it is "how representative."

Success Criteria: The Named Yes

The success criteria are the primary endpoint of the trial, and they must be named, measurable, and written before the pilot runs. "The tool was useful" is not a success criterion; it is a feeling, and feelings are exactly what a demo produces. A success criterion is a specific metric, threshold, and measurement method, such that two reasonable people looking at the results would agree on whether it passed. The discipline is to write the number you would need to see to say yes in advance, when you have no stake in the answer yet, because the moment the results are in you will be tempted to lower the bar to justify the effort spent.

Good field-vision success criteria are operational, not technical. The vendor will offer technical metrics (capture coverage, image resolution, model accuracy) because those are easy to win, but you care about whether the tool changes how the job runs. Tie the criteria to a decision the tool should improve: does percent-complete-by-trade match a manual audit within an acceptable tolerance, so the PM can trust it for pay-app review; does it surface rework or missed installation early enough to act before the next trade covers it; does it cut the time the super spends walking and documenting by a defined amount. Each is a number you can set a threshold on and measure at the end.

Write the criteria as a small set, three to five, each with a metric, a threshold, and a method, and decide in advance how many must pass for an overall yes. This is where the program's verification habit pays off: you are defining the standard the deliverable will be checked against before the deliverable exists, so the check is a gate, not a negotiation. A POC with named, pre-registered success criteria can be wrong and can be argued, but it cannot be fudged, because the standard was fixed before anyone knew the result.

A proof of concept that cannot fail is not a proof of concept, it is a demo with a budget. Write the success criteria and the sunset criteria before the first capture, so the pilot returns a decision you cannot fudge after the fact.

Sunset Criteria: Knowing When to Kill It

The part almost every pilot skips, and the part that separates a real POC from a year of dashboards, is the sunset criteria: the conditions under which you stop the pilot, including stopping early. Success criteria tell you when to say yes. Sunset criteria tell you when to say no, and when to walk away before the 60 days are up because the answer is already clear. A medical trial has stopping rules for both efficacy and futility: you stop early if the drug obviously works, and stop early if it obviously will not, because continuing is a waste. Your POC needs the futility rule even more than the efficacy rule, because the gravitational pull of a running pilot is to continue.

Sunset criteria come in a few flavors and you should write all of them. The hard sunset: the date the POC ends regardless (60 days), at which point a decision is made and the tool is adopted, rejected, or sent to a defined extension with a new question. The early-kill sunset: conditions that end the pilot before the date, such as the vendor failing to integrate, capture compliance falling below a usable threshold so there is no data to judge, a security or data-handling problem surfacing, or a clear early read that the tool will not hit the success criteria. And the per-tool sunset in a competitive POC: if one platform is clearly losing at the midpoint, you sunset it and concentrate the field team on the contender that might win.

The reason sunset criteria matter so much is organizational, not analytical. A pilot with no kill switch becomes a zombie: it consumes subscription cost and field attention, accumulates defenders who do not want their work declared a failure, and crowds out the next experiment. Writing them in advance gives the strategy lead the authority to stop, because the stop was agreed before anyone was attached to the outcome. The hardest sentence in vendor management is "we are killing this," and the only way to say it cleanly is to have written, before you started, exactly the conditions under which you would. A POC without a sunset is the year-long pilot that proved nothing.

Data Handling: Who Holds the Captures

A field-vision POC pours your project's visual reality into a vendor's platform: progress imagery, site conditions, sometimes faces of workers, sometimes proprietary details of the building and the means and methods. Before the first capture, the POC has to settle the data-handling terms, because a pilot is still a live project and the data is real. This is the program's data gate applied to vendor selection: you do not let consequential data leave your control on an unwritten understanding, and a POC is not an exception because it is temporary. The questions are concrete. Who owns the captured data. Where is it stored and under whose security. Can the vendor use your imagery to train its models, and if so, is that disclosed and acceptable. What happens to the data when the POC ends, especially for the tool you do not pick.

The end-of-POC data question is the one most often missed and the one that bites. When you run two platforms and adopt one, the loser is sitting on 60 days of your project's imagery, and the SOW has to say what happens to it: certified deletion within a defined window, with no retained right to use it. The same goes for the winner if the pilot does not convert. Treat the data as something you are lending under terms, not giving away because the engagement is short, because "short" is the framing that gets data-handling waved through, and the imagery does not become less sensitive because the pilot was brief.

There are also live-project handling concerns the POC must respect regardless of which tool wins. If the capture includes workers, there may be notice or consent obligations and a privacy posture the firm has to hold. If the imagery shows the owner's facility or a secure site, the owner's confidentiality terms flow down into the pilot. The strategy lead does not get to ignore these because the deployment is a test. Settling the terms up front is what makes the POC a controlled instrument rather than an uncontrolled leak, and it belongs in the SOW in plain language.

Contract Terms: Pricing the Experiment

The contract terms turn the POC from a handshake into an instrument you can stop cleanly. A POC is a bounded experiment and the contract has to be bounded to match, so the dollars gate applies: you are committing the firm's money, and the terms must protect its ability to walk away with the decision and without a tail. The fixed POC fee should be exactly that, fixed and capped, with no auto-renewal into a full subscription if no one acts, because the silent rollover is how the year-long pilot gets funded. POC pricing should be separated from production pricing so the experiment's cost is not entangled with the cost of the thing you are deciding whether to buy.

The terms also have to encode the exits the sunset criteria define. If you can kill a tool at the midpoint, the contract has to let you stop paying for it at the midpoint, so the early-kill sunset is a contractual right, not just an analytical decision. The SOW should define what each party owes: what the vendor provides (the platform, field support, integration help, defined deliverables and reports), what you provide (site access, capture cadence, a point of contact), and the milestones and dates. It should name the production pricing the POC would convert to if it succeeds, so the success decision is not ambushed by a price you never agreed to, and it should be explicit that running the POC creates no obligation to buy.

One more term protects the integrity of the comparison itself. Because two vendors are competing on the same job, the SOW should hold them to the same conditions: same scope, access, cadence, reporting, and evaluation criteria, disclosed to both, so neither can claim the comparison was rigged and you cannot be talked into special conditions for the favorite. Fairness in a competitive POC is a contract term, not a courtesy, because the decision is only as defensible as the even-handedness of the test that produced it. Contract terms are how you make the experiment safe to run and safe to stop, the precondition for it being a decision instrument at all.

The Applied Problem: The SOW for a 60-Day Two-Platform Field-Vision POC

Here is the exercise. Write the statement of work for a 60-day proof of concept of two competing field-vision platforms, for example OpenSpace versus Buildots, running on the same job. The SOW is the artifact that turns the design discipline of this lesson into a document a vendor signs and a strategy lead can defend, and it has to contain every part of the instrument: scope, success criteria, sunset criteria, data-handling terms, and contract terms, written so the POC returns a clear yes or no on the named date.

Build the SOW in sections that map to the lesson. Open with the objective stated as a decision: this POC exists to decide whether to adopt a field-vision platform and, if so, which one, measured against named criteria over 60 days. Define the scope as the same-job control: the specific phase, floors, and trades both tools cover, the capture cadence, who captures, and the explicit statement that both platforms run on the identical slice. Write the success criteria as three to five named metrics, each with a threshold and a measurement method (percent-complete-by-trade matching a manual audit within a stated tolerance; rework surfaced before cover within a stated lead time; superintendent walking time reduced by a stated amount), and state how many must pass for a yes and how the tools are scored against each other.

Then write the parts that make it a decision instrument rather than a demo. Specify the sunset criteria: the hard 60-day end with a forced decision, the early-kill conditions (failed integration, capture compliance below a usable floor, a security or data finding, a clear early miss), and the per-tool midpoint sunset that drops the losing platform. Specify the data-handling terms: ownership of the captures, storage and security, whether project imagery may train models, worker and owner privacy obligations, and certified deletion at POC end for the tool not selected. Specify the contract terms: a fixed capped POC fee with no auto-renewal, separated from production pricing, the named production pricing if it converts, the same-conditions fairness clause binding both vendors, and no obligation to buy. The deliverable is one SOW any partner could read and see exactly how the firm will get a defensible yes or no in 60 days, and the lasting product is a reusable template for every future competitive pilot.

Key Takeaways

  • A POC is a decision instrument, not a demo. A demo is designed by the vendor to impress under favorable conditions; a POC is designed by you to return a clear yes or no on a named question, on a named date, against criteria written before the pilot runs.
  • Define success against a written standard before you start, exactly as the program's verification discipline defines design intent before checking a deliverable. Write the criteria after seeing the results and you have run a demo and rationalized it, because nothing could have failed.
  • Run the competing tools on the same job: same building, trades, schedule, super, and cadence, so only the platform varies. Piloting OpenSpace and Buildots on different projects compares the projects, not the tools. The slice must be representative of the work you would actually deploy on.
  • Success criteria are the named yes: three to five operational metrics, each with a threshold and a measurement method, tied to a decision the tool should improve (trustworthy percent-complete, early rework detection, reduced documentation time), with a pre-set rule for how many must pass.
  • Sunset criteria are the named no and the kill switch: a hard 60-day end with a forced decision, early-kill conditions (failed integration, unusable capture compliance, a security finding, a clear early miss), and a per-tool midpoint sunset to drop a losing platform. A pilot without a sunset becomes a zombie that proves nothing.
  • Settle data-handling terms before the first capture: ownership of the captures, storage and security, whether imagery trains models, worker and owner privacy, and certified deletion at POC end for the tool not selected. The data gate does not relax because the engagement is short.
  • Write contract terms that make the experiment safe to run and safe to stop: a fixed capped POC fee with no auto-renewal, separated from production pricing, exits that match the sunset criteria, named production pricing if it converts, a same-conditions fairness clause binding both vendors, and no obligation to buy.
  • The named artifact is the SOW for a 60-day POC of two competing field-vision platforms on the same job. It turns the design discipline into something a vendor signs and a strategy lead can defend, and becomes a reusable template for every future competitive pilot.