BuildOptix Solutions · AI Model Development Brief Confidential · September 2026
Developer Brief · Version 1.0

AI Model Development

What we need built, what we will supply, and the order to build it in.

This document is for the developer or vendor who will build the models behind the BuildOptix platform. The platform UI is complete and every screen that consumes a model already exists, so the interfaces below are not proposals — they are the contracts the product is already written against.

Ten models are described. They are not equal: three of them gate everything else, and one of those is a prerequisite for all the rest. The sequence at the end of this document matters more than the model list.

1 · What already exists

The platform is a multi-tenant building management and IoT system covering 25 sites across 9 tenants, with roughly 4,000 monitored devices on BACnet/IP, Modbus TCP, MQTT, KNX and LoRaWAN. A canonical registry holds every tenant, site, building, system and point, and every module derives from it. The following screens are built and waiting on model output.

Screen Currently shows Needs model
Anomalies & FDDRule-derived faultsM1
Savings & ROI ledgerBaseline inferred from site healthM2
Sensor TrustHeuristic scoringM3
Tariff & DemandStatic tariff arithmeticM4, M9
AI action ledgerSample actions, approve/revert flow liveM5
GOC alarm groupingManual groupingM6
Alarm RationalisationDerived from alarm countsM7
Field App work ordersGenerated from alarm stateM8
AI CopilotScripted answers with provenanceM10

2 · The models

Each block states where the output surfaces, what it consumes, the response shape the UI is written against, and the condition under which we will accept it. The acceptance criteria are the important part — a model that cannot meet them cannot be put in front of a client.

P1 M1 Fault detection & diagnosis

Feeds  Anomalies, FDD, GOC alarm wall, automatic work-order creation.

Consumes  15-minute point history per device (temperature, pressure, flow, kW, valve position, run status), equipment metadata including design conditions and capacity, weather.

{ siteId, equipmentId, faultClass, confidence, firstSeen, evidencePoints[], estimatedCostImpact }

Accepted when  Precision ≥ 0.85 against an engineer-labelled historical fault set, and every output cites the specific points that triggered it. An uncited fault cannot be displayed — the UI has no place to put it.

P1 M2 Energy baseline & measurement

Feeds  Savings & ROI ledger, Client Portal savings claim, monthly client reports. This is the model that gets invoiced against, so it carries the most commercial risk of the ten.

Consumes  At least 12 months of pre-intervention interval consumption, weather (degree days), an occupancy or production driver, and the intervention log from the action ledger.

{ siteId, period, baselineKwh, actualKwh, avoidedKwh, confidenceInterval, cvrmse, nmbe, attribution[] }

Accepted when  IPMVP Option C is satisfied: CV(RMSE) ≤ 25% on monthly data and NMBE within ±0.5%. Outside those bounds the platform reports the saving as estimated rather than verified, and the model must return the statistics so the platform can make that call itself.

P1 M3 Sensor validation

Feeds  Sensor Trust, and through it the baseline quality gate and the AI autonomy gate. Build this first: every other model trains and infers on data this one grades.

Consumes  Raw point streams, reference meter readings where a site has them, device metadata and expected reporting interval.

{ pointId, status: ok | frozen | drifting | gap | dead, driftPct, lastGoodAt, confidence }

Accepted when  Injected faults are detected within one hour: a frozen value, a 3% drift, and a 10% interval gap. Drift must be returned as null rather than zero where a site has no reference to compare against — the platform distinguishes unmeasured from good.

P2 M4 Load & demand forecast

Feeds  Tariff & Demand. Indian HT tariffs set the demand charge from the single highest 15-minute average in the month, at 1.5× above the contracted ceiling, so a missed peak prediction has a direct rupee cost.

Consumes  15-minute kW history, weather forecast, calendar and occupancy, equipment start schedule.

{ siteId, horizon, forecastKw[], peakRiskWindow, breachProbability, confidence }

Accepted when  MAPE ≤ 8% at a four-hour horizon, and recall ≥ 0.9 on demand-ceiling breaches. Recall matters more than precision here: a false alarm costs an operator two minutes, a missed peak costs a month of penalty.

P2 M5 Setpoint & sequencing optimiser

Feeds  AI action ledger. Runs in Approve mode by default: the model drafts, a human commits. Per-site autonomy is set in AI Governance and the model must honour it rather than assume it.

Consumes  Current plant state, M4 forecast, tariff structure, comfort and safety constraints, equipment performance curves.

{ siteId, actions[{ pointId, currentValue, proposedValue, holdUntil, revertTo }], expectedKwh, expectedRupees, rationale, confidence }

Accepted when  No proposal ever falls outside the comfort or safety envelope, every action carries a revert value, and realised savings reconcile with M2 after the fact. A recommendation whose saving cannot later be measured is not a recommendation.

P2 M6 Alarm root-cause grouping

Feeds  Global Operations Centre alarm wall.

Consumes  Live alarm stream with timestamps, equipment topology from the published SLD drawings, historical co-occurrence.

{ groupId, rootAlarmId, childAlarmIds[], causeHypothesis, confidence }

Accepted when  At least 70% of multi-alarm events collapse into one group, and no genuinely independent alarm is ever suppressed as a child. The second condition is absolute — a suppressed real alarm is a safety failure, not a quality metric.

P3 M7 Alarm rationalisation classifier

Feeds  Alarm Rationalisation, measured against EEMUA 191.

Consumes  90 days of alarm history per point, equipment state, shelving history.

{ pointId, class: chattering | standing | duplicate | genuine, activations24h, standingDays, suggestedFix }

Accepted when  Agreement ≥ 0.8 with engineer classification on a labelled sample, and the suggested fix is specific enough to action — an on-delay value, a deadband width, a suppression condition, not "review this point".

P3 M8 Predictive maintenance

Feeds  Field App work orders, planned maintenance scheduling.

Consumes  Runtime hours, vibration, filter differential pressure, motor current trends, service and replacement history.

{ equipmentId, failureMode, predictedWindow, confidence, recommendedAction, partsRequired[] }

Accepted when  Lead time is at least 7 days — shorter than that and a technician cannot be scheduled or a part ordered, which makes the prediction worthless — with a false-positive rate ≤ 20%.

P3 M9 Tariff optimiser

Feeds  The recoverable-cost panel in Tariff & Demand: load shifting, staggered starts, power-factor correction, DG versus grid changeover.

Consumes  DISCOM tariff schedule per site (ToD bands, demand contract, PF incentive), M4 forecast, deferrable load inventory, DG fuel cost.

{ siteId, levers[{ action, expectedRupees, confidence, constraint }], blendedRate }

Accepted when  Modelled cost reconciles with the actual DISCOM bill within 5%. Until it does, the recommendations cannot be shown to a client, because the first thing they will do is compare.

P4 M10 Operations copilot

Feeds  The AI Copilot tab. Retrieval over the registry, point history, action ledger and site documents.

Consumes  Natural-language question plus the retrieval corpus, scoped to the asking user's site permissions.

{ answer, citations[{ sourceType, sourceId, value, asOf }], confidence, refused }

Accepted when  Every number in an answer traces to a tag, record or document, and the model refuses rather than estimates when it cannot cite. Tenant scoping is enforced server-side, not by prompt.

3 · What we supply

Assemble this before the developer starts. Most AI engagements lose their first month waiting on data access, and the items marked as gaps below are the ones that need a decision from us rather than a request to IT.

Item Detail Status
Site registry exportTenants, sites, buildings, systems, protocols, device countsReady
Point / tag manifestPer site, from published SLD drawings — includes equipment topologyReady
Historian export15-minute intervals, 24 months where available, per pointTo confirm
Alarm history12 months with acknowledgement and clearance timestampsTo confirm
Service historyWork orders, parts replaced, equipment failuresTo confirm
WeatherHistorical and forecast per city — third-party feed, we procureTo procure
Tariff schedulesPer DISCOM: MSEDCL, BESCOM, DHBVN, TANGEDCO, APSPDCL, WBSEDCLTo compile
Labelled fault setEngineer-labelled historical faults — needed to accept M1 at allGap
Reference meter readingsFor drift detection — only some sites have a referenceGap
Comfort & safety envelopesPer system, per site — the hard limits M5 may never crossGap

4 · Non-negotiable constraints

These are enforced by the platform and the model must be built to them. They are not preferences — the governance module already implements each one, and a model that ignores them will be rejected at the gate rather than at review.

  1. Approve mode is the default. Models draft, humans commit. Full autonomy is granted per site in AI Governance and the model reads that setting rather than assuming it.
  2. Every output carries provenance. A number without a traceable source cannot be displayed anywhere in the platform. This applies equally to a fault, a saving and a copilot answer.
  3. Every action carries a revert path. The action ledger offers one-click revert on all committed actions, so the model must return the prior value alongside the proposed one.
  4. Data quality gates action. Below a trust score of 55 no model may act on a site and no saving is claimed. Below 75 savings are reported as estimated. The platform enforces this; the model should not attempt to work around a low score.
  5. Model version and drift are tracked. Every inference records the model version that produced it, and the governance page shows drift against the validation baseline. Retraining without a version bump is not permitted.
  6. Tenant isolation is server-side. A model serving one tenant must be incapable of reading another's data, enforced at the data layer rather than by prompt or filter.

5 · How to move forward

The single most important decision is not which model to build first — it is to not start with a model at all. Phase 0 produces no AI, and skipping it is the most common way these programmes fail: models get trained on unvalidated point data, perform well in testing, and then disagree with the client's own meter in month one.

Phase 0 · Weeks 1–4 · Data foundation
Historian export and access, tag normalisation across the five protocols, weather feed procured, tariff schedules compiled. Deliverable: M3 sensor validation running on live data. No other model starts until this one is grading the inputs.
Phase 1 · Weeks 5–12 · The two claims
M2 baseline and M1 fault detection. These make the commercial claim and the operational claim respectively — they are what a client is actually buying. Runs in parallel with engineers labelling the historical fault set.
Phase 2 · Weeks 13–20 · Acting on it
M4 forecast then M5 optimiser, in Approve mode on two pilot sites only. Measure realised savings with M2 before widening. This is the phase where the platform stops reporting and starts doing, so it earns the longest pilot.
Phase 3 · Week 21 onward · Operator quality of life
M6 grouping, M7 rationalisation, M8 predictive maintenance, M9 tariff, M10 copilot. Valuable but none of them gate revenue, and each is easier once the first five are producing labelled operational history.

6 · Decisions we owe them

Each of these blocks work if left open. Answer them before the kick-off rather than during it.

Recommended first step
Commission Phase 0 as a separate, fixed-scope engagement before signing for model development. It is four weeks, it is the cheapest phase, and its output — validated, normalised, graded data — is worth having whether or not the same developer builds the models afterwards. It also tells you, for a small sum, whether your data can support the claims the platform makes.