This document is for the developer or vendor who will build the models behind the BuildOptix platform. The platform UI is complete and every screen that consumes a model already exists, so the interfaces below are not proposals — they are the contracts the product is already written against.
Ten models are described. They are not equal: three of them gate everything else, and one of those is a prerequisite for all the rest. The sequence at the end of this document matters more than the model list.
The platform is a multi-tenant building management and IoT system covering 25 sites across 9 tenants, with roughly 4,000 monitored devices on BACnet/IP, Modbus TCP, MQTT, KNX and LoRaWAN. A canonical registry holds every tenant, site, building, system and point, and every module derives from it. The following screens are built and waiting on model output.
| Screen | Currently shows | Needs model |
|---|---|---|
| Anomalies & FDD | Rule-derived faults | M1 |
| Savings & ROI ledger | Baseline inferred from site health | M2 |
| Sensor Trust | Heuristic scoring | M3 |
| Tariff & Demand | Static tariff arithmetic | M4, M9 |
| AI action ledger | Sample actions, approve/revert flow live | M5 |
| GOC alarm grouping | Manual grouping | M6 |
| Alarm Rationalisation | Derived from alarm counts | M7 |
| Field App work orders | Generated from alarm state | M8 |
| AI Copilot | Scripted answers with provenance | M10 |
Each block states where the output surfaces, what it consumes, the response shape the UI is written against, and the condition under which we will accept it. The acceptance criteria are the important part — a model that cannot meet them cannot be put in front of a client.
Feeds Anomalies, FDD, GOC alarm wall, automatic work-order creation.
Consumes 15-minute point history per device (temperature, pressure, flow, kW, valve position, run status), equipment metadata including design conditions and capacity, weather.
{ siteId, equipmentId, faultClass, confidence, firstSeen, evidencePoints[], estimatedCostImpact }
Accepted when Precision ≥ 0.85 against an engineer-labelled historical fault set, and every output cites the specific points that triggered it. An uncited fault cannot be displayed — the UI has no place to put it.
Feeds Savings & ROI ledger, Client Portal savings claim, monthly client reports. This is the model that gets invoiced against, so it carries the most commercial risk of the ten.
Consumes At least 12 months of pre-intervention interval consumption, weather (degree days), an occupancy or production driver, and the intervention log from the action ledger.
{ siteId, period, baselineKwh, actualKwh, avoidedKwh, confidenceInterval, cvrmse, nmbe, attribution[] }
Accepted when IPMVP Option C is satisfied: CV(RMSE) ≤ 25% on monthly data and NMBE within ±0.5%. Outside those bounds the platform reports the saving as estimated rather than verified, and the model must return the statistics so the platform can make that call itself.
Feeds Sensor Trust, and through it the baseline quality gate and the AI autonomy gate. Build this first: every other model trains and infers on data this one grades.
Consumes Raw point streams, reference meter readings where a site has them, device metadata and expected reporting interval.
{ pointId, status: ok | frozen | drifting | gap | dead, driftPct, lastGoodAt, confidence }
Accepted when Injected faults are detected within one hour: a frozen value, a 3% drift, and a 10% interval gap. Drift must be returned as null rather than zero where a site has no reference to compare against — the platform distinguishes unmeasured from good.
Feeds Tariff & Demand. Indian HT tariffs set the demand charge from the single highest 15-minute average in the month, at 1.5× above the contracted ceiling, so a missed peak prediction has a direct rupee cost.
Consumes 15-minute kW history, weather forecast, calendar and occupancy, equipment start schedule.
{ siteId, horizon, forecastKw[], peakRiskWindow, breachProbability, confidence }
Accepted when MAPE ≤ 8% at a four-hour horizon, and recall ≥ 0.9 on demand-ceiling breaches. Recall matters more than precision here: a false alarm costs an operator two minutes, a missed peak costs a month of penalty.
Feeds AI action ledger. Runs in Approve mode by default: the model drafts, a human commits. Per-site autonomy is set in AI Governance and the model must honour it rather than assume it.
Consumes Current plant state, M4 forecast, tariff structure, comfort and safety constraints, equipment performance curves.
{ siteId, actions[{ pointId, currentValue, proposedValue, holdUntil, revertTo }], expectedKwh, expectedRupees, rationale, confidence }
Accepted when No proposal ever falls outside the comfort or safety envelope, every action carries a revert value, and realised savings reconcile with M2 after the fact. A recommendation whose saving cannot later be measured is not a recommendation.
Feeds Global Operations Centre alarm wall.
Consumes Live alarm stream with timestamps, equipment topology from the published SLD drawings, historical co-occurrence.
{ groupId, rootAlarmId, childAlarmIds[], causeHypothesis, confidence }
Accepted when At least 70% of multi-alarm events collapse into one group, and no genuinely independent alarm is ever suppressed as a child. The second condition is absolute — a suppressed real alarm is a safety failure, not a quality metric.
Feeds Alarm Rationalisation, measured against EEMUA 191.
Consumes 90 days of alarm history per point, equipment state, shelving history.
{ pointId, class: chattering | standing | duplicate | genuine, activations24h, standingDays, suggestedFix }
Accepted when Agreement ≥ 0.8 with engineer classification on a labelled sample, and the suggested fix is specific enough to action — an on-delay value, a deadband width, a suppression condition, not "review this point".
Feeds Field App work orders, planned maintenance scheduling.
Consumes Runtime hours, vibration, filter differential pressure, motor current trends, service and replacement history.
{ equipmentId, failureMode, predictedWindow, confidence, recommendedAction, partsRequired[] }
Accepted when Lead time is at least 7 days — shorter than that and a technician cannot be scheduled or a part ordered, which makes the prediction worthless — with a false-positive rate ≤ 20%.
Feeds The recoverable-cost panel in Tariff & Demand: load shifting, staggered starts, power-factor correction, DG versus grid changeover.
Consumes DISCOM tariff schedule per site (ToD bands, demand contract, PF incentive), M4 forecast, deferrable load inventory, DG fuel cost.
{ siteId, levers[{ action, expectedRupees, confidence, constraint }], blendedRate }
Accepted when Modelled cost reconciles with the actual DISCOM bill within 5%. Until it does, the recommendations cannot be shown to a client, because the first thing they will do is compare.
Feeds The AI Copilot tab. Retrieval over the registry, point history, action ledger and site documents.
Consumes Natural-language question plus the retrieval corpus, scoped to the asking user's site permissions.
{ answer, citations[{ sourceType, sourceId, value, asOf }], confidence, refused }
Accepted when Every number in an answer traces to a tag, record or document, and the model refuses rather than estimates when it cannot cite. Tenant scoping is enforced server-side, not by prompt.
Assemble this before the developer starts. Most AI engagements lose their first month waiting on data access, and the items marked as gaps below are the ones that need a decision from us rather than a request to IT.
| Item | Detail | Status |
|---|---|---|
| Site registry export | Tenants, sites, buildings, systems, protocols, device counts | Ready |
| Point / tag manifest | Per site, from published SLD drawings — includes equipment topology | Ready |
| Historian export | 15-minute intervals, 24 months where available, per point | To confirm |
| Alarm history | 12 months with acknowledgement and clearance timestamps | To confirm |
| Service history | Work orders, parts replaced, equipment failures | To confirm |
| Weather | Historical and forecast per city — third-party feed, we procure | To procure |
| Tariff schedules | Per DISCOM: MSEDCL, BESCOM, DHBVN, TANGEDCO, APSPDCL, WBSEDCL | To compile |
| Labelled fault set | Engineer-labelled historical faults — needed to accept M1 at all | Gap |
| Reference meter readings | For drift detection — only some sites have a reference | Gap |
| Comfort & safety envelopes | Per system, per site — the hard limits M5 may never cross | Gap |
These are enforced by the platform and the model must be built to them. They are not preferences — the governance module already implements each one, and a model that ignores them will be rejected at the gate rather than at review.
The single most important decision is not which model to build first — it is to not start with a model at all. Phase 0 produces no AI, and skipping it is the most common way these programmes fail: models get trained on unvalidated point data, perform well in testing, and then disagree with the client's own meter in month one.
Each of these blocks work if left open. Answer them before the kick-off rather than during it.