THE SITUATION
A multi-building pharmaceutical campus runs GMP and other critical environmental spaces. Temperature, humidity, and room pressurization must hold continuously — an excursion is not discomfort, it is a deviation, an investigation, and potentially lost product. The building automation system was commissioned years ago and has drifted ever since. Operations runs reactively: they learn a system has failed when an alarm fires inside a qualified space — the worst possible place to find out.
Documentation, investigation, and risk to product.
The first signal arrives inside the qualified space.
Waiting for the break-down is the most expensive plan.
THE CHALLENGE
An ambiguous prose sequence, reinterpreted by a programmer, diverging over time.
Frozen controllers, overrides, and drift, with no fleet-level observability.
BACnet interoperability is partial. MS/TP and broadcast domains fail physically — and silently.
Out-of-spec operation, leaks, stuck dampers and valves, degrading toward failure.
The ground truth was never formalized
The designer’s intent lives in the OPR, the Basis of Design, and the control sequences — pages of prose. The controls contractor reinterprets that prose into controller code. Every RFI, revision, and shortcut widens the gap. Setpoints get hard-coded, sequences get simplified, and no one keeps a machine-checkable record of what the system was supposed to do. Over years, the code and the intent quietly diverge — and nothing flags it.
Bottom line: conformance drifts unchecked
The network can’t be trusted
BACnet promises interoperability; in practice it is partial. MS/TP trunks and broadcast domains fail physically and silently — a Who-Is / I-Am storm, a duplicate device instance, a priority-array conflict. The front end still looks connected while critical assets drop off the network. The result is poor network health that no one is watching until a controller stops responding.
Bottom line: poor network health
FAILURE DOMAIN 03
The endpoints can’t be trusted
The controllers are the endpoints the entire system trusts — yet nothing watches the watchers. Controllers go offline, slow down under load, accumulate manual overrides, and drift out of calibration. Without fleet-level health telemetry, a frozen controller can hold a stale value for weeks while the space slowly leaves spec and no alarm ever fires.
Bottom line: no controller-health telemetry
The equipment can’t be trusted
Even with perfect code and a healthy network, the equipment itself degrades. Chillers and heat pumps run outside manufacturer parameters; valves leak; dampers stick; sensors and actuators break. Reactive O&M waits for the failure — and in a critical space, that failure is an excursion first and a dead asset soon after.
Bottom line: reactive O&M ends in asset failure
THE SOLUTION
We formalize the sequence of operations into custom, machine-checkable AFDD rules — the majority written for this facility — and continuously verify the running code against the design intent.
Network- and device-health monitoring surfaces silent MS/TP dropouts, broadcast storms, and priority conflicts before a critical asset falls off the map.
Fleet-level controller-health telemetry flags offline, slow, overloaded, overridden, and drifting controllers before their stale values reach a qualified space.
Physics-based AFDD — refrigeration-cycle residuals layered with AI — catches out-of-spec operation and degradation on chillers, heat pumps, and AHUs while there is still time to act.
A BAS logic/data fault reporting physically impossible cooling capacity — silently corrupting control decisions. Caught by residual analysis, not by any built-in alarm.
Early refrigerant-side degradation on an air-source heat pump, visible in the thermodynamics long before it would become a mechanical failure.
Per-circuit R-454B saturation diagnostics surfaced heat-exchanger approach drift — an early indicator of fouling or a charge issue.
Continuous, monitoring-based commissioning reduces reliance on outside commissioning authorities and MEP re-design engagements.
Continuous optimization surfaces ECMs and keeps equipment running at its design efficiency instead of drifting away from it.
Fault isolation and root-cause analysis cut the diagnostic labor otherwise spent chasing alarms, overrides, and drift.
Note: facility-specific savings quantification is in progress — these are the value categories a MelRok engagement is measured against.