Remote Monitoring Setup for a Maintenance Customer
Purpose
This standing instruction enrolls a maintenance customer's equipment in remote monitoring only after every alert it will produce has a written response class, a named owner and a stated response window. Monitoring that emits alerts nobody has agreed to act on is not a service, it is a liability with a notification sound.
The failure runs in one direction. Thresholds are left at platform defaults, defaults fire on a heat wave, alerts pile into an inbox, someone mutes the channel because it is mostly noise, and the one that mattered lands in a muted channel three weeks later. The customer meanwhile believes you are watching, because that is what monitoring means to someone outside this trade. That gap between what the platform emits and what the shop actually promised is where this procedure works.
Safety actions that gate this procedure
- A remote command starts equipment with nobody standing at it. No setpoint change, mode force or test is issued from the office without confirming first that nobody is working on that system, and the confirmation is logged with the command.
- Never generate a test alert by tripping a safety, forcing a fault or overriding a protective point. Driving a limit, a float or a pressure switch from a keypad reaches the same hazard as moving it by hand. Use the platform's own test function or a non-equipment path such as briefly removing network access.
- Where a step measures on running equipment, the electrical and combustion gates in the maintenance SOPs apply in full, including a personal CO monitor for any run that fires a burner.
Scope
Covers enrolling residential and light-commercial equipment in a manufacturer or platform monitoring portal under a maintenance agreement: consent, capability check, baseline capture, threshold setting, alert routing, loop testing and periodic review.
Does not cover the connected thermostat's own setup and customer account, or fault code investigation once an alert has fired, each owned by its own SOP. Does not cover writing the maintenance agreement itself, or pricing it.
Roles and handoffs
| Role | Owns | Hands off |
|---|---|---|
| Service manager | The agreement clause defining alert classes, response windows and hours | The written clause, before any enrollment happens |
| Setup tech | Steps 1 to 6, including the measured baseline | The monitoring record and the baseline sheet, filed to the equipment |
| Dispatcher | Receiving alerts in the queue and executing the response class | A logged action per alert, which is what the quarterly review counts |
Procedure
Confirm the agreement says what monitoring is, what each alert class earns, and in what hours, before anything is enrolled. Acceptance: a signed clause naming the response classes and their windows, plus the customer's written consent to access their equipment data and their route to revoke it. Wrong looks like enrolling because the platform offered it, which creates an unwritten promise of continuous watching. Stop rule: no clause and no enrollment; monitoring without a stated response is worse than none, because it transfers the customer's attention to you. Hazard: none here, paperwork at the office.
Confirm the equipment can actually report what the agreement promises. Acceptance: a written list of the alert types the platform emits for this exact equipment and control combination, from the platform's documentation rather than its marketing page. Wrong looks like promising runtime monitoring on a control that reports only faults. Stop rule: a promised alert the platform does not emit for this equipment is struck from the clause before enrollment, not explained away later. Hazard: none here, document work at the truck.
Capture a measured baseline on a real visit, at a recorded outdoor condition. Acceptance: daily runtime, supply and return temperatures with the drop or rise computed, external static, and the date and outdoor temperature they were taken at. Wrong looks like enabling thresholds against a platform default, which describes an average machine and not this one. Stop rule: no baseline and thresholds stay disabled; an alert with no baseline behind it cannot be tuned, only tolerated. Hazard: these are measurements on running equipment, so gauges connected before start-up, probes clear of the blower wheel, and a CO monitor on for any heat-side run.
Set every threshold from the baseline, and write the alert-to-action map before you enable a single alert. Acceptance: each enabled alert carrying a class - same-business-day contact, add to next visit, or informational - plus a named owner and the hours that owner is watching. Wrong looks like every available alert switched on because the toggles were there. Stop rule: an alert nobody owns is disabled rather than enabled, because an unowned alert trains the whole queue to be ignored. Hazard: none here, it is configuration in a portal with no command authority exercised.
Decide and document what your shop may command remotely, and what it may not. Acceptance: a written authority line stating whether the shop may change setpoints or modes from the portal, who may do it, and the confirmation required first. Wrong looks like a portal with full command authority and no rule, which eventually means a technician's hands in a blower compartment while an office change starts it. Stop rule: no written authority and the account is set to read-only. Hazard: this step defines who can start equipment from a distance, so the rule that no command is issued without confirming the system is clear goes in the authority line itself, not in a memory.
Test the whole loop end to end before you tell the customer it is live. Acceptance: one alert generated by a permitted method, received by the named owner inside the response window the agreement states, and the response executed and logged. Wrong looks like a test that confirms the alert was emitted and never checks that a human received and acted on it, which tests the platform and not the shop. Stop rule: a loop failing at any hop is fixed and re-tested before the customer is told monitoring is active. Hazard: the test is generated by the platform's test function or by removing network access, never by tripping a safety or forcing a fault.
Review at a stated interval and count what each alert class produced. Acceptance: a quarterly count of alerts by class, how many produced an action, and how many produced none, with a retune or a disable decision recorded for any class that mostly produces nothing. Wrong looks like a review that reads the log and changes nothing, after which the queue gets muted informally by the people who work it. Stop rule: a class with a high no-action rate is retuned or turned off at the review, not left enabled on the theory that it might catch something. Hazard: none here, it is a review at a desk with the log open.
The record this produces
A monitoring record filed to the equipment and referenced from the agreement:
- The agreement clause reference, the consent record, and the revocation route given
- The alert types this equipment actually emits, from the platform documentation
- The baseline: runtime, supply and return with the drop or rise, static, date, outdoor temperature
- Every enabled alert with its threshold, the baseline figure it came from, its class, its owner and the watching hours
- The remote command authority line, and who holds it
- The loop test: alert generated, method used, time received, owner, response logged
- Each quarterly review: alerts by class, actions taken, no-action count, retune or disable decision
The no-action count is the field that keeps the system honest. It is the only number that shows a threshold is producing noise, and it is the number a shop is least inclined to write down.
Worked pass: split system on a residential maintenance agreement, cooling season enrollment
Step 1: agreement clause signed. Same-business-day contact for fault and offline alerts, next-visit handling for filter and runtime alerts, informational for everything else. Consent recorded with the revocation route shown to the customer at the table.
Step 2: platform documentation for this control and equipment combination lists fault, offline, filter runtime and daily compressor runtime. Whole-house humidity is not emitted by this combination, so the humidity line is struck from the clause before enrollment rather than promised.
Step 3: baseline taken on a measured visit at an outdoor temperature of 88 F. Daily compressor runtime 5.2 hours. Supply 56 F against a return of 76 F, so 76 minus 56 leaves a 20 F drop. External static 0.42 in w.c. Date and outdoor temperature recorded beside every figure.
Step 4: runtime threshold set at 40 percent above baseline, so 5.2 hours times 1.4 gives 7.28 hours, entered as 7.3 hours, and gated to days whose outdoor average is within a stated band of the 88 F baseline day. Fault and offline alerts classed same-business-day and owned by the dispatch queue. Filter and runtime classed next-visit, owned by the same queue with a different tag. Watching hours written out.
Step 5: account set so the shop may read everything and may change setpoints only with the customer on the phone, and only by the dispatcher. Written into the authority line with the confirmation requirement inside it.
Step 6 FAILS. A test offline alert is generated at 4:50 pm by taking the device off the network for 20 minutes. The platform emits it, but the notification routes to the shop's general inbox rather than the dispatch queue, and it is opened the next morning. The agreement promises same-business-day contact for offline alerts, and nobody contacted the customer that day. Stop rule taken: the customer is not told monitoring is live, and the routing is corrected before anything else.
The failure is worth naming precisely, because the platform did its part. The alert existed, on time, with the right content. What did not exist was a path from that alert to a person who had agreed to act on it inside the promised window, and that is the half a platform cannot supply.
Re-test the following morning: alert generated at 9:10 am by the same method, lands in the dispatch queue at 9:14 am, four minutes later. The dispatcher logs it and reaches the customer at 9:40 am, inside the same-business-day window with a wide margin. Loop closed, and the customer is told monitoring is live only at that point.
First quarterly review: 11 alerts total. Two same-day class, both real - a condensate float and a genuine offline event after the customer changed routers. Six filter alerts, all handled at the next visit. Three runtime alerts, of which one produced a truck roll for a dirty condenser coil and two fired during a heat wave with nothing wrong. That is three truck rolls from 11 alerts and a no-action count of two, all of them in the runtime class. Decision recorded: the runtime threshold keeps its 7.3 hour value but its outdoor-temperature gate is narrowed, because both no-action alerts came from days well outside the baseline band the threshold was meant to be compared within.
Checking this pass: step 3's acceptance names runtime, supply and return with the computed drop, static, date and outdoor temperature, and the run prints all six. Step 4's threshold shows its derivation inline rather than stating 7.3 alone. The review counts add: two plus six plus three is 11, and the truck rolls are the two same-day alerts plus one runtime alert, which is three. Step 6's acceptance requires receipt by the named owner inside the stated window, and the failing run missed that criterion even though the alert itself was emitted correctly, which is why the stop rule fires on receipt rather than on emission.
References
- Platform or manufacturer monitoring documentation for the exact control and equipment combination, the authority on which alerts are emitted and what each one measures
- The shop's own maintenance agreement clause defining alert classes, response windows and watching hours, which every acceptance above is measured against
- See related: the smart thermostat setup and customer handover SOP, which owns account ownership and remote access; the communicating equipment fault code investigation SOP, which takes over once a fault alert fires; the annual HVAC maintenance SOP, which supplies the baseline visit