Article
Robot fleet monitoring and heartbeats
Fleet monitoring means each deployed site continuously reports whether it is working, rather than waiting for somebody to complain. Hybot projects send heartbeats carrying internet quality, frontend, backend and WebSocket status and vendor API reachability, charted over the last hour so a degrading site is visible before it fails outright.
Written by Hybot operations lead, Deployment and operations lead, Hyrcan-Tech · · 5 min read
The point of monitoring is the phone call you do not get
A venue that has to ring you to report a problem has already had the problem. Monitoring exists to invert that: the site tells you it is degrading, and you act before the restaurant's Friday service becomes the incident report.
What a Hybot project reports
Every project heartbeats to the Central Platform with:
- Internet quality, measured as client-to-project latency.
- Frontend, backend and WebSocket status — three separate services, because they fail separately and a site with a healthy backend and a dead WebSocket behaves very strangely indeed.
- Vendor API reachability — whether the robot manufacturer's API is answering, which is a dependency you do not control but must be able to see.
Each is charted across the last hour with weak, medium and strong bands, with a configurable auto-refresh.
Why the hour matters more than the moment
Intermittent problems are the normal kind, and a single current-status indicator is blind to them. A site that flickers between strong and weak every few minutes reads as healthy at any instant you happen to look.
That flicker is also the most common explanation for the complaint that arrives as "the robots are being weird". The hour-long chart makes the pattern obvious; a green dot does not.
Redundancy you can see
Server nodes report primary and secondary status with Up/Down and Active/Standby. Redundancy that silently absorbs a failure is doing its job, but it is also hiding the fact that you have used up your margin. Seeing that a site is running on its secondary is early warning, not an alarm.
Ranking sites, not just listing them
Per-project risk analytics rank which sites need attention first. With one deployment a list is fine. With thirty, a list is a wall of green in which the one degrading site is indistinguishable, and ordering by risk is the difference between monitoring and merely displaying.
Monitoring and commands belong together
Seeing that something is wrong is only half of it. Hybot's remote commands are queued centrally, collected on the robot's next heartbeat and acknowledged back, so an operator can tell whether a corrective action was actually applied rather than merely sent. The command channel uses the same heartbeat that carries the health data.
Where to go next
Cluster hub: Intelligent robotics.
Frequently asked questions
What does a heartbeat actually contain?
- Per project: internet quality measured as client-to-project latency, the status of the frontend, backend and WebSocket services, and whether the robot vendor's API is reachable. Each is charted across the last hour with weak, medium and strong bands rather than shown as a single instantaneous value.
Why chart an hour instead of showing current status?
- Because intermittent failure is the common case and a snapshot hides it. A site flickering between strong and weak every few minutes looks fine at any single moment, and that pattern is exactly the one that produces unexplained robot behaviour on the floor.
What about server redundancy?
- Server nodes report as primary and secondary with Up or Down and Active or Standby status. Knowing that a site is running on its secondary node is useful information well before anything visibly breaks for the venue's staff.
Where this fits
This page is part of Intelligent robotics: coordinating a fleet. If you are working through the topic in order, these are the neighbouring pages.
Multi-robot coordination, explained
Two robots are more than twice the problem of one. The failure modes that appear at the second unit, and what a coordination layer has to do about them.
How robot task assignment algorithms work
Nearest-robot, first-come, and weighted scoring compared — why the dispatch rule you choose decides whether a fleet balances its work or exhausts one unit.
Staged update rollout for robot fleets
Why pushing one release to every site at once turns a small regression into a coordinated outage, and how per-project rollout changes the risk profile of shipping.
Take it further
If a question here applies to a venue you actually run, the specifics matter more than the general case.