Article
Staged update rollout for robot fleets
Staged rollout means a software release reaches some deployed sites before the rest, so a regression is discovered at one location rather than everywhere simultaneously. Hybot applies updates selectively per project, which keeps the blast radius of any change to a single venue while the release is being observed in real conditions.
Written by Hybot operations lead, Deployment and operations lead, Hyrcan-Tech · · 5 min read
The asymmetry that makes staging worth it
A regression that reaches one site is an incident. The same regression reaching forty sites during the same lunch service is a different category of event, and the difference is not the bug — it is the rollout policy.
Web software gets away with simultaneous deploys because rolling back is fast and the failure is a page that does not load. A robot fleet fails differently: a venue mid-service with machines behaving oddly cannot simply refresh, and the operational recovery involves people on a floor rather than a revert.
How Hybot does it
Update rollout is selective per project. A release goes to a chosen site, is observed, and then goes wider. Fleet management is where that selection happens, alongside per-project health.
Because each project runs on its own subdomain with its own database, "one site takes this release" is a genuine boundary rather than a feature flag that might leak.
Choosing the first site
This is an operational judgement rather than a technical one, and the criteria are mundane:
- Somewhere a problem will be noticed quickly. An engaged operator who will actually tell you is worth more than any amount of instrumentation.
- Somewhere it matters least. Lower stakes, more tolerant pattern of use.
- Somewhere with good visibility. Health reporting you already trust.
Notably absent: "somewhere small". A tiny site may not exercise the change at all, in which case a clean run there proves nothing.
What to watch, and for how long
The same heartbeat data used day to day — internet quality, service status and vendor reachability, charted across the last hour.
A useful discipline is to wait for one full busy period before widening. Most robot software problems are load-shaped: they appear when the queue is deep and several units are competing, which is exactly the condition a quiet Tuesday afternoon does not reproduce.
The relationship to fleet resilience
Staged rollout is the same instinct that makes Hybot skip a misconfigured robot at startup with a warning instead of refusing to start, and allow robots to be added or removed without a restart. In each case the design goal is that one bad thing stays one bad thing — the coordination article covers that principle more broadly.
Where to go next
Cluster hub: Intelligent robotics.
Frequently asked questions
Why not update every site at once?
- Because a regression that reaches every site simultaneously is a coordinated outage rather than an incident. Staging means the worst case is one venue having a bad shift, which is recoverable, instead of an entire estate failing during the same lunch service.
How is a site chosen to go first?
- Usually one with a tolerant operating pattern and good visibility — somewhere a problem will be noticed quickly and matters least if it happens. That is an operational judgement about the venue, not a technical property of the software.
What do you watch after a release?
- The same per-project heartbeat used day to day: internet quality, frontend, backend and WebSocket status and vendor API reachability, charted over the last hour. A release that degrades a site usually shows there before anyone in the venue articulates what is wrong.
Where this fits
This page is part of Intelligent robotics: coordinating a fleet. If you are working through the topic in order, these are the neighbouring pages.
Multi-robot coordination, explained
Two robots are more than twice the problem of one. The failure modes that appear at the second unit, and what a coordination layer has to do about them.
Robot fleet monitoring and heartbeats
How a deployed site reports its own health — heartbeats, latency, service status and node redundancy — so problems are found before a customer reports them.
How robot task assignment algorithms work
Nearest-robot, first-come, and weighted scoring compared — why the dispatch rule you choose decides whether a fleet balances its work or exhausts one unit.
Take it further
If a question here applies to a venue you actually run, the specifics matter more than the general case.