Article
Multi-robot coordination, explained
Coordinating several robots is harder than running one because the machines compete for jobs, for space and for charging points. A coordination layer has to decide which robot takes each task, keep that decision current as conditions change, and make sure a single misconfigured or offline unit does not stop the rest of the fleet working.
Written by Hybot technical lead, Technical lead, Hyrcan-Tech · · 6 min read
One robot is a device; two is a system
With a single robot there is no allocation problem. Every job goes to the only machine available, and the interesting questions are all about navigation.
Add a second and three new problems appear at once, none of which existed before:
- Who takes this job? A choice now has to be made, continuously.
- What about the space? Two machines in one corridor is a different situation from one.
- What about the chargers? A shared resource with a queue.
None of these are hardware problems. They are all decisions, which is why the coordination layer is where the value sits once a venue passes one unit.
The failure mode nobody expects
Bad allocation does not look like a crash. It looks like both robots being busy in the same third of the venue while a section goes unserved, or one unit doing three-quarters of the shift's distance while the other idles near its charger.
The venue reads this as "the robots aren't helping much". Nothing has failed. The allocation was simply poor, repeatedly, for four hours.
That is why how tasks are assigned is the single most consequential design decision in a fleet, and why Hybot re-scores the whole queue every two seconds rather than deciding once at task creation.
Central decisions rather than peer negotiation
Hybot's robots do not negotiate with each other. They report state — position, battery, status — to the management system, and the system allocates.
This is a deliberate simplification. Peer negotiation is academically more interesting and operationally much harder to debug: when a job goes to a surprising robot, a central decision can be explained, while an emergent one has to be reconstructed. An operator standing in a venue at service time needs the first kind.
Degrading one unit at a time
A fleet has to survive its own members. Two Hybot behaviours exist for exactly this:
- A misconfigured robot is skipped at startup with a warning. The fleet comes up; that unit does not. The alternative — one bad configuration preventing every robot from starting — turns a small mistake into a lost morning.
- Robots can be added, removed and reloaded without a server restart, so taking a unit out for inspection is not an outage.
Similar reasoning applies to releases, which is why update rollout is staged per project rather than pushed everywhere simultaneously.
Knowing what the fleet is doing
Coordination is only as good as the state it is based on. Position and status stream live; per-robot counters record completed tasks and distance travelled, daily and lifetime. Fleet monitoring covers how a site reports its own health.
Where to go next
Cluster hub: Intelligent robotics.
Frequently asked questions
Why is the second robot the hard one?
- Because with one robot there is no allocation question — it takes every job. The moment there are two, something must choose, and choosing badly is worse than not choosing at all: both units can converge on the same area while another part of the venue goes unserved.
What happens when one robot fails?
- In Hybot a misconfigured robot is skipped at startup with a warning rather than preventing the fleet from coming up, and robots can be added, removed or reloaded without restarting the system. One unit out of service should not become a whole venue out of service.
Do robots need to talk to each other?
- Not directly in this design. They report state to the management system and it makes the allocation decisions centrally, which is simpler to reason about and to debug than peer negotiation, and it means the operator can see why a job went where it did.
Where this fits
This page is part of Intelligent robotics: coordinating a fleet. If you are working through the topic in order, these are the neighbouring pages.
How robot task assignment algorithms work
Nearest-robot, first-come, and weighted scoring compared — why the dispatch rule you choose decides whether a fleet balances its work or exhausts one unit.
Robot fleet monitoring and heartbeats
How a deployed site reports its own health — heartbeats, latency, service status and node redundancy — so problems are found before a customer reports them.
Staged update rollout for robot fleets
Why pushing one release to every site at once turns a small regression into a coordinated outage, and how per-project rollout changes the risk profile of shipping.
Take it further
If a question here applies to a venue you actually run, the specifics matter more than the general case.