How to tell a real outage from a slow night

Restaurant operators try to answer this with the order count and it cannot be answered with the order count, because a slow night and a dark listing both produce a small number. Three tests do separate them, and none of them needs a data scientist. Whether the storefront can currently take an order is a binary that anyone can read. Your other branches on the same platform in the same city are a control group you already own. And the two situations have different shapes in time, one a step and the other a slope.

Why can order volume not settle this on its own?

Because it is the wrong kind of variable. Order count is a rate with genuine variance built into it, driven by weather, football, school holidays, a competitor’s promotion and the ordinary noise of a Tuesday. Any single evening’s total is compatible with a wide range of causes, and the range is widest at exactly the volumes where the question gets asked. A branch that normally takes forty orders and takes eleven has told you almost nothing.

Worse, the two situations converge as the evening goes on. An outage that started at seven and a genuinely quiet night both end at midnight with a low daily total, and by then the distinguishing evidence has gone. This is why teams that rely on the count end up arguing from memory the next morning, and why the argument is never settled. The count is a lagging summary of something that needed to be observed while it was happening.

What is the binary that actually separates them?

Whether a customer could have placed an order at that moment. Unlike volume, this has only two values, it does not depend on demand, and it is observable from outside the business by anybody with the app open. A quiet night is a listing that was orderable and not chosen. An outage is a listing that could not be chosen.

One reading is not enough, because platform front ends flicker and a single failure proves little. A confirmed reading is a negative observation repeated, which is how our own measurement works: a listing is checked at roughly ten minute intervals and a negative result is confirmed by a repeat check before it counts as an incident. That standard is worth borrowing even for a manual check on one branch. Open the app twice, five minutes apart, from an address inside the catchment, before deciding anything.

What baseline should the answer be tested against?

Your own, from a month when nothing unusual happened, and failing that the published ones. In our July 2026 UAE panel, listings were unavailable for 1.68 percent of the hours they said they would trade, and a typical listing lost roughly 9.5 trading hours over the month. Mean interruption length ran between 1 hour 50 minutes and 3 hours 34 minutes depending on the platform.

Two things follow for tonight’s question. Interruptions are long enough to be caught while they are running, which makes the binary test practical rather than theoretical. And they are frequent enough that a specific branch going dark in a given month is unremarkable, so an operator whose instinct says this never happens to us is usually making a statement about detection rather than about incidence. How those figures are bounded, and what they exclude, is at our methodology.

How do you use your own estate as a control group?

As a control for demand, which is the one variable you cannot otherwise hold still. Demand on a given Tuesday evening is a property of the city rather than of your branch, so sites sharing that city and that hour are exposed to the same weather, the same football match and the same school holiday, and differ only in themselves. That turns tonight’s low number from an unanswerable question into a comparison that has an answer, and it costs nothing, because you already own both sides of it.

The strength of the comparison depends entirely on picking siblings honestly. Two branches four hundred metres apart share the conditions. A branch in another emirate does not, and adding it puts back exactly the demand variation the control was meant to remove. Two or three of the nearest sites is the right size, and a control group assembled from whichever branches happen to be reporting is not a control group.

The sequence for turning that comparison into a diagnosis, one platform against all platforms and one branch against several, is written out step by step on orders have stopped coming in on one delivery app, and there is no reason to run it twice.

What does the shape of the drop tell you?

An outage is a step and a slow night is a slope. Orders do not taper when a listing goes dark, they stop inside a single clock hour and then resume in a similarly abrupt way when it returns. A quiet evening declines and recovers gradually and rarely produces a completely empty half hour during service. Plotting the day in half hour buckets, which any point of sale can do, makes the difference visible without any statistics at all.

There is a third shape worth recognising, which is zero from the opening time onward. That is usually not an outage at all but a failure to open, where the listing never came online after its stated opening time. It is common enough to be worth naming: in our July 2026 UAE panel there were 3,759 such events, of which 2,906 never came online within the check window and 853 opened late. A morning with no orders at all is more often a listing that never opened than a morning with no customers.

Which evidence disappears if you wait until the morning?

Almost all of it. The storefront state is a live value, not a log, so the fact that a listing was unorderable at 19:40 exists only if somebody observed it at 19:40. The platform’s own state field is equally transient, and several of the useful ones are only meaningful in the moment, such as Careem’s distinction between “Offline” and “Outlet Closed”, or the reason a Talabat closure carried out of its published set including TOO_BUSY_KITCHEN and TECHNICAL_PROBLEM.

By the following morning the branch is open, the portal shows a healthy status, and the only surviving artefact is a low order count that nobody can attribute. That is the exact situation this page exists to prevent, and the fix is not analysis, it is capture. Whatever is observed during the incident, in a screenshot with a timestamp, is worth more than any amount of reconstruction afterwards.

What test should a chain run tonight, in order?

Two readings and a photograph, in that order, because tonight’s job is evidence rather than diagnosis. Try to reach checkout from an address inside the branch’s catchment. Repeat five minutes later so the negative is confirmed rather than asserted. Then capture the customer view and the portal view together with a visible clock, and write down the time you first looked as well as the time you captured, since those two are rarely the same and the gap between them is the part everybody forgets to record.

That takes a few minutes and produces something a platform has to answer rather than an impression it can absorb. It also settles the argument before it starts, because a confirmed negative taken from the customer side is not a claim about how busy the evening felt and cannot be answered with one. Doing it continuously, across every branch and every platform, at an interval short enough to catch an incident before it ends, is what Kitchain (kitchain.co) automates, and the reason to automate it is that this evidence has a shelf life measured in hours rather than days.

Related

Start Monitoring



    No credit card. No integrations.
    We'll configure your first location and confirm within 24h.
    Request a Demo

    Book a personalized walkthrough of Kitchain Products.



      We'll get back to you within 24 hours.