What a chain should log about every delivery app outage
Most restaurant chains log outages as free text in a group chat, which is enough to remember an evening and useless a month later. A record that can be aggregated and escalated needs a fixed set of fields, and the ones that get left out are always the same three: the listing rather than the branch, the platform’s own name for the state, and anything captured from the customer side while the incident was still running. Everything else can be reconstructed afterwards. Those three cannot.
Is the record about a branch, a platform, or a listing?
A listing, meaning one branch on one platform, and getting this wrong makes the whole month uncountable. A branch trading on four apps can be dark on one and healthy on three at the same moment, and a record filed against the branch cannot express that. It also cannot be added up, because two incidents at the same branch on different platforms are two separate commercial events with two separate counterparties.
This is the same reason our own measurement uses the listing as its unit. Availability differs per listing in practice, so aggregating first and analysing second destroys the only distinction that leads anywhere. File one record per listing per incident, always, even when four listings went dark together in what everyone in the building experienced as one event.
Which timestamps are load bearing, and which are decoration?
Five carry weight and the rest are context. The stated opening time for that listing on that day, because everything is measured inside declared hours. The first confirmed unavailable observation. The last confirmed unavailable observation. The first confirmed orderable observation afterwards. And the moment somebody in the business became aware, which is almost never the same as the first of those.
| Field | Why it earns its place |
|---|---|
| Listing (branch plus platform) | The only unit that can be aggregated or escalated |
| Stated trading hours that day | Downtime outside declared hours is not downtime |
| First and last confirmed unavailable | Defines the duration everything else is derived from |
| Detection time | The gap to the first unavailable reading is your own response cost |
| Platform state name, verbatim | Decides who can lift it and therefore who to contact |
| Who could lift it | Splits self service from escalation before anyone tries |
| Customer side capture | The only evidence that survives the night |
| Promotion running at the time | Spend against a listing nobody could order from |
| Interruption or failure to open | Two different failures with different causes |
The gap between the first unavailable reading and the detection time is the single most useful number a chain can produce about itself, because it is entirely within its own control. The duration of the incident is largely not.
Which state name goes in the record, and in whose words?
The platform’s, verbatim, because the vocabulary decides the next action. Careem separates the state named “Offline”, which a branch can lift, from the state named “Outlet Closed”, whose tooltip reads “To reactivate your outlet, please reach out to Careem”. Deliveroo’s Forced Closure “can only be enacted by Deliveroo” and an attempt to lift it returns “This closed period can only be updated by Deliveroo.” Jahez runs visibility as a second axis with the values “Visible”, “Hidden” and “Partially Visible”. Snoonu labels the result of its own pause “Orders Paused”.
Recording that the branch was down collapses all of those into one useless category. Recording the exact string preserves the one fact that determines whether this is a five second fix by a shift manager or a week of account management. Where the platform also captures a reason, capture the same value: Talabat publishes a closure reason set including TOO_BUSY_KITCHEN, TOO_BUSY_NO_DRIVERS, TECHNICAL_PROBLEM and HOLIDAY_SPECIAL_DAY, and Careem requires one to be chosen, prompting “Select reason for taking outlet {{ selectedStatus }}”.
What has to be captured from the customer side, and when?
A screenshot of the storefront as a customer sees it, taken from an address inside the branch’s catchment, with the time visible, while the incident is running. Ideally paired with a screenshot of the merchant portal taken within the same minute. That pair is the whole evidentiary value of the record, because a portal showing a healthy branch beside an app showing an unorderable one is a contradiction that has to be explained rather than a complaint that can be absorbed.
The timing is not negotiable. Storefront state is a live value rather than a log, so nothing about 19:40 exists unless somebody looked at 19:40. Reconstructing it the next morning is impossible on every platform we have examined, and this is the field that separates a record a platform engages with from a record it does not.
What context makes a record readable a month after nobody remembers it?
Four items, all cheap at the time and unrecoverable later. Which promotion was live on that listing, since paid placement charged by time keeps running against a listing nobody can reach. What the schedule said for that day, because a narrowed schedule is a common hidden cause. Whether a menu publish, credential change or integration deployment happened that day, since those cluster with incidents. And whether the platform tablet was in use alongside the integration, which is a known source of conflicting order acceptance.
Add a plain sentence of what was done and when, in order. Not a narrative, a sequence: noticed, checked the app, checked the portal, called the branch, contacted support, reopened. A month of those sequences shows where the time actually goes, and it is almost never where the team assumes.
Which fields make a month of records add up to something?
The ones that are the same on every row. Duration in minutes, so lost hours can be totalled per listing and per platform. A cause class from a fixed short list rather than free text. Whether the incident was an interruption inside trading hours or a failure to open at the start of them, since those have different fixes. And who could lift the state, which sorts a month into the part you control and the part you escalate.
That last split is what turns a log into a management tool. Our July 2026 UAE panel recorded 3,759 failures to open across the month, of which 2,906 never came online within the check window, which is a class of failure that a chain filing everything as “offline” would never see as a category at all. The definitions behind those counts are at our methodology.
What should a chain do with the month once it has one?
Three things, and only the first is about platforms. Take the total lost hours per platform into the account management conversation, with dated readings rather than assertions. Rank your own branches by detection gap rather than by downtime, because that is the number your operation owns. And check whether one platform is producing states that only it can lift, which is a commercial fact rather than an operational one.
Maintaining this by hand across a large estate is where it usually fails, because the fields that matter most are the ones that must be captured in the first minutes. Kitchain (kitchain.co) records the listing, the confirmed timestamps and the customer side reading automatically, which leaves a chain to supply only the parts that live inside its own building: what was done, by whom, and how long it took to notice.