Does being offline hurt your rating on a delivery app?
No delivery platform we can verify publishes availability as an input to its star rating, so restaurant chains will not find downtime anywhere inside the formula. What downtime does instead is change the arithmetic around it. A dark listing stops collecting the new scores that would have diluted the old ones, it produces a small crop of cancellation reviews rather than none at all, and it puts the branch at risk on the rating floors that gate promotion. On Just Eat, time offline is scored directly, but as a separate status rather than as a star.
Do any delivery apps put availability into the rating formula?
Not in anything they publish. Deliveroo defines its number as customer scores and nothing else, stating that “The rating shown on the restaurant list is an average of the last 400 ratings you have received”. Uber Eats is narrower still and excludes anything that did not complete, saying the metric “is based on your average customer rating, weighted for the past 90 days” and counts “only ratings submitted from completed orders”. An order that never existed because the storefront was dark cannot appear in either. The rating is a record of orders that happened, and downtime is a record of orders that did not.
That sets the ceiling on what an outage can do. A four hour interruption does not subtract stars. It removes a slice of trading from the sample the rating is built on, and the rest of the effect is second order. Operators who go looking for a direct penalty and find nothing conclude that downtime is a rating non event, which is the wrong conclusion drawn from the right observation. The penalty is real, it is applied somewhere other than the star.
Does a dark listing lose stars, or does it simply stop moving?
It stops moving, and whether that helps or hurts depends entirely on which way the number was travelling. This is the part that surprises people. A branch on a bad run that goes offline for a day freezes a bad average in place. A branch that has just fixed a problem and is climbing loses the days that would have carried the climb.
The freeze behaves differently depending on how the window is defined, and the two common definitions pull in opposite directions. Deliveroo’s window is a count of the last 400 ratings, so a dark listing adds nothing to the count and the same old scores stay on the card. Uber Eats weights a rolling 90 days, so the calendar keeps advancing whether or not you trade and old scores age out on their own. On a count window, downtime preserves your history. On a calendar window, downtime spends it. Two branches of one brand, one on each platform, come out of the same closure with the number moving in different directions.
Which reviews does an outage actually produce?
Very few, and that is precisely what makes them dangerous. A customer who opens the app and does not see your branch writes nothing, because there is nothing to write about. A customer who placed an order that was then rejected or cancelled writes something, and what they write is about your restaurant rather than about the platform. Snoonu’s rejection dialogue makes the mismatch visible in its own vocabulary: the reasons a branch selects include “Branch Closed”, explained as “The store is not operating at the time of the order”, and “Kitchen/Staff Overload”, explained as “The restaurant is too busy to fulfill the order”. Both are recorded against the restaurant.
So the review load from an outage is small in count and hostile in tone, and it lands in a window that has stopped receiving ordinary five star traffic to balance it. That is the mechanism by which a two hour incident can move a rating more than a busy fortnight of average service. It is not that the outage was scored. It is that the ordinary traffic which normally absorbs a handful of complaints was not there.
Where is time offline scored directly, and what does it decide?
On Just Eat, and it decides a status rather than a rating. The Local Legend criteria are published, they must be held “for two consecutive quarters”, and among them sit both “Customer food review rating of 3+ stars” and “Time offline at 10% or below”. Those are two separate tests inside one gate. A restaurant can hold its stars perfectly and still fail on availability, and the published benefits of the status include “Priority in giveaways, competitions and marketing campaigns”, which is commercial rather than cosmetic.
The two quarter requirement is the sharp edge. A quarter is not a window you can repair in a fortnight, and there is no partial credit. A chain with a branch that lost a bad week in month one of a quarter has already spent that quarter, and the criterion runs across the next one as well. This is the only place in the platforms we have examined where downtime is expressed as a numeric threshold that a restaurant can be measured against, which is a good argument for holding an availability figure per branch even if you never intend to file a claim about it.
Can a rating drop then cause more downtime?
It can, and this is the loop that makes the two numbers a single problem rather than two. Rating floors gate access to the tools that generate volume. Deliveroo requires that “Star rating must be 3.8 or more” among its Marketer criteria, reassesses shortly before each month begins, and removes Adverts and Offers “for a full month” from a site that misses. Careem is blunter for paid campaigns: “Only outlets rated 4.0 and above can run ads.” Foody sets its own floor for paid placement, requiring that a partner “Have an average rating of Users over 3.5” before it can buy a Recommended Position.
The loop closes on noon Food, where the contract turns behaviour into an availability consequence directly. Its supplemental terms define a “Brand Matter” as an event causing concern for the brand, “including, but not limited to, high cancellation or non-acceptance rates (as determined by Noon Food)”, and a Brand Matter is a named ground for suspension under the clause headed “j. Suspension of Noon Food Services”. An outage that produces rejections rather than a clean closure therefore feeds a metric that can itself end in a suspension, which is more downtime.
Why does the rating damage arrive weeks after the outage?
Because every window on every platform is long relative to an incident, and none of them are the month you are reporting on. A Deliveroo count of 400 empties at whatever rate the branch trades. A 90 day weighting on Uber Eats carries a bad Friday into the following quarter. Just Eat’s median can sit unmoved through the whole episode and then step a full point when the ordering of the underlying scores finally changes.
The practical consequence is a reporting trap. The month in which the outages happened usually shows a healthy rating, because the damage has not surfaced yet, and the month in which the rating drops usually shows clean availability, because the outages are over. Anyone reading the two numbers month by month in the same table will conclude that they are unrelated. They are related with a lag, and the lag is set by the platform rather than by you.
How should a chain read a rating next to an availability record?
By date, on the same branch and the same platform, rather than by month in a summary. The question worth asking is not what the rating is, but what the branch was doing during the days that the current window covers. That requires an availability history at branch level going back further than the rating window, which is roughly a quarter on most of these platforms.
Across our July 2026 UAE panel a typical listing lost about 9.5 trading hours in the month, and the mean Deliveroo interruption ran 3 hours 10 minutes, which on a count based window is roughly a service worth of scores that never entered the sample. A rating history is only interpretable next to an availability history for the same dates, and that outside view is what Kitchain (kitchain.co) records from the customer side. How the availability figures are defined and bounded is set out in our methodology.
One caution about causation before anyone builds a target on this. We measure availability and we read published rating rules. We have not measured the size of the rating effect of an outage, and no platform publishes one, so the mechanisms above are documented while the magnitude is not.