Scraped Airbnb income estimates land within 10-20% of reality for professionally-run listings and 30-50% off for the long tail, because they infer bookings from review counts. Airbnb does not publish per-listing revenue to anyone. Every figure you will find — from Inside Airbnb, from commercial data vendors, and from us — is a model built on public signals, and the size of the error is predictable enough to correct for once you know how the model works.
The detail
Inside Airbnb documents its method openly, which is why it is the right thing to build on. Its San Francisco Model estimates bookings as reviews divided by an assumed review rate of 50%, multiplies by the city's average length of stay, and caps the result at 70% occupancy. Each of those three steps introduces a known, directional bias. The review-rate assumption understates listings whose guests review less often — corporate and repeat stays especially — and overstates listings with very few reviews, where a single review can imply two bookings that never happened. The 70% cap means the method structurally cannot show you a listing running 85%, which is exactly the listing you most want to identify.
Commercial vendors improve on this by scraping the booking calendar and classifying nights as available, booked or blocked. That is a better signal, but it cannot distinguish an owner blocking dates for personal use from a genuine booking, so it inflates occupancy for part-time hosts. It also cannot see the rate actually paid — only the rate displayed — which diverges from the achieved rate whenever a host offers weekly or monthly discounts, and those discounts are commonly 10-25%.
There is also a definitional trap worth naming: some datasets include the cleaning fee in revenue and some do not, and some report gross of the platform fee while others report net. A 15% difference between two sources is often not disagreement at all, just two different definitions of the word revenue.
The numbers that matter
Where this stops holding
The honest limitation is that no public dataset, ours included, sees an actual payout. Airbnb has never published per-listing revenue, and everything downstream is inference from reviews, calendars and displayed prices. That means we can characterise the direction and rough size of the error but cannot eliminate it, and the error is largest for exactly the listings that look most attractive — new, high-rate, low-review-count properties where a handful of reviews drive the whole estimate. The right posture is to read every income figure as a range, underwrite the bottom of that range, and treat anything above it as upside rather than plan.
Sources
Get this answered for your address
A HostPal Invest report runs the real occupancy, nightly rate, RevPAR, regulation risk and a buy / wait / avoid verdict for one specific property or drawn area, in any of 117 markets — not a national average. Street-Level reports from £29.