Valuation · 7 min read · updated

Automated home valuation: how the number is built, and when the model must stay silent

A human appraiser takes up to a week and charges 150,000 to 1 million soum. A model answers in a second, for free. The speed difference is obvious — the useful question is when the fast number can be trusted.

A cloud of data points with a trend line, and an empty dashed area where there is no data

Where the number comes from

An automated valuation model — AVM — follows the same logic a person does: find comparable properties, look at what they sold for, adjust for the differences. The difference is volume: a person holds five comparables in their head, a model looks at thousands.

Which leads to the first thing worth understanding: the quality of the answer is set by what the model sees, not by how clever it is. On dirty data any architecture is wrong — just confidently, and on every property at once.

≤ 12%
target MdAPE error across Tashkent
104,235
listings come down to about 62,000 properties after cleaning and merging duplicates
100
observations a year — below that a district gets no figure at all

Deduplication matters more than the model

The main work in housing valuation is not training. It is merging duplicates.

The same apartment sits in open data as several listings: the owner or the same agent reposts it, and a neighbouring agent lists it too. The prices differ; sometimes the floor areas do too. To a model this looks like several different properties that somehow cost different amounts in the same building — and it faithfully learns that spread as “normal market noise”.

The order of magnitude: of 104,235 listings, 76,862 remain after cleaning, and merging brings them down to about 62,000 properties. Roughly one listing in five turns out to be a repeat. Any metric computed on unmerged data always looks better than the truth: the model is not predicting a price, it is recognising another listing of the same flat.

How we counted. 101,663 is the archive of closed for-sale housing listings from one large open listings platform in Uzbekistan, which we downloaded on 22 August 2026. Together with the active listings that makes 104,235 records. After dropping unusable records (no floor area or coordinates, area outside 15–400 m², room count outside 1–8, price per m² outside $150–6,000, outside Tashkent), 76,862 remained. Duplicates were merged by a single rule: same district, room count, floor and number of storeys, area within 1.5 m², points no more than 200 m apart, prices less than 15% apart (25% if the descriptions are similar), listed at the same time or within 45 days of each other. That left about 62,000 properties. We checked 30 random groups of three or more listings against the originals by hand, and all 30 were the same apartment. The 35,000–55,000 range was an estimate we wrote into the specification before measuring, and the measurement did not confirm it.

Validation split by time only

The second thing that decides the fate of the number is how the data was split into training and validation.

A random split produces pretty figures that mean nothing: listings from the same month land on both sides, and inside a month the market barely moves. The model is effectively peeking.

An honest split is by time only: train on the past, validate on the future. The figures come out worse. But those are the conditions the product actually runs in — nobody knows tomorrow’s price.

Two models instead of one, and prices in dollars

Two more decisions that off-the-shelf AVMs usually leave unsaid.

New builds and resale are modelled separately. They are different markets: a new build’s price is tied to construction stage and the developer’s price list, resale is tied to condition, floor and haggling. One shared model averages two different mechanics and misses on both.

Prices are computed in dollars. The soum moved over the observation period, and a model trained on soum prices learns the exchange rate instead of the price of housing. That is the most galling failure mode: the metrics look decent while the product measures currency.

The refusal rule

A good AVM has to know how to stay silent.

If a district has fewer than a hundred observations in a year, there is nothing for the statistics to stand on. The honest answer is “not enough data”, not a number with a confidence interval as wide as the market.

A product that always answers passes off a guess as a calculation whenever the data is thin — and you have no way of knowing whether your property is one of those cases.

This is awkward to sell: clients want a figure for every address. But refusing to answer on thin data is exactly what separates an instrument from a number generator. Our target error is no worse than 12% MdAPE across Tashkent, and it only means anything alongside the refusal rule: without it, any metric can be improved by inventing confidence where there is none.

What to ask any valuation vendor

Five questions that quickly reveal what you are dealing with.

  1. How do you merge duplicates? If the answer is “our data is clean”, nobody looked for duplicates.
  2. How was the train/validation split done? “Randomly” means the metric cannot be trusted.
  3. Which currency is the price modelled in? And if it is soum — what was done about exchange-rate drift.
  4. Can the model decline to answer? If it answers every address, ask what it does in a district with five transactions a year.
  5. What data was the error measured on? Error on the training set is not an error.

None of these questions requires a technical background. All five require a straight answer from the vendor.

What we don’t know

We are not claiming automated valuation replaces the official kind. By law, a valuation is mandatory when state-owned property is involved in a deal, and otherwise when the value is disputed, including in mortgage lending (Article 11 of the Law on Appraisal Activity). The mortgaged property is valued by agreement between the parties or by an appraiser (Article 10 of the Law on Mortgage). In practice banks require an appraiser’s report for a mortgage, and a model has no such status and cannot acquire one.

Its use is elsewhere: before you go to the appraiser, who charges from 150,000 to 1 million soum for an apartment by published price lists and takes two to seven working days over the report. Knowing whether a property lands in range, not spending a week on something that would never clear as collateral, seeing that an asking price is inflated — that is enough to pay for the second you waited.

The rules above come from the specification of our own valuation engine, EstIQ, and the duplicate figures from our own measurement on 22 August 2026. The product has no paying customers yet, and we are not going to pretend otherwise. EstIQ’s own site shows different figures: it now collects listings from two platforms across the whole country, while this is a snapshot of one platform in Tashkent. EstIQ publishes its actual error and the areas where it refuses to name a price on its accuracy page (in Russian).

Read next

Show your process — I will send back an automation map in 48 hours

Six questions about how things are set up at your place today. The output is a diagram: what can be taken off people, in what order and what it costs.