Where our data comes from
What we collect, what we do not have, and how to judge whether our figures are reliable.
What we collect
Our figures are built from live retail diamond listings - the diamonds actually offered for sale, at the prices actually asked, by the retailers we track. We currently track in the region of 500,000 listings, refreshed continuously.
For each listing we record the attributes that drive price: shape, carat weight, colour, clarity, cut, polish, symmetry, fluorescence, measurements, certifying laboratory, and the asking price. Those attributes are what let us compare like with like rather than averaging across diamonds that have little in common.
We also keep a daily snapshot, which is what makes the price history on this site possible. That archive currently runs to 16 months of daily observations across ten shapes, natural and lab-grown, broken out by carat band and quality band.
What we do not have
This is the more useful half of the page, and the half most data pages leave out.
- We do not have transaction prices. We see what a diamond is listed at, not what it finally sold for. Negotiation, promotions and trade-ins are invisible to us. Treat our figures as the asking-price market.
- We do not have the whole market. Our coverage is the retailers we have a data relationship with. A diamond sold somewhere we do not track does not exist in our figures, however good its price.
- We do not physically inspect diamonds. We rely on the certificate. Where a lab is more or less strict than another, that variation flows into our data too.
- We do not cover custom, estate or private-sale diamonds. Those markets price differently and are excluded entirely.
- Our history has gaps. Where our snapshot did not run, the series breaks and the chart says so. We do not interpolate across missing days to make a line look continuous.
How the data becomes a price
Collecting listings is the easy part. Turning them into a number that means something is where the judgement lives - which quality band to publish, why a median rather than an average, why shapes cannot share a price curve, and why a thin sample is more dangerous than no sample at all.
All of that is set out in the methodology, including an error we found in our own index and the change we made because of it.