Of a model's three inputs, hazard comes from a science, vulnerability from an engineering discipline, and exposure comes from a file the insurer has built itself. It is the only one of the three it fully controls, and it is by far the one that degrades its results most. An excellent model fed a mediocre exposure returns a mediocre figure, presented with all the model's authority. No downstream treatment repairs this, and that is why exposure quality is an underwriting subject and not an IT one.
The first attribute is location, and it comes in resolution levels that must be named. A coordinate taken on the building is the finest level. A geocoded postal address is close to it, within a few tens of meters. A postcode reduced to its centroid is something else entirely: on a large commune, the distance between the centroid and the actual site reaches several kilometers, which, for a flood peril whose intensity varies over a hundred meters, amounts to not knowing where the risk is. The model, for its part, will compute without complaint.
The effect of that aggregation is not a symmetric noise that would cancel out over a large number of contracts, and that is the subtlety worth understanding. Grouping at the centroid smooths intensities: genuinely highly exposed sites receive an average intensity that is too low, protected sites receive one that is too high. On the portfolio average the error may seem to offset; in the TAIL it does not offset, because the tail is made precisely of the most exposed sites, the ones the smoothing pulled back toward the average. An aggregated portfolio therefore reports an extreme loss that is too low, and it reports it systematically.
The second attribute is value, and it carries two independent traps. The first is the basis: a net book value, an old purchase value and a reinstatement value do not describe the same amount, and the gap commonly reaches a factor of two on an old industrial building. The second is the perimeter: building alone, building and contents, with or without business interruption. A model applying a damage ratio to a value that is not the one the curve was calibrated on produces a proportional error, invisible and not random.
The third attribute is the description of the asset: construction class, year, number of stories, occupancy, presence of a basement, height of the first floor. These fields are rarely filled in completely, and what is missing is not left blank by the model: it applies a default value, usually the regional average. A portfolio with half its construction fields absent is therefore half modeled as an average regional building, which erases exactly the differences one was trying to measure.
One must add the way protections and deductibles are transmitted, because that is where whole amounts are lost without anyone noticing. A site protected by a flood wall whose existence is not in the file is modeled without that wall. A statutory or contractual deductible not transmitted gives a gross loss treated as if it were net. An ignored per-site sublimit lets the model climb beyond what the contract can pay. None of these gaps shows in the output: the table comes out formatted, with its return periods, and nothing in it flags what was not said.
The conduct to adopt comes down to three cheap habits. Exposure quality is measured before risk is measured: share of sites geocoded at the building, share of construction fields recorded, basis of values, and those three rates are published beside every result. Sensitivity is tested by deliberately degrading part of the portfolio to the centroid, so as to know what aggregation costs on this particular portfolio. And data improvement is treated as an underwriting project with an order of priority, starting with the sites that weigh most in the tail, because correcting ten thousand addresses of no consequence costs as much as correcting the hundred that decide.
An insurer takes over a commercial portfolio of 1.9 billion euros in insured values, introduced by a broker. The file transmitted holds 3,100 sites, of which 2,250 are geocoded to the address and 850 attached to their postcode centroid, the latter representing 640 million in values. The construction class is recorded for 61% of sites. Values are declared at reinstatement value for buildings, but the business interruption column is empty across the whole file although 1,400 contracts carry one in the wording. The model returns a two-hundred-year flood loss of 74 million. The committee must decide on the takeover. What is to be said about that figure?
The 74 million figure is not approximate, it is biased, and the file's three defects all push the same way, downward. The 850 centroid sites carry a third of the values and are modeled on a smoothed intensity: that smoothing pulls back toward the average the genuinely most exposed sites, and those are the ones that make the tail of the distribution, that is, precisely the two-hundred-year number the committee is looking at. The error is therefore not symmetric and does not cancel out over the number of sites. The 39% of missing construction classes are replaced by the regional average, which erases the very sensitivity one pays a model for and makes the output less discriminating than a zone rate. And the empty business interruption column means the model returns a property burden only, while the portfolio carries interruption on 1,400 contracts: on a flood peril, where stoppages run for weeks, that is the term that can dominate the burden, and here it is absent by construction rather than by decision. What has to be asked before deciding is counted in working days, not in euros. First an address-level geocoding of the 850 sites, starting with the largest by value, since they decide the tail. Then a measurement of what aggregation costs on this portfolio, obtained by deliberately degrading a sample of well-located sites to the centroid: that is the only way to state an order of magnitude for the bias rather than allude to it. Finally the recovery of the business interruption amounts, failing which the figure presented does not answer the question asked. Deciding on 74 million as it stands would mean deciding on a measurement whose direction of error is known and whose size is not.
- 01Exposure is the only one of the three inputs the insurer fully controls, and it is the one that degrades its results most.
- 02Attaching sites to a postcode centroid does not add noise: it smooths intensities and understates the tail systematically.
- 03A book value where the curve expects a reinstatement value produces a proportional error, invisible and not random.
- 04A missing construction field is not left blank: the model applies a regional average, erasing the sensitivity one pays to obtain.
- 05Protections, deductibles and sublimits not transmitted never show in the output: the table comes out formatted and does not flag what was left unsaid.