The emissions attached to a model's training compute, whose magnitude depends more on the local electricity mix than on the size of the model.
Training a large model consumes an amount of electricity that converts into emissions according to the electricity mix of the place and the hour, so two identical training runs can differ by a substantial factor depending on the region chosen. That finding shifts the question, since an operator can cut its footprint more by choosing where to compute than by reducing what it computes. A second distinction is indispensable and often omitted: training is a one-off cost while inference is repeated at every use, so cumulative inference exceeds training as soon as a model is widely deployed, moving attention from announcements to daily operation. For insurance the significance is not that of a peril but of a transition and reputational risk: a company publishing a reduction pathway while simultaneously deploying compute-intensive services must account for the gap, and its indirect emissions include those of the compute provider it uses.
The paper Carbon Emissions and Large Neural Network Training, published in 2021 by researchers at Google and Berkeley, estimated the emissions of training the GPT-3 model at around 550 tonnes of carbon dioxide equivalent, and showed that the choice of location and electricity mix changed that figure by an order of magnitude.
training emissions, empreinte carbone des modèles, coût environnemental de l'IA