The Open Dataset

Open publication

All donations together form an open dataset that we will publish and then update continuously. It will be freely available to researchers and to anyone interested, for example for:

  • Simulating and validating future power grids (the core of SAFEr Grid)
  • Load modeling and forecasting methods
  • Appliance signature detection (NILM) and flexibility analyses

The time series stay unchanged

The measurements are not summarized, but published at the resolution in which they reach us. Aggregated load curves already exist, and they are too coarse for the research questions behind the data donation.

The anonymization therefore works on the information attached to a time series.

Metadata and k-anonymity

Every time series comes with descriptive information: country, postcode, number of people in the household, living area, construction year of the building, appliance type, manufacturer, and model. We treat all of it as quasi-identifiers: features that name nobody on their own, but in combination can lead to a single household, for example the only four-person household with a particular heat pump in a small municipality.

Before every publication we therefore generalize this information until every combination that occurs is shared by at least k households (k-anonymity). A postcode becomes a larger region, a construction year a construction period, a rare appliance model the bare appliance type. Features that would still stand alone after that are dropped.

Identifiers and self-chosen appliance names are removed beforehand. After that, the metadata narrows things down to a group of at least k households at best. We document the procedure and the chosen value of k together with the dataset.

Residual risk in the load profile

The time course of consumption and generation reveals habits, such as when someone is at home. These patterns are part of the measurements themselves. Removing them would require a coarsening that noticeably reduces the scientific value of the dataset.

Without metadata that can be matched to anyone, such a pattern describes an unknown household. Assigning it to a specific home would require already knowing that home’s electricity consumption. A residual risk remains, since a load profile itself can carry features that make a household recognizable.

Further limits apply earlier on: camera, microphone, presence, and motion data are never read out, and no location more precise than the optional postcode is ever collected. Which appliances take part at all is your choice (Which data is transmitted). Until publication, we treat all raw data as personal data under the GDPR (privacy policy).

Access for project partners and verified researchers

Our partners in the SAFEr Grid consortium (two further RWTH institutes and Aalborg University) can obtain pseudonymized data under a data use agreement, with clear purpose limitations, without contact data, and no right to redistribute. For research questions that require finer metadata, such as the full postcode for regional analyses, we plan controlled access for verified researchers under the same conditions.