What an auditor, or Sekit's evidence engine, asks for.
AI dataset record: provenance, permitted use and quality
For each dataset an AI system uses, the record of where it came from, what the company is allowed to do with it (licence, consent, restrictions), whether it contains personal data, and the quality and preparation checks run before using it (cleaning, labelling, data split).
From the Sekit evidence catalog
In practice
For each dataset an AI system uses, whether it is a training set for a custom model or the documents fed to a vendor's retrieval feature, the company needs a record of where the data came from, what it is licensed or consented to be used for, and whether it contains personal data. An auditor pulls the AI dataset provenance record and checks the licence or consent basis matches how the data is used in practice, and that a quality or cleaning step is documented before the data reached the AI system. The typical gap is a dataset pulled from a public source with no record of its licence terms at all.
Common gaps
A dataset used to fine-tune or ground an AI feature has no recorded licence or consent basis, only a note that it came from the web.
Personal data feeding an AI system was never flagged as such, so the data protection team never reviewed it.
The dataset record exists, but nobody documented the cleaning step applied, or the labelling and data-split steps used when the company trains a model itself.
Questions your auditor will ask
What is your AI system allowed to do with this dataset?
The dataset provenance record states the licence or consent basis and any restrictions for every dataset the AI system uses.
Does this dataset contain personal data?
The record flags personal data content explicitly, feeding the same data inventory used for the company's other personal data assets.
What quality checks ran on the data before the AI system used it?
The dataset record documents cleaning and de-duplication of the source documents, plus labelling and data-split steps where the company trains or tunes a model itself.
Where regulation demands it
GDPR Article 30.1 requires a record of processing activities that names personal data in use, which for AI extends to datasets that train or ground a model.
Related controls
Via the shared Sekit CSF topic, not the framework's own index.