What Data Should AI Use
What data should AI use? ==AI should use data that is relevant, representative, accurate, secure, lawfully obtained, and appropriate for the specific context in which the system will operate.== [:cite[1]{ln=1}] [:c...
What data should AI use? ==AI should use data that is relevant, representative, accurate, secure, lawfully obtained, and appropriate for the specific context in which the system will operate.== [:cite[1]{ln=1}] [:cite[2]{ln=1}] [:cite[3]{ln=1}] Good AI data should be: Relevant to the local context. Because training data shape what an AI system can do, the data should reflect the population, language, environment, and conditions where the system will be used.[:cite[1]{ln=1}] [:cite[4]{ln=4}] Representative and inclusive. Data should not systematically overrepresent certain demographic, cultural, or epidemiological groups while excluding others, since unrepresentative data can produce biased results.[:cite[2]{ln=1}] Accurate and high quality. Data should be sufficiently complete, timely, consistent, and reliable for training, testing, and monitoring; incomplete or delayed administrative records can weaken AI systems.[:cite[5]{ln=3}] Digitized, structured, and interpretable. Useful datasets need more than machine readable files: they should include context specific metadata and standardized formats that make their meaning clear to machines and users.[:cite[1]{ln=2}] [:cite[6]{ln=2}] Available at the necessary scale. The volume and quality of data should match the requirements of the intended model and task.[:cite[1]{ln=3}] Collected and used with privacy protections. Personal data should be processed for specific, legitimate purposes, with appropriate consent, authorization, data minimization, and security safeguards.[:cite[7]{ln=2}] [:cite[8]{ln=2}] Shareable under clear governance rules. Data sharing arrangements should specify what may be shared, with whom, and under what conditions; technical systems should also support secure interoperability.[:cite[9]{ln=5}] [:cite[9]{ln=6}] Available in local languages and sectors. Countries developing locally useful AI systems need data in local languages and sector specific datasets, particularly where existing digital data are scarce.[:cite[6]{ln=5}] [:cite[10]{ln=2}] Auditable and properly documented. Data provenance, access rules, limitations, and intended uses should be recorded so that people can evaluate how the data affect the system’s results.[:cite[6]{ln=2}] [:cite[11]{ln=3}] Not all data should be openly available. Sensitive information may require restricted access, privacy enhancing technologies, synthetic data, or federated learning so that organizations can gain useful insights without exposing individuals’ raw data.[:cite[12]{ln=1}] [:cite[13]{ln=1}] [:cite[13]{ln=4}] ==The central rule is: use the minimum data necessary for a clearly defined purpose, but make sure it is sufficiently representative and reliable for that purpose.== [:cite[8]{ln=2}] [:cite[14]{ln=2}] [:cite[1]{ln=2}]