How Data Workers Shape Datasets: The Role of Positionality in Data Collection and Annotation for Computer Vision
Data workers play a key role in the big data industry. Clients hire data workers to collect and annotate data with human identity concepts, like demographic categories or clothing items. Often, such workers are treated as computational—they are expected to quickly and objectively conduct their work, with the goal of having unbiased datasets for training and evaluating models. Computer vision is especially interested in fair and impartial data due to biases and unethical practices in the field. However, far from impartial, data workers imbue computer vision data with "biases" beyond correct versus incorrect answers. Data workers embed their own specific positional perspectives about identity concepts in both collection and annotation processes. Through interviews and ethnographic observations of data workers (both freelance and business process outsourcing (BPO) employees), we show how worker positionality influences decisions during data work. We also show how unintended outcomes, generally portrayed as "biases," occur when positionality is not explicitly considered in client instructions. We discuss how employing a lens of positionality in data work reveals the gulfs between data worker perspectives and client expectations, which are colored by a web of positional actors beyond isolated data workers. We propose positional (il)legibility as an approach to data work that embraces the reality of positionality in classification practices that the lens of "bias" fails to appropriately account for.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)