Emerging Data Practices: Data Work in the Era of Large Language Models
Authors
Generative AI (Text, Image, Music, Video)AI Ethics, Fairness & AccountabilityAlgorithmic Fairness & BiasData Scientists & AnalystsAI/ML Researchers & EngineersHCI Researchers
Research Background and Problem
- Problems and Challenges: The authors identify significant challenges in constructing and managing datasets required for large language models (LLMs) amidst their rapid development and widespread application. These challenges include data uncertainty, the demand for large-scale data, ethical and legal issues surrounding generative AI, and the rapid iteration in development processes.
- Importance: Data is the foundation of AI development, and the capabilities of LLMs heavily depend on data. This not only affects model performance but also raises societal concerns such as bias, privacy issues, intellectual property disputes, and other social risks. Therefore, understanding data processing practices is crucial for designing more responsible and secure AI systems.
- Research Motivation and Related Work:
- LLMs require vast amounts of data of diverse types, yet many definitions and standards under current technological and legal frameworks lack consensus.
- Academic research in related fields emphasizes that a comprehensive study of practitioners' data practices can provide critical insights into addressing the risks of generative AI.
Solution
- Proposed Approach:
- The authors conducted semi-structured interviews with 25 practitioners involved in different development stages at a multinational IT company to explore their data processing practices.
- The study focuses on data practices across the following three development stages:
- Defining and operationalizing data governance.
- Data processing during pre-training and fine-tuning phases.
- Post-training use of synthetic data for model generation and evaluation.
- Innovations:
- Provides empirical analysis of data processing changes throughout the LLM development lifecycle.
- Reveals how practitioners navigate uncertainties by relying on contextual resources (e.g., regulations, expert guidance) to drive model development.
- Explores how the unique characteristics of LLMs influence data practices and their future technical and societal impacts.
- Implementation Steps:
- Recruit practitioners with data processing experience and design a tailored interview protocol.
- Use inductive thematic analysis to categorize and extract challenges and dependency mechanisms related to data practices.
- Integrate findings with cutting-edge issues in the human-computer interaction (HCI) field to propose research opportunities that support practitioners.
Research Outcomes
- Specific Findings:
- Describes the transitions and key decision points in data processing practices across different LLM development stages.
- Highlights how practitioners address uncertainties, such as through standardized processes, expert collaboration, and leveraging LLM technologies to achieve goals.
- Elucidates the profound impact of LLMs' unique characteristics (e.g., large-scale demands and ethical ambiguities) on data practices.
- Advantages Over Existing Solutions:
- Provides a deeper analysis of the complexities and critical issues in generative AI data work, such as ethical implications and the societal consequences of technical decisions.
- Offers actionable and practical recommendations to support practitioners in managing large-scale, uncertain, and complex data challenges.
- Experimental or Evaluation Results: Breaks down typical data practices in each development stage, uncovering how uncertainties influence operations during these stages.
- Limitations and Future Directions:
- Due to the complexity of data work, there may be unexplored details and practices.
- The study focuses on a multinational company, leaving out practices in smaller enterprises or other industries.
- Future research could explore cross-industry comparisons to enhance a comprehensive understanding of the societal impacts of generative AI. Researchers could also investigate tools and frameworks that support practitioners in improving data transparency and traceability.
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- What are the main data governance and processing challenges in large language model (LLM) development?Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
- How are data practices at different development stages influenced by LLM characteristics?Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
- How do practitioners handle uncertainty in data processing to advance LLM development?Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
lightbulb
Practical Problems
1- Practitioners struggle to manage the large-scale, complex data required for LLM development.Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
- 71%
Can AI Be a Moral Victim? The Role of Moral Patiency and Ownership Perceptions in Ethical Judgments of Using AI-Generated Content
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 71%
Interface Support for Evaluating Disability Bias in AI Generated Images
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 67%
The Landscape and Gaps in Open Source Fairness Toolkits
CHI '21· AI Ethics, Fairness & Accountability +1
- 67%
Towards Fairness in Practice: A Practitioner-Oriented Rubric for Evaluating Fair ML Toolkits
CHI '21· AI Ethics, Fairness & Accountability +1
- 67%
Jury Learning: Integrating Dissenting Voices into Machine Learning Models
CHI '22· AI Ethics, Fairness & Accountability +1
- 67%
Capable but Amoral? Comparing AI and Human Expert Collaboration in Ethical Decision Making
CHI '22· AI Ethics, Fairness & Accountability +1
- 67%
Out of Context: Investigating the Bias and Fairness Concerns of "Artificial Intelligence as a Service"
CHI '23· AI Ethics, Fairness & Accountability +1
- 67%
“It is currently hodgepodge”: Examining AI/ML Practitioners’ Challenges during Co-production of Responsible AI Values
CHI '23· AI Ethics, Fairness & Accountability +1
- 67%
STILE: Exploring and Debugging Social Biases in Pre-trained Text Representations
CHI '24· AI Ethics, Fairness & Accountability +1
- 67%
A Design Space of Behavior Change Interventions for Responsible Data Science
IUI '25· AI Ethics, Fairness & Accountability +1
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3714069
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), AI Ethics, Fairness & Accountability, Algorithmic Fairness & Bias
work
Professions
Data Scientists & Analysts, AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers