Emerging Data Practices: Data Work in the Era of Large Language Models

Generative AI (Text, Image, Music, Video)AI Ethics, Fairness & AccountabilityAlgorithmic Fairness & BiasData Scientists & AnalystsAI/ML Researchers & EngineersHCI Researchers

Research Background and Problem

  • Problems and Challenges: The authors identify significant challenges in constructing and managing datasets required for large language models (LLMs) amidst their rapid development and widespread application. These challenges include data uncertainty, the demand for large-scale data, ethical and legal issues surrounding generative AI, and the rapid iteration in development processes.
  • Importance: Data is the foundation of AI development, and the capabilities of LLMs heavily depend on data. This not only affects model performance but also raises societal concerns such as bias, privacy issues, intellectual property disputes, and other social risks. Therefore, understanding data processing practices is crucial for designing more responsible and secure AI systems.
  • Research Motivation and Related Work:
    • LLMs require vast amounts of data of diverse types, yet many definitions and standards under current technological and legal frameworks lack consensus.
    • Academic research in related fields emphasizes that a comprehensive study of practitioners' data practices can provide critical insights into addressing the risks of generative AI.

Solution

  • Proposed Approach:
    • The authors conducted semi-structured interviews with 25 practitioners involved in different development stages at a multinational IT company to explore their data processing practices.
    • The study focuses on data practices across the following three development stages:
      1. Defining and operationalizing data governance.
      2. Data processing during pre-training and fine-tuning phases.
      3. Post-training use of synthetic data for model generation and evaluation.
  • Innovations:
    • Provides empirical analysis of data processing changes throughout the LLM development lifecycle.
    • Reveals how practitioners navigate uncertainties by relying on contextual resources (e.g., regulations, expert guidance) to drive model development.
    • Explores how the unique characteristics of LLMs influence data practices and their future technical and societal impacts.
  • Implementation Steps:
    1. Recruit practitioners with data processing experience and design a tailored interview protocol.
    2. Use inductive thematic analysis to categorize and extract challenges and dependency mechanisms related to data practices.
    3. Integrate findings with cutting-edge issues in the human-computer interaction (HCI) field to propose research opportunities that support practitioners.

Research Outcomes

  • Specific Findings:
    • Describes the transitions and key decision points in data processing practices across different LLM development stages.
    • Highlights how practitioners address uncertainties, such as through standardized processes, expert collaboration, and leveraging LLM technologies to achieve goals.
    • Elucidates the profound impact of LLMs' unique characteristics (e.g., large-scale demands and ethical ambiguities) on data practices.
  • Advantages Over Existing Solutions:
    • Provides a deeper analysis of the complexities and critical issues in generative AI data work, such as ethical implications and the societal consequences of technical decisions.
    • Offers actionable and practical recommendations to support practitioners in managing large-scale, uncertain, and complex data challenges.
  • Experimental or Evaluation Results: Breaks down typical data practices in each development stage, uncovering how uncertainties influence operations during these stages.
  • Limitations and Future Directions:
    • Due to the complexity of data work, there may be unexplored details and practices.
    • The study focuses on a multinational company, leaving out practices in smaller enterprises or other industries.
    • Future research could explore cross-industry comparisons to enhance a comprehensive understanding of the societal impacts of generative AI. Researchers could also investigate tools and frameworks that support practitioners in improving data transparency and traceability.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/188816/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3714069
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), AI Ethics, Fairness & Accountability, Algorithmic Fairness & Bias
work
Professions
Data Scientists & Analysts, AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers