Nurturing Capabilities: Unpacking the Gap in Human-Centered Evaluations of AI-Based Systems

AI-Assisted Decision-Making & AutomationAlgorithmic Fairness & BiasHCI ResearchersCognitive Scientists

Research Background and Issues

  • What problems or challenges did the authors identify?
    This study identified two major issues: first, the potential of artificial intelligence (AI) to support frontline caregiving work under high workload conditions has not been thoroughly evaluated; second, existing human-centered evaluation methods fail to adequately consider the broader social needs and professional aspirations of caregiving workers, leading to a "sociotechnical gap" between technology and social needs.

  • Why is this issue important?
    Caregiving work is highly stressful, and high workloads can negatively impact service quality and worker well-being. As service scales expand, the absence of effective technological support could further exacerbate the burden on workers. Moreover, the promotion and deployment of current AI technologies may overlook opportunities for caregivers to achieve their personal and professional goals, necessitating the optimization of technology to enhance their capabilities and freedoms.

  • Research Motivation and Related Work:
    This study is grounded in Amartya Sen's "Capability Approach," which emphasizes essential human freedoms and achievements. It systematically evaluates how AI technologies can help caregivers improve work efficiency in high-risk health domains (e.g., public health) while achieving long-term professional development goals. Existing research primarily focuses on the technical performance of AI systems (e.g., accuracy, recall rates) while neglecting the broader impacts of technology on its users.


Solution

  • What methods or solutions did the authors propose?
    The authors proposed a solution based on the "Design-Based Implementation Research" (DBIR) methodology, developing an intent classification system using artificial intelligence (GPT-4 and its optimized version GPT-4o) to separate non-medical messages in high-workload caregiving contexts. The research framework evaluates the potential of AI systems to reduce workload while uncovering the social and technical factors influencing AI model selection.

  • What are the innovative aspects of this solution?

    1. Introduction of the Capability Approach: Expands the evaluation scope of AI's role from traditional human capital (efficiency, skills) to the enhancement of individual capabilities and freedoms.
    2. Multi-Stage Deployment and Evaluation: Employs a phased deployment strategy with real-time adjustments, enhancing the reliability of AI models through human feedback and model refinement.
    3. Incorporation of Validation Workers: To reduce AI model errors, the solution integrates a "validation worker" mechanism (i.e., secondary human verification of AI outputs) to meet the high accuracy demands of public health.
  • What are the implementation steps and key technologies used?

    1. Preliminary Research and Needs Analysis: Methods such as observation, interviews, and focus group discussions were used to understand the actual needs and current conditions of caregiving workers.
    2. AI Model Selection and Development: An intent classification system was built using the open-source GPT-4o model to filter and label medical and non-medical information for the target audience.
    3. Multi-Stage Deployment:
      • Phase 1: The AI classification model runs in the background to collect performance data.
      • Phase 2: AI prediction labels are displayed in the user interface without affecting message distribution.
      • Phase 3: The system is fully deployed to filter non-medical messages, with validation workers handling AI model outputs.
    4. Real-Time Feedback and Iterative Optimization: Error prediction logs were analyzed, and continuous communication with users was conducted to improve input data and workflows for the model.

Research Outcomes

  • What specific outcomes were achieved?
    The AI system significantly reduced the non-medical workload of caregiving staff, improving their response efficiency to medical-related messages. Additionally, the system design highlighted the value of "validation workers," enhancing the controllability of high-risk information (e.g., medical errors).

  • What advantages does it have compared to existing solutions?

    1. The AI deployment in this study integrates traditional performance metrics with social adaptability factors, embedding design constraints within sociotechnical systems (e.g., financial costs, usability, domain-specific adaptability).
    2. The research emphasizes an evaluation framework aimed at expanding human capabilities, addressing workers' long-term professional development needs rather than merely focusing on efficiency improvements.
  • What were the experimental or evaluation results?

    1. Workload Reduction: The average daily message workload of caregiving workers (MSEs) decreased from 60% to 40%, achieving record-high service performance during system implementation (e.g., recorded response ticket counts).
    2. Error Control: The AI model achieved a final error rate of 1%, with validation workers promptly identifying and correcting medical messages missed by the model.
    3. Challenges: Some new burdens were shifted to validation workers (TT), highlighting the need to balance fairness when redistributing workloads.
  • Limitations and Future Directions:

    1. Limitations:
      • Certain non-medical messages (e.g., patient gratitude messages) hold emotional value for caregivers but were filtered out by the AI system, neglecting their psychological significance.
      • The increased burden on validation workers necessitates further optimization of workload distribution and welfare structures.
    2. Future Directions:
      • Develop multi-level, multidimensional evaluation metrics (technical, social, individual capability expansion).
      • Explore closer collaboration between validation workers and AI to reduce wasted effort and optimize human-AI interaction efficiency.
      • Conduct long-term assessments of AI's impact on healthcare service expansion and the professional development of grassroots workers.

Conclusion

This study, through a human-centered evaluation based on the Capability Approach, highlights the potential and directions for improvement of artificial intelligence in enhancing the professional capabilities of caregiving workers. It comprehensively demonstrates how to introduce and refine AI technologies in resource-constrained, high-risk public health domains, achieving a holistic evaluation of AI effectiveness across social needs, technical performance, and individual well-being. This provides a new paradigm for future AI deployment in high-risk fields while emphasizing the importance of workers' professional aspirations and social welfare.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/189556/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713278
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
AI-Assisted Decision-Making & Automation, Algorithmic Fairness & Bias
work
Professions
HCI Researchers, Cognitive Scientists
article
Content Status
Full text indexed
hub
Related Papers
8 related papers