The Promises and Perils of using LLMs for Effective Public Services

Human-LLM CollaborationAI-Assisted Decision-Making & AutomationParticipatory DesignGovernment Officials & Civil ServantsSocial WorkersChild Welfare Workers

Paper Title

The Promises and Perils of using LLMs for Effective Public Services

Publication Info

  • Topic area: Application of AI and language models in public sector services, focusing on child welfare.
  • Keywords: LLMs, child welfare, public sector, BERTopic, LocalLLM, case management, social work, AI ethics, participatory design, thematic analysis.

Background and Problem

  • Problem / challenge: Governments are increasingly adopting AI tools like LLMs to improve public services, but their feasibility, limitations, and potential harms in high-stakes domains like child welfare remain underexplored.
  • Significance: Effective use of AI in child welfare could improve service delivery, reduce unnecessary family involvement, and optimize caseworker efforts, addressing critical societal needs.
  • Motivation and related work: Prior research has explored computational text analysis in public sector contexts, including predictive algorithms and topic modeling. However, gaps remain in understanding how LLMs can support nuanced, context-dependent decision-making in child welfare systems.

Solution

  • Proposed approach: Application of BERTopic and LocalLLMs (Llama 3.1) to analyze child welfare casenotes and Service Plans, aiming to track case progress and identify thematic connections.
  • Novelty:
    1. Collaborative study between HCI researchers and child welfare practitioners to investigate AI applications in operational challenges.
    2. Thematic analysis of Service Plans and Regular casenotes to identify shared and unique themes.
    3. Evaluation of LocalLLMs for tracking progress and identifying thematic trends in child welfare cases.
    4. Roadmap for participatory design of AI tools tailored to public sector needs.
  • Procedure and key techniques:
    • Preprocessing and anonymizing child welfare documents (Service Plans and Regular casenotes).
    • Applying BERTopic for thematic clustering of text data.
    • Using LocalLLMs to identify Service Plan Activity relevance in casenotes.
    • Manual validation of LocalLLM outputs and thematic analysis of Activity-relevant and irrelevant casenotes.

Results

  • Concrete findings:
    • BERTopic identified 60 topics in Regular casenotes, 38 in Service Plan Objectives, and 82 in Activities, grouped into high-level themes like parenting, health, and child custody.
    • LocalLLMs achieved moderate agreement with manual labels for Activity relevance, with Cohen’s kappa ranging from 0.604 (short cases) to 0.402 (long cases).
    • Activity-relevant casenotes decreased in extremely long cases, reflecting emergent child welfare concerns not captured in existing Service Plans.
  • Advantage over baselines:
    • LLMs demonstrated potential for granular thematic analysis and diagnostic insights compared to traditional predictive models.
    • BERTopic provided nuanced semantic clustering superior to LDA and LMF.
  • Experiments / evaluation:
    • Dataset: 720 cases with 52,748 Regular casenotes and 1,213 Service Plans from a Canadian child welfare agency.
    • Metrics: Agreement rates, Cohen’s kappa, false positive/negative rates.
    • Manual validation and thematic tracking across normalized case timelines.
  • Limitations and future work:
    • Findings may not generalize to other child welfare systems or public sectors.
    • Regular casenotes reflect worker perceptions, which may include biases or omissions.
    • Future work should explore additional LLMs, refine thematic categories, and validate findings across diverse datasets.

Summary

This study explored the application of BERTopic and LocalLLMs to analyze child welfare casenotes and Service Plans, revealing thematic connections and challenges in tracking case progress. While LLMs showed promise for diagnostic insights and thematic analysis, they struggled with complex cases requiring discretionary judgments grounded in social work practices. The findings underscore the need for human-centered, participatory approaches to design AI tools that support rather than supplant caseworker decision-making. Future research should focus on refining methodologies, expanding datasets, and addressing systemic constraints in public sector AI applications.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/223484/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790297
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
8 authors
sell
Subtopics
Human-LLM Collaboration, AI-Assisted Decision-Making & Automation, Participatory Design
work
Professions
Government Officials & Civil Servants, Social Workers, Child Welfare Workers
article
Content Status
Full text indexed
hub
Related Papers
1 related papers