Agent-Supported Foresight for AI Systemic Risks: AI Agents for Breadth, Experts for Judgment

Explainable AI (XAI)AI-Assisted Decision-Making & AutomationAI Ethics, Fairness & AccountabilityParticipatory DesignUser Research Methods (Interviews, Surveys, Observation)AI/ML Researchers & EngineersHCI ResearchersSociologists & Anthropologists

Paper Title

Agent-Supported Foresight for AI Systemic Risks: AI Agents for Breadth, Experts for Judgment

Publication Info

  • Topic area: Foresight methods for identifying systemic risks in AI technologies.
  • Keywords: AI systemic risks, foresight, in-silico agents, Futures Wheel, hybrid workflows, Technology Readiness Levels (TRLs), systemic consequences, large language models (LLMs), risk evaluation, participatory methods.

Background and Problem

  • Problem / challenge: Systemic risks of AI technologies are difficult to foresee due to cognitive limitations, lack of long-term thinking, and the abstract nature of such risks. Existing foresight methods are constrained by near-term focus and limited scalability.
  • Significance: Anticipating systemic risks is crucial for responsible AI governance, especially for emerging technologies with potentially far-reaching societal impacts.
  • Motivation and related work: Prior work in foresight and AI risk assessment has explored speculative design, participatory methods, and taxonomies but often lacks scalability and fails to address systemic risks for early-stage technologies. Recent advances in LLMs offer new opportunities to augment foresight processes.

Solution

  • Proposed approach: A hybrid foresight workflow combining in-silico agents with human expertise to identify systemic risks using the Futures Wheel method.
  • Novelty:
    1. Development of a scalable pipeline for generating systemic risks with in-silico agents.
    2. Empirical evaluation of agent-generated risks against human-ideated risks across four AI use cases.
    3. Introduction of a structured rubric for evaluating systemic risks based on specificity, novelty, usability, applicability, and diversity.
    4. Proposal of a hybrid governance workflow integrating agents for breadth and humans for contextual judgment.
  • Procedure and key techniques:
    1. Selection of four AI use cases across Technology Readiness Levels (TRLs): Chatbot Companion (TRL 9), AI Toy (TRL 7), Griefbot (TRL 5), and Death App (TRL 2).
    2. Implementation of the Futures Wheel method with six in-silico agents using the Plurals framework to generate cascading consequences.
    3. Classification of consequences into risks and benefits, followed by deduplication to create non-redundant risk lists.
    4. Evaluation of risks by domain experts and leaders using a structured rubric and annotation cards.
    5. Comparison of agent-generated risks with human-ideated risks in both human-only and human-plus-AI conditions.

Results

  • Concrete findings:
    • Agents generated 27–47 unique systemic risks per use case, with 75–93% judged systemic by domain experts.
    • Risks were rated as likely (mean likelihood ≥ 3.57) and severe (mean severity ≥ 3.51).
    • Agent-generated risks were broader and more systemic than human-ideated risks, especially for low-TRL use cases.
  • Advantage over baselines:
    • Agents consistently produced more systemic risks than humans (e.g., 21–32 vs. 2–10 per use case).
    • Agent-generated risks were more novel and detailed for speculative use cases but harder to engage with compared to human-ideated risks.
  • Experiments / evaluation:
    • Three studies involving 290 domain experts and 7 domain leaders evaluated risks across four AI use cases.
    • Metrics included systemic scope, likelihood, severity, specificity, novelty, usability, and applicability.
    • Human-plus-AI collaboration matched or exceeded agent-only outputs in volume but focused on narrower, less systemic risks.
  • Limitations and future work:
    • Limited diversity in risk categories (e.g., few environmental risks).
    • Dependence on specific LLMs; future work could explore alternative models and prompts.
    • Need for longitudinal and intersectional methods to capture lived experiences and long-term impacts.

Summary

This study introduces a hybrid foresight workflow combining in-silico agents and human expertise to identify systemic risks of novel AI technologies. Using the Futures Wheel method, agents broadened the search space, generating numerous systemic risks across four AI use cases with varying technological maturity. Human evaluators provided contextual grounding, emphasizing emotional depth and cultural nuances. The findings highlight the complementary strengths of agents and humans in systemic risk assessment, with implications for AI governance and anticipatory decision-making. Future work should address topical biases, integrate intersectional perspectives, and refine hybrid workflows for broader applicability.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222389/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790712
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Explainable AI (XAI), AI-Assisted Decision-Making & Automation, AI Ethics, Fairness & Accountability, Participatory Design
work
Professions
AI/ML Researchers & Engineers, HCI Researchers, Sociologists & Anthropologists
article
Content Status
Full text indexed
hub
Related Papers
4 related papers