Forage: Understanding LLM-facilitated sensemaking of conversation data
Authors
Paper Title
Forage: Understanding LLM-facilitated sensemaking of conversation data
Publication Info
- Topic area: LLM-enabled tools for exploratory analysis of unstructured conversation data.
- Keywords: LLMs, RAG systems, exploratory search, sensemaking, conversation analysis, thematic analysis, interpretive facets, user studies, AI in journalism, AI in government.
Background and Problem
- Problem / challenge: Existing retrieval-augmented generation (RAG) systems and LLMs are widely used for information retrieval and synthesis but lack robust evaluation in open-ended, subjective sensemaking tasks, such as analyzing conversation data. There is also concern about potential homogenization and bias in LLM-generated insights.
- Significance: Understanding how LLMs can support exploratory sensemaking is critical for domains like journalism, government, and research, where unstructured data analysis is central. This has implications for efficiency, insight generation, and mitigating biases in knowledge work.
- Motivation and related work: Prior work in exploratory search and sensemaking has identified opportunities for technology to support iterative and interpretive processes. However, the integration of LLMs into these workflows remains underexplored, particularly in contexts requiring subjective interpretation. This paper builds on theories of sensemaking and exploratory search, addressing gaps in understanding user needs and LLMs' role in subjective analysis.
Solution
- Proposed approach: Forage, a RAG-based sensemaking tool, facilitates exploratory analysis of conversation data by combining retrieval and LLM-generated synthesis. Wild Forage extends this by introducing perspective-driven interpretive facets to generate multiple thematic interpretations.
- Novelty:
- Development of Forage, a RAG-based system for exploratory sensemaking of conversation data.
- Introduction of Wild Forage, a design provocation that generates multiple interpretations of data along specified axes (e.g., political orientation).
- Empirical insights into user needs and practices through multi-part user studies.
- A taxonomy of user query types in LLM-enabled exploratory search.
- Comparative study of LLM-enabled sensemaking versus traditional text search.
- Procedure and key techniques:
- Forage employs a four-stage pipeline: retrieval of relevant speaker turns, ranking by similarity, LLM synthesis with citations, and re-sorting based on thematic salience.
- Wild Forage uses LLM prompts to generate thematic analyses from specified perspectives (e.g., Democrat vs. Republican).
- User studies involved semi-structured interviews, think-aloud protocols, and iterative design changes informed by user feedback.
Results
- Concrete findings:
- Forage supported thematic generation (36% of queries) and synthesis (15%), with users favoring LLM-generated insights over pure retrieval.
- Wild Forage revealed differences in linguistic framing and citation focus across interpretive facets (e.g., Democrat vs. Republican).
- In comparative studies, LLM-enabled Forage was judged easier to use, less mentally demanding, and more effective at generating new ideas than text search.
- Advantage over baselines:
- Forage provided structure and surfaced novel insights compared to manual or search-only methods.
- Wild Forage enabled users to reflect on biases by presenting alternative thematic framings.
- Experiments / evaluation:
- User studies with 27 participants from NPR, the City of Durham, a European football club, and a US political campaign.
- Comparative study with five trained sensemakers analyzing conversation data using both search-only and LLM-enabled Forage.
- Computational analysis of linguistic and citation differences in Wild Forage's perspective-driven facets.
- Limitations and future work:
- Risks of misrepresentation and stereotyping in generating interpretive facets.
- Limited generalizability to non-US contexts and other languages.
- Challenges in validating LLM-generated outputs and ensuring appropriate use.
- Future work includes exploring computational methods for identifying axes of interpretive difference and studying human interaction with multiple perspectives.
Summary
This paper introduces Forage, a RAG-based tool for exploratory analysis of conversation data, and Wild Forage, a design provocation for generating multiple thematic interpretations. User studies across journalism, government, and research contexts demonstrate Forage's utility in structuring insights and surfacing novel themes. Comparative studies confirm that LLM-enabled Forage outperforms text search in ease of use and idea generation. Wild Forage highlights the potential of perspective-driven facets to counteract bias and enrich sensemaking, though risks of misrepresentation remain. These findings inform the design of LLM-powered tools for subjective, open-ended analysis tasks.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)