Learning from AVA: Early Lessons from a Curated and Trustworthy Generative AI for Policy and Development Research
Authors
Paper Title
Learning from AVA: Early Lessons from a Curated and Trustworthy Generative AI for Policy and Development Research
Publication Info
- Topic area: Generative AI for policy and development research
- Keywords: Generative AI, epistemic humility, policy research, development, citation verifiability, reasoned abstention, multilingual AI, evidence-based synthesis, Humble AI, World Bank Reports
Background and Problem
- Problem / challenge: General-purpose LLMs often produce misinformation and lack mechanisms for verifiable, evidence-based outputs, which is critical for policy and development research.
- Significance: Ensuring trustworthy and verifiable AI outputs is essential for professionals in policy and development to make informed decisions and save time.
- Motivation and related work: Prior work on LLMs has not adequately addressed the need for epistemic humility or the ability to decline unsupported queries. This paper builds on these gaps by introducing a specialized AI system tailored for evidence-based policy research.
Solution
- Proposed approach: AVA (AI + Verified Analysis), a generative AI platform utilizing a curated library of over 4,000 World Bank Reports to provide evidence-based syntheses with multilingual support.
- Novelty:
- Implementation of citation verifiability, allowing claims to be traced back to original sources.
- Introduction of reasoned abstention, where the system declines unsupported queries with explanations and redirections.
- A multi-agent pipeline designed for evidence-based query handling.
- Design guidelines for creating specialized, ecosystem-aware AI systems.
- Procedure and key techniques:
- AVA uses a curated library of World Bank Reports as its knowledge base.
- It employs a multi-agent pipeline to process user queries and generate evidence-based responses.
- Mechanisms for citation verifiability and reasoned abstention are integrated to ensure trustworthiness and clarity.
Results
- Concrete findings: Sustained engagement with AVA was associated with saving 2.4–3.9 hours per week for users.
- Advantage over baselines: AVA provided a specialized “evidence engine” with calibrated trust through institutional provenance and page-anchored citations, addressing limitations of general-purpose LLMs.
- Experiments / evaluation:
- Evaluated in-the-wild with over 2,200 participants from 116 countries.
- Methods included log analysis, surveys, and 20 interviews.
- Limitations and future work: The paper does not specify detailed limitations but suggests further exploration of ecosystem-aware AI and refinement of design guidelines.
Summary
AVA is a generative AI platform designed for policy and development research, leveraging a curated library of World Bank Reports to provide evidence-based, multilingual syntheses. Its key contributions include mechanisms for citation verifiability and reasoned abstention, which enhance trust and usability. Evaluated with over 2,200 users globally, AVA demonstrated time savings of 2.4–3.9 hours weekly and was perceived as a reliable “evidence engine.” The work advances the concept of Humble AI and provides design insights for creating trustworthy, specialized AI systems.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)