Take It, Leave It, or Fix It: Measuring Productivity and Trust in Human-AI Collaboration
Authors
Although recent developments in generative AI have greatly enhanced the capabilities of conversational agents such as Google's Bard or OpenAI's ChatGPT, it's unclear whether the usage of these agents aids users across various contexts. To better understand how access to conversational AI affects productivity and trust, we conducted a mixed-methods, task-based user study, observing 76 software engineers (N=76) as they completed a programming exam with and without access to Bard. Effects on performance, efficiency, satisfaction, and trust vary depending on user expertise, question type (open-ended "solve" questions vs. definitive "search" questions), and measurement type (demonstrated vs. self-reported). Our findings include evidence of automation complacency, increased reliance on the AI over the course of the task, and increased performance for novices on “solve”-type questions when using the AI. We discuss common behaviors, design recommendations, and impact considerations to improve collaborations with conversational AI.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How does GenAI affect users' productivity (e.g., performance, efficiency, satisfaction) in programming tasks?Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
- How has users' trust in GenAI changed over time and across task contexts?Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
- What behavioral differences do experts and novices show when using GenAI-assisted tools?Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
Practical Problems
1- When programming with GenAI, users struggle to identify erroneous suggestions, and misplaced trust reduces efficiency.Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
- 100%
Discovering the Syntax and Strategies of Natural Language Programming with Generative Language Models
CHI '22· Generative AI (Text, Image, Music, Video) +1
- 100%
"What It Wants Me To Say": Bridging the Abstraction Gap Between End-User Programmers and Code-Generating Large Language Models
CHI '23· Generative AI (Text, Image, Music, Video) +1
- 100%
D-Twins: Your Digital Twin Designed for Real-Time Boredom Intervention
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 80%
Method for Exploring Generative Adversarial Networks (GANs) via Automatically Generated Image Galleries
CHI '21· Generative AI (Text, Image, Music, Video) +2
- 80%
MUD: Towards a Large-Scale and Noise-Filtered UI Dataset for Modern Style UI Modeling
CHI '24· Generative AI (Text, Image, Music, Video) +1
- 80%
Ivie: Lightweight Anchored Explanations of Just-Generated Code
CHI '24· Generative AI (Text, Image, Music, Video) +2
- 80%
Generative AI Uses and Risks for Knowledge Workers in a Science Organization
CHI '25· Generative AI (Text, Image, Music, Video) +2
- 80%
How the Role of Generative AI Shapes Perceptions of Value in Human-AI Collaborative Work
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 80%
Exploring Empty Spaces: Human-in-the-Loop Data Augmentation
CHI '25· Generative AI (Text, Image, Music, Video) +2
- 80%
GenComUI: Exploring Generative Visual Aids as Medium to Support Task-Oriented Human-Robot Communication
CHI '25· Generative AI (Text, Image, Music, Video) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)