Live in the Loop: Rapid Run-time Feedback for Prompts
Authors
Paper Title
Live in the Loop: Rapid Run-time Feedback for Prompts
Publication Info
- Topic area: AI-assisted programming with live feedback mechanisms.
- Keywords: LLMs, live programming, run-time feedback, code generation, exploratory programming, microversioning, prompt refinement, user study, AI assistants, IDE extensions.
Background and Problem
- Problem / challenge: Current AI coding assistants like GitHub Copilot lack mechanisms to provide immediate run-time feedback for generated code, leading to inefficiencies in refining prompts and validating code suggestions.
- Significance: Addressing the "gulf of envisioning" and "gulf of evaluation" in AI-assisted programming can improve the accuracy and intentionality of code generation, reducing errors and iterations.
- Motivation and related work: Prior work on live programming environments and tools like LEAP has demonstrated the benefits of immediate feedback for evaluating code changes. However, these systems focus on code-level interactions rather than run-time results, leaving gaps in prompt refinement and exploration of multiple solutions.
Solution
- Proposed approach: ReFiQ (“Result-first Queries”), a live programming tool that prioritizes run-time results over code and displays multiple results for a single prompt.
- Novelty:
- Presentation of run-time values before code suggestions.
- Support for exploring multiple interpretations of prompts simultaneously.
- Integration of live programming mechanisms like probes to anchor interactions at the run-time data level.
- Encouragement of intentional prompt refinement through immediate feedback.
- Procedure and key techniques:
- Users interact with probes to view live run-time values.
- Prompts are issued directly referencing run-time data.
- Multiple code suggestions are generated and executed automatically, displaying their run-time results first.
- Users select desired behavior based on run-time effects and progressively disclosed code changes.
- The system supports microversioning to track and revert changes.
Results
- Concrete findings:
- ReFiQ reduced manual code iterations by 49% compared to GitHub Copilot.
- Participants issued 9% more prompts and spent 56% more time refining prompts in ReFiQ.
- Programs created with ReFiQ had significantly fewer errors (p = 0.022) and higher output quality.
- Advantage over baselines:
- Faster feedback loops for prompt refinement.
- Higher trust in correctness and output of generated code.
- Better exploration of solution space through multiple run-time results.
- Experiments / evaluation:
- Exploratory user study with 8 participants using a within-subjects design.
- Tasks: OCR preprocessing and topography segmentation.
- Metrics: task completion, output quality, interaction logs, surveys, and thematic analysis.
- Limitations and future work:
- Limited validity outside visual domains like image processing.
- Small sample size (n=8) affects statistical power.
- Lack of support for undo/redo of prompts and backtracking across iterations.
- Future work could explore integrating prompts as first-class programming artifacts and enhancing explanations for code suggestions.
Summary
ReFiQ introduces a result-first approach to AI-assisted programming, enabling users to refine prompts and validate code suggestions through immediate run-time feedback. The tool demonstrated reduced manual iterations, higher output quality, and improved exploration of solution space compared to GitHub Copilot. While the study highlighted trade-offs in code comprehension and generalization, ReFiQ's design advances live programming workflows by prioritizing run-time results and supporting intentional decision-making. Future work could address limitations in backtracking and expand applicability to non-visual domains.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
DALL: Data Labeling via Data Programming and Active Learning Enhanced by Large Language Models
CHI '26· Human-LLM Collaboration +3
- 100%
Data-Prompt Co-Evolution: Growing Test Sets to Refine LLM Behavior
CHI '26· Human-LLM Collaboration +3
- 88%
Interview-Informed Generative Agents for Product Discovery: A Validation Study
CHI '26· Human-LLM Collaboration +3
- 88%
"Shall We Dig Deeper?": Designing and Evaluating Strategies for LLM Agents to Advance Knowledge Co-Construction in Asynchronous Online Discussions
CHI '26· Human-LLM Collaboration +3
- 88%
Just-In-Time Objectives: A General Approach for Specialized AI Interactions
CHI '26· Human-LLM Collaboration +3
- 88%
TurnStyle: A Framework for Analyzing Human Conversational Behaviors to Predict Success in LLM-Assisted Tasks
CHI '26· Human-LLM Collaboration +3
- 75%
Perspectra: Choosing Your Experts Enhances Critical Thinking in Multi-Agent Research Ideation
CHI '26· Human-LLM Collaboration +3
- 75%
"Here, Let Me Help": An Empirical Study of User Interventions in Human–Web Agent Collaboration
CHI '26· Human-LLM Collaboration +3
- 75%
Designing Staged Evaluation Workflows for LLMs: Integrating Domain Experts, Lay Users, and Model-Generated Evaluation Criteria
CHI '26· Human-LLM Collaboration +3
- 75%
The Invisible Mentor: Inferring User Actions from Screen Recordings to Recommend Better Workflows
CHI '26· Human-LLM Collaboration +3
Based on Jaccard similarity of research subtopics & professions (≥60%)