PromptHive: Bringing Subject Matter Experts Back to the Forefront with Collaborative Prompt Engineering for Educational Content Creation
Authors
Research Background and Problem
-
What problems or challenges did the authors identify?
With the proliferation of large language models (LLMs), generating high-quality educational content has become increasingly feasible. However, effectively integrating the expertise of Subject Matter Experts (SMEs) into content creation to ensure accuracy and applicability remains a significant challenge. Additionally, designing interfaces that support rapid iteration and experimentation to help experts optimize prompt engineering is another issue that needs to be addressed. -
Why is this problem important?
Existing research has demonstrated the critical role of domain experts in guiding LLM outputs to generate professional content. Education is one of the most promising application areas for AI. By optimizing prompt engineering, the quality of educational content can be improved, the content creation process can be accelerated, and domain experts can retain control and trust in the system. -
Research Motivation and Related Work
The authors' previous research highlighted the complexity of prompt engineering and its impact on the quality of LLM outputs, while also pointing out the unpredictability of generative AI tools and the high metacognitive demands of prompt design. Furthermore, the integration of domain experts and collaborative prompt iteration has not been sufficiently explored in the existing literature.
Solution
-
What methods or solutions did the authors propose?
The authors designed a collaborative prompt creation interface called PromptHive, which connects domain knowledge with generative AI and supports rapid iteration and prompt optimization.
PromptHive employs a four-stage workflow: loading textbook data, creating and iterating prompts, sharing effective prompts, and incorporating multi-level prompt design at both textbook and course levels to achieve highly adaptable content generation. -
What are the innovative aspects of this solution?
- Unlike traditional prompt engineering tools, PromptHive is deeply embedded into experts' existing workflows, supporting rapid experimentation through randomized buttons and a "Prompt Draft Board."
- It provides a "Shared Prompt Library" to facilitate collaborative prompt design, enabling SMEs to draw inspiration from and contribute ideas across courses.
- It includes a backend logging engine to capture how users evolve prompts during collaborative iterations, introducing a new perspective on "dynamic prompt iteration" to the research field.
-
What are the implementation steps and key technologies used?
- Load textbook content: Import structured data sources via links.
- Create and test prompts: Use the Prompt Draft Board to generate multiple prompt versions and conduct model output experiments.
- Optimize and iterate: Quickly adjust content by comparing different prompt versions and their outputs.
- Share effective prompts: Save the best designs in the "Prompt Library" for use and optimization by other experts.
PromptHive also supports structured outputs from GPT-4 and integrates self-consistency prompting techniques to reduce inconsistencies in AI-generated content.
Research Outcomes
-
What specific outcomes were achieved?
- In the field of mathematics, PromptHive significantly reduced experts' cognitive workload (NASA-TLX scores dropped from 55.17 to 26.73) and shortened the time required to generate useful content from "several months" to "a few hours," achieving a 30-fold efficiency improvement.
- The average System Usability Scale (SUS) score was 89/100, indicating high levels of user satisfaction and trust.
-
What advantages does it have compared to existing solutions?
- PromptHive enables non-AI experts to create content matching the quality of manually crafted materials while significantly reducing labor input.
- Through collaboration, SMEs demonstrated greater creativity and lower cognitive burdens during the process of sharing and customizing prompts.
-
What were the experimental or evaluation results?
- Learning gain experiments showed that prompts generated using PromptHive were equally effective in learning outcomes compared to human-written content (average learning gains of 8.13% and 7.47%, respectively, with no statistically significant difference).
- Compared to manual prompt engineering, PromptHive significantly reduced the time required and the frequency of repetitive tasks.
-
Limitations and Future Directions
- The quality of prompts generated by PromptHive depends on the underlying model (e.g., GPT-4), which carries risks such as bias and hallucinations, necessitating cautious deployment in large-scale educational settings.
- Although experts felt that PromptHive preserved domain control, future iterations could expand its functionality to support direct editing of generated content and increase expert involvement in "system-level prompt" design.
- The current research primarily focuses on mathematics; future studies could explore its applicability to other subjects, educational systems, and even other domains.
Conclusion
PromptHive is an innovative collaborative prompt engineering tool that enables the rapid creation of high-quality educational content while allowing experts to maintain a high degree of control. By facilitating collaborative optimization of the prompt iteration process and significantly improving work efficiency, it demonstrates a new model of human-AI collaboration. It also provides a reference pathway for designing "human-centered yet highly automated" educational tools. Future research could further optimize system functionality and application scenarios by expanding subject applicability and deepening expert involvement.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can domain experts' knowledge be effectively integrated into generative AI content-creation workflows?Category: LLM Learning Scaffolding and Reflection SupportSimilar questionsarrow_forward
- How can interfaces supporting rapid iteration and experimentation help domain experts optimize prompt engineering?Category: LLM Learning Scaffolding and Reflection SupportSimilar questionsarrow_forward
- Can PromptHive significantly reduce domain experts' cognitive load and improve efficiency of educational content generation?Category: LLM Learning Scaffolding and Reflection SupportSimilar questionsarrow_forward
Practical Problems
1- Domain experts struggle to participate efficiently in AI-generated educational content; prompt design is time-consuming and complex.Category: LLM Learning Scaffolding and Reflection SupportSimilar questionsarrow_forward
- 67%
AutoPBL: An LLM-powered Platform to Guide and Support Individual Learners Through Self Project-based Learning
CHI '25· Human-LLM Collaboration +2
- 60%
Computational Interaction: Theory and Practice
CHI '18· Programming Education & Computational Thinking +1
- 60%
Developing a Community of Practice to Support Global HCI Education
CHI '18· Programming Education & Computational Thinking +1
- 60%
Cody: An AI-Based System to Semi-Automate Coding for Qualitative Research
CHI '21· Human-LLM Collaboration +1
- 60%
PaperWeaver: Enriching Topical Paper Alerts by Contextualizing Recommended Papers with User-collected Papers
CHI '24· Human-LLM Collaboration +1
- 60%
CodeAid: Evaluating a Classroom Deployment of an LLM-based Programming Assistant that Balances Student and Educator Needs
CHI '24· Human-LLM Collaboration +1
- 60%
Evaluating Large Language Models on Academic Literature Understanding and Review: An Empirical Study among Early-stage Scholars
CHI '24· Human-LLM Collaboration +1
- 60%
Understanding the Role of Large Language Models in Personalizing and Scaffolding Strategies to Combat Academic Procrastination
CHI '24· Human-LLM Collaboration +1
- 60%
Teach AI How to Code: Using Large Language Models as Teachable Agents for Programming Education
CHI '24· Human-LLM Collaboration +1
- 60%
Playing Dumb to Get Smart: Creating and Evaluating an LLM-based Teachable Agent within University Computer Science Classes
CHI '25· Human-LLM Collaboration +1
Based on Jaccard similarity of research subtopics & professions (≥60%)