Cody: An AI-Based System to Semi-Automate Coding for Qualitative Research
Authors
Document Title
Cody: An AI-Based System to Semi-Automate Coding for Qualitative Research
Document Information
- Subject Areas: Human-Computer Interaction (HCI), Qualitative Data Analysis (QDA), AI-Assisted Tool Design
- Keywords: Qualitative research, qualitative coding, rule-based coding, supervised machine learning, user-centered design, artifact design
Research Background and Issues
-
Problems or Challenges:
- The coding process in qualitative research is highly time-consuming and repetitive, especially for large datasets.
- Existing qualitative data analysis systems (QDAS) have limited machine learning functionalities and lack interactivity and transparency, hindering the adoption of automated coding techniques.
- There is a lack of studies on the interaction between qualitative researchers and AI-assisted tools, leading to issues of trust and ineffective utilization of these tools.
-
Importance:
- Qualitative research is widely used to answer "what," "how," and "why" questions, and the quality of coding determines the reliability and applicability of subsequent theory development.
- As dataset sizes grow, effective qualitative coding tools are crucial for ensuring consistency and reducing workload.
-
Research Motivation and Related Work:
- Literature suggests that interactive coding tools designed with user-centered principles are more likely to be accepted by researchers.
- AI technology holds great potential in qualitative coding but currently faces challenges in usability and user trust.
- The academic community calls for the development of interactive, transparent, and user-friendly AI-assisted tools that integrate rule definition and machine learning model training.
Solution
-
Method or Solution:
- Proposed and designed an interactive AI system named Cody, which semi-automates qualitative coding through rule definition and supervised machine learning.
- Cody allows users to interactively define and modify coding rules and extends manual coding to unseen data.
-
Innovations:
- Integrates rule-based coding with machine learning model training, directly incorporating rule definition into the coding workflow.
- Provides code suggestions and explanations, with a streamlined and transparent interface to alleviate user concerns about system complexity.
-
Implementation Steps and Techniques:
- Defined six system design requirements and developed Cody based on these, including support for unit analysis selection, rule definition and modification, transparent model training, and suggestion explanations.
- The rule generator creates initial coding rules based on semantic similarity and Levenshtein distance.
- Supervised learning using a logistic regression model is employed to update data classification in real-time based on manual annotations.
- Introduced a learning method designed to address the "cold start problem," including the generation of artificial negative examples.
Research Outcomes
-
Specific Outcomes:
- Cody helps users define coding rules, encourages reflection on coding methods, and improves coding consistency.
- Automated suggestions assist users in identifying uncovered text segments and support iterative refinement of coding rules.
- Enhances transparency and structure in qualitative coding, making the coding process easier for third parties to understand.
-
Advantages Compared to Existing Solutions:
- Compared to traditional QDAS tools (e.g., MAXQDA), Cody improves coding consistency and quality (Krippendorff’s Alpha increased from 0.085 to 0.33).
- Provides multi-level support, including rule suggestions for seen data and machine learning predictions for unseen data.
-
Experimental or Evaluation Results:
- Participants found Cody beneficial for repetitive data coding tasks, facilitating learning of coding rules and providing an overall view of the document.
- Evaluation revealed user interest in rule definition and iteration processes, while interest in machine learning suggestions was lower; however, ML suggestions showed potential in enhancing coding rules.
-
Limitations and Future Directions:
- The small dataset size limits comprehensive evaluation of the tool's utility; long-term field evaluations in real research scenarios are recommended.
- The quality of the machine learning model is still constrained by the number of coding samples, suggesting further optimization of model training strategies.
- Multi-user collaborative coding and the impact of different coding styles on tool usage were not considered, requiring further research on the tool's performance and effects in collaborative scenarios.
- Future studies could explore the use of technologies like eye-tracking to identify unannotated coding segments and evaluate how to better encourage users to explore coding rules and automated suggestions.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can an interactive AI system semi-automate the coding process in qualitative research?Category: Transparency, Auditability, and Trust Calibration MechanismsSimilar questionsarrow_forward
- How can rule definition and supervised machine learning be effectively combined to improve qualitative coding efficiency and consistency?Category: Transparency, Auditability, and Trust Calibration MechanismsSimilar questionsarrow_forward
- How do users' trust and transparency in AI-assisted tools affect their acceptance in qualitative data analysis?Category: Transparency, Auditability, and Trust Calibration MechanismsSimilar questionsarrow_forward
Practical Problems
1- Qualitative research data coding is time-consuming and repetitive, and existing tools struggle to balance transparency and efficiency.Category: Transparency, Auditability, and Trust Calibration MechanismsSimilar questionsarrow_forward
- 100%
ReviewFlow: Intelligent Scaffolding to Support Academic Peer Reviewing
IUI '24· Human-LLM Collaboration +1
- 80%
Are We On Track? AI-Assisted Active and Passive Goal Reflection During Meetings
CHI '25· Human-LLM Collaboration +2
- 80%
Large Language Models in Qualitative Research: Uses, Tensions, and Intentions
CHI '25· Human-LLM Collaboration +2
- 80%
IdeaSynth: Iterative Research Idea Development Through Evolving and Composing Idea Facets with Literature-Grounded Feedback
CHI '25· Human-LLM Collaboration +2
- 80%
Underreporting of AI Use: The Role of Social Desirability Bias
CHI '26· Human-LLM Collaboration +2
- 80%
LLM-based In-situ Thought Exchanges for Critical Paper Reading
IUI '26· Human-LLM Collaboration +2
- 80%
DiscipLink: Unfolding Interdisciplinary Information Seeking Process via Human-AI Co-Exploration
UIST '24· Human-LLM Collaboration +2
- 75%
How to Write CHI Papers -- Second Edition
CHI '18· User Research Methods (Interviews, Surveys, Observation)
- 75%
3rd Early Career Development Symposium
CHI '18· User Research Methods (Interviews, Surveys, Observation)
- 75%
Introduction to Human-Computer Interaction
CHI '18· User Research Methods (Interviews, Surveys, Observation)
Based on Jaccard similarity of research subtopics & professions (≥60%)