The Language of Approval: Identifying the Drivers of Positive Feedback Online
Authors
Paper Title
The Language of Approval: Identifying the Drivers of Positive Feedback Online
Publication Info
- Topic area: Linguistic drivers of positive feedback in online communities
- Keywords: Reddit, causal inference, linguistic analysis, community feedback, positive reinforcement, predictive modeling, community guidelines, online governance, social computing, HCI
Background and Problem
- Problem / challenge: Despite the centrality of positive feedback mechanisms in online communities, the causal linguistic drivers of such feedback remain unclear, and existing community guidelines often fail to reflect these drivers.
- Significance: Understanding these drivers can inform better community design, moderation workflows, and user guidance systems, improving community health and engagement.
- Motivation and related work: Prior work has focused on surveys and descriptive studies, which are limited by aspirational biases or correlational insights. There is a lack of causal understanding of how linguistic attributes influence positive feedback, limiting actionable interventions.
Solution
- Proposed approach: A three-part methodology combining causal inference, predictive modeling, and guideline audits to identify and leverage linguistic drivers of positive feedback on Reddit.
- Novelty:
- Application of quasi-experimental causal inference to isolate linguistic effects from confounding factors.
- Development of predictive models to identify high-quality posts in real time using linguistic features.
- Audit of community guidelines to reveal gaps between empirical findings and existing rules.
- Procedure and key techniques:
- Analyzed 11M posts from 100 subreddits using a selection-on-observables framework.
- Extracted 100+ linguistic features (e.g., LIWC, semantic styles, toxicity) and controlled for author reputation, timing, and community norms.
- Built global and local predictive models using XGBoost to detect high-quality posts.
- Compared findings with prior surveys and audited subreddit guidelines for alignment with empirical drivers.
Results
- Concrete findings:
- Broad, discussion-generating posts had ≈43% higher odds of high scores, while clear, readable writing increased odds by ≈40%.
- Question-heavy posts (≈30% lower odds) and toxic language (≈5% lower odds) were penalized.
- Specific linguistic features like future focus and causal framing positively influenced feedback.
- Predictive models achieved AUC scores of 0.654 (global) and 0.726 (local), outperforming existing prosociality-based approaches by 12%.
- Advantage over baselines: The curated linguistic feature set outperformed existing prosociality models in predicting high-quality posts, demonstrating stronger alignment with community reward patterns.
- Experiments / evaluation:
- Evaluated causal effects using logistic regression with fixed effects and risk stratification.
- Tested predictive models on global and subreddit-specific datasets.
- Audited guidelines from five high-performing subreddits to assess alignment with empirical findings.
- Limitations and future work:
- Temporal generalizability is limited to a five-month window.
- Visual content and cross-community transferability were not addressed.
- Sensitivity to unmeasured confounders was analyzed but remains a limitation for weaker effects like toxicity.
- Guidelines audit covered only five subreddits, requiring broader analysis.
Summary
This study identifies causal linguistic drivers of positive feedback in online communities using 11M Reddit posts. Broad, clear, and discussion-generating posts were rewarded, while question-heavy and toxic posts were penalized. Predictive models based on these findings outperformed existing approaches, achieving high AUC scores. An audit revealed a "policy-practice gap," as community guidelines often fail to teach empirically-supported strategies. These results suggest actionable paths for designing real-time guidance tools, improving moderation workflows, and revising community guidelines to better reflect what drives positive feedback. Future work should explore temporal stability, cross-platform generalization, and ethical considerations in deploying such systems.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 83%
Timing Matters: Designing Effective Corrections for Short-Form Video Misinformation
CHI '26· Misinformation & Fact-Checking +1
- 71%
Collab: Fostering Critical Identification of Deepfake Videos on Social Media via Synergistic Annotation
CHI '26· Deepfake & Synthetic Media Detection +2
- 71%
Influencers vs. Legacy Media on Instagram: Effects on Perceived Credibility and Following Intention
CHI '26· Social Platform Design & User Behavior +2
- 67%
Microblog Analysis as a Programme of Work
CHI '18· Social Platform Design & User Behavior +1
- 67%
Assessing enactment of content regulation policies: A post hoc crowd-sourced audit of election misinformation on YouTube
CHI '23· Content Moderation & Platform Governance +1
- 63%
Simple changes to content curation algorithms affect the beliefs people form in a collaborative filtering experiment
CHI '26· Social Platform Design & User Behavior +3
Based on Jaccard similarity of research subtopics & professions (≥60%)