Campus AI vs. Commercial AI: Comparing How Students and Employees Perceive their University’s LLM Chatbot vs. ChatGPT
Authors
Paper Title
Campus AI vs. Commercial AI: Comparing How Students and Employees Perceive their University’s LLM Chatbot vs. ChatGPT
Publication Info
- Topic area: User perceptions of customized LLMaaS chatbots versus commercial LLM chatbots in academic settings.
- Keywords: LLMaaS, ChatGPT, user trust, privacy, hallucinations, sustainability, academic AI, chatbot customization, trust calibration, human-AI interaction.
Background and Problem
- Problem / challenge: Existing studies focus on technical adaptations of LLMaaS chatbots but neglect user perceptions, particularly in comparison to commercial alternatives like ChatGPT.
- Significance: Understanding user perceptions is critical for universities adopting AI systems to ensure trust, privacy, and appropriate use while minimizing risks like hallucinations and privacy concerns.
- Motivation and related work: Prior research highlights the increasing adoption of LLM chatbots in academia and the importance of trust and privacy in AI systems. However, the impact of user-facing customizations on perceptions and behavior remains underexplored. This study addresses this gap by comparing a university’s customized LLMaaS chatbot to ChatGPT.
Solution
- Proposed approach: A survey-based field study comparing user perceptions of a university’s customized LLMaaS chatbot and ChatGPT, focusing on trust, privacy, hallucinations, and sustainability-aware AI usage.
- Novelty:
- Demonstrates higher trust, perceived privacy, and fewer perceived hallucinations for the customized LLMaaS chatbot compared to ChatGPT.
- Theorizes the role of customization cues, such as branding and interface design, in shaping user perceptions based on the Trustworthiness Assessment Model (TrAM).
- Extends the concept of trust calibration to include privacy and hallucinations in academic AI use cases.
- Provides practical recommendations for designing and deploying LLMaaS chatbots to align user perceptions with system capabilities.
- Procedure and key techniques:
- Conducted a survey with 526 participants (students and employees) at a German university.
- Focused on a subsample of 116 participants who regularly used both the university chatbot and ChatGPT for within-subject comparisons.
- Measured trust, privacy concerns, perceived hallucinations, cautious behavior, and sustainability-aware AI usage.
- Conducted exploratory benchmark evaluations (TruthfulQA and HaluEval) to assess hallucination tendencies and detection performance.
Results
- Concrete findings:
- Trust in the customized LLMaaS chatbot was significantly higher (M = 3.60) than in ChatGPT (M = 3.08, p < 0.001, d = 0.73).
- Perceived privacy was greater for the customized chatbot (M = 2.38) compared to ChatGPT (M = 3.61, p < 0.001, d = -1.15).
- Fewer hallucinations were perceived in the customized chatbot (M = 3.23) than in ChatGPT (M = 3.85, p < 0.001, d = -0.72).
- No significant differences in cautious behavior toward hallucinations or sustainability-aware AI usage.
- Benchmark evaluations showed the customized chatbot hallucinated more frequently (50% vs. 38% in TruthfulQA) but had slightly better hallucination detection performance (F1-score: 0.7591 vs. 0.7397 in HaluEval).
- Advantage over baselines:
- The customized LLMaaS chatbot was perceived as more trustworthy and privacy-friendly, despite using the same underlying LLM technology as ChatGPT.
- Institutional branding and interface design likely influenced user perceptions positively.
- Experiments / evaluation:
- Survey-based field study with paired t-tests for within-subject comparisons.
- Exploratory benchmark evaluations (TruthfulQA and HaluEval) to assess hallucination tendencies and detection performance.
- Limitations and future work:
- Self-selection bias in the subsample of dual-system users.
- Reliance on self-reported measures; future studies should include behavioral data.
- Limited generalizability to non-European contexts or domains outside academia.
- Need for experimental designs to isolate the effects of specific customization cues.
Summary
This study investigates how students and employees perceive a university’s customized LLMaaS chatbot compared to ChatGPT. The customized chatbot was associated with higher trust, greater perceived privacy, and fewer perceived hallucinations, despite benchmark evidence showing higher hallucination rates. These differences highlight the influence of customization cues, such as branding and interface design, on user perceptions. The findings emphasize the importance of aligning user perceptions with system capabilities to ensure calibrated trust and appropriate use. Future research should explore causal relationships between customization cues and user perceptions and extend the findings to other contexts and domains.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
Understanding and Supporting Peer Review Using AI-reframed Positive Summary
CHI '25· Human-LLM Collaboration +2
- 71%
OmniQuery: Contextually Augmenting Captured Multimodal Memories to Enable Personal Question Answering
CHI '25· Human-LLM Collaboration +2
- 71%
Invisible Saboteurs: Sycophantic LLMs Mislead Novices in Problem-Solving Tasks
CHI '26· Human-LLM Collaboration +2
- 71%
The Impact of Response Latency and Task Type on Human-LLM Interaction and Perception
CHI '26· Human-LLM Collaboration +2
- 71%
Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
CHI '26· Human-LLM Collaboration +2
- 71%
A Multimodal Investigation of Controllability and Cognitive Load in Interactive Machine Learning
IUI '26· Human-LLM Collaboration +2
- 67%
Understanding the Role of Large Language Models in Personalizing and Scaffolding Strategies to Combat Academic Procrastination
CHI '24· Human-LLM Collaboration +1
- 67%
SummAct: Uncovering User Intentions Through Interactive Behaviour Summarisation
CHI '25· Human-LLM Collaboration +1
- 67%
The Influence of Curiosity Traits and On-Demand Explanations in AI-Assisted Decision-Making
IUI '25· Explainable AI (XAI) +1
- 63%
DiLLS: Interactive Diagnosis of LLM-based Multi-agent Systems via Layered Summary of Agent Behaviors
CHI '26· Human-LLM Collaboration +3
Based on Jaccard similarity of research subtopics & professions (≥60%)