Campus AI vs. Commercial AI: Comparing How Students and Employees Perceive their University’s LLM Chatbot vs. ChatGPT

Human-LLM CollaborationExplainable AI (XAI)AI-Assisted Decision-Making & AutomationUniversity Professors & ResearchersSoftware Engineers & DevelopersHCI Researchers

Paper Title

Campus AI vs. Commercial AI: Comparing How Students and Employees Perceive their University’s LLM Chatbot vs. ChatGPT

Publication Info

  • Topic area: User perceptions of customized LLMaaS chatbots versus commercial LLM chatbots in academic settings.
  • Keywords: LLMaaS, ChatGPT, user trust, privacy, hallucinations, sustainability, academic AI, chatbot customization, trust calibration, human-AI interaction.

Background and Problem

  • Problem / challenge: Existing studies focus on technical adaptations of LLMaaS chatbots but neglect user perceptions, particularly in comparison to commercial alternatives like ChatGPT.
  • Significance: Understanding user perceptions is critical for universities adopting AI systems to ensure trust, privacy, and appropriate use while minimizing risks like hallucinations and privacy concerns.
  • Motivation and related work: Prior research highlights the increasing adoption of LLM chatbots in academia and the importance of trust and privacy in AI systems. However, the impact of user-facing customizations on perceptions and behavior remains underexplored. This study addresses this gap by comparing a university’s customized LLMaaS chatbot to ChatGPT.

Solution

  • Proposed approach: A survey-based field study comparing user perceptions of a university’s customized LLMaaS chatbot and ChatGPT, focusing on trust, privacy, hallucinations, and sustainability-aware AI usage.
  • Novelty:
    1. Demonstrates higher trust, perceived privacy, and fewer perceived hallucinations for the customized LLMaaS chatbot compared to ChatGPT.
    2. Theorizes the role of customization cues, such as branding and interface design, in shaping user perceptions based on the Trustworthiness Assessment Model (TrAM).
    3. Extends the concept of trust calibration to include privacy and hallucinations in academic AI use cases.
    4. Provides practical recommendations for designing and deploying LLMaaS chatbots to align user perceptions with system capabilities.
  • Procedure and key techniques:
    1. Conducted a survey with 526 participants (students and employees) at a German university.
    2. Focused on a subsample of 116 participants who regularly used both the university chatbot and ChatGPT for within-subject comparisons.
    3. Measured trust, privacy concerns, perceived hallucinations, cautious behavior, and sustainability-aware AI usage.
    4. Conducted exploratory benchmark evaluations (TruthfulQA and HaluEval) to assess hallucination tendencies and detection performance.

Results

  • Concrete findings:
    • Trust in the customized LLMaaS chatbot was significantly higher (M = 3.60) than in ChatGPT (M = 3.08, p < 0.001, d = 0.73).
    • Perceived privacy was greater for the customized chatbot (M = 2.38) compared to ChatGPT (M = 3.61, p < 0.001, d = -1.15).
    • Fewer hallucinations were perceived in the customized chatbot (M = 3.23) than in ChatGPT (M = 3.85, p < 0.001, d = -0.72).
    • No significant differences in cautious behavior toward hallucinations or sustainability-aware AI usage.
    • Benchmark evaluations showed the customized chatbot hallucinated more frequently (50% vs. 38% in TruthfulQA) but had slightly better hallucination detection performance (F1-score: 0.7591 vs. 0.7397 in HaluEval).
  • Advantage over baselines:
    • The customized LLMaaS chatbot was perceived as more trustworthy and privacy-friendly, despite using the same underlying LLM technology as ChatGPT.
    • Institutional branding and interface design likely influenced user perceptions positively.
  • Experiments / evaluation:
    • Survey-based field study with paired t-tests for within-subject comparisons.
    • Exploratory benchmark evaluations (TruthfulQA and HaluEval) to assess hallucination tendencies and detection performance.
  • Limitations and future work:
    • Self-selection bias in the subsample of dual-system users.
    • Reliance on self-reported measures; future studies should include behavioral data.
    • Limited generalizability to non-European contexts or domains outside academia.
    • Need for experimental designs to isolate the effects of specific customization cues.

Summary

This study investigates how students and employees perceive a university’s customized LLMaaS chatbot compared to ChatGPT. The customized chatbot was associated with higher trust, greater perceived privacy, and fewer perceived hallucinations, despite benchmark evidence showing higher hallucination rates. These differences highlight the influence of customization cues, such as branding and interface design, on user perceptions. The findings emphasize the importance of aligning user perceptions with system capabilities to ensure calibrated trust and appropriate use. Future research should explore causal relationships between customization cues and user perceptions and extend the findings to other contexts and domains.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/221846/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790622
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Human-LLM Collaboration, Explainable AI (XAI), AI-Assisted Decision-Making & Automation
work
Professions
University Professors & Researchers, Software Engineers & Developers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers