X6.04.3Longitudinal field evaluation of social robotsdesignresearch

Judging whether a social robot succeeds requires evaluating it over a long deployment, not a short trial

Aliases: longitudinal HRI · diary study · long-term evaluation

What it is

Judging whether a social robot is actually accepted and delivers long-term value requires a longitudinal field evaluation — placing the robot in a real use environment and observing it for weeks to years. Studies that measure only first contact or a few days of data systematically overstate acceptance.

Why it happens

This is the methodological consequence of the novelty effect producing inflated early data. Since the novelty effect inflates engagement and satisfaction early in a deployment, any study that measures only the first few days is mostly capturing the novelty effect itself rather than the product's durable value. That is a confounding-variable problem in the methodological sense, not something a bigger sample fixes — the measurement window itself is drawn from the wrong phase, and no amount of extra sample collected inside the novelty period changes that. There is a second confound underneath it: a cross-sectional comparison — different users measured at the same point in time — cannot separate "people simply differ from each other" from "the same person changed over time," and easily conflates two unrelated causes into one conclusion. And as long as researchers are still present to remind participants or to make maintenance visits, use may be sustained by that extra attention; once the study ends and the reminders and visits stop, natural use can drop sharply to a lower level.

Studying it

This fact is itself an account of how to study the question. Longitudinal deployments typically place a robot in homes, schools, or care institutions for weeks to years, tracking how use patterns change through usage logs, periodic questionnaires, scheduled interviews, and diary studies. Compared to lab studies, this approach runs into a few concrete difficulties: attrition, where participants drop out partway through and later data volume collapses; uncontrolled context, since household routines, seasons, and changes in who lives in the home all affect use but are hard to log and rule out; and cost and duration that are far higher than a lab study — which is the actual reason long-term deployments are rarer, not a lack of methodological know-how. A common sampling strategy is denser early logging (weekly) shifting to sparser later logging (monthly), which models the adaptation process itself rather than arbitrarily discarding the first few weeks as "contaminated by novelty" — that discarded window is exactly where the information about the decay rate lives, and is analytic material rather than noise.

Where it stops holding

Results from a longitudinal deployment are highly context-dependent: home deployments differ from institutional ones, and a single user having exclusive access differs from multiple users sharing one robot, so a trajectory measured in one setting cannot simply be transplanted onto another. When budget or time genuinely rules out a full longitudinal deployment, the reporting should at minimum flag that "this result is based on N days of data and may be influenced by the novelty effect, and does not represent long-term adoption," rather than presenting short-term data as a long-term conclusion for product decisions. Research-provided hardware, typically free with prompt on-call repair, also limits external validity, so retention measured under that support cannot be taken as a forecast for what an ordinary buyer would experience; short controlled experiments remain irreplaceable for isolating a single causal mechanism and are not simply superseded by longitudinal deployment — the two answer different questions.

Applying it

Decisions about whether to keep investing in a robot or scale up its deployment should draw on longitudinal data rather than launch-trial or press-demo data. When a full longitudinal deployment is not feasible, a fallback check is a "novelty removed" comparison: let a subset of users interact with the robot unsupervised, with no extra prompting, for two to three weeks — long enough for the novelty effect to plausibly have decayed — before collecting their evaluation, rather than gathering feedback on launch day and treating it as a long-term conclusion. Triangulate logs against experience interviews, measured benefit, and bystander accounts (family, staff) instead of letting one interaction-count metric stand for "success" — the three often disagree about whether use has been abandoned or merely scaled back, and that disagreement is a signal worth chasing, not noise to average away.

Related

  • Same group: X6.04.1 Early enthusiasm for a social robot does not predict whether people keep using it · X6.04.2 A social robot that provides no lasting practical value gets abandoned once the novelty fades
  • Nearby: X6.05 Use by special populations · X4.07 Operator situation awareness
  • Search terms: longitudinal field evaluation · diary study · long-term HRI · novelty effect

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/X6.04.3