Interaction safety
Longitudinal evaluation
Longitudinal evaluation is the assessment of an AI system by observing the same people using it over an extended period, rather than by scoring isolated responses. It is the method that can see cumulative effects — trust, dependency, attachment, influence — which snapshot benchmarks are structurally unable to detect.
Also called long-term AI evaluation, in-the-wild evaluation
Benchmarks answer a specific question well: given this input, how good is the output. They are reproducible, comparable, and cheap. They are also blind to everything that takes time to happen, which is most of what matters about a system somebody talks to every day.
Longitudinal work asks a different question: what happens to this person, and to this relationship, over months of contact. It surfaces the slow effects — a trust level that drifted upward without being tested, a pattern of agreement that eroded someone’s willingness to be challenged, a reliance that displaced something else — none of which appear in any individual exchange.
The difficulty is real. Longitudinal study is slow, expensive, hard to control, and involves the most sensitive material anyone produces. It also cannot be done credibly without access to genuine ongoing relationships, which is why so little of it exists relative to how much these systems are used.
Why it matters
The effects that most warrant attention in relational AI are precisely the ones the dominant evaluation methods cannot see. Without longitudinal work, the field is measuring what is convenient rather than what matters.
Where TAICU stands
TAICU operates HUG partly to make this kind of study possible: a real system, in real use, over real time, treated as a site for forming better research questions rather than as evidence in itself.
Common questions
- Why do AI benchmarks miss so much?
- They score responses in isolation. Cumulative effects — how trust develops, what reliance displaces, how a manner shapes someone over months — do not exist at the level of a single response and cannot be measured there.
- What does longitudinal AI research require?
- Access to genuine ongoing relationships over long periods, careful handling of unusually sensitive material, and a tolerance for slow, expensive work that produces fewer clean comparisons than benchmarking does.