Interaction safety

Interaction safety is the study of harms that arise from the ongoing relationship between a person and an AI system rather than from any single output. It covers influence, dependency, attachment, misplaced trust, and emotional failure — risks that only become visible across time and in context.

Also called relational AI safety, safety in human-AI interaction

Most AI safety work evaluates outputs: whether a model will produce dangerous content, whether it can be jailbroken, whether it is accurate. That work is necessary and it is not sufficient for systems people live alongside, because the harms specific to those systems are not located in any individual response.

A system can pass every content evaluation and still, over six months, make someone less able to tolerate disagreement, more isolated, or more confident in beliefs nobody has tested. Each response along that path was safe. The trajectory was not, and nothing that looks at responses one at a time can see it.

This requires different methods. The unit of analysis is the relationship rather than the exchange; the timescale is months rather than turns; and the relevant evidence often lives in what people do outside the conversation, which makes it correspondingly harder to collect responsibly.

Why it matters

Relational systems are being deployed to millions of people while the evaluation apparatus for their characteristic risks barely exists. The gap is not in willingness to measure but in method.

Where TAICU stands

Interaction safety is one of TAICU’s three research pillars. The lab’s position is that these risks cannot be evaluated from benchmarks alone and require studying real relationships over time, which is a large part of why it operates a system of its own.

Common questions

How is interaction safety different from AI safety?
Conventional AI safety evaluates outputs — whether a model produces harmful content. Interaction safety examines harms that emerge from an ongoing relationship, such as dependency, misplaced trust, and influence, none of which are visible in a single response.
Why can’t benchmarks measure interaction safety?
Benchmarks score responses in isolation. The harms here are properties of a trajectory across months, where every individual response can be entirely defensible while the overall pattern is not.

All terms in the research glossary