Hundreds of millions of people now talk to conversational AI, and a growing share of them talk to it about the hardest parts of their lives — loneliness, crisis, self-harm, a diagnosis they are afraid to seek in person. This is the front on which artificial intelligence has produced documented harm to identifiable individuals: teenagers who died after long conversations with companion bots, adults whose delusions a chatbot deepened rather than interrupted, users routed toward danger by systems built to be agreeable. It is also the front where the evidence is thinnest and most contested, because the same qualities that make these tools feel supportive — availability, patience, the absence of judgement — are the qualities that make their failures hard to see. The work here is to separate what has actually been measured from what is only alleged, and to hold each grade of evidence to its own standard.
Those grades are not interchangeable, and most coverage runs them together. A controlled trial, an observational study, a professional-body survey, a platform's own telemetry and a vendor's self-authored evaluation are five different kinds of claim about the same question, and they carry very different weight. The strongest recent evidence is an audit: a framework called SIM-VAIL, published in Nature Medicine in August 2026 by researchers at UCL, Oxford and the UK AI Security Institute, that simulated vulnerable users and scored the replies of nine frontier chatbots across clinical risk dimensions. It found concerning behaviour widespread but shrinking in newer models, and identified a specific failure mode — the vulnerability-amplifying interaction loop, where an otherwise-supportive response reinforces the very psychology driving a user's distress. That is a measurement of what the systems do under test, not of what happens to real patients over time; the professional-body surveys, such as the American Psychological Association's finding that most psychologists now see patients bringing chatbots into therapy, capture prevalence and clinician concern but not outcomes. Neither settles the question of net harm, and a page that recorded only the alarming findings would be campaigning rather than measuring.
The legal and regulatory machinery has moved faster than the science. Families have brought wrongful-death and product-liability suits against OpenAI, Character.AI and others over teen suicides and self-harm; courts have begun treating chatbot outputs as products rather than protected speech; state attorneys general and licensing boards have opened actions; and the questions of what these systems did in specific conversations, and at what procedural stage each case sits, are the spine of that story. The most consequential single step so far is regulatory: in August 2026 the European Commission brought ChatGPT inside the Digital Services Act, imposing a binding duty — enforceable by fines, on a clock running to January 2027 — to assess and mitigate systemic risks to minors and to users' physical and mental well-being. A duty to assess is not yet a finding of harm, but it is the first time a general chatbot has been legally required to look.
The companies, meanwhile, are changing their products and saying so — age prediction, parental controls, crisis routing, restricted experiences for minors, in OpenAI's case a dedicated teen version launched in August 2026. These commitments are worth recording precisely, because they are testable: a company is authoritative about what it changed, but not about whether the change worked, and child-safety advocates have been quick to warn that an announcement is not a safeguard until it is shown to run. That distinction — between what was done, what was promised, what was found, and what was merely alleged — is the whole of this subject's reliability. What remains unresolved is the largest question of all: whether, across a population that now numbers in the billions of conversations, these tools are on balance helping the people who turn to them in their worst moments, or harming them. The honest answer today is that nobody has measured it, and the instruments to do so are only now being built.