The Number Was Specific.
The Country Was a Guess.
Researchers at Wroclaw Medical University tested twenty-nine AI chatbot agents (five general-purpose, twenty-four built specifically for mental-health support) against standardized, simulated suicide-crisis scenarios. Of the twenty-three that attempted to provide an emergency number, fewer than one in four initially supplied one appropriate for the researcher’s location. Among the agents that initially supplied a number inappropriate for the region, most answered as though the person on the other end were calling from the United States.
A team of researchers at Wroclaw Medical University in Poland built a set of escalating crisis scenarios, modeled on the Columbia-Suicide-Severity-Scale used across clinical psychiatry, and ran them against twenty-nine different AI chatbot agents: five familiar general-purpose systems (GPT-4o mini, Gemini 2.0 Flash, DeepSeek-v1, LeChat, and Llama 3.1 8B) and twenty-four apps marketed specifically as mental-health support tools. Most agents responded as though the exchange required escalation: twenty-four recommended professional assistance, and twenty-five advised contacting a hotline or emergency number. The harder test was whether the contact information actually fit the researcher’s location: when a distressed person asked for help, did the agent point them toward the emergency number for the country they were actually in?
Twenty-three of the twenty-nine agents attempted to provide an emergency number of some kind; six never offered one at all. Of those twenty-three, the number was correct and appropriate for the researcher’s region on the first attempt in only 21.74 percent of cases: five agents out of twenty-three. The researchers’ own paper states the pattern plainly: “Most agents that initially provided a number inappropriate to the region assumed that the user was based in the United States.” Seventeen of the full twenty-nine eventually reached the right number, but only after two supplementary prompts — the second of which was the researchers stating outright, “I am in Poland.”
This isn’t only a data problem. After both supplementary prompts, seventeen of all twenty-nine agents provided correct emergency contact information. At the initial regional-accuracy measurement, only five of the twenty-three agents that attempted to provide a number were correct for the researcher’s location. The second supplementary prompt did not hint at the missing context. It supplied it outright: “I am in Poland.” Before location was supplied, the tested systems frequently produced an answer without establishing whether its regional context matched the researcher’s. In this test, most agents that initially supplied region-inappropriate information assumed a United States context.
This isn’t simply a story about a chatbot producing an incorrect answer. It is a story about systems that often produced emergency contact information without first establishing whether that information applied to the person in front of them. Producing a number is not the same as establishing that it applies to the person asking.
What changed when location was supplied
The measured failure was regional appropriateness. The paper reports that most agents supplying an initially inappropriate number assumed a United States location. It does not disclose enough response-level data to determine whether every incorrect number was otherwise valid, current, or operational.
A locale signal was a couple of direct questions away. The country had not been established. Once the researcher stated it directly, seventeen of twenty-nine agents provided the correct emergency contact information.
Before the researcher supplied a location, most agents that provided an inappropriate number assumed a United States context. The study measured whether the information was region-appropriate; it did not separately score the confidence or hedging of each response. Filed as an Agentics Contextual Ambiguity specimen, one of the five clinical characteristics of Digital Derangement Syndrome™: a specific answer produced before the context determining its applicability had been established.
What speed doesn’t fix
The researchers caught the pattern through a standardized protocol: escalating simulated scenarios, a fixed prompt sequence, and twenty-nine systems evaluated against the same criteria. The boundary matters. This was one English-language test of free or free-trial versions selected beginning in November 2024 — not a natural conversation, a clinical trial, or a test of the products as they operate today. The authors also could not confirm whether every application’s responses were generated solely by AI, by rules, or by some combination of the two.
Outside a controlled test, a user gets one conversation — and may have no way to know which unstated context shaped the system’s answer.
This record has named the gap between recognition and correctness before, in registries and citation indexes and search results. Here it’s stated in its plainest possible terms: a system can produce highly specific information and still misapply it, because it never established the context that determined whether the answer fit.
If you or someone you know is in crisis: in the United States, call or text 988. Outside the United States, findahelpline.com lists crisis lines by country — a directory, not a guess.
Pichowicz, W., Kotas, M., Piotrowski, P., “Performance of mental health chatbot agents in detecting and managing suicidal ideation,” Scientific Reports 15, 31652 (2025) — peer-reviewed and open access; its publication fee was paid from a Wroclaw Medical University reserve under payment reference REZD.Z505.25.001, and the authors declared no competing interests. Figures verified against the full text: twenty-three of twenty-nine agents attempted to provide an emergency number; the number was correct for the researcher’s region on the first attempt in “only 21.74% (n = 5)” of those cases; “Most agents that initially provided a number inappropriate to the region assumed that the user was based in the United States”; seventeen agents (58.62%) eventually reached a correct number only “after both supplementary prompts,” the second of which stated the country directly — doi.org/10.1038/s41598-025-17242-4.