Health, Safety / AI news for Malaysia
OpenAI releases MentalHealthBench to test how chatbots handle mental health
People already confide in chatbots, so the quality of their answers matters. OpenAI's open benchmark grades replies against expert rubrics, but it tests synthetic conversations and is not proof that any app is safe.

In brief
- OpenAI released MentalHealthBench on 23 September 2026, an open benchmark for how AI systems respond in realistic mental health conversations.[1]
- It was built with more than 80 licensed psychologists and psychiatrists from 22 countries, who wrote the rubrics used to grade answers.[1]
- The benchmark uses synthetic conversations, not real chats, and OpenAI says ChatGPT is not a substitute for therapy or professional care.[1]
What was released
OpenAI has released an open benchmark that measures how AI systems respond when people bring mental health concerns to a chatbot. MentalHealthBench, published on 23 September 2026, covers everything from everyday stress and relationship trouble to emergencies that need real-world help.[1]
OpenAI says more than one billion people use ChatGPT each week. It argues that earlier tests focused mostly on emergencies and broad pass or fail rules, leaving a gap in how models handle the wider range of conversations.[1]

How the benchmark was built
OpenAI says it created synthetic conversations with privacy-preserving techniques, so they reflect real usage patterns without exposing real chats. Some include background about the user, such as a recent death in the family, to see whether a model uses that context. Scenarios cover adults, teens aged 13 to 17, caregivers and clinicians.[1]
Each conversation was reviewed by at least three experts from a group of more than 80 licensed psychologists and psychiatrists in 22 countries, speaking 19 languages. They wrote grading criteria weighted from minus 10 to plus 10. Only criteria agreed by two experts and not contradicted by a third were kept.[1]

How answers are graded
An automated grader, OpenAI's GPT-5.6 Sol model, scores each response against the expert criteria. Positive points reward helpful behaviour, such as asking the right question, while negative points penalise harmful replies. The score can be split into ten behaviours defined by the experts, including safety, seeking context and respecting the user's choices.[1]
OpenAI says results show steady improvement in newer models from several providers, especially in seeking context. OpenAI built the benchmark, designed the grading and uses its own model as the grader. That makes independent tests of the results important.[1]
Where users and experts differed
Separately, OpenAI asked 44 adults from 16 countries who had used AI for emotional support to rate responses to non-urgent conversations. They valued practical next steps and tone more than the experts did. The experts put more weight on gathering context and carefully reading unclear situations.[1]
OpenAI says this user study did not change the scoring, which rests on expert consensus. It also points to crisis hotline suggestions inside ChatGPT, a Trusted Contact feature and a teen version of ChatGPT. The company states plainly that ChatGPT does not replace therapy or professional care.[1]
The Malaysian picture
The Institute for Public Health's National Health and Morbidity Survey 2023 estimated that about one million people in Malaysia aged 16 and above, or 4.6 per cent, have depression. The number doubled from 2019, about half had thoughts of self-harm, and younger age groups were more likely to be affected.[2]
A benchmark cannot tell a Malaysian parent whether a particular app is safe, and OpenAI has not published results by language or country. It does give local health researchers, universities and app builders an open tool to test models against their own needs.[1][2]
Why Malaysia should care
Depression affects about one million Malaysians aged 16 and above, the national health survey estimates. Some will turn to chatbots first. OpenAI has not said whether Malaysian languages or scenarios are in the benchmark.
Chatbot users
Answers are improving, but only on synthetic tests.[1]
Practical move: Use AI for support, not in place of a professional.
App and HR tech builders
An open yardstick for wellbeing features.[1]
Practical move: Run the benchmark before launching emotional-support chat.
What Malaysians can do now
- If you or someone you know is in crisis, contact emergency services or a trained counsellor directly rather than relying on a chatbot.
- Companies adding AI chat to HR or customer apps should test how it responds to distress before launch and set a clear hand-off to a human.
- Parents should check which AI apps their teenagers use and switch on any teen or safety settings those apps provide.
What we still do not know
What we still do not know
- Whether the benchmark includes Malay, Chinese or Tamil conversations.
- The exact scores behind OpenAI's charts, and whether independent groups can reproduce them.
- Whether ChatGPT's crisis hotline suggestions list Malaysian services.
Sources
- 1.Introducing MentalHealthBench OpenAI, 23 September 2026
- 2.National Health and Morbidity Survey (NHMS) 2023: Non-Communicable Diseases and Healthcare Demand - Key Findings Institute for Public Health, Ministry of Health Malaysia
- 3.File:Interior of Sri Petaling Line train (230906).jpg Wikimedia Commons, 6 September 2023
- 4.File:Jabatan Kerja Sosial Perubatan Hospital Kuala Lumpur, Kuala Lumpur 20230427 125118.jpg Wikimedia Commons, 27 April 2023
- 5.File:Permai Hospital Johor Bahru.jpg Wikimedia Commons, 8 February 2016


