Hello CMB team,
I am a physician from Turkey building MedFailBench, an open source clinical AI safety benchmark based on synthetic clinician authored cases.
Project: https://github.com/goktugozkanmd/medical-ai-failure-atlas
Your CMB work is one of the clearest fits I found for a China Turkey medical AI safety collaboration because it is rooted in Chinese medical benchmark design rather than translated general evaluation.
I want to build a shared Turkish, English, and Chinese clinical safety benchmark, run reproducible evaluations across medical and frontier models, and aim for a coauthored paper.
I can bring clinician authored safety cases, Turkish clinical wording risk, a safety taxonomy, and an open source evaluation repo.
No patient data is involved. This is research infrastructure for safer medical LLM evaluation, not clinical advice or deployment.
Would your team be open to a short call or async discussion about a joint pilot?
Best,
Goktug Ozkan
Hello CMB team,
I am a physician from Turkey building MedFailBench, an open source clinical AI safety benchmark based on synthetic clinician authored cases.
Project: https://github.com/goktugozkanmd/medical-ai-failure-atlas
Your CMB work is one of the clearest fits I found for a China Turkey medical AI safety collaboration because it is rooted in Chinese medical benchmark design rather than translated general evaluation.
I want to build a shared Turkish, English, and Chinese clinical safety benchmark, run reproducible evaluations across medical and frontier models, and aim for a coauthored paper.
I can bring clinician authored safety cases, Turkish clinical wording risk, a safety taxonomy, and an open source evaluation repo.
No patient data is involved. This is research infrastructure for safer medical LLM evaluation, not clinical advice or deployment.
Would your team be open to a short call or async discussion about a joint pilot?
Best,
Goktug Ozkan