Health, Healthcare, and Wellbeing / Medical Simulation Training and Communication Support

How consistent are LLM evaluation results with domain expert judgments on complex tasks?

Similar questions

Related papers