간호학생이 제출한 과제가 너무 완벽하다. 문장은 매끄럽고, 인용도 적절하다. 하지만 교수는 안다. 이건 학생이 쓴 게 아니다. ChatGPT가 쓴 것이다. 그렇다면 어떻게 평가해야 할까?
호주 울런공대학교 연구팀은 생성형 AI 시대에 간호·조산 교육의 평가 방식을 어떻게 재설계해야 하는지를 다룬 토론 논문을 발표했다. 이 연구는 동료 평가 문헌, 전문가 의견, 정책 문서를 종합한 서술적 문헌고찰(narrative review) 방식으로 진행되었다.
연구팀은 세 가지 접근 경로를 제시했다. 첫 번째 경로(Lane One)는 GenAI가 복제하기 어려운 과제를 설계하는 것이다. 임상 실습 관찰, 구두 발표, 대면 시험, 시뮬레이션 평가 등이 이에 해당한다. 고차원적 사고와 실제 수행 능력을 요구하는 과제는 AI가 대신할 수 없다.
두 번째 경로(Lane Two)는 인간-AI 협업을 수용하는 것이다. 학생이 AI를 사용하는 과정을 투명하게 공개하고, AI 결과물을 비판적으로 평가하는 능력을 키우는 방식이다. 여기서는 결과물보다 과정과 평가적 판단(evaluative judgement)이 중요해진다.
세 번째 경로(Lane Three)는 하이브리드 방식이다. 정해진 경계 내에서 조건부 AI 사용을 허용하는 것이다. 예를 들어, 초안 작성에 AI를 쓰되 최종 수정과 논리 전개는 학생이 직접 수행해야 한다는 식이다.
연구팀이 가장 강조한 점은 구조적 변화의 필요성이다. 단순히 AI 사용 금지라는 규칙을 공표하는 것(담론적 변화)은 실효성이 없다. AI가 학생의 산출물을 복제할 수 있는 상황에서 규칙만으로 억제하는 것은 불가능에 가깝다. 대신 평가의 구조 자체를 변경해야 한다. 프로그램 수준에서 평가를 체계적으로 재설계하고, 허용 가능한 AI 사용 범위를 명확히 해야 한다.
이 논문의 의미는 간호·조산 교육이 직면한 현실적 딜레마를 구체적 해법과 함께 제시했다는 점이다. 다만 토론 논문이므로 실험적 검증은 없었고, 제시된 전략의 효과성은 향후 실제 적용을 통해 확인해야 한다.
간호교육 기관과 교수진이 실천할 수 있는 구체적 방법은 다음과 같다. 면접 평가와 시뮬레이션을 확대하고, AI 사용 과정을 명시하도록 요구하는 과제를 도입하며, 환자 안전 원칙이 평가의 핵심에 자리잡도록 보장하는 것이다. 결국 AI 시대에도 간호사의 핵심 역량은 환자를 안전하게 돌보는 능력이며, 평가는 그 능력을 검증해야 한다.
📖 *Refining Assessment Design and Practice in Nursing and Midwifery Education in the Era of Generative AI: A Discussion Paper (토론 논문, 서술적 문헌고찰)* |
논문 원문
※ 이 기사는 의학 논문을 바탕으로 작성되었습니다. 개인 건강 상태에 따라 다를 수 있으니 전문의와 상담하세요.
A nursing student submits an assignment that is suspiciously perfect. The prose is polished, citations are spotless. The professor knows — ChatGPT wrote this. How should it be graded?
A research team from the University of Wollongong in Australia published a discussion paper addressing exactly this dilemma. Through a narrative review drawing on peer-reviewed literature, expert commentary, and policy documents, the team examined how nursing and midwifery education should redesign assessment in the generative AI era.
The authors proposed three lanes of assessment redesign. Lane One involves creating GenAI-resistant tasks that demand higher-order thinking: clinical observation, oral presentations, in-person exams, and simulation assessments. These are tasks AI cannot replicate because they require physical presence and real-time reasoning.
Lane Two embraces human-AI collaboration. Students are required to disclose their AI use transparently, and the focus shifts to their ability to critically evaluate AI outputs — what the authors call evaluative judgement. The process matters more than the final product.
Lane Three is a hybrid approach, permitting conditional AI use within defined boundaries. For example, students might use AI for initial drafting but must independently handle revisions and logical development.
The team's most forceful argument was for structural rather than discursive change. Simply posting an AI use prohibited rule — a discursive change — is ineffective when AI can replicate student outputs. Instead, the mechanics of assessment itself must be redesigned at the program level, with clear definitions of acceptable AI use.
The paper's significance lies in confronting a real dilemma with concrete strategies. As a discussion paper, however, it lacks experimental validation; the effectiveness of each approach awaits empirical testing.
For nursing education institutions, practical steps include expanding oral and simulation-based assessments, introducing assignments that require explicit disclosure of AI use, and ensuring patient safety principles remain central to all evaluations. Ultimately, the core competency of nurses — safely caring for patients — must be what assessments verify, regardless of AI's presence.
📖 *Refining Assessment Design and Practice in Nursing and Midwifery Education in the Era of Generative AI: A Discussion Paper (Discussion paper, narrative review)* |
PubMed
※ This article is based on a medical research paper. Individual health conditions may vary; please consult a healthcare professional.