This study compares human and AI assessments of "register awareness" in high-stakes mixed-genre writing. By analyzing student scripts from a national-level EFL examination, the research reveals limitations in traditional rubrics and highlights the irreplaceable role of human judgment in evaluating complex rhetorical shifts that generative AI currently struggles to address.