Human-in-the-loop evaluation 通过让真人评判 AI agent 的输出与行为来检查其表现。测试者不只看 automated score,还邀请 user、domain expert 或 crowd worker 观察任务、标注答案、标记错误,并评价 clarity、fairness 或 safety。他们的反馈揭示数字 alone 难以发现的问题,如 hidden bias、confusing language 或对人感觉不对的行动。团队研究这些笔记、调整 model 并再跑一轮,重复直至 agent 达到质量与信任目标。混合 human judgment 与 data 可得到更准确、有用且日常更安全的系统。