数据评测开发工程师
Why Work at Lenovo
We are Lenovo. We do what we say. We own what we do. We WOW our customers.
Lenovo is a US$83 billion revenue global technology powerhouse, ranked #153 in the Fortune Global 500, and serving millions of customers every day in 180 markets. Focused on a bold vision to deliver Smarter Technology for All, Lenovo has built on its success as the world’s largest PC company with a full-stack portfolio of AI-enabled, AI-ready, and AI-optimized devices (PCs, workstations, smartphones, tablets), infrastructure (server, storage, edge, high performance computing and software defined infrastructure), software, solutions, and services. Lenovo’s continued investment in world-changing innovation is building a more equitable, trustworthy, and smarter future for everyone, everywhere. Lenovo is listed on the Hong Kong stock exchange under Lenovo Group Limited (HKSE: 992) (ADR: LNVGY).
This transformation together with Lenovo’s world-changing innovation is building a more inclusive, trustworthy, and smarter future for everyone, everywhere. To find out more visit www.lenovo.com, and read about the latest news via our StoryHub.
Description and Requirements
岗位职责
1. 评测体系搭建 建立 AI Agent 端到端评测体系,制定覆盖问答质量、多轮对话一致性、工具调用准确率、任务完成率、稳定性、安全性与用户体验的多维评测标准,形成可量化、可复用的评分量表。
2. 评测数据集建设与管理 构建并持续维护高质量 Golden Dataset:完成样本采集、清洗、标注、分层、去重与版本管理;覆盖标准问答、多步推理、边界 Case 与对抗样本,保障数据集的代表性、场景覆盖度与区分度。
3. 自动化评测流水线与基建 开发自动化评测 Pipeline,融合规则校验、代码断言、LLM-as-a-Judge 与人工专家评审多路评分;接入 CI/CD 实现版本发布前自动回归;打通"离线评测 → 自动化回归 → 线上监控"全链路。
4. 竞品 Benchmark 对标 对标 Claude Code、OpenCode 等主流 Coding Agent 产品,定期开展横向 Benchmark 评测,输出能力差距与优化方向。
5. 问题归因与迭代闭环 建立失败案例归因体系,对幻觉、误判、漏答、工具调用失败、执行中断、结果不稳定等问题做聚类分析;输出评测报告,推动算法、工程、产品团队完成 Top 问题专项治理。
6. 离线-线上效果对齐 探索用户线上行为、业务结果与离线指标的相关性,持续提升离线评测对真实用户体验的预测与解释能力。
岗位要求
1、1年以上大模型/算法评测、自动化测试开发或 AI 产品质量经验
2、熟悉Python,能独立开发评测脚本与流水线框架;熟练使用 SQL 做数据处理与实验统计分析
3、熟悉至少 1 种主流评测框架(RAGAS、DeepEval、TruLens、AgentBench、PromptFoo、PawBench 等),4、了解 SWE-Bench、WebArena、GAIA、ToolBench 等学术 Benchmark
5、掌握主流评测方法论:准确率/召回率、成对对比、评分量表、LLM-as-a-Judge、人工评测
6、理解大模型、Agent、RAG、Function-Calling 与 Prompt 工程基本原理
7、具备评测数据集完整构建经验,熟悉标注流程、样本分层与质量控制
加分
1、有自动化评测平台、大规模评测流水线建设经验
2、有复杂任务 E2E 评测或工具调用 Agent 评测经验
3、熟悉 LangSmith、LangFuse 等 LLM 可观测性工具,能做全链路故障归因
4、有离线评测指标与线上业务指标对齐的实践经验
5、较强的结构化分析与报告撰写能力,能从失败样本中提炼共性问题并给出可落地建议
6、优秀的跨团队协作能力,能驱动算法、工程、产品、QA 多方落地问题治理

