模型后训练算法专家

LenovoApplyPublished 1 days agoFirst seen 19 hours ago
Apply

Why Work at Lenovo

We are Lenovo. We do what we say. We own what we do. We WOW our customers. 

Lenovo is a US$83 billion revenue global technology powerhouse, ranked #153 in the Fortune Global 500, and serving millions of customers every day in 180 markets. Focused on a bold vision to deliver Smarter Technology for All, Lenovo has built on its success as the world’s largest PC company with a full-stack portfolio of AI-enabled, AI-ready, and AI-optimized devices (PCs, workstations, smartphones, tablets), infrastructure (server, storage, edge, high performance computing and software defined infrastructure), software, solutions, and services. Lenovo’s continued investment in world-changing innovation is building a more equitable, trustworthy, and smarter future for everyone, everywhere. Lenovo is listed on the Hong Kong stock exchange under Lenovo Group Limited (HKSE: 992) (ADR: LNVGY). 

This transformation together with Lenovo’s world-changing innovation is building a more inclusive, trustworthy, and smarter future for everyone, everywhere. To find out more visit www.lenovo.com, and read about the latest news via our StoryHub.

Description and Requirements

岗位职责:
1. 负责大语言模型的后训练全流程研发,包括SFT、RLHF、DPO、GRPO等主流后训练算法的落地与优化,提升模型在特定场景下的效果与对齐度。
2. 设计并搭建高效的后训练数据处理、训练、评测全链路pipeline,保障训练流程的稳定性、可复现性与高效率。
3. 针对模型幻觉、偏见、安全对齐等核心问题,研发针对性的后训练优化方案,持续提升模型的可靠性与合规性。
4. 负责大模型Agent能力的后训练对齐研发,包括Agentic Workflow的指令遵循、工具调用、多步推理、任务规划、长程记忆等核心能力的后训练优化,提升模型在Agent场景下的任务完成率与稳定性。
5. 探索并落地大模型前沿后训练技术,包括Process Reward Model(PRM)、Outcome Reward Model(ORM)、Agentic RLHF、多模态对齐、Constitutional AI、自洽性对齐等前沿算法,持续提升模型的通用能力与Agent任务执行能力。
6. 设计Agent场景下的后训练评测体系,包括多步任务、工具调用、长程规划、反事实推理等维度的评测基准与自动化评测流程,保障模型Agent能力的持续迭代。
7. 跟进业界前沿的大模型后训练技术与算法,结合业务场景进行创新落地,输出技术方案、专利与技术文档。
8. 与算法、工程、产品团队协同,完成模型迭代的全流程落地,支撑业务场景的效果目标达成。

岗位要求:
1. 精通大语言模型后训练相关技术,深入理解SFT、RLHF、DPO、GRPO、Agentic RLHF等主流对齐算法的原理与工程实现,有相关项目落地经验。
2. 熟练掌握PyTorch、DeepSpeed、Megatron-LM、vLLM等主流训练与推理框架,具备大模型分布式训练、显存优化、性能调优、推理加速的实战经验。
3. 具备扎实的机器学习、自然语言处理基础,熟悉大模型预训练、微调、评测、部署的全流程技术体系。
4. 深入理解大模型Agent的核心技术体系,有Agent能力后训练、工具调用对齐、多步推理优化相关的项目经验,熟悉AgentBench、ToolBench等主流Agent评测基准。
5. 具备优秀的问题分析与解决能力,能够快速定位训练过程中的效果、性能、稳定性问题,并给出有效解决方案。
6. 具备良好的沟通协作能力与文档撰写能力,能够清晰输出技术方案与研发文档,有较强的团队协作意识。