AI 开发工程师-大模型/端侧推理

LenovoApplyPublished 1 days agoFirst seen 1 days ago
Apply

Why Work at Lenovo

We are Lenovo. We do what we say. We own what we do. We WOW our customers.

Lenovo is a US$83 billion revenue global technology powerhouse, ranked #153 in the Fortune Global 500, and serving millions of customers every day in 180 markets. Focused on a bold vision to deliver Smarter Technology for All, Lenovo has built on its success as the world’s largest PC company with a full-stack portfolio of AI-enabled, AI-ready, and AI-optimized devices (PCs, workstations, smartphones, tablets), infrastructure (server, storage, edge, high performance computing and software defined infrastructure), software, solutions, and services. Lenovo’s continued investment in world-changing innovation is building a more equitable, trustworthy, and smarter future for everyone, everywhere. Lenovo is listed on the Hong Kong stock exchange under Lenovo Group Limited (HKSE: 992) (ADR: LNVGY). 

To find out more visit www.lenovo.com and read about the latest news via our StoryHub.

Description and Requirements

岗位职责

  1. 负责 LLM、多模态模型的工程落地,完成模型转换、部署、推理链路开发与性能调优;
  2. 深入理解 Transformer 等主流模型结构,剖析 Prefill/Decode 推理流程,针对 KV Cache、注意力机制、采样策略开展优化;
  3. 使用 PyTorch、ONNX Runtime、llama.cpp 等框架完成模型封装、量化、算子调试,解决内存占用、推理时延、吞吐瓶颈;
  4. 搭建模型评测基准,对比不同硬件、推理引擎下模型表现,输出优化方案;
  5. 协同算法、客户端、内核团队推进 AI 功能集成,支撑端侧 / 本地 AI 推理方案落地;
  6. 跟踪前沿模型架构、推理加速技术,沉淀工程化工具链。

岗位要求

  1. 本科及以上,计算机、人工智能、电子信息相关专业;
  2. 深入理解 Transformer 基础架构,熟悉大模型前向推理完整流程,掌握 Prefill、Decode 阶段运行原理,了解 KV Cache、上下文管理、Token 生成机制;
  3. 熟悉模型结构、推理方式;熟悉主流模型推理范式,了解 INT4/INT8 量化、模型压缩、计算图优化等加速手段;
  4. 熟练 Python,具备 C/C++ 开发能力优先;熟悉 PyTorch、ONNX、GGUF 模型格式;
  5. 有 llama.cpp、ONNX Runtime、TensorRT、vLLM 任一推理框架实战经验;
  6. 具备端侧 / 本地部署经验,了解 CPU/XPU/NPU 异构推理调度优先;
  7. 良好问题排查能力,能够定位推理精度异常、性能瓶颈;具备优秀跨团队协同意识。