Agent Infra架构师
Why Work at Lenovo
We are Lenovo. We do what we say. We own what we do. We WOW our customers.
Lenovo is a US$83 billion revenue global technology powerhouse, ranked #153 in the Fortune Global 500, and serving millions of customers every day in 180 markets. Focused on a bold vision to deliver Smarter Technology for All, Lenovo has built on its success as the world’s largest PC company with a full-stack portfolio of AI-enabled, AI-ready, and AI-optimized devices (PCs, workstations, smartphones, tablets), infrastructure (server, storage, edge, high performance computing and software defined infrastructure), software, solutions, and services. Lenovo’s continued investment in world-changing innovation is building a more equitable, trustworthy, and smarter future for everyone, everywhere. Lenovo is listed on the Hong Kong stock exchange under Lenovo Group Limited (HKSE: 992) (ADR: LNVGY).
This transformation together with Lenovo’s world-changing innovation is building a more inclusive, trustworthy, and smarter future for everyone, everywhere. To find out more visit www.lenovo.com, and read about the latest news via our StoryHub.
Description and Requirements
岗位职责
1. 设计并实现面向异构算力设备(移动端、边缘盒子、桌面端)的轻量化推理引擎,覆盖 Android/Linux/Windows 多平台,支持动态模型加载、内存优化、算力感知调度。
2. 针对不同设备算力(ARM NEON、x86 AVX、Mobile GPU),设计模型量化策略(INT4/INT8/FP16),适配主流开源模型(Qwen、Llama、Mistral、Phi等),实现推理延迟 < 500ms、内存占用 < 4GB 的生产级标准。
3. 构建端侧 Agent 执行引擎(AgentCore Harness),实现工具调用、多轮对话、上下文管理、流式输出、离线运行能力,与TianxiAI形成端边云协同闭环。
4. 设计端侧模型蒸馏回灌机制,支持云端大模型能力包向端侧增量更新,实现版本一致性校验、能力感知路由、断点续传与灰度发布。
5. 建立端侧推理性能基准体系(RTF、首token延迟、内存峰值、功耗),输出端侧模型选型指南,推动模型压缩工具链、自动化评测、CI/CD 流水线落地。
岗位要求
1、5年以上 AI 系统或推理引擎开发经验,深度参与过至少一个端侧推理框架的核心模块设计(ONNX Runtime / MNN / NCNN / MLX / llama.cpp / ExecuTorch)
2、精通 C++17 及以上版本,具备扎实的多线程与高并发编程能力,熟悉并发队列、无锁/低锁设计、锁竞争分析与性能优化;熟悉 Rust 者优先
3、深入理解大模型推理的核心瓶颈,能够从显存带宽、计算能力、PCIe 传输、CPU-GPU 页交换等维度定位性能问题,理解 Prefill 与 Decode 阶段在计算密集、访存密集及并行特征上的差异
4、有生产环境端侧模型部署经验,熟悉模型转换工具链(ONNX / CoreML / TFLite / MLC-LLM)
5、深入掌握 KV-Cache 底层原理及工程实现,熟悉 PagedAttention、逻辑块与物理块管理、页表映射、显存池等关键机制,能够针对长上下文和高并发场景优化缓存利用率
6、熟悉大模型推理请求调度策略,包括连续批处理(Continuous Batching)、动态批处理、预取与请求抢占,能够在吞吐、时延、公平性和资源利用率之间进行系统性权衡
7、具备 GPU 显存管理与优化经验,熟悉显存池化、显存碎片治理、CPU-GPU 页交换与 Offload 机制,能够设计受限算力和受限内存设备上的模型装载与运行方案
8、熟悉推理服务性能度量与调优方法,能够围绕吞吐量(TPS)、首 Token 时延(TTFT)、单 Token 解码时延(TPOT)建立基准测试、性能剖析和持续优化体系

