AI Data Infra Senior Researcher

Lenovo•Published 1 days ago•First seen 16 hours ago

Why Work at Lenovo

We are Lenovo. We do what we say. We own what we do. We WOW our customers.

Lenovo is a US$83 billion revenue global technology powerhouse, ranked #153 in the Fortune Global 500, and serving millions of customers every day in 180 markets. Focused on a bold vision to deliver Smarter Technology for All, Lenovo has built on its success as the world’s largest PC company with a full-stack portfolio of AI-enabled, AI-ready, and AI-optimized devices (PCs, workstations, smartphones, tablets), infrastructure (server, storage, edge, high performance computing and software defined infrastructure), software, solutions, and services. Lenovo’s continued investment in world-changing innovation is building a more equitable, trustworthy, and smarter future for everyone, everywhere. Lenovo is listed on the Hong Kong stock exchange under Lenovo Group Limited (HKSE: 992) (ADR: LNVGY).

This transformation together with Lenovo’s world-changing innovation is building a more inclusive, trustworthy, and smarter future for everyone, everywhere. To find out more visit www.lenovo.com, and read about the latest news via our StoryHub.

Description and Requirements

岗位职责:

  • 研究和应用机器学习、深度学习、统计学习及时间序列方法,对大规模 telemetry、metrics、logs、events 和系统运行数据进行建模,支持异常检测、趋势分析、change-point detection、性能预测、容量预测、故障预测和根因分析。
  • 研究面向复杂基础设施的数据科学方法,包括多变量时间序列分析、表征学习、异常模式识别、因果分析及预测建模,并将算法应用于实际的大规模分布式系统。
  • 设计和开发面向 AI 工作负载的大规模数据基础设施,支持模型训练、Fine-tuning、Embedding、RAG 和在线推理过程中的数据采集、处理、存储、索引、检索和高吞吐数据供给。
  • 研究大规模并发 GPU 环境中的数据处理和数据供给问题,优化 CPU、GPU、内存、存储和网络之间的数据路径,降低 GPU data starvation、I/O bottleneck 和数据移动开销,提高 GPU 集群整体利用率。
  • 设计和优化高并发数据处理架构,支持多 GPU、多节点及大规模训练和推理 workload 下的数据预处理、shuffle、batching、prefetch、cache、数据并行和分布式数据访问。
  • 设计和优化大规模实时数据处理链路,覆盖数据摄取、消息系统、流式计算、批处理、分析存储和数据服务,提高系统吞吐、延迟、可靠性及资源利用效率。
  • 开展 Flink 等分布式流处理系统的设计和性能优化,解决 stateful processing、event time、watermark、checkpoint、backpressure、exactly-once、数据倾斜及故障恢复等问题。
  • 设计和优化 Kafka / Flink / ClickHouse 或类似技术体系下的端到端实时数据链路,支持高吞吐数据摄取、实时计算和低延迟分析。
  • 研究 ClickHouse 等列式分析系统中的数据组织和查询优化,包括 partitioning、sorting、compression、indexing、materialized view、distributed query 和 high-throughput ingestion。
  • 研究 Data Lake、Lakehouse、OLAP 和 Streaming 系统在 AI 工作负载中的协同架构,优化 Parquet、Arrow、Iceberg 等数据格式及存储系统的数据访问效率。
  • 参与 GPU / 加速器、CPU、内存、存储、网络、运行时和 AI 框架之间的跨层性能分析,使用 profiling、benchmark 和 production telemetry 定位复杂性能瓶颈。
  • 与研究、工程、架构和产品团队合作,将数据科学算法、数据基础设施和系统优化技术转化为 Enterprise AI 和 Personal AI 产品能力。

任职要求:

  • 计算机科学、人工智能、数据科学、计算机工程、电子工程、应用数学或相关专业硕士或博士学位,或具备同等实践经验。
  • 7 年及以上 Machine Learning、Deep Learning、Data Science、AI Data Infrastructure、Distributed Systems、Machine Learning Systems 或大规模数据平台相关研发经验。
  • 具备扎实的数据科学、机器学习和深度学习基础,熟悉监督学习、无监督学习、表示学习、时间序列分析、异常检测、预测建模及模型评估方法。
  • 熟悉 PyTorch、TensorFlow、JAX 或其他主流机器学习框架,具有模型训练、推理、数据分析或算法工程化经验。
  • 具备大规模分布式数据系统设计和开发能力,熟悉 streaming、batch processing、distributed storage、OLAP 和数据服务等技术。
  • 熟悉 Kafka、Flink、Spark、ClickHouse 或类似技术,并能够分析大规模实时数据链路中的吞吐、延迟、状态管理、backpressure 和故障恢复问题。
  • 熟悉 Flink 或类似 stateful streaming framework,理解 event time、watermark、window、state、checkpoint 和 exactly-once 等关键机制。
  • 熟悉 ClickHouse 或类似列式分析系统,理解 columnar storage、partitioning、sorting、compression、高吞吐写入和低延迟查询。
  • 理解大规模 GPU 训练和推理中的数据处理需求,能够分析 storage、network、CPU preprocessing 与 GPU execution 之间的数据供给和性能瓶颈。
  • 具备扎实的软件开发能力,熟练掌握 Python、C++、Java、Go、Rust 等一种或多种编程语言,并具备复杂系统性能分析和问题定位能力。

Preferred Qualifications

  • 具有大规模 telemetry 或 time-series 数据分析经验,并在异常检测、预测分析、根因分析或系统智能诊断方面具有实际项目经验。
  • 具有大规模 GPU 集群、分布式训练或推理系统的数据 pipeline、data loader、distributed storage 或高性能数据访问优化经验。
  • 具有 Kafka / Flink / ClickHouse 大规模生产环境设计、开发和性能调优经验。
  • 具有 Flink runtime、state backend、checkpoint、backpressure 或 distributed state 深入实践经验。
  • 具有 ClickHouse MergeTree、partition / order key、materialized view、distributed table、query optimization 或 ingestion tuning 经验。
  • 熟悉 CUDA、ROCm、NCCL、RDMA、GPUDirect Storage 或其他 GPU、网络和高性能数据访问相关技术。
  • 具有 Vector Search、Embedding Pipeline、RAG 数据处理或大规模非结构化数据平台经验。
  • 在 CCF 推荐 A 类会议或期刊发表过机器学习、数据挖掘、数据库、分布式系统、系统或相关领域论文者优先。
  • 具有高质量专利、重要开源项目贡献,或相关技术在商业产品和生产环境中产生实际影响者优先。