Staff/Sr. ML Infrastructure / Platform Engineer

Trend Micro•Published 1 days ago•First seen 1 days ago

Join Trend ‧ Join New Generation

趨勢科技 - 全球雲端資安領航者 / 全亞洲最大軟體公司 / 企業版圖橫跨五大洲 / 趨勢全球研發基地在台灣 
===============================================================

About the Role

We are building a production-grade, GPU-accelerated LLM serving platform that powers multiple AI products at enterprise scale. You will be responsible for designing, building, and operating the infrastructure that serves large language models — from raw Kubernetes cluster management to multi-GPU inference optimization and autoscaling. 

Required Qualifications

Model Serving & Inference

  • Operate multi-model LLM serving infrastructure 
  • Tune autoscaling policies to balance GPU cost and latency SLAs 

Kubernetes & GPU Infrastructure

  • Operate production K8s clusters with NVIDIA GPU nodes 
  • Handle GPU node lifecycle: NVIDIA driver setup 

Infrastructure as Code

  • Write and maintain Terraform/Terragrunt modules for AWS/GCP cloud 
  • Package platform components and model deployments as Helm charts 
  • Manage multi-environment configurations 

Observability & Performance

  • Maintain monitoring stack: Prometheus, Grafana, 
  • Build dashboards for GPU utilization, KV cache occupancy, TTFT/ITL latency, and cost per token 
  • Set up alerting for SLA violations and OOM events 

Bonus Skills

These are not required, but candidates with these skills will stand out. 

  • LoRA / PEFT fine-tuning workflows 
  • MLflow for experiment tracking, model registry, and automated adapter deployment 
  • Experience building LoRA adapter CI/CD pipelines (training → registry → serving) 
  • Experience with alternative inference frameworks such as SGLang or NVIDIA NIM, including deep Parameter Tuning for Continuous Batching, KV Cache management, and Speculative Decoding. 

===============================================================
連結智慧 守護世界 --- Connected Intelligence for Securing a Connected World