Software - Software Engineer, Compiler (Kernel Optimization)
About FuriosaAI
FuriosaAI builds high-performance, high-efficiency AI compute for the Inference Era. Founded in 2017 by veteran semiconductor and AI algorithm engineers, Furiosa operates globally with offices in Korea and Silicon Valley, along with a compiler-focused R&D lab in Lisbon.
Our vision is to make AI computing sustainable, enabling access to powerful AI for everyone on Earth. We solve the AI hardware energy and operational cost crisis at the architectural level, rather than through brute force, building the world's first truly AI-native compute platform to unlock the full potential of artificial intelligence for every enterprise.
About the Job
The compiler plays a central role in FuriosaAI's mission to build high-performance, energy-efficient AI systems. Modern deep learning models are evolving rapidly and becoming increasingly diverse, making compilation a challenging problem. Transforming these models into efficient executable programs requires careful reasoning about complex transformations while preserving program meaning and structure.
In this role, you will own the performance of critical AI kernels integrated into Furiosa-LLM, FuriosaAI's software stack for large language model serving. You will analyze end-to-end serving workloads, identify performance-critical bottlenecks, and implement highly optimized kernels using TCL (Tensor Contraction Language), FuriosaAI's programming language for kernel optimization. Because real-world serving workloads are inherently dynamic, you will develop scheduling, specialization, and algorithmic techniques that map them efficiently onto compiler abstractions and execution environments optimized for static workloads.
Responsibilities
- Analyze end-to-end LLM serving workloads and identify performance bottlenecks that can be addressed through kernel-level optimization.
- Design, implement, and optimize high-performance kernels in TCL for critical operations in Furiosa-LLM.
- Develop algorithmic techniques for efficiently supporting dynamic serving workloads.
- Integrate, benchmark, and validate optimized kernels across representative models, input shapes, and serving scenarios.
- Collaborate with compiler and serving teams to improve compiler capabilities and ensure optimized kernels work effectively within the production software stack.
Minimum Qualifications
- BS in Computer Science, Artificial Intelligence, Electrical Engineering, or a related field.
- Experience in developing low-level or performance-critical software.
- Experience in analyzing performance bottlenecks using profiling, benchmarking, and hardware performance characteristics.
- Understanding of parallel computation, memory hierarchies, and data movement on modern architectures.
Preferred Qualifications
- MS or PhD in Computer Science, Artificial Intelligence, Electrical Engineering, or a related field.
- Experience in optimizing high-performance kernels on AI accelerators (e.g., GPU, TPU) for AI products.
- Hands-on knowledge of kernel optimization techniques such as tiling, computation scheduling, operator fusion, memory layout transformation, vectorization, pipelining, and data movement optimization.
- Experience reasoning about trade-offs among parallelism, memory bandwidth, on-chip memory capacity, compute utilization, and synchronization overhead.
- Experience with accelerator programming or domain-specific kernel languages such as Triton, cuTile, or Pallas.
Why Join FuriosaAI
The defining bottleneck of the AI era is building the right hardware and software stack to run it at global scale. Furiosa is solving this challenge holistically from the ground up.
With our flagship chip, RNGD, in mass production today and our next-generation platform in development with Broadcom, we are proving that full-stack, tensor-native compute is the future of AI infrastructure. This is a pivotal moment to join our team, right as we accelerate our global expansion.
At Furiosa, you will:
Solve AI’s Most Urgent Challenge. Help build the high-performance, energy-efficient inference hardware and software required to fulfill the promise of advanced AI.
Pioneer Full-Stack Co-Design. Work with teams that are architecting solutions from silicon up through the compiler (featuring innovations like Tensor Contraction Language and Virtual ISA) and serving frameworks.
Ship Real-World Silicon, Software, and Solutions. Turn breakthrough technology into commercial deployment. RNGD is in mass production with TSMC and running live enterprise workloads for global leaders like LG AI Research and Samsung SDS.
Partner With the Industry's Best. Collaborate across an elite global ecosystem that includes TSMC, Broadcom, SK Hynix, and GUC.
Do Your Life’s Best Work. Join a brilliant, low-ego, mission-driven team in a high-trust environment that values autonomy, intellectual curiosity, and shared ambition.

