NPU Compiler Engineer

Pedestal 

📍 Neihu District, Taiwan 🇹🇼

full-time
junior
80000
Posted —

Key Skills

C++PythonCompilerONNXTensorFlow

Industry

SemiconductorAI

Job Description

About Pedestal

Pedestal is a fast-growing NPU IP startup developing flexible, scalable NPU core hardware for next-generation AI accelerators. Our expertise spans ultra-low-power architecture, near-memory computing, and high-efficiency NoC design. Since our founding in 2025, we have secured commercial customers, achieved successful silicon tape-outs, received prestigious industry awards, and been selected for multiple government-funded programs. Join us to shape the future of AI hardware.



Role Overview

In this role, you will bridge the gap between high-level neural network frameworks and custom NPU hardware, driving performance, power efficiency, and seamless execution.



Key Responsibilities

  • Compiler Framework & Development: Construct compiler frameworks (utilizing open-source infrastructure like TVM or MLIR) and build an NPU compiler to translate high-level neural network graphs (ONNX, TensorFlow, PyTorch) into optimized hardware execution binaries.
  • Graph Optimization: Perform computational graph optimizations—including operator fusion, kernel selection, and memory layout transformations—to optimize inference latency, power efficiency, and memory footprint.
  • Performance Profiling: Analyze compiled workloads to identify system bottlenecks and drive continuous optimization.
  • System Integration: Integrate the compiler toolchain with runtime libraries, drivers, and the broader AI software stack.


Basic Qualifications

  • Education: Master’s degree or above in Electrical Engineering, Computer Science, or a related field.
  • Programming: Strong proficiency in C, C++, and Python.
  • Compiler Fundamentals: Solid understanding of compiler construction, Intermediate Representations (IR), and code generation techniques.
  • AI Frameworks: Familiarity with modern neural network frameworks (ONNX, TensorFlow, PyTorch) and their graph representations.
  • Performance Tuning: Practical experience in performance profiling, bottleneck analysis, and code optimization techniques.
  • Experience: 0~5 years of work experience.


Preferred Qualifications

  • AI/ML Knowledge: Familiarity with AI/ML fundamentals and workloads, especially Computer Vision, Signal Processing, and LLM/VLM/MLLM architectures.
  • Industry Awareness: Understanding of modern AI hardware architectures, NPU technology, and domain trends.
  • Mindset: Highly adaptable with a strong results-driven and execution-oriented approach.


Expected Base Pay Range (NTD)

NT$ 80,000 – NT$ 250,000 monthly