Back to jobs

AI/ML Compiler Developer (NPU Acceleration)

Hyderabad, India

Pay
Salary not listed in the saved posting
Work setup
Unconfirmed
Employment
Unconfirmed
Apply at Advanced Micro Devices

What you’ll work on

Full posting
  • Design and implement highly optimized C++ kernel library for NPU/GPU.

  • Work closely with hardware engineers to understand the architecture of VLIW vector core units such as MAC, GeMM, and non-linear functions.

  • Develop CPU models for the ML operators in C++/ Python to validate accuracy.

From the employer’s posting
Kernel Development: Design and implement highly optimized C++ kernel library for NPU/GPU. Collaborate with the research and software teams to integrate these kernels into the existing software stack.
Vector Processor Optimization: Work closely with hardware engineers to understand the architecture of VLIW vector core units such as MAC, GeMM, and non-linear functions. Develop vectorized code that leverages SIMD (Single Instruction, Multiple Data) and ILP (instruction level parallelism) for maximum performance.
Testing and Validation: Develop CPU models for the ML operators in C++/ Python to validate accuracy. Write unit tests and integration tests to ensure correctness and reliability.

What you’ll bring

All qualifications

Preferred experience

  • Good understanding of SIMD/Tensor/VLIW processor architecture to exploit parallelism.
  • Experience with vectorized programming (SIMD) and parallel computing.
  • Familiarity with machine learning frameworks (e.g., TensorFlow, PyTorch) is a plus.
  • Experience with silicon bring-up and pre-silicon validation on Emulation platforms is a plus.
Qualification wording
Good understanding of SIMD/Tensor/VLIW processor architecture to exploit parallelism.
Experience with vectorized programming (SIMD) and parallel computing.
Familiarity with machine learning frameworks (e.g., TensorFlow, PyTorch) is a plus.
Experience with silicon bring-up and pre-silicon validation on Emulation platforms is a plus.

Tools in this posting

  • C
  • Python
  • PyTorch
  • TensorFlow
  • C++
Source — Tool mentions in context
THE ROLE: We are looking for a dynamic, energetic Lead / Staff Software Engineer to join our growing team in AI (Artificial Intelligence) group. In this role, the individual will be responsible for developing AI/ML specific C/C++ kernels and dataflow schedules for AMD Ryzen Processors built on XDNA Neural Processor Units (NPU) to map LLMs, Stable Diffusion networks on NPU. As a C++ Kernel Developer, you will play a crucial role in designing, optimizing, and implementing machine learning kernels specifically tailored for vector processors. Your work will directly impact the efficiency, speed, and accuracy of our machine learning models. KEY RESPONSIBILITIES:
PREFERRED EXPERIENCE: - Excellent C/C++ and Python coding skills - Good understanding of SIMD/Tensor/VLIW processor architecture to exploit parallelism.
- Testing and Validation: - Develop CPU models for the ML operators in C++/ Python to validate accuracy. - Write unit tests and integration tests to ensure correctness and reliability.
- Experience with vectorized programming (SIMD) and parallel computing. - Familiarity with machine learning frameworks (e.g., TensorFlow, PyTorch) is a plus. - Experience with silicon bring-up and pre-silicon validation on Emulation platforms is a plus.
- Kernel Development: - Design and implement highly optimized C++ kernel library for NPU/GPU. - Collaborate with the research and software teams to integrate these kernels into the existing software stack.

Job description

View original posting ↗



ADVANCE YOUR CAREER. ADVANCE THE WORLD. 

At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future. 

 

Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career.




MTS SOFTWARE DEVELOPMENT ENGINEER 

 

THE ROLE:  

We are looking for a dynamic, energetic Lead / Staff Software Engineer to join our growing team in AI (Artificial Intelligence) group. In this role, the individual will be responsible for developing AI/ML specific C/C++ kernels and dataflow schedules for AMD Ryzen Processors built on XDNA Neural Processor Units (NPU) to map LLMs, Stable Diffusion networks on NPU. As a C++ Kernel Developer, you will play a crucial role in designing, optimizing, and implementing machine learning kernels specifically tailored for vector processors. Your work will directly impact the efficiency, speed, and accuracy of our machine learning models. 

  

KEY RESPONSIBILITIES: 

  • Kernel Development: 
  • Design and implement highly optimized C++ kernel library for NPU/GPU. 
  • Collaborate with the research and software teams to integrate these kernels into the existing software stack. 
  • Vector Processor Optimization: 
  • Work closely with hardware engineers to understand the architecture of VLIW vector core units such as MAC, GeMM, and non-linear functions. 
  • Develop vectorized code that leverages SIMD (Single Instruction, Multiple Data) and ILP (instruction level parallelism) for maximum performance. 
  • Performance Profiling and Tuning: 
  • Profile and analyze the performance of existing kernels. 
  • Identify bottlenecks and optimize critical sections for better throughput. 
  • Testing and Validation: 
  • Develop CPU models for the ML operators in C++/ Python to validate accuracy. 
  • Write unit tests and integration tests to ensure correctness and reliability. 
  • Validate kernel performance across different hardware platforms. 
  • Documentation and Collaboration: 
  • Document design specs for new kernels and the performance improvements. 
  • Follow coding guidelines, use tools like git to maintain code and create pull-requests, and documentation. 
  • Collaborate with cross-functional teams, including machine learning researchers and software engineers. 

 

PREFERRED EXPERIENCE:  

  • Excellent C/C++ and Python coding skills 
  • Good understanding of SIMD/Tensor/VLIW processor architecture to exploit parallelism. 
  • Experience with vectorized programming (SIMD) and parallel computing. 
  • Familiarity with machine learning frameworks (e.g., TensorFlow, PyTorch) is a plus. 
  • Experience with silicon bring-up and pre-silicon validation on Emulation platforms is a plus. 
  • Knowledge of low-level hardware details (cache hierarchy, memory access patterns) is desirable. 
  • Excellent problem-solving skills and a passion for performance optimization. 

 

ACADEMIC CREDENTIALS: 

BS/Masters/PhD degree in Computer Science, Electrical Engineering, or a related field with around 10/8/5 year experience respectively.

 

#LI-NR1




Benefits offered are described:  AMD benefits at a glance.

 

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law.   We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

 

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position.  AMD’s “Responsible AI Policy” is available here.

 

This posting is for an existing vacancy.

Your next step

  • Have your CV and examples of relevant work ready.
  • Check the listed location, eligibility and core experience before starting.
  • Ask the employer about the salary range before committing time to the process.

Complete your application on global-external-amd.icims.com. The employer’s form will show what is required.

Already applied? Track this application

Source & posting history

View original posting ↗

Source notes

Source excerpts

Selected passages from the saved posting. Check the full description for conditions and exceptions.

Pay

No pay amount identified in the saved description.

Location & working pattern

Hyderabad, India

Working pattern and location restrictions need checking in the full posting.

Work authorization

No clear work-authorization passage found. Eligibility is unconfirmed.

Status in our records
Active
First seen by us
Sep 1, 2026
Recorded sightings
174
Last seen by us
Oct 8, 2026

These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.

Report an error

See how this role fits your experience

Add your resume to compare the role’s scope, tools and requirements with your experience.

Find answers in the posting

AI
How answers work

AI selects complete passages from this posting. Check them for conditions and exceptions.

Uses this posting and your question. No profile needed.