Opens an external site
- Employment type
- Full-time
- Experience level
- Senior · 5+ years
- Minimum education
- Bachelor’s degree
- Apply by
- Oct 5, 2027
- Posting language
- English
- Working hours
- 40 hours per week
Job summary
Architect, develop, and scale AI fleet management solutions across global regions while driving performance optimization and reliability. Influence the long-term architectural evolution of AI compute solutions and leverage AI to enhance organizational productivity.
Job details
ADVANCE YOUR CAREER. ADVANCE THE WORLD. At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future. Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career. THE ROLE: AMD is looking for an AI solutions systems Engineer who is passionate about complex and innovative AI solutions, AI infrastructure, fleet management and automation solutions. You will be a member of a core team of incredibly talented industry specialists and will work with the very latest hardware and software technology. THE PERSON: The ideal candidate should be passionate about software engineering, system design, automation and possess leadership skills to drive sophisticated issues to resolution. Able to communicate effectively and work optimally with different teams across AMD. KEY RESPONSIBILITIES: Work with AMD’s architecture specialists and hybrid cloud/edge developers to architect, develop and scale AI fleet management solutions spanning multiple global regions Drive performance optimization and reliability based on industry best practices Leverage AI to enhance developer and organizational productivity and solution capabilities Influence roadmap and long-term architectural evolution of cutting-edge AI compute solutions PREFERRED EXPERIENCE: Extensive experience building and operating distributed systems or large-scale infrastructure platforms Strong proficiency in Python and JS based stacks (eg, Angular, React etc) Experience designing and operating containerized systems (Docker, Kubernetes) and managing compute clusters (e.g., Slurm or similar schedulers) Proven experience with public cloud platforms (AWS, Azure, or GCP) and/or hybrid or edge infrastructure Strong knowledge of SQL databases Deep understanding of Linux systems, administration, and scripting General understanding in platform control, fleet management, IT and data center infrastructure Familiar with AI dev tools such as Cursor or Claude Demonstrated experience leading complex architectural initiatives across teams Experience building AI-driven or agent-based systems in production environments Contributions to open-source infrastructure projects Experience with authentication and security Experience with DevOps practices, CI/CD pipelines, and infrastructure-as-code tools (Terraform, Ansible, etc.) PREFERRED ACADEMIC CREDENTIALS: Bachelor or master's degree in Electrical/Computer Engineering, Mathematics, Computer Science or an equivalent preferred LOCATION: Austin, Texas or Markham, Ontario #LI-DNI Benefits offered are described: AMD benefits at a glance. AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process. AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here. This posting is for an existing vacancy.
What you’ll do
Architect, develop, and scale AI fleet management solutions across global regions while driving performance optimization and reliability. Influence the long-term architectural evolution of AI compute solutions and leverage AI to enhance organizational productivity.
Requirements
Requires extensive experience in building distributed systems, large-scale infrastructure, and containerized environments. Candidates should possess strong proficiency in Python, JavaScript, and cloud platforms, along with a degree in a relevant engineering or technical field.
Listed skills
- Kubernetes · Preferred
- SQL · Preferred
- CI/CD · Preferred
- Docker · Preferred
- JavaScript · Preferred
- Python · Preferred
Other relevant skills
Identified from the job description. Confirm important requirements above.
- Python
- JavaScript
- Docker
- Kubernetes
- Distributed Systems
- Cloud Computing
- SQL
- Linux Administration
- Fleet Management
- Infrastructure as Code
- CI/CD
- DevOps
- System Design
- Automation
- Performance Optimization
- Security
- Influencing Skills
- Hybrid Cloud Computing
- Edge Intelligence
- Infrastructure as Code (IaC)
- JavaScript (Programming Language)
- Amazon Web Services
- Angular (Web Framework)
- Systems Engineering
- Authentications
- Microsoft Azure
- Management
- Communication
- Computer Science
- Computer Engineering
- Data Center Infrastructure Efficiency
- Leadership
- Innovation
- Python (Programming Language)
- Mathematics
- Public Cloud
- Ansible
- Software Engineering
- SQL (Programming Language)
- Systems Design
- Scripting
- Reliability
- React.js (Javascript Library)
- Slurm (Batch Scheduling Software)
- Terraform
- Docker (Software)
- Artificial Intelligence Infrastructure
- Claude AI
Job areas
- Technology
- Software
- Engineering
- Data & Analytics
- Fleet Management Specialist
- Site Reliability Engineer
- Software and Applications Developers and Analysts Not Elsewhere Classified
- Computer and Information Research Scientists
More jobs from AMD
Software Development Engineer (Chip Product Security)
- Hybrid
- Markham, ON
- Posted Oct 10, 2026
Senior Firmware Engineer
- On-site
- Markham, ON
- Posted Oct 9, 2026
Systems Design Engineer (New Grad)
- On-site
- Markham, ON
- Posted Oct 9, 2026
Principal ASIC Design Engineer
- Hybrid
- Markham, ON
- Posted Oct 9, 2026
