Please click on the Apply to verify the status of jobs posted more than 15 days ago, as they may have expired. Similar Jobs
Job Description
Position Title: AI/HPC System Engineer
Location: San Jose, CA (Onsite)
Description
We are hiring an AI/HPC System Engineer to build and operate the compute infrastructure supporting our HPC and AI development workloads. This role deploys, automates, and maintains GPU clusters across on-premise and cloud environments, delivering reliable, scalable, and cost-efficient compute for engineering and R&D teams.
Responsibilities:
• GPU/HPC infrastructure: Build, configure, and operate GPU and HPC clusters across compute, storage, and networking; support capacity planning, performance tuning, and optimization for AI training, inference, and compute-intensive workloads
• Hybrid cloud infrastructure: Deploy and maintain compute environments spanning on-premise and public cloud, and contribute to modernization and scaling initiatives for HPC/AI infrastructure
• Automation and observability: Implement infrastructure-as-code, provisioning automation, monitoring, and alerting, and drive improvements in resource utilization and efficiency
• AI platform support: Deploy, integrate, and support LLM APIs, coding assistants, and AI/agent platforms used by internal engineering teams
• Operations and collaboration: Troubleshoot and resolve infrastructure issues, document standards and runbooks, and work with relevant stakeholders to support day-to-day IT operations
Qualifications: Looking to get Placed? Try our Placement Guarantee Plan
• Bachelors degree in Computer Science, Engineering, or a related technical field
• 3+ years of hands-on experience in IT infrastructure, cloud, platform engineering, or HPC
• Hands-on experience with Linux-based infrastructure and public cloud environments such as AWS, Azure, or Google Cloud Platform
• Experience deploying or operating GPU/HPC environments, including workload scheduling or orchestration platforms such as Kubernetes or Slurm
• Experience with infrastructure automation, monitoring, troubleshooting, and performance optimization
• Solid understanding of compute, storage, networking, and container technologies; experience with AI/ML infrastructure or workloads is a plus
• Strong collaboration and communication skills, with the ability to work across engineering and IT teams
Skills
Cloud InfrastructureAi/mlAiGoogle CloudMlLlmIf a job posting appears fraudulent, asks for payment, contains misleading information, or violates our guidelines, please report it immediately. Our team will review it promptly, Jobaaj does not charge any fee from the applicants.
About Company
Xoriant, a leading technology services and consulting company, specializes in product engineering and technology solutions. With a commitment to innovation and excellence, Xoriant empowers clients across various industries by delivering cutting-edge solutions that enhance competitiveness and drive growth. Xoriant careers offer professionals the opportunity to work on transformative projects, leveraging the latest technologies in data analytics, cloud computing, cybersecurity, and more. As part of the Xoriant team, employees engage in a culture of learning, collaboration, and innovation, making significant contributions to clients' success. Xoriant careers are pathways to personal and professional growth, where talent meets opportunity in the tech world.
Important dates & deadlines?
Application Deadline
26 Oct 26, 04:15 PM IST
Similar Jobs
View All

