Application Reliability Engineer – Data & AI

Department Icon Data Science Analytics & Machine Learning
149+ Applicants
Posted: 3 weeks ago
5-7 years
Pune, Maharashtra
work from office

Posted: 3 weeks ago
|
Applicants: 149+
Job Description
About Company
Similar Jobs
Please verify your account first! Send OTP

Please click on the Apply to verify the status of jobs posted more than 15 days ago, as they may have expired. Similar Jobs

Job Description

Application Reliability Engineer - Data & AI

- - - - - - - - - - - -

The Data and AI Application Reliability Engineer you are responsible for owning the post-deployment health, latency, and performance of productionized AI application and Medallion architectures followed for applications.

Acting as the problem-solver for technical issues affecting data pipelines, databases, and deployed AI/ML models, ensuring continuous operation and high user satisfaction.

Partner with core Feature teams to perform root-cause analysis, optimize PySpark jobs, and reduce system debt

Build automated observability dashboards and self-healing mechanisms for automated failure recovery

As a part of Job you are required to balance your responsibilities between Proactive Engineering & Automation and Operational Reliability/ Production Health.

Key Responsibilities

Analyze and refactor resource-intensive PySpark jobs, queries, and API endpoints to optimize cost, execution speed, and compute efficiency.

Develop automated recovery routines, DAG rerun triggers, and data quality checks to minimize manual intervention.

Partner closely with Feature Teams and Architects to establish strict Definition of Done (DoD) standards and production readiness gates for new deployments.

Design, implement, and maintain real-time monitoring and alerting frameworks for Medallion architecture pipelines, feature stores, and AI Applications.

Lead technical resolution for high-priority production incidents, conducting thorough post-mortems to eliminate recurring failure patterns.

Technical Expertise:

 Azure Data Engineering Stack

Proficiency in Databricks , Python, and PySpark.

Azure (ADLS Gen2, Azure Data Factory, Key Vault, Azure DevOps)

Hands-on experience with Medallion architecture.

 Cloud and DevOps Fundamentals

Understanding of cloud computing concepts and Services, specifically Microsoft Azure.

Good handson experience on Python.

 Good to Have Technical Abilities :

Looking to get Placed? Try our Placement Guarantee Plan

Understanding of PowerBI reportswill be a plus.

DevOps & CI/CD Fundamentals

AI & Machine Learning Fundamentals

Data Science Fundamentals

Generative AI & LLM Fundamentals

Behaviour

Problem Solver:Ability to reverse-engineering complex system behavior and tracking down bugs.

Automation-First:An instinct to automate repetitive tasks .

Clear Communicator: Ability to explain technical root causes to non-technical stakeholders clearly.
Relevant work experience - 5 yrs

Skills

PythonData ScienceMachine LearningAi/mlAiMlLlmGenerative AiDatabricks

If a job posting appears fraudulent, asks for payment, contains misleading information, or violates our guidelines, please report it immediately. Our team will review it promptly, Jobaaj does not charge any fee from the applicants.

About Company

Michelin, a leading tire manufacturer, develops and sells tires for various vehicles, including cars, trucks, motorcycles, and bicycles. They are also involved in tire services, and offer a range of solutions for mobility. Michelin's commitment to innovation and sustainability is evident in their research and development efforts.

Important dates & deadlines?

Application Deadline

26 Oct 26, 02:34 PM IST

Similar Jobs

View All
Loading...
Bag Logo
Jobaaj
Don't Miss out any Updates

Subscribe now for the latest job alerts
and never miss an update

Job Alert
Google hiring for Specific Roles Apply Now!
1 min ago
New Opportunity
Amazon is hiring freshers Apply Now!
5 min ago
Featured Jobs
Microsoft opening 50+ positions Apply Now!
10 min ago

Application Reliability Engineer – Data & AI

Share with