Please click on the Apply to verify the status of jobs posted more than 15 days ago, as they may have expired. Similar Jobs
Job Description
If working in an environment that encourages you to innovate and excel, not just in professional but personal life, interests you- you would enjoy your career with Quantiphi!
Role: Lead Support Engineer
Experience Level: 8-15 Years
Work location: Trivandrum
As a Application Lead for the Production Engineering (PE) and Service Delivery (SD) stream, you will take ownership of the reliability, scalability, and availability of critical healthcare AI platforms. You will move beyond day-to-day ticket resolution to define the Support Strategy, focusing on Site Reliability Engineering (SRE) principles.
You will lead a team of engineers, guiding them through complex incident resolutions, orchestrating cloud automation, and bridging the gap between Data Engineering, ML Engineering, and Operations. You will serve as the primary escalation point and technical architect for support operations, ensuring high availability for multi-geography projects.
Role & Responsibilities-
Technical Leadership & Strategy
- Lead and mentor senior and junior support engineers; conduct code reviews, technical sessions, and manage 24/7 shift rotations.
- Implement SRE best practices by defining and tracking SLOs, SLIs, and Error Budgets.
- Act as Major Incident Manager and primary escalation point, leading war rooms, RCAs, and corrective actions.
- Architect and enforce IaC standards using Terraform/CloudFormation to ensure reproducible, self-healing environments.
- Design and maintain monitoring, logging, and alerting systems for proactive issue detection.
- Ensure zero-downtime releases, automated rollbacks, and mature CI/CD pipelines.
- Partner with Solution Architects and engineering teams to improve system supportability and maintainability.
- Introduce modern cloud-native tools to enhance automation and operational efficiency.
Cloud & Infrastructure (AWS/GCP)
- Compute & Serverless: EC2/GCE, Lambda/Cloud Functions, autoscaling with custom metrics.
- Networking: VPCs, Transit Gateways, Load Balancers, DNS, VPN/Direct Connect.
- Security: IAM (least privilege), KMS, WAF, secure public endpoints.
- Storage: S3/GCS lifecycle policies, EBS/Persistent Disk optimization.
- Strong development experience in Python (Flask/Django/FastAPI) or Java (Spring Boot).
- Proven ability in deep code debugging, performance tuning, and memory leak analysis.
- Hands-on experience with PRs, code reviews, and automated testing.
- Microservices using REST, gRPC, Kafka/SQS/PubSub with resiliency patterns.
- Distributed tracing and observability using OpenTelemetry.
- Kubernetes (EKS/GKE): cluster upgrades, ingress, autoscaling (HPA/VPA), security policies.
- Experience with Helm charts and Kubernetes Operators.
- Terraform modules, remote state management, multi-environment setups.
- Ansible for configuration management and OS hardening.
- CI/CD with Jenkins, Gitflow/Trunk-based strategies.
- Artifact and container registry management (Docker, Artifactory, Nexus).
- PostgreSQL performance tuning, replication, and DR strategies.
- NoSQL databases (Cassandra/MongoDB) and caching layers (Redis/Memcached).
Skills
OperationsConfiguration ManagementOperational EfficiencyService DeliveryProductionDeliverySupport EngineerIf a job posting appears fraudulent, asks for payment, contains misleading information, or violates our guidelines, please report it immediately. Our team will review it promptly, Jobaaj does not charge any fee from the applicants.
About Company
Quantiphi, founded in 2013, is headquartered in McLean, Virginia and operates in AI-driven solutions. It operates as private, with a global workforce of approximately 1000+, with offices in Multiple locations globally including India, US, UK, Singapore.
Important dates & deadlines?
Application Deadline
14 Sep 26, 06:09 PM IST
Similar Jobs
View All




