Job Description
- CI/CD for ML: Build and maintain automated pipelines for model training, testing, and deployment.
- Infrastructure: Containerize and orchestrate model-serving infrastructure (Docker, Kubernetes) at scale.
- Monitoring: Set up monitoring for model performance, data/concept drift, latency, and system health.
- Versioning & Reproducibility: Manage model/experiment versioning and reproducibility (MLflow, DVC, or similar).
- Cloud & GPU Management: Manage scalable, cost-efficient GPU/cloud infrastructure for training and inference workloads.
- Strong hands-on experience with Docker and Kubernetes
- CI/CD tooling (GitHub Actions, Jenkins, GitLab CI, or similar)
- Experience with model-serving frameworks (Triton Inference Server, TorchServe, or similar)
- Cloud platform experience (AWS/Azure/GCP), especially GPU infrastructure
- Python for automation and tooling
- Experience with experiment tracking and model registries (MLflow, DVC, Weights & Biases)
- Monitoring/observability tooling (Prometheus, Grafana, or similar)
- Production experience with Triton Inference Server: model repositories, ensembles, dynamic batching, instance groups
Looking to get Placed? Try our Placement Guarantee Plan
- GPU operations on Kubernetes: NVIDIA GPU Operator, MIG/time-slicing, node pools, driver/CUDA version management
- Hands-on with TensorRT / ONNX Runtime conversion and performance profiling (Nsight, perf_analyzer) gRPC and streaming service patterns; load testing tools (Locust, k6)
- Strong Linux, networking and debugging fundamentals for bare-metal environments
- Docker, Kubernetes, CI/CD, Python, cloud infrastructure (AWS/Azure/GCP)
- Experience deploying LLM/NLP inference services specifically (batching, quantization, low-latency serving)
- Familiarity with vLLM, Ollama, or similar for local LLM hosting
- Infrastructure-as-code experience (Terraform, Helm)
- Experience with NVIDIA NVCF, NeMo or NIM deployments
- Serving TTS/ASR models where time-to-first-byte matters
- Log/trace stacks (Loki, OpenTelemetry, ELK) and on-call tooling
Skills
LinusGCPCICDKubernetesAzureMonitoring ToolsPythonDockerAWSPythonGRPCNLPLLMAzureKubernetesIf a job posting appears fraudulent, asks for payment, contains misleading information, or violates our guidelines, please report it immediately. Our team will review it promptly, Jobaaj does not charge any fee from the applicants.
About Company
Important dates & deadlines?
Application Deadline
24 Nov 26, 06:25 PM IST
Similar Jobs
View AllDon't Miss out any Updates
Subscribe now for the latest job alerts
and never miss an update

