Junior Big Data Hadoop And Spark Pipeline Engineer
Job Description
What role will you play in the team
As a Junior Big Data Hadoop and Spark Pipeline Engineer, you will serve as a foundational pillar within our core Data Science, Analytics & Machine Learning department in Kolkata. Your primary focus will be designing, developing, and maintaining high-throughput distributed data pipelines. You will collaborate closely with senior data architects and machine learning engineers to ingest, clean, and process massive volumes of structured and unstructured enterprise datasets, ensuring seamless data flow for downstream analytical models and predictive algorithms.
What you will do
You will write and optimize complex Apache Spark jobs, manage Hadoop distributed file systems, and configure Apache Kafka event streams to handle real-time data ingestion. Your daily routines will involve monitoring cluster performance, troubleshooting pipeline failures, writing efficient ETL workflows using Talend and Hive, and implementing robust data governance policies. Furthermore, you will curate structured database stores utilizing PostgreSQL, MySQL, and MongoDB to support rapid querying by data science teams.
Key responsibility
- Develop and maintain scalable batch and real-time data processing pipelines using Apache Spark, Hadoop, and Apache Kafka.
- Design and implement efficient ETL and ELT data integration workflows leveraging tools like Talend and Apache NiFi.
- Optimize large-scale data queries and database schemas across relational and NoSQL platforms including PostgreSQL, MySQL, and MongoDB.
- Monitor distributed computing clusters, diagnose performance bottlenecks, and tune resource allocation for maximum throughput.
- Collaborate with machine learning engineers to format, clean, and deliver enterprise datasets required for feature engineering and model training.
- Implement rigorous data validation, cleansing protocols, and error-handling mechanisms within automated data ingestion pipelines.
- Maintain comprehensive technical documentation of data architecture blueprints, pipeline dependencies, and operational runbooks.
Required Qualification and Skills
- Bachelor's degree in Computer Science, Information Technology, Statistics, or a closely related quantitative discipline.
- 0 to 1 year of hands-on professional or academic experience in big data processing, database management, and pipeline engineering.
- Proficiency in core programming languages including Python and SQL for data manipulation and querying.
- Demonstrated practical knowledge of Big Data Technologies such as Hadoop, Spark, and Apache Kafka.
- Familiarity with database administration and querying concepts across PostgreSQL, MySQL, and MongoDB.
- Understanding of ETL tools like Talend, Informatica, or Apache NiFi for data movement orchestration.
- Strong analytical problem-solving skills, ability to work collaboratively in agile teams, and excellent written communication abilities.
Looking to get Placed? Try our Placement Guarantee Plan
Benefits Included
- Competitive entry-level annual salary package with performance-based appraisal reviews conducted semi-annually.
- Comprehensive health, medical, and life insurance coverage extending to immediate family members.
- Robust professional development stipend allocated for certifications in big data engineering and cloud platforms.
- Flexible working hours with modern ergonomic workstation setups provided for optimal productivity.
- Generous paid time off, casual leave policies, and dedicated mental health wellness days.
- Regular internal technical workshops, hackathons, and mentorship programs led by industry veterans.
A Day in the Life
Your day begins by reviewing real-time monitoring dashboards for your Apache Spark and Hadoop clusters running overnight batch loads. You spend the morning debugging a slow-running ETL pipeline and optimizing SQL queries for a PostgreSQL data warehouse. After a quick stand-up with the data science team, you dedicate the afternoon to writing Python scripts for automated data validation and testing a new Apache Kafka streaming topic before its deployment to staging.
Skills
PythonSQLHadoopSparkApache KafkaPostgreSQLMySQLMongoDBTalendApache NifiIf a job posting appears fraudulent, asks for payment, contains misleading information, or violates our guidelines, please report it immediately. Our team will review it promptly, Jobaaj does not charge any fee from the applicants.
About Company
Jobaaj is India’s premier job portal, offering a comprehensive suite of services that include recruitment, training, placements, and career management solutions. With a dedicated focus on the Accounting, Finance, Data, and Consulting sectors, Jobaaj bridges the gap between skilled professionals and dynamic opportunities. Through an innovative approach and a commitment to excellence, Jobaaj ensures that job seekers and employers alike find the perfect match to drive mutual success.
Guided by a vision to create a close-knit professional community, Jobaaj aspires to help one crore individuals secure their dream jobs by 2030. By combining industry expertise with in-house tools like AI-powered portals, Resume Builder, and LinkedIn Optimizer, the platform provides candidates with the resources needed to stand out in an increasingly competitive job market. Jobaaj also offers mentorship programs and skill-enhancement training, equipping individuals with the knowledge and experience required to excel in their chosen fields.
Under the leadership of an experienced team, Jobaaj operates as a holistic ecosystem that goes beyond traditional job portals. From guiding students in identifying their ideal career domains to offering hands-on internships and facilitating placements, Jobaaj delivers a seamless journey for professionals at every stage. This unique approach, supported by years of firsthand industry experience, ensures personalized solutions that cater to the specific needs of clients and candidates.
Ranked as one of India’s top job portals, particularly in the Finance and Accounting sectors, Jobaaj stands out as a trusted platform for both job seekers and recruiters. With a mission to revolutionize the recruitment landscape, Jobaaj continues to expand its connected ecosystem, creating value for its community and shaping the future of employment in India.
Important dates & deadlines?
Application Deadline
25 Oct 26, 09:00 PM IST
Similar Jobs
View AllDon't Miss out any Updates
Subscribe now for the latest job alerts
and never miss an update

