Junior Big Data And Spark Pipeline Developer

Department Icon Data Science Analytics & Machine Learning
149+ Applicants
Posted: 7 hours ago
0-1 years
Bangalore
Work From Office
₹5,00,000 - ₹7,50,000

Posted: 7 hours ago
|
Applicants: 151+
Job Description
About Company
Similar Jobs
Please verify your account first! Send OTP

Job Description

What role will you play in the team

You will play a foundational role in our core Data Science, Analytics & Machine Learning infrastructure team in Bangalore. Your primary objective is to design, construct, and maintain robust data pipelines using distributed computing frameworks. You will bridge the gap between raw unstructured data sources and machine learning model training datasets, ensuring optimal performance, scalability, and high reliability across all data engineering workflows within enterprise operations.

What you will do

You will actively write, optimize, and deploy distributed data processing jobs using Apache Spark, Hadoop, and SQL. Collaborating closely with senior data engineers and machine learning scientists, you will ingest massive volumes of transactional logs and streaming data. Your day-to-day work involves troubleshooting pipeline bottlenecks, refining ETL logic, managing MongoDB and PostgreSQL databases, and validating data integrity to feed downstream predictive models and analytics dashboards seamlessly.

Key responsibility

  • Design, develop, and maintain large-scale distributed data pipelines utilizing Apache Spark, Hadoop, and advanced SQL querying techniques to process multi-terabyte datasets daily.
  • Integrate diverse data streams from enterprise transactional systems into centralized data lakes and data warehouses using optimized ETL tools and custom Python scripts.
  • Monitor, debug, and optimize existing data workflows running on cluster environments to reduce job execution time and minimize resource compute costs.
  • Collaborate with machine learning engineers to structure feature stores and deliver clean, transformed training data for production-grade predictive models and neural networks.
  • Implement rigorous data quality checks, schema validations, and automated error-handling mechanisms across all ingestion nodes to prevent data corruption.
  • Manage database schema changes, indexing strategies, and query performance tuning across PostgreSQL, MySQL, and MongoDB data storage layers.
  • Document pipeline architectures, data lineage maps, and operational runbooks to ensure seamless knowledge transfer across the core analytics engineering division.

Required Qualification and Skills

  • Bachelor's degree in Computer Science, Statistics, Data Science, Information Technology, or a highly quantitative engineering discipline from a recognized university.
  • Demonstrated foundational knowledge of Big Data Technologies, specifically Apache Spark, Hadoop distributed file systems, and distributed computing principles.
  • Strong programming proficiency in Python and advanced SQL for complex data manipulation, aggregation, and relational database management systems.
  • Hands-on experience working with modern database systems including PostgreSQL, MySQL, and NoSQL databases like MongoDB.
  • Familiarity with ETL concepts, data warehousing architecture, version control systems such as Git, and Linux command-line environments.
  • Excellent analytical thinking, structured problem-solving abilities, and effective communication skills to collaborate within cross-functional technical teams in Bangalore.

Benefits Included

  • Competitive annual compensation package with performance-linked bonuses and annual appraisal reviews based on technical milestone achievements.
  • Comprehensive health, medical, and life insurance coverage extending to immediate family members from day one of employment.
  • Generous paid time off, casual leaves, maternity and paternity leave policies, and dedicated mental health wellness days.
  • Continuous learning allowance for professional certifications, cloud computing credentials, advanced data engineering courses, and technical conference attendance.
  • State-of-the-art office workspace located in Bangalore equipped with ergonomic seating, high-performance computing hardware, and collaborative recreation zones.

A Day in the Life

Your day begins by checking cluster health metrics and reviewing automated pipeline execution logs from overnight batch runs in Bangalore. You spend the morning optimizing a slow-running Spark SQL join that processes customer transaction histories. After a quick sync with the data science squad, you write custom Python data transformation scripts, update PostgreSQL index configurations, and push your code through CI/CD pipelines before wrapping up with a technical retrospective.

Skills

PythonSQLDatabase ManagementPostgreSQLMySQLMongoDBBig Data TechnologiesHadoopSpark

If a job posting appears fraudulent, asks for payment, contains misleading information, or violates our guidelines, please report it immediately. Our team will review it promptly, Jobaaj does not charge any fee from the applicants.

About Company

Jobaaj is India’s premier job portal, offering a comprehensive suite of services that include recruitment, training, placements, and career management solutions. With a dedicated focus on the Accounting, Finance, Data, and Consulting sectors, Jobaaj bridges the gap between skilled professionals and dynamic opportunities. Through an innovative approach and a commitment to excellence, Jobaaj ensures that job seekers and employers alike find the perfect match to drive mutual success.

Guided by a vision to create a close-knit professional community, Jobaaj aspires to help one crore individuals secure their dream jobs by 2030. By combining industry expertise with in-house tools like AI-powered portals, Resume Builder, and LinkedIn Optimizer, the platform provides candidates with the resources needed to stand out in an increasingly competitive job market. Jobaaj also offers mentorship programs and skill-enhancement training, equipping individuals with the knowledge and experience required to excel in their chosen fields.

Under the leadership of an experienced team, Jobaaj operates as a holistic ecosystem that goes beyond traditional job portals. From guiding students in identifying their ideal career domains to offering hands-on internships and facilitating placements, Jobaaj delivers a seamless journey for professionals at every stage. This unique approach, supported by years of firsthand industry experience, ensures personalized solutions that cater to the specific needs of clients and candidates.

Ranked as one of India’s top job portals, particularly in the Finance and Accounting sectors, Jobaaj stands out as a trusted platform for both job seekers and recruiters. With a mission to revolutionize the recruitment landscape, Jobaaj continues to expand its connected ecosystem, creating value for its community and shaping the future of employment in India.

Read More

Important dates & deadlines?

Application Deadline

30 Sep 26, 09:00 PM IST

Similar Jobs

View All
Loading...
Bag Logo
Jobaaj
Don't Miss out any Updates

Subscribe now for the latest job alerts
and never miss an update

Job Alert
Google hiring for Specific Roles Apply Now!
1 min ago
New Opportunity
Amazon is hiring freshers Apply Now!
5 min ago
Featured Jobs
Microsoft opening 50+ positions Apply Now!
10 min ago

Junior Big Data And Spark Pipeline Developer

Share with