Junior Big Data And Apache Kafka Stream Processing Engineer
Job Description
What role will you play in the team
You will play a critical foundational role within our enterprise data infrastructure engineering team, focusing on ingestion pipelines and real-time event streaming systems. Your primary objective will be to design, develop, and maintain robust data streaming architectures utilizing Apache Kafka and distributed computing frameworks. You will bridge the gap between raw data generators and downstream machine learning models by ensuring reliable, high-throughput message delivery and stream processing across massive enterprise datasets.
What you will do
You will actively build and optimize distributed data ingestion pipelines, configure Kafka brokers, topics, and consumer groups, and integrate big data ecosystems with relational and non-relational database management systems. Your day-to-day engineering duties will involve writing high-performance transformation scripts, monitoring stream latency, debugging distributed pipeline bottlenecks, and collaborating closely with data scientists to deliver clean, production-ready streaming features for predictive analytical models.
Key responsibility
- Design, implement, and maintain real-time data pipelines utilizing Apache Kafka, ZooKeeper, and Spark Streaming components.
- Develop scalable ETL and data ingestion jobs to process terabytes of incoming structured and unstructured telemetry data.
- Configure and manage Kafka producers, consumers, and cluster parameters to ensure fault-tolerance and high availability.
- Integrate big data storage layers with PostgreSQL, MySQL, and Hadoop distributed file systems for long-term retention.
- Monitor pipeline performance, identify processing bottlenecks, and tune JVM garbage collection and network buffers.
- Write comprehensive unit and integration tests for all streaming components using automated CI/CD frameworks.
- Collaborate with cross-functional teams of data architects and machine learning engineers to align stream schemas.
Required Qualification and Skills
Candidates must possess a Bachelor's degree in Computer Science, Information Technology, or a related quantitative discipline with hands-on technical competencies. Essential technical requirements include strong proficiency in Python and SQL for data manipulation. Hands-on experience with Apache Kafka, Hadoop, and Spark Big Data Technologies is mandatory. Working knowledge of relational databases such as PostgreSQL and MySQL, alongside experience with ETL Tools like Talend or Informatica, is strictly required. Strong debugging skills, understanding of distributed systems architecture, and familiarity with Linux environment management are vital.
Benefits Included
Looking to get Placed? Try our Placement Guarantee Plan
- Comprehensive health, medical, and life insurance coverage for employees and immediate family members.
- Competitive performance-based annual bonuses linked directly to engineering milestone achievements.
- Sponsored professional certifications in Big Data Engineering, Apache Kafka, and Cloud Infrastructure.
- Flexible work arrangements with standard work from office protocols and ergonomic workspace allowances.
- Generous paid time off, maternal and paternal leave policies, and annual wellness stipend packages.
A Day in the Life
Your typical workday begins by reviewing stream processing dashboards and investigating overnight Kafka cluster lag alerts. You will spend the morning writing Python scripts to parse incoming JSON payloads into structured PostgreSQL tables. In the afternoon, you will participate in an agile sprint planning session, refactor a Spark streaming transformation job with senior architects, and test a newly deployed data ingestion pipeline in a staging environment.
Skills
PythonSQLPostgreSQLMySQLBig Data TechnologiesHadoopSparkApache KafkaETL ToolsTalendIf a job posting appears fraudulent, asks for payment, contains misleading information, or violates our guidelines, please report it immediately. Our team will review it promptly, Jobaaj does not charge any fee from the applicants.
About Company
Jobaaj is India’s premier job portal, offering a comprehensive suite of services that include recruitment, training, placements, and career management solutions. With a dedicated focus on the Accounting, Finance, Data, and Consulting sectors, Jobaaj bridges the gap between skilled professionals and dynamic opportunities. Through an innovative approach and a commitment to excellence, Jobaaj ensures that job seekers and employers alike find the perfect match to drive mutual success.
Guided by a vision to create a close-knit professional community, Jobaaj aspires to help one crore individuals secure their dream jobs by 2030. By combining industry expertise with in-house tools like AI-powered portals, Resume Builder, and LinkedIn Optimizer, the platform provides candidates with the resources needed to stand out in an increasingly competitive job market. Jobaaj also offers mentorship programs and skill-enhancement training, equipping individuals with the knowledge and experience required to excel in their chosen fields.
Under the leadership of an experienced team, Jobaaj operates as a holistic ecosystem that goes beyond traditional job portals. From guiding students in identifying their ideal career domains to offering hands-on internships and facilitating placements, Jobaaj delivers a seamless journey for professionals at every stage. This unique approach, supported by years of firsthand industry experience, ensures personalized solutions that cater to the specific needs of clients and candidates.
Ranked as one of India’s top job portals, particularly in the Finance and Accounting sectors, Jobaaj stands out as a trusted platform for both job seekers and recruiters. With a mission to revolutionize the recruitment landscape, Jobaaj continues to expand its connected ecosystem, creating value for its community and shaping the future of employment in India.
Important dates & deadlines?
Application Deadline
01 Oct 26, 09:00 PM IST
Similar Jobs
View All




