Job Title: Data Engineering InternRemote - Night shift (5:30 pm to 2:30 am)As a
Data Engineering Intern, you will have the opportunity to work closely with the data engineering team and gain practical experience in
designing, developing, and
maintaining data pipelines and
infrastructure. You will contribute to the development and
optimization of data storage,processing, and retrieval systems,
enabling efficient and
accurate data analysis. This internship will provide you with hands-on experience in the field of
data engineering and
exposure to various tools, technologies, and best practices.
ResponsibilitiesDesigning and developing data pipelines: Collaborate with the team to design and
implement efficient and
scalable data pipelines, including
data ingestion, transformation, and
storage.
Data infrastructure management: Assist in setting up and
maintaining data infrastructure components such as
databases, data warehouses, and
data lakes. Monitor and
optimize performance to ensure
data availability and
reliability.Data integration: Integrate data from various sources and formats, ensuring data quality and consistency. Develop data integration workflows and scripts to automate data processing tasks.
Data transformation and modeling: Implement data transformation processes to convert raw data into structured formats suitable for analysis. Work on data modeling activities to support data analysis and reporting requirements.
Data quality and validation: Implement data validation and quality checks to ensure accuracy, completeness, and consistency of data. Identify and resolve data quality issues in collaboration with the data engineering team.
Documentation and collaboration: Document technical specifications, data flows, and processes. Collaborate with cross-functional teams, including data analysts and data scientists, to understand data requirements and support their analytical needs.
Performance optimization: Identify opportunities for performance improvement and optimization of data processing and storage systems. Propose and implement solutions to enhance system efficiency and scalability.
Stay updated with industry trends: Keep abreast of the latest developments, tools, and techniques in data engineering and related fields. Share knowledge and contribute to the teams learning and growth.
RequirementsCurrently pursuing a bachelors or masters degree in Computer Science, Engineering, or a related field.
Strong programming skills in languages such as Python, Java, or Scala.Familiarity with
databases (SQL, NoSQL) and
data modeling concepts.Knowledge of
data processing frameworks and
technologies such as
Apache Spark, Apache Hadoop, Git, Jupyter Notebooks, Docker,Splunk, Elasticsearch / Logstash / Kibana (ELK Stack) or Apache Kafka.
Understanding of
data integration and
ETL (Extract, Transform, Load) processes.
Proficiency in working with
Linux/Unix environments and
command-line tools.Excellent
problem-solving and
analytical skills, with
attention to detail.
Good communication and collaboration skills.
Self-motivated and able to work independently as well as in a team.
Previous experience with
cloud platforms (e.g., AWS, Azure, Google Cloud) is a plus.