- Experience
- 2–5 yrs
- Salary
- —
- Openings
- 1
- Posted
- 1 day ago
- Work mode
- In office
- Education
- B.C.A., B.Sc., B.Tech, or B.E. in any specialization
- Eligibility
- Candidates with B.C.A., B.Sc., B.Tech, or B.E. degrees in any specialization are eligible to apply.
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Role Overview and Responsibilities
- Design and implement scalable data pipelines leveraging PySpark technology.
- Create ETL workflows and perform data transformation tasks to handle large datasets.
- Process extensive structured and unstructured data collections.
- Develop Spark applications optimized for enhanced performance.
- Utilize SQL and Spark SQL for effective data querying and processing operations.
- Maintain high standards of data quality, ensuring scalability and reliability of data solutions.
- Identify and resolve issues encountered in data pipeline operations.
- Collaborate effectively with cross-functional business and technical teams to align solutions with objectives.
Candidate Profile and Requirements
- Possess between 2 to 5 years of professional experience in Data Engineering roles.
- Demonstrated strong hands-on expertise with PySpark and Apache Spark frameworks.
- Proficient in programming with Python.
- Solid understanding and working knowledge of SQL.
- Experience working with ETL processes and Big Data technologies.
- Exposure or familiarity with Databricks, Hadoop, or Hive is advantageous.
- Excellent analytical and problem-solving skills.
Eligibility Criteria
Applicants should hold a B.C.A., B.Sc., B.Tech, or B.E. degree in any specialization.
Company Information
Location details: Floor No.: 3rd Cross, Building No./Flat No.: Infosys Limited, Buildings 44 and 97A, Electronic City, Bangalore, Karnataka, India.
Job Location
Noida, India.
Minimum education
Bachelor's Degree
Industry
IT Services & ConsultingSkills
Tools & software
Apache Spark
· 2 to 5 years required
Apache Hive
required
Apache Hadoop
required
Databricks
required
How they work
Teamwork & Collaboration
Problem Solving