Lead I - Data Engineering (Python/Pyspark, Azure Databricks)

Trivandrum · Onsite Contract 5-7yrs B.Tech/MCA Direct Client hiring Posted 5 months ago💰 Not disclosed
Apply now

Must Have Skills

nilPythonPySparkLinux Shell scriptingDatabricksDBT (Data Build Tool)Data Modeling, Quality & TestingGCP, Azure or AWSSDLCNIL

Preferred Skills

nilPythonPySparkLinux Shell scriptingDatabricksDBT (Data Build Tool)Data Modeling, Quality & TestingGCP, Azure or AWSSDLCNIL


Experience: 6 – 10 Years

Employment Type: Full-Time / Contract (as applicable)

Shift: As per project requirements

Job Summary

You will lead and contribute to the design, development, optimization, and delivery of large-scale data platforms and pipelines. This role combines hands-on engineering with team leadership, code quality ownership, SDLC governance, and strong cross-functional collaboration to deliver robust, scalable and production-ready data solutions using modern cloud stacks (Databricks, DBT, Python, PySpark).

Mandatory Skills

  1. Python for data engineering & automation
  2. PySpark for scalable distributed data processing
  3. Linux Shell scripting for data workflows & automation
  4. Databricks platform & Spark ecosystem
  5. DBT (Data Build Tool) for modular data transformations
  6. Data Modeling, Quality & Testing expertise
  7. Experience implementing cloud data tech on GCP, Azure or AWS
  8. Strong SDLC knowledge & code quality practices (version control, reviews)

These skills are core to building resilient data processing ecosystems and pipelines at scale.

Key Responsibilities

Engineering & Delivery

  1. Interpret solution designs and build data pipelines aligned to business requirements.
  2. Develop, debug, test, and deploy ETL/ELT workflows using Python, PySpark, and Databricks.
  3. Build reusable, efficient data transformation layers using DBT (models, tests, macros).
  4. Implement scalable data architectures using cloud services (e.g., Data Lake, Unity Catalog).
  5. Optimize data models, storage formats (e.g., Delta Lake), and pipeline performance.
  6. Ensure data correctness, reliability, and quality assurance through testing and monitoring.
  7. Maintain documentation of design, data flows, pipeline logic, and best practices.

SDLC & Quality Engineering

  1. Follow engineering standards, enforce coding best practices, and conduct code reviews.
  2. Create and manage unit tests, integration tests, and pipeline validation.
  3. Ensure timely deliveries per schedule with minimal production defects.
  4. Utilize version control (GitHub/GitLab) and CI/CD pipelines for deployments.

Leadership & Collaboration

  1. Provide engineering direction, mentorship, and technical coaching to junior data engineers.
  2. Collaborate with architects, data scientists, product owners, and stakeholders.
  3. Conduct design discussions, technical demos, and solution walkthroughs with customers.
  4. Lead efforts to document patterns, templates, checklists, and engineering guidelines.

Project & Team Management

  1. Provide effort estimates, manage user stories, and track delivery progress.
  2. Manage defect root cause analysis (RCA), mitigation, and trend reduction practices.
  3. Support Agile ceremonies (Sprint planning, standups, retrospectives).
  4. Ensure team alignment with project goals and high performance expectations.

Technical Outputs Expected

  1. High-quality, production-ready code aligned with architectural designs.
  2. Code reviews, templates, standards, and reusable components.
  3. Functional & technical documentation (requirements, designs, tests, runbooks).
  4. Data pipeline orchestration and operational tooling for monitoring and error handling.
  5. Clear guidelines for development, testing, and release processes.

Performance Measures

  1. Engineering process adherence & SDLC compliance.
  2. On-time delivery of features with minimal defects.
  3. Number of production issues and post-release defects.
  4. Quality of documentation and code review results.
  5. Effectiveness of automation and test coverage.

Knowledge & Skills Examples

  1. Deep understanding of data architectures, distributed processing, and data modeling.
  2. Competence in Python, PySpark, SQL, and data transformation frameworks.
  3. Experience with data orchestration tools (Airflow, Databricks Jobs).
  4. Familiarity with cloud platforms (GCP / Azure / AWS) and storage technologies.
  5. Strong analytical, problem-solving, and communication skills.

Soft Skills & Leadership Traits

  1. Ability to break down complex problems and propose logical solutions.
  2. Mentor and guide team members effectively.
  3. Maintain high motivation and positive team dynamics.
  4. Communicate clearly with technical and non-technical stakeholders.
  5. Drive customer satisfaction and build confidence through quality delivery.

Certifications & Continuous Learning

  1. Relevant certifications such as Databricks Certified Data Engineer, DBT Fundamentals, Cloud provider certifications (AWS/Azure/GCP) are beneficial.
  2. Encourage and complete domain-specific certifications and learning paths.


Apply to this job

Required — upload a file or paste the text