Lead I - Data Engineering (Python/Pyspark, Azure Databricks)
Trivandrum · Onsite Contract 5-7yrs B.Tech/MCA Direct Client hiring Posted 5 months ago💰 Not disclosed
Must Have Skills
nilPythonPySparkLinux Shell scriptingDatabricksDBT (Data Build Tool)Data Modeling, Quality & TestingGCP, Azure or AWSSDLCNIL
Preferred Skills
nilPythonPySparkLinux Shell scriptingDatabricksDBT (Data Build Tool)Data Modeling, Quality & TestingGCP, Azure or AWSSDLCNIL
Experience: 6 – 10 Years
Employment Type: Full-Time / Contract (as applicable)
Shift: As per project requirements
Job Summary
You will lead and contribute to the design, development, optimization, and delivery of large-scale data platforms and pipelines. This role combines hands-on engineering with team leadership, code quality ownership, SDLC governance, and strong cross-functional collaboration to deliver robust, scalable and production-ready data solutions using modern cloud stacks (Databricks, DBT, Python, PySpark).
Mandatory Skills
- Python for data engineering & automation
- PySpark for scalable distributed data processing
- Linux Shell scripting for data workflows & automation
- Databricks platform & Spark ecosystem
- DBT (Data Build Tool) for modular data transformations
- Data Modeling, Quality & Testing expertise
- Experience implementing cloud data tech on GCP, Azure or AWS
- Strong SDLC knowledge & code quality practices (version control, reviews)
These skills are core to building resilient data processing ecosystems and pipelines at scale.
Key Responsibilities
Engineering & Delivery
- Interpret solution designs and build data pipelines aligned to business requirements.
- Develop, debug, test, and deploy ETL/ELT workflows using Python, PySpark, and Databricks.
- Build reusable, efficient data transformation layers using DBT (models, tests, macros).
- Implement scalable data architectures using cloud services (e.g., Data Lake, Unity Catalog).
- Optimize data models, storage formats (e.g., Delta Lake), and pipeline performance.
- Ensure data correctness, reliability, and quality assurance through testing and monitoring.
- Maintain documentation of design, data flows, pipeline logic, and best practices.
SDLC & Quality Engineering
- Follow engineering standards, enforce coding best practices, and conduct code reviews.
- Create and manage unit tests, integration tests, and pipeline validation.
- Ensure timely deliveries per schedule with minimal production defects.
- Utilize version control (GitHub/GitLab) and CI/CD pipelines for deployments.
Leadership & Collaboration
- Provide engineering direction, mentorship, and technical coaching to junior data engineers.
- Collaborate with architects, data scientists, product owners, and stakeholders.
- Conduct design discussions, technical demos, and solution walkthroughs with customers.
- Lead efforts to document patterns, templates, checklists, and engineering guidelines.
Project & Team Management
- Provide effort estimates, manage user stories, and track delivery progress.
- Manage defect root cause analysis (RCA), mitigation, and trend reduction practices.
- Support Agile ceremonies (Sprint planning, standups, retrospectives).
- Ensure team alignment with project goals and high performance expectations.
Technical Outputs Expected
- High-quality, production-ready code aligned with architectural designs.
- Code reviews, templates, standards, and reusable components.
- Functional & technical documentation (requirements, designs, tests, runbooks).
- Data pipeline orchestration and operational tooling for monitoring and error handling.
- Clear guidelines for development, testing, and release processes.
Performance Measures
- Engineering process adherence & SDLC compliance.
- On-time delivery of features with minimal defects.
- Number of production issues and post-release defects.
- Quality of documentation and code review results.
- Effectiveness of automation and test coverage.
Knowledge & Skills Examples
- Deep understanding of data architectures, distributed processing, and data modeling.
- Competence in Python, PySpark, SQL, and data transformation frameworks.
- Experience with data orchestration tools (Airflow, Databricks Jobs).
- Familiarity with cloud platforms (GCP / Azure / AWS) and storage technologies.
- Strong analytical, problem-solving, and communication skills.
Soft Skills & Leadership Traits
- Ability to break down complex problems and propose logical solutions.
- Mentor and guide team members effectively.
- Maintain high motivation and positive team dynamics.
- Communicate clearly with technical and non-technical stakeholders.
- Drive customer satisfaction and build confidence through quality delivery.
Certifications & Continuous Learning
- Relevant certifications such as Databricks Certified Data Engineer, DBT Fundamentals, Cloud provider certifications (AWS/Azure/GCP) are beneficial.
- Encourage and complete domain-specific certifications and learning paths.
Apply to this job
Required — upload a file or paste the text
