Skills highlighted in blue are preferred key skills
Design dependable pipelines from operational sources into analytical platforms. Implement transformations, business rules and reusable data workflows. Use distributed processing when data volume or complexity requires it. Manage scheduling, dependencies, retries, logging and pipeline monitoring. Diagnose failed jobs and data-quality issues and implement lasting fixes. Work with analysts and scientists to shape useful curated datasets. Bring strong sql and practical cloud or distributed-data engineering experience.
Opportunity for a PySpark Data Pipeline Developer to contribute to a growing technology, analytics, enterprise applications, or sales function.
Client Company