
Data Engineer, PySpark & ML Infrastructure
Overview
You will build and support ML infrastructure, feature stores, and large-scale data pipelines using PySpark, Python, and SQL. You will join a team focused on data engineering, not model development. Your work will power machine learning workflows for a Google Cloud environment. This role centers on infrastructure and pipeline reliability, offering the chance to shape data systems from day one.
What You'll Do6
- 1Build scalable data pipelines using PySpark and Python to process large datasets
- 2Design and manage feature stores for machine learning applications, ensuring data consistency and availability
- 3Optimize SQL queries and BigQuery workloads for performance and cost efficiency
- 4Ship reliable data ingestion and transformation jobs for ML infrastructure
- 5Debug and resolve pipeline failures, maintaining data quality and timeliness
- 6Collaborate with data scientists to understand feature requirements and deliver robust data solutions
Requirements6
- 15+ years in data engineering with Python and SQL for large-scale data processing
- 23+ years hands-on with PySpark for ETL and data transformation
- 3Experience building and maintaining BigQuery datasets and queries
- 4Track record of developing and managing data pipelines in Google Cloud Platform
- 5Strong understanding of ML infrastructure concepts, including feature stores
- 6Proven ability to write efficient, tested code for production environments
Salary Insight
Salary not disclosed in listing
Location
Required Skills
Similar open positions
Explore active roles that match your skills and interests.

Finoit Inc.
VerifiedSenior Software Engineer Data Infrastructure
Design and lead development of scalable data pipelines for AI/ML platforms. Own design and implementation of distributed systems using Python and cloud services. Drive improvements in data quality and visualization. Lead cross-functional teams to deliver high-performance solutions.
Apex Systems
VerifiedData Engineer - Spark & Scala
As a Data Engineer, you'll design and build ETL pipelines that feed a data lake supporting Spark and Scala workloads. Your pipelines will power anomaly detection models and AI agents, handling a growing portfolio of network data sources. Part of the Apex Systems team in Denver, Colorado, you'll work onsite in Greenwood Village with remote flexibility. This role focuses on building robust and scalable data infrastructure to ensure data quality and availability for downstream consumers.

AgreeYa Solutions
VerifiedML Engineer, Databricks & Python
You will own the machine learning engineering lifecycle for a Dallas-based client, designing and deploying scalable data pipelines. You will work with Python, PySpark, and Databricks to process streaming and batch data, collaborating with data scientists and software engineers. You will build real-time features using Kafka and Snowflake, ensuring low-latency model serving. This 18-month contract offers the chance to tackle complex data modeling challenges, including slowly changing dimensions in MongoDB and PostgreSQL.

PRIMUS Global Services Inc.
VerifiedBig Data Machine Learning Engineer, Spark & Python
You will design and build large-scale data processing pipelines and machine learning platforms using Apache Spark, Python, and SQL. You will work within a distributed systems team at PRIMUS Global Services Inc., collaborating on feature extraction and data analysis. This role requires deep experience with Hadoop, Snowflake, and related big data tools, and offers the chance to solve complex data challenges on an onsite team in Sunnyvale, CA.

Vertex Solutions Inc.
VerifiedSystems Manager, Google Cloud Platform Data Engineering
You lead the hands-on technical team that designs, builds, and maintains enterprise-grade data pipelines powering the Enterprise Data & Analytics Platform (EDAP). Your team owns the core stages of the data lifecycle on Google Cloud Platform, from transfer and ingestion to transformation, curation, and exposure. You manage a highly skilled group of engineers, setting technical direction and ensuring delivery at scale. This role combines deep GCP expertise with team leadership to drive reliable, high-performance data infrastructure.

Delviom LLC
VerifiedData Engineer, AWS & Python
Data Engineer will own the build of scalable data pipelines on AWS, transforming raw data into analytics-ready assets. You will join a platform team of 6 engineers, working with Python, PySpark, Snowflake, and dbt to unify data for product and business units. This contract role offers direct ownership of cloud infrastructure and data flow, with a focus on Terraform and CI/CD automation. You will bring 5+ years of experience and a track record of shipping production data systems.