Octans Group LLC
Octans Group LLCVerified Source

Data Engineer, PySpark & ML Infrastructure

Onsite · San Jose, California
Posted August 13, 2026
contract

Overview

You will build and support ML infrastructure, feature stores, and large-scale data pipelines using PySpark, Python, and SQL. You will join a team focused on data engineering, not model development. Your work will power machine learning workflows for a Google Cloud environment. This role centers on infrastructure and pipeline reliability, offering the chance to shape data systems from day one.

What You'll Do6

  • 1Build scalable data pipelines using PySpark and Python to process large datasets
  • 2Design and manage feature stores for machine learning applications, ensuring data consistency and availability
  • 3Optimize SQL queries and BigQuery workloads for performance and cost efficiency
  • 4Ship reliable data ingestion and transformation jobs for ML infrastructure
  • 5Debug and resolve pipeline failures, maintaining data quality and timeliness
  • 6Collaborate with data scientists to understand feature requirements and deliver robust data solutions

Requirements6

  • 15+ years in data engineering with Python and SQL for large-scale data processing
  • 23+ years hands-on with PySpark for ETL and data transformation
  • 3Experience building and maintaining BigQuery datasets and queries
  • 4Track record of developing and managing data pipelines in Google Cloud Platform
  • 5Strong understanding of ML infrastructure concepts, including feature stores
  • 6Proven ability to write efficient, tested code for production environments

Salary Insight

Salary not disclosed in listing

Location

Typeonsite
LocationSan Jose, California

Required Skills

pysparkpythonsqlbigquerygoogle cloud platform
Share:

Similar open positions

Explore active roles that match your skills and interests.

Finoit Inc.

Finoit Inc.

1d agoSan Francisco, Californiapayroll

Senior Software Engineer Data Infrastructure

Design and lead development of scalable data pipelines for AI/ML platforms. Own design and implementation of distributed systems using Python and cloud services. Drive improvements in data quality and visualization. Lead cross-functional teams to deliver high-performance solutions.

180K–270K
PythonAWSGoogle Cloud Platform+3 more

Apex Systems

13h agoDenver, Coloradopayroll

Data Engineer - Spark & Scala

As a Data Engineer, you'll design and build ETL pipelines that feed a data lake supporting Spark and Scala workloads. Your pipelines will power anomaly detection models and AI agents, handling a growing portfolio of network data sources. Part of the Apex Systems team in Denver, Colorado, you'll work onsite in Greenwood Village with remote flexibility. This role focuses on building robust and scalable data infrastructure to ensure data quality and availability for downstream consumers.

146K–166K
sparkscalasql+2 more
AgreeYa Solutions

AgreeYa Solutions

1d agoDallas, Texascontract

ML Engineer, Databricks & Python

You will own the machine learning engineering lifecycle for a Dallas-based client, designing and deploying scalable data pipelines. You will work with Python, PySpark, and Databricks to process streaming and batch data, collaborating with data scientists and software engineers. You will build real-time features using Kafka and Snowflake, ensuring low-latency model serving. This 18-month contract offers the chance to tackle complex data modeling challenges, including slowly changing dimensions in MongoDB and PostgreSQL.

Competitive salary
PythonPySparkDatabricks+12 more
PRIMUS Global Services Inc.

PRIMUS Global Services Inc.

1d agoSan Jose, Californiacontract

Big Data Machine Learning Engineer, Spark & Python

You will design and build large-scale data processing pipelines and machine learning platforms using Apache Spark, Python, and SQL. You will work within a distributed systems team at PRIMUS Global Services Inc., collaborating on feature extraction and data analysis. This role requires deep experience with Hadoop, Snowflake, and related big data tools, and offers the chance to solve complex data challenges on an onsite team in Sunnyvale, CA.

40K–45K
pythonapache sparksql+2 more
Vertex Solutions Inc.

Vertex Solutions Inc.

14h agoNew York, New Yorkpayroll

Systems Manager, Google Cloud Platform Data Engineering

You lead the hands-on technical team that designs, builds, and maintains enterprise-grade data pipelines powering the Enterprise Data & Analytics Platform (EDAP). Your team owns the core stages of the data lifecycle on Google Cloud Platform, from transfer and ingestion to transformation, curation, and exposure. You manage a highly skilled group of engineers, setting technical direction and ensuring delivery at scale. This role combines deep GCP expertise with team leadership to drive reliable, high-performance data infrastructure.

140K–190K
google cloud platformbigquerypython+2 more
Delviom LLC

Delviom LLC

13h agoChicago, Illinoiscontract

Data Engineer, AWS & Python

Data Engineer will own the build of scalable data pipelines on AWS, transforming raw data into analytics-ready assets. You will join a platform team of 6 engineers, working with Python, PySpark, Snowflake, and dbt to unify data for product and business units. This contract role offers direct ownership of cloud infrastructure and data flow, with a focus on Terraform and CI/CD automation. You will bring 5+ years of experience and a track record of shipping production data systems.

50K–55K
pythonpysparksnowflake+2 more