
Data Platform Engineer, Lakehouse & Cloud Infrastructure
Overview
You will architect and operate a cloud-based data lakehouse platform serving analytics teams and downstream applications across AWS, GCP, and Azure. You will build scalable data infrastructure on open table formats like Delta Lake and Iceberg, powered by distributed compute from Spark and Trino. You will join a platform engineering squad of 8, collaborating daily with data scientists and analysts to ensure reliable, high-performance data access. This role stands out for its greenfield design of a multi-cloud lakehouse, paired with an on-prem Kubernetes migration path.
What You'll Do10
- 1Build and ship the core lakehouse storage layer using Delta Lake, Iceberg, and Hudi across AWS, GCP, and Azure within your first 90 days.
- 2Design and deploy containerized data services on Kubernetes with Helm charts, handling multi-cluster orchestration and service mesh.
- 3Develop data ingestion pipelines with Apache Airflow and Kafka, ensuring exactly-once semantics and backfill capabilities.
- 4Optimize query performance across Spark and Trino by tuning file formats, partitioning, and catalog integration with Hive Metastore.
- 5Drive the migration of on-prem Hadoop workloads to the cloud, refactoring jobs and building hybrid data access patterns.
- 6Own observability for the platform, implementing metrics, logs, and tracing using Prometheus and Grafana.
- 7Ship infrastructure as code with Terraform, managing state across multiple cloud providers.
- 8Drive cross-functional alignment by defining data contract standards and SLAs for downstream consumers.
- 9Lead chaos engineering and recovery drills to ensure platform resilience during peak event streams.
- 10Mentor two junior engineers on cloud-native data patterns and code review practices.
Requirements10
- 15+ years building data platforms on AWS, GCP, or Azure, with at least 2 years on Kubernetes in production.
- 2Strong proficiency in Python or Scala for data pipeline development.
- 3Hands-on experience with Spark, Trino, or Presto for large-scale data processing (PB+ scale).
- 4Deep knowledge of Delta Lake, Iceberg, or Hudi and their transaction models.
- 5Experience with Airflow, Kafka, or similar streaming and orchestration tools.
- 6Proficiency in Terraform or Pulumi for infrastructure automation.
- 7Understanding of Hadoop ecosystem and migration strategies to cloud object stores.
- 8Familiarity with Hive Metastore or AWS Glue Catalog for schema management.
- 9BS in Computer Science, Engineering, or equivalent practical experience.
- 10Ability to work hybrid from Austin, TX or Sunnyvale, CA with 2 days onsite.
Salary Insight
Salary not disclosed in listing
Similar open positions
Explore active roles that match your skills and interests.
Apex Systems
VerifiedPlatform Engineer, AWS & Data Lake Infrastructure
Design, build, and maintain the AWS infrastructure underpinning a large-scale data lake and agent runtime environments. You will own the stability, scalability, and security of production environments enabling data science and agentic AI workloads. Collaborate with a Network Analytics team on a project transitioning from legacy systems. This role stands out for its direct impact on reliable CI/CD pipelines and modern cloud-native architecture.

ISite Technologies Inc
VerifiedData Engineer, AWS Lakehouse Migration
Own the end-to-end migration of on-prem Data Lake to AWS Lakehouse. You will join the datastore-migration Factory team in Dallas, TX, working hybrid onsite. Refactor legacy ETL pipelines, migrate scheduling, and execute large-scale data transfers. This role offers direct impact on cloud transformation with AWS technologies.

Rivago infotech inc
VerifiedData Architect (Databricks), Lakehouse Platform
You will lead enterprise-scale data platform modernization, migrating from AWS EMR, Apache NiFi, and legacy ETL frameworks to the Databricks Lakehouse Platform. You will architect and implement scalable Lakehouse solutions using Delta Lake, Unity Catalog, and Databricks Workflows. You will design end-to-end data pipelines for batch, streaming, and CDC, and govern data across the organization. This role centers on technical leadership and hands-on architecture, shaping the data strategy for a long-term project.

Anblicks
VerifiedLead Databricks Engineer, Lakehouse & ETL Pipelines
You will own the Lakehouse platform architecture for Anblicks in Dallas, designing and implementing end-to-end data solutions on Databricks. You will define best practices and build scalable ETL/ELT pipelines that power Spark workloads. Working closely with stakeholders, you will shape technical strategy and deliver production-grade data systems. This role centers on driving performance, cost, and scalability across a growing data estate.

Headway Tek Inc
VerifiedDatabricks Architect, Enterprise Lakehouse Platform
You will own the Databricks platform architecture end-to-end, designing and implementing next-generation data platforms on the Lakehouse paradigm. You will drive scalability, governance, performance, and cost optimization across enterprise data ecosystems, partnering with data engineering, data science, and business teams. This role goes beyond pipeline development, offering leadership in platform strategy and adoption at scale. You will define the roadmap for Unity Catalog, Delta Lake, and Spark workloads in a hybrid environment.

Oscar Technology
VerifiedData Engineer, Cloud & IoT Pipelines
You will design and build cloud data pipelines handling high-volume manufacturing and IoT data, using Python, SQL, and modern lakehouse architectures. You'll join a hands-on engineering squad focused on reliable data delivery and scalable solutions. This role stands out for its direct impact on production data systems from day one, with no legacy constraints.