IlluminaVerified Source

Staff Data Engineer at Illumina

142K–212K
Onsite · San Diego, California
Posted August 10, 2026
payroll

Overview

Lead a senior data engineering role designing and scaling data products on cloud lakehouses while impacting global health equity. You will own end-to-end analytics and AI/ML solutions driving life-changing discoveries.

What You'll Do11

  • 1Partner across business AI and platform teams translating domain needs into scalable data products
  • 2Design and build end-to-end data pipelines on Databricks and Snowflake following medallion architecture
  • 3Develop reusable Python frameworks for ingestion transformation and publishing with robust data modeling
  • 4Create reusable frameworks in Python for data ingestion transformation and validation using functional and OOP paradigms
  • 5Design lakehouse data models applying system design judgment to build performant distributed pipelines with Spark Delta Lake and dbt
  • 6Embed data quality governance using Unity Catalog for lineage RBAC masking and PII handling
  • 7Implement monitoring alerting and root cause analysis for critical datasets ensuring SLA adherence
  • 8Adopt AI in daily workflows to accelerate development testing and optimization
  • 9Act as technical leader setting standards leading code reviews and mentoring engineers globally
  • 10Collaborate with SAP Manufacturing and Quality teams to deliver GxP compliant solutions
  • 11Mentor engineers on advanced cloud platforms and data engineering best practices

Requirements11

  • 110+ years of professional data engineering experience building and scaling products on Databricks or Snowflake
  • 2Strong proficiency in Python including functional and object oriented programming
  • 3Advanced SQL and expertise in relational dimensional and lakehouse data modeling
  • 4Solid understanding of distributed systems and system design for large scale processing
  • 5Hands-on experience with open table formats Delta Lake and Apache Iceberg
  • 6Experience with Spark and modern ELT tooling such as dbt
  • 7Familiarity with data observability governance security and compliance practices
  • 8Proven adoption of AI in data and analytics engineering workflows
  • 9Solid software engineering foundation including Git REST APIs JSON and CI/CD on cloud platforms
  • 10Bachelor's degree in Computer Science Data Science or related field
  • 11Bonus qualifications: Domain knowledge of SAP Manufacturing Quality data processes and GxP 21 CFR Part 11 experience

Salary Insight

$142 - $212k per year

Location

Typeonsite
LocationSan Diego, California

Required Skills

PythonSQLData ModelingSparkDelta LakeApache IcebergdbtSnowflakeDatabricksAWSGitREST APIs
Share:

Similar open positions

Explore active roles that match your skills and interests.

100 Eli Lilly and Company

100 Eli Lilly and Company

20d agoIndianapolis, Indianapayroll

Senior Data Engineer Lakehouse Architecture at Lilly

We seek a Sr. Principal Data Engineer to architect and lead large-scale Lakehouse solutions driving modern data platform evolution. This role involves implementing unified data architectures and mentoring a team of 3-5 engineers. Success is measured by delivering AI-ready data pipelines and shaping technical direction.

132K–244K
DatabricksSnowflakeLakehouse architecture+5 more
Apex Systems

Apex Systems

1d agoAtlanta, Georgiapayroll

Data Engineering Lead, DevOps & Data Platforms

You own the core foundation of the Data Intelligence Platform Enablement Framework, building governance-as-code, infrastructure-as-code, CI/CD automation, and observability. You lead a team of engineers and collaborate with data scientists and analysts to ship a reliable, secure platform. Your work directly enables faster, compliant data product delivery across the organization. This role offers technical leadership and the chance to shape platform standards from day one.

146K–156K
governance-as-codeinfrastructure-as-codeCI/CD automation+3 more

Matterworks

1d agoBoston, Massachusettspayroll

Senior Software Engineer Data Platform

Lead design and scaling of data contracts and infrastructure. Own pipelines serving millions of petabytes. Build scalable label enrichment systems and interfaces for AI and chemistry tools. Ensure operational excellence and data quality at scale.

Competitive salary
PythonSQLKubernetes+11 more

Ambiencehealthcare

11h agoSan Francisco, Californiapayroll

Senior Data Engineer, Healthcare AI & Snowflake

You'll own the trust layer for Ambience's AI platform, building and running the pipelines that ingest, validate, and transform clinical and operational data at scale. The metrics you deliver power our AI story and drive decisions across clinical, engineering, and business teams. You'll partner with product managers, engineers, and clinicians to turn raw data into governed, self-service analytics. This role stands out because you'll have end-to-end ownership of data quality and credibility, directly impacting healthcare outcomes. We're a fast-growing startup backed by top investors, with a hybrid office in San Francisco.

Competitive salary
snowflakepythonsql+2 more
100 Eli Lilly and Company

100 Eli Lilly and Company

21d agoSan Francisco, Californiapayroll

Advisor - Data Architect Data Foundry

Lead Data Architect at Lilly's Data Foundry drives AI-native drug discovery by designing scalable data infrastructure. Transform raw scientific data into actionable insights for global health impact. Shape architecture for discovery scientists and autonomous AI agents.

152K–244K
DatabricksSnowflakeSpark+7 more
3202 St. Jude Medical Business Services, Inc.

3202 St. Jude Medical Business Services, Inc.

7d agoSan Francisco, Californiapayroll

Lead Data & Solutions Architect at Abbott

We seek a Lead Data & Solutions Architect to design and modernize enterprise data ecosystems enabling AI and analytics. This role drives cloud-native data platforms using Databricks AWS Azure and supports data governance and security. You will collaborate with cross-functional teams to deliver scalable solutions for healthcare data initiatives.

114K–228K
DatabricksDelta LakeAWS+5 more