University of California San FranciscoVerified Source

Distributed Data Pipeline Architect - UCSF Savic Lab

Onsite · San Francisco, California
Posted August 2, 2026
payroll

Overview

The Savic Lab seeks a solution-minded Distributed Data Pipeline Architect to design and operate a research-grade scale-distributed data architecture. This role drives advanced analytics and accelerates drug development through modern data infrastructure. The ideal candidate will lead the transition from traditional data storage to interconnected secure solutions.

What You'll Do10

  • 1Design Build Lead Ship Own Scale Debug Drive
  • 2Develop data algorithms Perform computations Statistical analyses Interpret results
  • 3Implement streaming compute frameworks Manage distributed object storage
  • 4Apply NoSQL Cassandra DynamoDB MongoDB
  • 5Ensure HIPAA NIST compliance Secure data pipelines
  • 6Create data mesh environments Semantic web modeling
  • 7Integrate streaming Kafka Pulsar AWS Kinesis
  • 8Utilize Apache Spark Flink Beam
  • 9Collaborate with clinical teams across institutions
  • 10Lead migration from legacy storage models

Requirements8

  • 15+ years hands-on experience data engineering distributed systems enterprise platforms
  • 2Proven track record architecting production-grade distributed architectures
  • 3Direct experience large-scale data ingestion multi-tenant databases event-driven flows
  • 4Exceptional communication skills presenting complex concepts to executives
  • 5Bachelor's degree in Computer Computational Data Science or Domain Sciences
  • 6Familiarity with NIH schemas BIDS OMOP LOINC SNOMED CT FHIR
  • 7Master's degree preferred in related fields
  • 8Knowledge of HIPAA NIST compliance governance security frameworks

Salary Insight

Salary not disclosed in listing

Location

Typeonsite
LocationSan Francisco, California

Required Skills

Data EngineeringDistributed SystemsApache KafkaApache SparkNoSQL DatabasesClinical Data StandardsHIPAA ComplianceCommunication SkillsBachelor's Degree in Computer Science or related field
Share:

Similar open positions

Explore active roles that match your skills and interests.

2755 Barclays Services Corpor

2755 Barclays Services Corpor

20d agoNew York, New Yorkpayroll

Senior Data Engineer AVP - Cloud Data Pipelines & Warehouses

Lead the design and execution of enterprise-grade Java data pipelines and warehouses to ensure accurate accessible secure data. Build and maintain robust systems for ingestion transformation and storage across large volumes. Drive optimization of pipeline performance and collaborate with cross-functional teams. Shape data governance and compliance while mentoring technical staff. Seeking a leader who can influence decision making and foster a culture of excellence.

190K–195K
JavaSparkKafka+10 more
Wise Skulls Corp.

Wise Skulls Corp.

17h agoSan Jose, Californiacontract

Software Engineer (C++/Java) Data Platform

Own scalable high throughput data processing platforms using C++ and Java. Lead ship of core infrastructure components. Drive performance improvements across large scale systems. Shape architecture decisions for future growth. This role differs by focusing on low latency streaming solutions.

Competitive salary
C++JavaDistributed Systems+2 more
Finoit Inc.

Finoit Inc.

11h agoSan Francisco, Californiapayroll

Senior Software Engineer Data Infrastructure

Design and lead development of scalable data pipelines for AI/ML platforms. Own design and implementation of distributed systems using Python and cloud services. Drive improvements in data quality and visualization. Lead cross-functional teams to deliver high-performance solutions.

180K–270K
PythonAWSGoogle Cloud Platform+3 more
TECHNEPTUNE CONSULTING INC

TECHNEPTUNE CONSULTING INC

15d agoSan Francisco, Californiapayroll

Data Solution Architect - Expertise in AWS and Data Platforms

Lead ownership of scalable data solutions at Technep Tuning Inc. in San Francisco. Design and deliver enterprise-grade data platforms leveraging advanced cloud technologies. This role focuses on architecting robust data ecosystems that drive business insights and operational efficiency.

Competitive salary
DatabricksAWSApache Spark+12 more
102 ICF Incorporated, LLC

102 ICF Incorporated, LLC

1d agoRemotepayroll

Senior Software Engineer CMS Data Pipelines

Design and lead development of Spark-based data pipelines for CMS scoring systems. Build ETL routines and data engineering solutions across Postgres Redshift and S3 Parquet. Collaborate with UI UX and quality teams to define data requirements. Ensure compliance with CMS standards while supporting clinicians. Must reside in the United States and work remotely from any U.S. location.

175K–184K
ScalaSparkSpark SQL+11 more
100 Eli Lilly and Company

100 Eli Lilly and Company

20d agoSan Francisco, Californiapayroll

Advisor - Data Architect Data Foundry

Lead Data Architect at Lilly's Data Foundry drives AI-native drug discovery by designing scalable data infrastructure. Transform raw scientific data into actionable insights for global health impact. Shape architecture for discovery scientists and autonomous AI agents.

152K–244K
DatabricksSnowflakeSpark+7 more