We use cookies. Find out more about it here. By continuing to browse this site you are agreeing to our use of cookies.
#alert
Back to search results
New

CMPTL & DATA SCI RSH SPEC 4 RP

University of California - San Francisco
127,995 - 209,475
United States, California, San Francisco
1700 4th Street (Show on map)
Aug 03, 2026

Involves a hybrid of multiple computational / data / Cyberinfrastrcture (CI) dominated fields such as, but not limited to, bioinformatics, geological information services (GIS), data analytics, and computational chemistry. Applies computational, computer science, data science and cyber infrastructure (CI) research and development principles, with relevant domain science knowledge, to perform research and technology integration and development. Responsibilities include research, design, development, analysis, operation, and support of high performance computing (HPC) and data science research, software, tools and hardware resources. Develops data algorithms and performs computations, statistical analyses, interpretation and reporting of research. This specialty / function exists for those positions whose primary responsibility is to do research and use computational and data science technology as a tool to accomplish the research.

The Preclinical Design and Clinical Translation of Regimens for Tuberculosis (PReDiCTR-TB) Consortium is a 21st-century, model-informed drug development (MIDD) platform integrating computational science, translational pharmacology, and global clinical insight to accelerate the design and delivery of transformative TB regimens.

Positioned at the intersection of data, biology, engineering, econometrics, and decision science, PReDiCTR-TB functions as a strategic intelligence engine for regimen development, linking preclinical evidence, synthetic experiments, mechanistic models, and clinical data into a unified predictive framework. By embedding quantitative systems pharmacology (QSP), AI-driven analytics, and probabilistic decision modeling throughout development, the consortium enables real-time prioritization of regimens with the highest probability of clinical success.

Rather than advancing compounds in isolation, PReDiCTR-TB takes a regimen-first, translation-driven approach that optimizes combinations, dosing strategies, and treatment durations through iterative simulation-validation cycles grounded in human-relevant biology. This framework helps reduce reliance on costly empirical experimentation while increasing translational fidelity from bench to bedside.

The consortium delivers:

  • Predictive insights that de-risk development earlier
  • Optimized trial designs informed by mechanistic and statistical rigor
  • Faster and more confident go/no-go decisions across the R&D lifecycle

PReDiCTR-TB is redefining how infectious disease regimens are designed, prioritized, and translated into clinical impact.

Custom Scope

Model-informed drug development (MIDD) and AI in Clinical Trials market is experiencing a massive shift, projected to skyrocket from $1.35 billion to $2.74 billion by 2030. Data and model-driven efficiencies in patient selection, drug repurposing, and real-time trial monitoring are fundamentally compressing drug development timelines.

At the UCSF Savic Lab, we aren't just reacting to this shift-we are driving it. We are moving past traditional, isolated data storage models to establish a modern, highly interconnected, and secure data & model infrastructure.

We are seeking a solution-minded, highly technical, and mission-driven Distributed Data Pipeline Architect. In this role, you will design, build, lead implementation, and operate a research-grade, scale-distributed data architecture. This framework will serve as the foundation for advanced analytics, multi-institution translational science, and 21st-century accelerated drug development decision support.

Department Overview

The Savic Lab is a global leader in model-informed drug development for infectious diseases and serves as a quantitative innovation hub for translational pharmacology, AI-enabled modeling, and next-generation regimen design. The laboratory conducts groundbreaking research across TB, HIV, malaria, pediatric infectious diseases, translational PK/PD, and systems pharmacology. Through its leadership role in PReDiCTR-TB, the lab collaborates with international partners to integrate computational science, mechanistic modeling, and clinical translation into actionable strategies that improve global health outcomes.

UC San Francisco seeks candidates whose experience, teaching, research, or community service has prepared them to contribute to our commitment to diversity and excellence. The University of California is an Equal Opportunity/Affirmative Action Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion.


  • Applies advanced HPC / data / CI research and development concepts to plan, design, develop, modify, debug, deploy and evaluate highly complex HPC (software and / or hardware) or data science or computational science or CI software and technologies or combination thereof. Analyzes existing highly complex software, scientific codes, data science / analytics codes / algorithms, and HPC related hardware or works to formulate logic for new and highly complex systems and devises new algorithms. Performs highly complex analysis and tests / debugs highly complex software and hardware. Applies highly complex programming principles. Initiates large and complex research projects with multi-institutional scope in HPC / data / CI areas. May involve collaboration with domain science experts.

    Architect the Ecosystem: Design and implement distributed data platforms, data fabrics, and data mesh capabilities that safely unify complex datasets.

  • Specifies, develops, implements and executes highly complex software and hardware research and development plans. Performs or directs highly complex HPC, computational and data modeling, performance and integration testing. Works with research communities to develop, implement, and optimize computational and data analysis / analytics software / tools / algorithms / research codes with broad applicability.

    Build Resilient Pipelines: Oversee the development of scalable data pipelines optimized for massive batch processing of pre-clinical and clinical data.

  • Initiates and contributes to HPC / data science / CI research proposals with partners internal and external to the institution, in collaboration with other researchers and PIs. May lead a proposal of small to moderate size as a PI.

    Enable Cross-Domain Sharing: Create secure, cross-institution data frameworks that align with NIH Data Management and Sharing (DMS) policies and FAIR principles.

  • Drive Bio-Informatics Decisions: Set the architectural direction for formatting and structuring data that directly feeds pharmacometrics, biostatistics, and AI/ML pipelines.
  • Understands and applies advanced research and development practices, community standards and department policies and procedures. May serve as technical lead for multiple research and development projects of moderate to broad scope.

    Lead and Translate: Convert high-level clinical research data goals into concrete technical roadmaps. Collaborate across software engineering, machine learning, cyber security, clinical pharmacology teams, and consortium participants.

  • Brief Stakeholders: Author technical documentation, lead architectural reviews, and deliver executive briefings to internal leadership and external consortium partners.

Required Qualifications

  • Experience: 5+ years of hands-on experience in data engineering, distributed systems, or enterprise platform architecture.
  • Systems Design: A proven track record of architecting, deploying, and maintaining production-grade distributed data architectures.
  • Data Processing: Direct experience handling large-scale data ingestion, multi-tenant databases, and event-driven data flows.
  • Communication: Exceptional communication and presentation skills, with a demonstrated ability to explain complex technical concepts to non-technical executive stakeholders.
  • Bachelor's degree in Computer / Computational / Data Science, or Domain Sciences with computer / computational / data specialization or equivalent experience.

Preferred Qualifications

  • Architectures: Data mesh environments, data fabric layers, and semantic web modeling.
  • Streaming & Compute: Apache Kafka, Pulsar, AWS Kinesis, Apache Spark, Flink, or Apache Beam.
  • Storage & Databases: Distributed object storage (S3-compatible API environments) and NoSQL systems (Cassandra, DynamoDB, MongoDB).
  • Clinical Data Standards: Familiarity with NIH-preferred schemas and ontologies (e.g., BIDS for imaging, OMOP common data models, LOINC, SNOMED CT, or FHIR transfer protocols).
  • Governance & Security: Data cataloging, automated data lineage tools, and strict data access control frameworks (HIPAA / NIST compliance).
  • Master's degree in Computer / Computational / Data Science, or Domain Sciences with computer / computational / data specialization preferred.


Required Qualifications

  • Experience: 5+ years of hands-on experience in data engineering, distributed systems, or enterprise platform architecture.
  • Systems Design: A proven track record of architecting, deploying, and maintaining production-grade distributed data architectures.
  • Data Processing: Direct experience handling large-scale data ingestion, multi-tenant databases, and event-driven data flows.
  • Communication: Exceptional communication and presentation skills, with a demonstrated ability to explain complex technical concepts to non-technical executive stakeholders.
  • Bachelor's degree in Computer / Computational / Data Science, or Domain Sciences with computer / computational / data specialization or equivalent experience.

Preferred Qualifications

  • Architectures: Data mesh environments, data fabric layers, and semantic web modeling.
  • Streaming & Compute: Apache Kafka, Pulsar, AWS Kinesis, Apache Spark, Flink, or Apache Beam.
  • Storage & Databases: Distributed object storage (S3-compatible API environments) and NoSQL systems (Cassandra, DynamoDB, MongoDB).
  • Clinical Data Standards: Familiarity with NIH-preferred schemas and ontologies (e.g., BIDS for imaging, OMOP common data models, LOINC, SNOMED CT, or FHIR transfer protocols).
  • Governance & Security: Data cataloging, automated data lineage tools, and strict data access control frameworks (HIPAA / NIST compliance).
  • Master's degree in Computer / Computational / Data Science, or Domain Sciences with computer / computational / data specialization preferred.
Applied = 0

(web-77cf7d65c7-wz29x)