Verified current Job

Principal Data Engineer (RWE)

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Principal Data Engineer (RWE) based in India.

Job Remote Full source details
Jobgether Source published Sep 22, 2026 Verified 11 minutes ago
✓ 100% verification score · Source: jobgether (lever) · Always confirm final requirements on the original source.
Complete source information imported The available role or programme description, requirements, benefits and source facts were imported from the public official endpoint and formatted for reading.
EmploymentFull-time
Work modeRemote / location-flexible

Overview

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Principal Data Engineer (RWE) based in India.

Full job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Principal Data Engineer (RWE) based in India. As a Principal Data Engineer, you will play a key role in building scalable data products that support Real World Evidence (RWE), observational research, epidemiology studies, dashboards, and other analytical outputs. You will transform complex, heterogeneous healthcare datasets into reusable, high-quality data models and pipelines. The role combines advanced data engineering with deep healthcare data expertise, including OMOP standards and clinical terminologies. You will collaborate closely with epidemiologists, statisticians, health economists, market access specialists, IT teams, and other technical stakeholders. You will also contribute to FAIR data principles and the development of AI-ready datasets for emerging generative AI use cases. This is an opportunity to work with large-scale patient-level data in a collaborative, globally connected data engineering environment.

Develop automated data processes for the ongoing generation of patient-level data products supporting dashboards, reports, studies, and other business needs. Transform raw datasets received from data vendors and partners into usable, structured data products for RWE studies, dashboards, and analytical outputs. Convert heterogeneous healthcare datasets into reusable data models that support observational research and epidemiology studies. Convert bespoke datasets, such as biomarkers and mutation data, into OMOP format where appropriate, while identifying residual data that cannot be standardized and determining how it can still support analysis. Build FAIR (Findable, Accessible, Interoperable, Reusable) data pipelines and semantic data engineering frameworks that improve healthcare data discoverability and interoperability. Develop AI-ready datasets and data products capable of supporting generative AI and other advanced analytics use cases. Engage with epidemiologists, statisticians, market access specialists, health economists, and other stakeholders to understand, scope, document, and translate business requirements into actionable technical data structures. Collaborate with the RWE programming team to develop data structures required for study outputs and provide technical support where data engineering expertise adds value. Liaise with IT teams to ensure inbound datasets from data partners are fit for their intended analytical purposes. Collaborate with technical teams from data and analytics software vendors, including Databricks, when required. Maintain clear documentation covering data flows, schemas, pipelines, and processes to support onboarding, troubleshooting, and auditing. Design and implement comprehensive testing, validation, and monitoring approaches to ensure the accuracy, reliability, and quality of data products. Troubleshoot issues related to data loading, extraction, transformation, and ETL processes. Collaborate with other members of the Data Engineering team, providing support and taking on additional workload when needed. Requirements Strong understanding of Real World Data (RWD) and Real World Evidence (RWE) concepts and their application to healthcare analytics. Ability to assess business requirements and recommend appropriate real-world healthcare datasets for analytical use cases. Deep understanding of healthcare data models, healthcare data ecosystems, and patient-level datasets. Strong expertise in OMOP Common Data Model v5.4 and v6, including extensions. Knowledge of healthcare standards and terminologies such as SNOMED CT, RxNorm, ICD-10, LOINC, and HCPCS/CPT. Strong experience building scalable ETL/ELT pipelines using Databricks, PySpark, Spark SQL, SQL, and Delta Lake. Experience working with large-scale healthcare and patient-level datasets and distributed data processing frameworks. Strong understanding of Semantic Data Engineering principles and experience developing FAIR-compliant data pipelines. Experience with cloud-based data platforms and large-scale distributed processing environments. Strong Power BI development and data modeling capabilities, with the ability to create reusable analytical datasets for dashboards and studies. Experience designing AI-ready datasets and analytics data products. Strong data profiling, validation, monitoring, and automated data quality framework experience. Understanding of healthcare data quality assessment methodologies. Excellent stakeholder management, communication, and collaboration skills, with the ability to translate complex business needs into effective technical solutions. Experience working with cross-functional and globally distributed teams. Exposure to one or more therapeutic areas such as Oncology, Respiratory, Immunology and Inflammation, or Infectious Diseases. Nice-to-have: working knowledge of R, sparklyR, R Shiny, observational research methodologies, OHDSI tools, and Azure Data Platform services. Benefits India-based opportunity with a flexible working environment. Opportunity to work with large-scale healthcare and patient-level datasets. Exposure to Real World Data, Real World Evidence, observational research, epidemiology, and healthcare analytics. Opportunity to contribute to FAIR data engineering and AI-ready data initiatives. Collaborative environment with cross-functional and globally distributed teams. Opportunities to work across data engineering, analytics, healthcare standards, and emerging AI use cases. Professional development through collaboration with experienced data engineering, programming, analytics, and healthcare specialists. Inclusive workplace culture that values diversity, integrity, honesty, and respect.

Tips for this job

Practical Job and Scholarship guidance. These tips do not replace official rules or create new eligibility requirements.

  1. Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
  2. Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
  3. Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
  4. Apply through the original employer or official recruitment destination shown on this page.

Verification notes

laptop-ats-crawler v3

Original authoritative source

Job and Scholarship is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.

jobgether (lever) ↗

Browse current Job and Scholarship listings from jobgether (lever) →

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books