Overview
Create high-quality datasets for training and evaluation; run experiments on new datasets (data ablations) to assess their impact and determine the most effective data Develop and maintain scalable data pipelines for data ingestion, pre-processing, filtering, and annotation Analyse real-world multimodal datasets to assess quality, diversity, relevance, and identify areas for improvement Build tools and workflows for dataset auditing, visualization, and versioning Collaborate with Safety, Ethics, and Governance teams to ensure datasets meet standards for quality, privacy, and responsible AI practices Bachelor's Degree in AI, Computer Science, Data Science, Statistics, Physics, Engineering, or a related technical field. Proficiency in Python. Experience with distributed data-processing frameworks such as Spark, Ray and workflow-orchestration tools such as Airflow. Experience with processin
Full job description
Full Job Description
Create high-quality datasets for training and evaluation; run experiments on new datasets (data ablations) to assess their impact and determine the most effective data Develop and maintain scalable data pipelines for data ingestion, pre-processing, filtering, and annotation Analyse real-world multimodal datasets to assess quality, diversity, relevance, and identify areas for improvement Build tools and workflows for dataset auditing, visualization, and versioning Collaborate with Safety, Ethics, and Governance teams to ensure datasets meet standards for quality, privacy, and responsible AI practices Bachelor's Degree in AI, Computer Science, Data Science, Statistics, Physics, Engineering, or a related technical field. Proficiency in Python. Experience with distributed data-processing frameworks such as Spark, Ray and workflow-orchestration tools such as Airflow. Experience with processing datasets at many petabyte scale and with trillion rows. Experience building datasets for training foundation models, including language or multimodal models. Strong experience in data analysis, data engineering, or both. Ability to communicate technical findings clearly and effectively to research, engineering, and product teams. Experience evaluating dataset quality and measuring the impact of data through controlled model-training experiments. Master's degree in Computer Science or a related technical field, or equivalent experience. Experience working with large-scale, real-world datasets that are unstructured or semi-structured Experience evaluating dataset quality and measuring the impact of data through controlled model-training experiments. Experience with multimodal data, such as image, video, or audio data.
Tips for this job
Practical JobOpportunity guidance. These tips do not replace official rules or create new eligibility requirements.
- Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
- Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
- Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
- Apply through the original employer or official recruitment destination shown on this page.
Verification notes
Verified from public schema.org JobPosting structured data on the official source page. The complete published description, responsibilities, requirements and benefits were normalized when present; unstated facts were not inferred.
JobOpportunity is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.
Apply through JobOpportunity →Browse current JobOpportunity listings from Microsoft Careers →