Overview
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Data Engineer (AI/ML) based in India.
Full job description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Data Engineer (AI/ML) based in India. As a Senior Data Engineer (AI/ML), you will build the data and AI infrastructure powering next-generation intelligent products and experiences. You will combine modern data engineering with Generative AI, working across LLMs, RAG, embeddings, vector search, and AI agents. The role involves designing scalable data platforms and production-grade pipelines for training, inference, evaluation, and retrieval workloads. You will collaborate closely with data scientists, ML engineers, software engineers, and product teams to turn innovative prototypes into reliable systems. Your work will help establish strong foundations for data quality, governance, observability, performance, and cost efficiency. This is an opportunity to shape AI-enabled data products while working with large-scale distributed and streaming technologies in a global environment.
Design and build AI/LLM data pipelines supporting training, inference, evaluation, embeddings, and retrieval workloads. Build production-grade RAG systems covering ingestion, chunking, embedding generation, indexing, retrieval, reranking, and context construction. Develop AI applications using LLMs, structured outputs, function and tool calling, and agentic workflows. Build and optimize semantic search and vector retrieval systems. Develop frameworks for LLM evaluation, monitoring, tracing, quality measurement, latency analysis, and cost optimization. Design scalable batch and streaming pipelines using Databricks, Apache Spark, Delta Lake, Snowflake, and Airflow. Build data products and platforms that make structured and unstructured enterprise data accessible to AI applications. Develop reliable ETL/ELT pipelines and optimize large-scale distributed workloads for performance and cost. Establish data quality, governance, lineage, security, and observability practices. Partner with ML and application engineering teams to transition AI prototypes into reliable, production-ready systems. Support large-scale data platforms, real-time processing, event-driven architectures, and complex orchestration workflows. Contribute to AI evaluation datasets and pipelines that measure quality, accuracy, relevance, latency, and cost. Monitor production AI systems, including token usage, model performance, failures, latency, and overall system health. Requirements: 5+ years of experience in data engineering, software engineering, distributed systems, or a related field. Strong programming skills in Python and/or Scala/Java, combined with advanced SQL capabilities. Hands-on experience with Databricks, Snowflake, Apache Spark, Delta Lake, and Airflow. Strong experience working with cloud-based data platforms and scalable data architectures. Practical experience building applications using LLMs or Generative AI. Strong understanding of RAG architectures, embeddings, vector databases, semantic search, and retrieval systems. Familiarity with prompting, structured outputs, tool calling, model evaluation, and other modern LLM concepts. Experience designing scalable, reliable, observable production data systems. Strong knowledge of large-scale data platforms, distributed processing, and complex data workflows. Experience with real-time and streaming architectures using technologies such as Kafka or Spark Structured Streaming. Experience designing low-latency pipelines and event-driven architectures. Strong experience with multi-stage ETL/ELT and data orchestration workflows using Airflow or similar platforms. Experience optimizing Spark or Databricks workloads through partitioning, clustering, caching, joins, and compute optimization. Experience supporting both batch and real-time AI/ML workloads. Experience with LLM/AI evaluation frameworks, automated evaluations, experimentation, quality metrics, and evaluation datasets. Familiarity with AI observability and tracing, including token usage, model performance, latency, failures, and production monitoring. Experience with LangGraph, LangChain, LlamaIndex, or similar AI orchestration frameworks is preferred. Experience with vector databases such as Qdrant, Pinecone, Weaviate, or Databricks Vector Search is preferred. Familiarity with Kafka, MLflow, Unity Catalog, Databricks Mosaic AI, or model-serving platforms is a plus. Strong understanding of distributed systems, cloud architecture, APIs, CI/CD, data governance, and production operations. Benefits: 100% remote position across India. Work from almost anywhere for up to 20 days per year. Generous paid vacation and time off for your birthday. Paid parental leave. Company-paid therapy sessions through SpringHealth. Company-paid Headspace subscription. Annual company-wide week off to support rest and well-being. Generous health insurance and pension fund. Tax optimization options. Development Dollars and leadership development opportunities. Access to thousands of on-demand learning resources. Paid volunteer time. Travel discounts. Employee Resource Groups. Quarterly team offsites. Global and collaborative working environment. Opportunities to work with large-scale data engineering, Generative AI, distributed systems, and modern AI infrastructure. Flexible collaboration across international teams and time zones, with local laws and regulations taken into consideration.
Tips for this job
Practical Job and Scholarship guidance. These tips do not replace official rules or create new eligibility requirements.
- Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
- Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
- Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
- Apply through the original employer or official recruitment destination shown on this page.
Verification notes
laptop-ats-crawler v3
Job and Scholarship is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.
jobgether (lever) ↗Browse current Job and Scholarship listings from jobgether (lever) →