Verified current Job

Senior Data Engineer (Web Scraping)

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Data Engineer (Web Scraping) based in India

Job Remote Full source details
Jobgether Source published Sep 30, 2026 Verified 1 hour ago
✓ 100% verification score · Source: jobgether (lever) · Always confirm final requirements on the original source.
Complete source information imported The available role or programme description, requirements, benefits and source facts were imported from the public official endpoint and formatted for reading.
EmploymentFull-time
Work modeRemote / location-flexible

Overview

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Data Engineer (Web Scraping) based in India

Full job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Data Engineer (Web Scraping) based in India. This is a senior, hands-on opportunity to own the next generation of reliable web-data acquisition within a modern data platform. You will design, build, and operate production-grade Python scrapers and scalable ingestion pipelines that support trusted research and data products. The role focuses on improving the maturity, reliability, and scalability of web-scraping capabilities across a growing data environment. You will establish reusable patterns for extraction, scheduling, storage, monitoring, validation, and failure handling. From investigating new data sources to deploying and supporting production workloads, you will have substantial ownership over the full engineering lifecycle. The role combines deep Python engineering with cloud infrastructure, data pipelines, observability, and practical problem-solving in a remote-first environment. You will work independently while collaborating with a distributed engineering team through code reviews, documentation, and structured development workflows.

Own the development, deployment, and ongoing operation of web-scraping and web-data ingestion pipelines. Design and establish scalable web-scraping frameworks with reusable patterns for extraction, scheduling, storage, monitoring, validation, and failure handling. Build, maintain, and improve reliable production scrapers for both new and existing data sources. Investigate websites and determine the most appropriate acquisition method, including APIs, direct HTTP requests, HTML parsing, browser automation, or third-party tooling. Evaluate build-versus-buy options for scraping infrastructure and external services, considering capabilities, reliability, cost, operational complexity, and risk. Ensure web-data acquisition activities appropriately account for internal policies, website terms, robots.txt, access restrictions, privacy, and intellectual-property considerations, escalating unclear situations when required. Diagnose and resolve scraping challenges related to website changes, dynamic content, authentication, sessions, rate limits, concurrency, and other operational constraints. Integrate scraping workloads into scalable data-platform and lakehouse architectures. Improve scheduling, monitoring, storage, validation, and operational support for scraping workloads. Use AI-assisted engineering tools where appropriate while maintaining a thorough understanding of, and accountability for, the code being delivered. Support production workloads through monitoring, debugging, maintenance, and continuous improvement. Contribute to a remote engineering environment through code reviews, documentation, ticket-based workflows, and knowledge sharing. Requirements Demonstrated professional experience building and operating production web-scraping systems at scale . Proven ability to independently take substantial scraping projects from initial investigation through implementation, deployment, and ongoing production support. Strong production-level Python engineering skills, with experience developing maintainable applications rather than standalone scripts. Hands-on experience with scraping technologies such as Requests/httpx, BeautifulSoup, Scrapy, Playwright, or Selenium . Strong practical understanding of HTTP, HTML, APIs, JavaScript-rendered websites, and browser/network behavior. Experience addressing common scraping challenges including pagination, authentication, sessions, retries, rate limiting, concurrency, and proxies. Strong understanding of data pipelines, data quality, and how collected data should be validated, stored, and consumed by downstream systems. Experience deploying, monitoring, and supporting production workloads in a cloud environment. Strong debugging, analytical, and problem-solving abilities, with the judgment to make effective engineering decisions independently. Comfortable working within a remote engineering team and participating in code reviews, documentation, and ticket-based development workflows. Experience with AWS is desirable. Familiarity with lakehouse or data-lake architectures, particularly Apache Iceberg , is a plus. Experience with PySpark or other distributed data-processing technologies is beneficial. Familiarity with Docker and containerized workloads is advantageous. Experience with Terraform or other infrastructure-as-code tools is a plus. Familiarity with Grafana or comparable observability platforms is desirable. Experience operating high-volume or distributed crawling systems is beneficial. Experience evaluating or operating commercial scraping, proxy, or browser-infrastructure services is a plus. Experience implementing automated scraper testing, canary runs, or source-drift detection is desirable. Exposure to legal, compliance, privacy, or data-governance processes related to web-data acquisition is advantageous. Strong ownership, autonomy, documentation, communication, and collaboration skills. Benefits Fully remote position within a remote-first technology team. Opportunity to take ownership of a critical web-data acquisition capability and influence its architecture and operating standards. Senior, hands-on role with substantial autonomy across investigation, engineering, deployment, and production support. Work on scalable data pipelines and modern lakehouse architectures supporting research and data products. Exposure to cloud infrastructure, distributed processing, observability, browser automation, APIs, and production scraping technologies. Opportunity to establish reusable engineering patterns and improve the reliability and scalability of data ingestion. Collaboration with a distributed engineering team through code reviews, documentation, and structured workflows. Environment that supports independent problem-solving, technical ownership, and continuous improvement. Fully remote setup available across the relevant distributed team environment.

Tips for this job

Practical JobOpportunity guidance. These tips do not replace official rules or create new eligibility requirements.

  1. Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
  2. Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
  3. Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
  4. Apply through the original employer or official recruitment destination shown on this page.

Verification notes

laptop-ats-crawler v3

Original authoritative source

JobOpportunity is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.

Apply through JobOpportunity →

Browse current JobOpportunity listings from jobgether (lever) →

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books