Senior Data & Python Software Engineer
Job description
About the role
The is responsible for designing, building, and maintaining high-performance web scraping systems and backend services that power web data extraction and brand protection initiatives. This role owns the end-to-end lifecycle of data-centric infrastructure, from raw ingestion to reliable serving, with a focus on robustness and efficiency. You will implement and maintain scraping-focused APIs and data services that support both internal products and external integrations, ensuring these interfaces are stable and performant. A core part of the position involves building reliable ingestion, processing, and storage workflows capable of handling large-scale web data volumes with integrity. You will handle the cleaning of web data and enforce strict data quality validation and governance across the entire ingestion, storage, and serving pipeline. The role requires optimizing scraping systems for performance, scalability, reliability, and cost efficiency to meet evolving business demands. You will monitor, debug, and improve scraping system reliability using observability tools, logging, metrics, and tracing to ensure operational excellence. Collaboration with product and engineering teams is essential to deliver features from initial design through full end-to-end production deployment and maintenance. You will take independent ownership of systems in production, driving iteration and performance management to ensure long-term success.
Key facts
What you'll do
- Design, build, and maintain high-performance web scraping systems and backend services for web data extraction and brand protection, ensuring scalability and resilience under varying loads.
- Implement and maintain scraping-focused APIs and data services for internal products and external integrations, focusing on clear contracts and efficient data delivery.
- Build reliable ingestion, processing, and storage workflows for large-scale web data, incorporating best practices for error handling and retry logic.
- Handle cleaning of web data and ensure data quality validation and governance across ingestion, storage, and serving to maintain consistency and accuracy.
- Optimize scraping systems for performance, scalability, reliability, and cost efficiency, identifying bottlenecks and implementing targeted improvements.
- Monitor, debug, and improve scraping system reliability using observability tools, logging, metrics, and tracing to enable proactive issue resolution.
- Collaborate with product and engineering teams to deliver features from design through full end-to-end production deployment, aligning technical solutions with business goals.
- Take independent ownership of systems in production, including maintenance, iteration, and performance management, to reduce manual intervention.
- Design and implement data models and database schemas that support efficient querying and long-term maintainability of ingested data.
- Evaluate, adopt, and maintain tools and frameworks that streamline data workflows, ensuring they integrate smoothly with existing infrastructure.
- Work with complex and dynamic web environments, developing strategies to handle anti-scraping measures and ensure reliable data capture.
- Define and enforce coding standards, documentation practices, and operational runbooks to support scalable and maintainable systems.
- Partner with cross-functional stakeholders to translate requirements into technical specifications and deliver robust data platforms.
- Continuously assess new technologies and methodologies to improve the efficiency and effectiveness of data ingestion and processing pipelines.
Requirements
- You need Experience with web scraping, including handling dynamic content and managing request rates.
- You need Strong SQL skills for writing complex queries and optimizing database performance.
- You need Strong Python experience, demonstrating proficiency in writing clean, maintainable, and efficient code.
- You need Experience with PostgreSQL or similar relational databases for designing schemas and ensuring data integrity.
- You need Experience designing and building scalable APIs and backend services using frameworks such as FastAPI or Django.
- You need Experience designing efficient, scalable data models and database schemas to support growing data requirements.
- You need Hands-on experience deploying and operating systems in the cloud, including AWS, GCP, or Azure.
- You need Experience working with Docker and containerized environments to ensure consistent deployments.
- University education in a technical field such as Computer Science, Engineering, or similar is preferred.
- You need + years of professional experience or 2+ years in an early-stage startup, showing ability to handle responsibility.
Nice to have
- It helps if you have Experience with workflow orchestration tools such as Airflow for managing complex data pipelines.
- It helps if you have Experience with browser-based automation tools like Playwright or Selenium for advanced scraping scenarios.
- It helps if you have Experience with DBT or analytics-focused data transformation workflows to ensure reliable data modeling.
- It helps if you have Experience building or operating high-concurrency systems and task queues to manage heavy workloads.
- It helps if you have Experience designing and deploying cloud-native workflows on AWS to leverage managed services.
- It helps if you have Familiarity with CI/CD pipelines and production deployment practices to enable rapid iteration.
- It helps if you have Experience working in a high-growth early-stage startup environment, adapting quickly to changing priorities.
Practical notes
Note: Design, build, and maintain high-performance web scraping systems and backend services for web data extraction and brand protection.
Implement and maintain scraping-focused APIs and data services for internal products and external integrations.
Build reliable ingestion, processing, and storage workflows for large-scale web data.
Handle cleaning of web data and ensure data quality validation and governance across ingestion, storage, and serving.
Optimize scraping systems for performance, scalability, reliability, and cost efficiency.
Monitor, debug, and improve scraping system reliability using observability tools, logging, metrics, and tracing.
Collaborate with product and engineering teams to deliver features from design through full end-to-end production deployment.
Take independent ownership of systems in production including maintenance, iteration, and performance management.
For , Experience with web scraping is required.