Software Engineer II, Data Engineer
Job description
About the role
This position offers the chance to architect and sustain the data infrastructure that drives critical advertising platforms at scale. You will transform fragmented raw inputs into coherent, reliable data products that fuel business intelligence and revenue operations. The role demands ownership of complex pipeline logic and collaboration with cross-functional teams to ensure data integrity and usability. You will be instrumental in establishing robust data governance frameworks that align with regulatory requirements. Success in this position requires balancing system performance with the delivery of actionable insights for stakeholders. The work involves solving intricate distributed systems problems while adhering to strict compliance standards. Your contributions will directly influence the reliability of data flows that power major revenue-generating products.
Key facts
What you'll do
- Architect and develop scalable backend services and data pipelines designed to process high-volume AdTech and ACR data streams efficiently.
- Construct and maintain internal platform tooling that enforces stringent Data Governance, Data Quality, and Data Compliance standards across the organization.
- Independently define scope, technical enhancements, and deliverables for features of limited complexity while providing accurate time estimates and resource planning.
- Troubleshoot, debug, and resolve complex bugs within distributed environments and large-scale workflows using advanced analytical and diagnostic techniques.
- Design and implement data processing workflows that optimize resource utilization and ensure high throughput and low latency system performance.
- Partner closely with product managers, data scientists, and engineering teams located in the US and India to align technical solutions with evolving business objectives.
- Evaluate and integrate modern data platforms and orchestration tools to streamline operations and improve the manageability of data workflows.
- Clearly communicate intricate technical concepts, including system designs, architecture trade-offs, and engineering decisions, to both technical and non-technical audiences.
- Develop and maintain comprehensive documentation for data pipelines, platform tools, and operational procedures to ensure knowledge transfer and system maintainability.
- Monitor production data pipelines proactively, identifying bottlenecks and potential failure points to maintain system reliability and data accuracy.
- Contribute to the planning and development of features that impact multiple components of the data ecosystem, ensuring cohesive system evolution.
- Implement data quality checks and validation frameworks to ensure the trustworthiness and consistency of data throughout its lifecycle.
- Collaborate with security teams to ensure that data handling practices comply with internal policies and external regulatory standards.
- Continuously assess emerging technologies and methodologies to improve the efficiency and scalability of the data platform.
Requirements
- Hold a Bachelor's Degree in Computer Science, Engineering, or a related technical field from an accredited institution.
- Possess over two years of professional software development experience, with a focus on Backend or Data Engineering roles in production environments.
- Demonstrate mastery of fundamental Data Structures and Algorithms and apply them to solve complex engineering problems.
- Exhibit a solid grasp of core Computer Science concepts, including Operating Systems, Computer Networks, and DBMS principles.
- Show a proactive, self-starter mindset with strong problem-solving and analytical skills to navigate ambiguous challenges.
- Possess the ability to communicate technical challenges effectively in both written and verbal formats to diverse stakeholders.
- Prove a proven ability to collaborate effectively across time zones, maintaining productivity and alignment with global teams.
- Have hands-on experience with Big Data processing frameworks, specifically Apache Spark (PySpark or Scala), in real-world scenarios.
- Demonstrate experience with modern data platforms like Databricks and orchestration tools like Apache Airflow for managing workflows.
- Show a background building, maintaining, and scaling complex ETL processes to meet evolving business demands.
- Have familiarity with cloud computing platforms such as AWS, GCP, or Azure and their associated data services.
- Understand distributed systems concepts deeply and possess Low-Level Design proficiency to architect scalable solutions.
- Comply with the requirement for daily overlap with US hours, specifically between 8 AM and 11 AM MST, to facilitate real-time collaboration.
- Maintain eligibility to work in India without the need for visa sponsorship, as this is a requirement for the position.
Nice to have
The preferred qualifications focus on advanced data engineering skills that enhance platform stability and scalability. Hands-on experience with Big Data processing frameworks, specifically Apache Spark (PySpark or Scala), is highly valued. Experience with modern data platforms like Databricks and orchestration tools like Apache Airflow is preferred. A background building, maintaining, and scaling complex ETL processes is advantageous. Familiarity with cloud computing platforms such as AWS, GCP, or Azure is considered a plus. Understanding of distributed systems concepts and Low-Level Design proficiency is strongly preferred.
Practical notes
This role requires daily overlap with US hours, specifically between 8 AM and 11 AM MST. Candidates must have employment eligibility in India. Current visa sponsorship is not available for this position.