Software Engineer III
Job description
About the role
You will own the design and implementation of backend data pipelines that ingest, validate, normalize, enrich, and persist high-volume raw datasets from diverse acquisition sources. You will own the development of ETL and ELT workflows and processing jobs using technologies such as Apache Airflow, Apache Beam, Google Dataflow, DataProc, Spark, Kafka, or Pub/Sub. You will own the implementation of new features in Java-based services and data processing applications that support data acquisition at scale. You will own work across batch and streaming architectures for scheduled, near-real-time, and event-driven data flows. You will own improvements to data quality through schema validation, deduplication, enrichment, monitoring, retries, and controlled backfills. You will own contributions to observability for pipeline health, throughput, latency, cost, and error rates. You will own collaboration with product managers, data teams, and platform teams to translate business requirements into reliable technical solutions. You will own the troubleshooting of pipeline failures, data anomalies, and performance bottlenecks in production environments.
Key facts
What you'll do
- Execute the work posted for within the Data Acquisition team.
- Follow the scope and guidelines outlined About the company, recognizing that ZoomInfo Technologies Inc. is a registered data broker in the United States.
- Design and build backend data pipelines that ingest, validate, normalize, enrich, and store high-volume raw data from multiple sources.
- Develop ETL and ELT workflows and processing jobs using technologies such as Apache Airflow, Apache Beam, Google Dataflow, DataProc, Spark, Kafka, or Pub/Sub.
- Implement new features in Java-based services and data processing applications that support data acquisition at scale.
- Work with batch and streaming architectures for scheduled, near-real-time, and event-driven data flows.
- Improve data quality through schema validation, deduplication, enrichment, monitoring, retries, and controlled backfills.
- Contribute to observability for pipeline health, throughput, latency, cost, and error rates.
- Collaborate with product managers, data teams, and platform teams to translate business requirements into reliable technical solutions.
- Help design, plan, and execute initiatives for next-generation data acquisition technologies, defining priorities and execution plans.
- Champion troubleshooting practices for pipeline failures, data anomalies, and performance bottlenecks in production.
- Operate production data pipelines and services with a focus on reliability, scalability, and performance.
- Partner with cross-functional stakeholders to identify opportunities for optimization and new capabilities.
- Ensure solutions adhere to data governance, security, and compliance standards relevant to data brokering.
- Leverage cloud infrastructure to build cost-effective and resilient data processing systems.
- Write clean, testable, and maintainable code that meets production quality standards.
- Participate in code reviews and contribute to the continuous improvement of engineering practices.
Requirements
You bring 3+ years of professional software engineering experience with strong Java proficiency. You have built and operated production data pipelines or ETL/ELT workflows in distributed environments. You understand streaming technologies like Kafka or Pub/Sub and batch processing, schema design, and data quality practices. You are experienced with cloud infrastructure, especially GCP services such as Dataflow, BigQuery, Pub/Sub, and Kubernetes on GKE. You can troubleshoot pipeline issues, optimize throughput, and balance quality with business impact. A bachelor's degree in computer science or a related field is required. You are legally authorized to work in the United States without sponsorship now or in the future.
Nice to have
Experience with Kubernetes on GKE or EKS, plus work in Snowflake, BigQuery, or similar query engines.
Skills & Tools
Java, Apache Airflow, Apache Beam, Google Dataflow, DataProc, Spark, Kafka, Pub/Sub, BigQuery, GCS, GKE.
Practical notes
Please What you'll do
ZoomInfo Technologies Inc. is a registered data broker in the United States. The company collects and sells access to its database of information about business people and companies to sales, marketing, and recruiting professionals.