Senior Software Engineer
Job description
About the role
ZoomInfo is where careers accelerate. You will own the design and operation of backend systems and data pipelines that acquire, transform, validate, and store large raw data sets from a wide range of sources. You will work across distributed processing, workflow orchestration, streaming and batch data flows, and cloud infrastructure in collaboration with product, data science, and data quality teams. The role is a senior individual-contributor position focused on technical depth, production execution, and high-quality data systems. You will translate complex data acquisition requirements into scalable, reliable systems that empower ZoomInfo customers. You will move fast, think boldly, and be surrounded by teammates who care deeply and celebrate wins.
Key facts
What you'll do
Design, build, and operate large-scale data acquisition pipelines that ingest, validate, transform, enrich, and store high-volume raw data.
Architect resilient ETL/ELT workflows for batch, streaming, scheduled, and event-driven data processing.
Develop production Java services and data processing applications for ingestion, orchestration, enrichment, deduplication, and delivery.
Build and improve systems using technologies such as Apache Airflow, Apache Beam, Spark, Google Dataflow, DataProc, Kafka, and Pub/Sub.
Define practical approaches for schema evolution, data contracts, data validation, backfills, replayability, and idempotent processing.
Improve reliability, performance, scalability, and cost efficiency across data acquisition pipelines and services.
Implement observability, monitoring, and alerting for pipeline health, throughput, latency, failure rates, and data quality metrics.
Work with product, data science, platform, and data quality teams to translate business needs into production-ready systems.
Contribute to technical designs, implementation plans, and system modernization efforts across Data Acquisition.
Collaborate with cross-functional stakeholders to prioritize work, refine requirements, and ensure timely delivery of robust data solutions.
Support on-call responsibilities for production data pipelines and respond to operational incidents with clear communication.
Champion best practices for data quality, testing, and documentation to enable long-term maintainability and scalability.
Partner with data platform teams to align data acquisition efforts with broader platform strategy and standards.
Drive continuous improvement by evaluating new tools, techniques, and architectural patterns to enhance data workflows.
Requirements
Must have a Bachelor's degree in Computer Science, Engineering, or a related field.
Must have 5+ years of professional software engineering experience with a strong focus on backend systems, data engineering, or distributed processing.
Must have proven experience building and operating production data pipelines at scale.
Must have deep proficiency with Java and object-oriented design.
Must have hands-on expertise with data processing and orchestration technologies such as Apache Beam, Apache Airflow, Spark, Google Dataflow, or DataProc.
Must have strong experience with streaming systems such as Apache Kafka, Google Pub/Sub, or similar technologies.
Must have a strong understanding of batch processing, streaming processing, data modeling, schema evolution, and data quality management.
Must have experience designing ETL/ELT workflows that process large volumes of structured and semi-structured data.
Must have experience with at least one cloud provider, preferably GCP.
Must have hands-on experience with cloud services such as BigQuery, GCS, GKE, Dataflow, DataProc, and Pub/Sub.
Must have experience operating production services with monitoring, logging, alerting, SLIs, and incident investigation.
Must be able to analyze and improve pipeline performance, reliability, resource utilization, and cost.
Must be able to write clean, maintainable production code and evaluate tradeoffs in system design.
Must have strong API, integration patterns, retries, backpressure, idempotency, and operational failure modes understanding.
Must be able to work effectively with product, data science, platform, and data quality teams.
Must be able to contribute to technical designs, implementation plans, and system modernization efforts.
Must be able to communicate clearly and collaborate with cross-functional teams in a fast-paced environment.
Must be able to travel occasionally if required for team meetings or company events.
Nice to have
Experience with additional programming languages such as Python or Scala.
Familiarity with containerization and orchestration tools such as Kubernetes.
Knowledge of data governance, privacy, and compliance considerations related to data processing.
Experience with infrastructure as code and CI/CD practices for data pipelines.
Practical notes
This role is full-time.
Travel may be required occasionally for team meetings or company events.
No specific visa sponsorship information is provided in the source.
No deadlines for application submission are specified in the source.