Data Engineer
Job description
About the role
OpenGov is looking for a high-ownership, solution-oriented Data Engineer to join our rapidly growing Data Platform team. You will work hands-on with dbt to model and transform data that powers decision-making across the organization. In this role, you will take end-to-end ownership of how analytical work gets shipped, ensuring version control, rigorous testing, and reliable deployment through CI/CD pipelines instead of ad hoc handoffs. You view data infrastructure as a product, treating data and pipelines with the same care as production software that is maintained, monitored, and iterated upon over time. You will partner closely with business analytics, product, and AI teams to deeply understand their needs and deliver data solutions that are well-structured, well-documented, and production-ready for high-stakes public sector use.
Key facts
What you'll do
You will partner with business stakeholders across Analytics, Operations, GTM, and G&A to gather requirements and clarify problem statements, translating them into precise data requirements and success metrics. You will design and build dimensional and analytical data models in Snowflake using dbt (Cloud/Core), establishing clear staging, intermediate, and mart layers that serve as a single source of truth for key business domains. You will develop, test, and maintain dbt models following best practices, writing comprehensive documentation and robust data tests to ensure correctness, consistency, and reliability over time. You will optimize SQL queries and dbt models for performance in Snowflake, leveraging clustering keys, thoughtful materialization strategies, and query profiling to deliver fast, efficient analytics. You will own the scheduling and orchestration of dbt jobs, implementing monitoring and observability practices that uphold strict reliability and SLA adherence for critical data pipelines. You will define and implement CI/CD workflows for dbt and pipeline deployments using GitHub Actions, including automated testing, linting, and safe promotion through development, staging, and production environments. You will work with the Data Platform team on AWS resource provisioning via Terraform to support scalable ingestion and pipeline infrastructure that meets growing public sector demands. You will implement data quality checks and freshness tests within dbt to proactively detect anomalies, preventing bad data from reaching consumers and supporting trustworthy insights. You will perform deep exploratory data analysis to validate source data, understand distributions and edge cases, and clearly communicate findings during scoping and build phases to reduce risk. You will collaborate with data analysts, product managers, and AI teams to ensure that data models and pipelines align with analytical use cases and downstream consumption patterns in a regulated environment.
Requirements
You hold a Bachelor's degree in Computer Science, Mathematics, Engineering, Statistics, or a closely related quantitative field that provides a strong foundation for analytical systems. You bring 4-6 years of professional experience in data engineering, backend engineering, or a similar role where you have delivered production-grade data solutions. You have hands-on data warehouse experience and deep proficiency with dbt, having built and maintained complex dbt projects in a cloud data warehouse such as Snowflake. You are experienced with DevOps practices and have built or maintained CI/CD pipelines, with a strong preference for GitHub Actions to automate testing, linting, and deployment of data pipelines. You understand dimensional modeling principles deeply, including star schema design, handling schema drift, and creating models that balance clarity, performance, and flexibility for evolving business needs. You are comfortable working with core AWS services like S3, Lambda, and IAM to integrate data sources and support secure, scalable pipeline execution. You possess mastery of SQL, writing advanced queries that include complex transformations, window functions, and performance tuning for large analytical workloads. You are proficient in Python, using it for data manipulation, pipeline scripting, automation, and lightweight tooling that supports data workflows. You are comfortable exploring raw data sets, identifying quality issues, diagnosing root causes, and clearly articulating insights to both technical and non-technical stakeholders. You treat data code as serious software, applying software engineering practices such as version control, peer review, testing, and thorough documentation throughout the lifecycle. You communicate effectively and collaborate well across technical and non-technical audiences, including public sector partners with varying levels of technical familiarity. You have experience using AI-assisted development tools such as Claude Code, Cursor, or Codex to accelerate engineering workflows and improve delivery speed without sacrificing quality.
Nice to have
You have some working exposure to Terraform or other Infrastructure as Code tools for provisioning and managing cloud resources in a controlled, auditable way. You have comfort conducting discovery sessions with business stakeholders, translating ambiguous analytical problems into well-structured data models and clear requirements. You have exposure to workflow orchestrators such as Airflow, Dagster, or Prefect, even if they are not part of your daily responsibilities, demonstrating an understanding of scheduling and dependency management. You are familiar with dbt Mesh and multi-database architectures, showing an interest in scalable data mesh concepts and managing shared semantic layers across diverse data platforms.