Senior Data & Analytics Engineer, Domain Enablement
Job description
About the role
You will architect and own the core semantic data layer that allows Shield AI business domains to operate with a single version of the truth on the Databricks lakehouse. In this hybrid position, you translate ambiguous business questions into durable data models, balancing the needs of finance, GTM, and operations teams with long-term platform scalability. You will design and govern the transition from raw Bronze tables to curated Silver and Gold assets, ensuring data quality, lineage, and performance meet enterprise standards. This role requires you to build reusable patterns and modular datasets that domain analysts can self-serve without constant engineering intervention. You will collaborate closely with data scientists and business stakeholders to validate model outputs and ensure analytical consistency across the organization. You will also act as a technical coach for domain teams, helping them understand data structures and best practices so they can extend datasets responsibly. This position is critical for moving Shield AI from ad hoc reporting to a scalable, governed analytics environment that supports rapid decision-making.
Key facts
What you'll do
Architect and deliver governed Silver and Gold datasets on Databricks that serve as the foundational layer for domain analytics and reporting.
Implement reusable semantic models and calculation patterns that standardize key business metrics such as financials, pipeline, and operational KPIs across G&A and GTM functions.
Partner with stakeholders in Accounting, Program Finance, RevOps, Marketing, and HR to translate complex requirements into clear data specifications and model designs.
Establish data quality frameworks and automated monitoring that detect anomalies, drifts, and breakdowns before they impact business decisions.
Optimize data structures and partitioning strategies to improve query performance and reduce costs across large and growing datasets.
Lead the definition of naming conventions, data dictionaries, and documentation standards to ensure clarity and consistency for both technical and non-technical users.
Develop and maintain reusable SQL and Python patterns that accelerate common analytical workflows and reduce redundant effort across domains.
Enable self-service analytics by building flexible domain datasets that business users can explore in BI tools without requiring deep SQL expertise.
Collaborate with data platform and infrastructure teams to ensure new domain datasets integrate cleanly with existing lakehouse architecture and security controls.
Champion best practices in data governance, including access controls, lineage tracking, and compliance with internal policies and external regulations.
Support the rollout of simulation and synthetic reality data capabilities by preparing analytics-ready datasets for Aechelon and related platforms.
Evaluate emerging tools and frameworks that can enhance the performance, reliability, or usability of domain data products over time.
Act as a technical liaison between engineering and business teams during incident investigations, data disputes, or model validation exercises.
Contribute to the long-term roadmap for domain data platforms, providing input on priorities, trade-offs, and resource requirements based on observed usage and feedback.
Requirements
You possess a Bachelor's degree in Computer Science, Engineering, Mathematics, or a related quantitative field, or equivalent practical experience.
You have proven experience building and maintaining analytics datasets in a modern lakehouse environment, with deep expertise in Databricks or a comparable data platform.
You demonstrate advanced SQL capabilities, including complex joins, window functions, CTEs, and performance tuning for large-scale analytical workloads.
You are proficient in Python or Scala for data transformation, UDF development, and pipeline logic within a data engineering context.
You have a strong understanding of data modeling principles, including dimensional modeling, conformed dimensions, and slowly changing dimensions for enterprise reporting.
You show mastery over data quality techniques, including anomaly detection, constraint validation, and reconciliation processes across distributed datasets.
You bring experience working with finance, operations, or GTM data sets, including general ledger concepts, revenue metrics, and pipeline analytics.
You communicate effectively with both technical and non-technical audiences, translating ambiguous problems into structured analytical approaches.
Nice to have
Experience with data virtualization, semantic layer tools, or metadata management platforms that support self-service analytics.
Familiarity with compliance frameworks relevant to defense and government environments, including ITAR, EAR, or export control considerations.
Hands-on work with synthetic data generation, simulation data pipelines, or digital twin concepts in prior roles.
Background in agile product delivery, where you have participated in sprint planning, retrospectives, and stakeholder demonstrations.
Practical notes
This is a fully remote position based on the information provided.
No specific working hours are defined in the source material, allowing for flexibility within a professional remote work arrangement.
No travel requirements are specified in the source material.
No visa sponsorship details are provided in the source material.
No application deadline is mentioned in the source material.