Data Platform Engineer
Job description
-automation.
About the role
Hadrian is constructing autonomous factories to reindustrialize America by merging AI, advanced software, robotics, and full-stack manufacturing. This integration helps aerospace and defense companies build rockets, satellites, aircraft, ships, and other mission-critical systems up to 10x faster and at significantly lower cost. As a Data Platform Engineer, you will own critical segments of the data flow from source to trusted dataset, managing ingestion, change data capture, streaming, lakehouse storage, orchestration, transformation, contracts, quality, and lineage. Some data sources may involve machines, PLCs, historians, and industrial protocols, where controls experience is advantageous but not required. Your work will function as a data-platform and distributed-systems role, ensuring new sources can be integrated easily, failures are straightforward to repair, and datasets remain reliable for Analytics, Data Science, Operations Research, and ML systems. You will play a key role in transforming machine signals, quality events, work orders, and application changes into trusted data that powers scheduling, ML, and operations for the future of American manufacturing.
Key facts
What you'll do
- Build the data backbone for autonomous factories, transforming machine signals, quality events, work orders, and application changes into trusted data used by scheduling, ML, and operations.
- Build ingestion, CDC, and streaming capabilities for transactional data, events, telemetry, and files; explicitly manage ordering, deletes, retries, replay, idempotence, and backpressure.
- Define versioned data and event contracts with upstream teams, supported by testing and service targets for freshness, completeness, and correctness.
- Model telemetry, quality events, work orders, and operational data into datasets with explicit grain, identity, time, provenance, and history.
- Own Dagster orchestration, dbt transformation, data CI/CD, backfills, lineage, observability, and offline feature datasets for ML.
- Collaborate with Manufacturing Operations and Infrastructure to acquire data from machines, PLCs, historians, OPC-UA, MTConnect, and MQTT sources as needed.
- Ensure data reliability and performance across distributed systems by implementing robust error handling, monitoring, and recovery processes.
- Partner with data scientists and analysts to optimize dataset structures, access patterns, and storage formats for analytical and machine learning workloads.
- Contribute to the design and evolution of the data platform architecture to support scalability, maintainability, and operational excellence.
- Participate in on-call rotations to address production incidents, diagnose root causes, and implement preventative improvements.
- Document data pipelines, schemas, contracts, and operational procedures to support knowledge sharing and system maintainability.
- Engage with the broader engineering community through code reviews, design discussions, and cross-team collaboration to uphold best practices.
Requirements
- Experience building and operating production data infrastructure or distributed data systems, including on-call ownership and recovery efforts.
- Strong production Python and advanced SQL and data-modeling skills, including incremental processing, temporal data, and schema evolution.
- Experience with Kafka or another event-streaming platform, plus CDC or other stateful incremental pipelines.
- Experience operating Snowflake, and with a lakehouse table format such as Iceberg, Delta, or Hudi, including expertise in partitioning and compaction.
- Experience with tools such as Dagster, Airflow, Argo, or Prefect; dbt or similar transformation frameworks; and Kubernetes or infrastructure as code.
- Strong judgment regarding contracts, failure modes, and the needs of downstream analytics, ML, and operational systems.
- Proven ability to work effectively in a fast-paced, high-growth manufacturing environment with evolving requirements and complex data sources.
- Commitment to writing clean, testable, and maintainable code, with a focus on reliability, observability, and performance.
Nice to have
- Experience running Snowflake and Iceberg together or designing a hybrid warehouse and lakehouse architecture.
- Production experience with PeerDB, Debezium, Flink, Spark Structured Streaming, Redpanda, Bufstream, or similar CDC and streaming systems.
- Proficiency with ClickHouse or another low-latency analytical database, including performance tuning and lifecycle management.
- Experience with industrial or edge data collection using OPC-UA, MTConnect, MQTT, historians, PLCs, or handling intermittently connected systems.
- Background in performance-sensitive data systems built with Go, Rust, or Scala; regulated-environment experience; or contributions to dbt, Dagster, Iceberg, or related open-source projects.
Practical notes
Full_time engagement.
Compensation range is $170,000 - $300,000.
Location in Los Angeles, CA.