Staff Data Platform Engineer
Job description
Staff Data Platform Engineer at Trm Labs.
About the role
This position owns the performance, reliability, and operational continuity of TRM's government cloud data infrastructure. You will be the second engineer on the StarRocks-backed serving layer, eliminating single points of failure and enabling resilient, compliant scaling. You will leverage AI-assisted tooling to accelerate root-cause analysis, enforce data pipeline integrity, and prevent incidents before they reach customers. Your work will directly support regulated investigative workloads that depend on real-time analytical clarity and strict compliance. You will operate in a high-stakes, high-availability environment where decisions are driven by data and evidence. The role demands hands-on distributed systems expertise paired with disciplined change management. You will collaborate closely with Forward Deployed Engineering and Product to align infrastructure capabilities with mission needs. Success is measured by uptime, audit readiness, and the speed at which production issues are resolved.
Key facts
What you'll do
- Assume ownership of the StarRocks serving layer in the government cloud, tuning queries and optimizing resource usage to sustain high throughput under investigative load.
- Drive reliability of data pipelines that feed public sector investigations, implementing AI code review workflows to catch issues early and ship changes safely in regulated environments.
- Reduce operational risk by becoming the second engineer capable of independently supporting the serving layer, ensuring continuity when the primary owner is unavailable.
- Use AI-assisted debugging, log analysis, and profiling tools to accelerate triage and convert multi-hour incidents into rapid, targeted fixes.
- Participate in on-call rotations to respond to production issues, applying runbooks, playbooks, and real-time telemetry to restore service quickly.
- Collaborate with compliance and security teams to implement configuration changes, retention policies, and access controls that satisfy government cloud requirements.
- Design and validate backup and restore procedures for critical OLAP workloads, ensuring recoverability within defined recovery time objectives.
- Work with product and FDE engineers to translate investigation workflows into scalable data models and query patterns that preserve platform parity across commercial and government clouds.
- Contribute to architecture decisions for the serving layer, balancing performance, cost, and regulatory constraints with pragmatic engineering tradeoffs.
- Document operational procedures, incident responses, and system behaviors to accelerate onboarding and sustain knowledge across the distributed team.
- Support migration efforts that move data services into the government cloud, validating integrity, performance, and audit readiness at each step.
- Mentor engineers new to distributed OLAP systems by pairing on query optimization, log analysis, and debugging techniques in regulated contexts.
- Coordinate with SRE and platform teams to align infrastructure health checks, monitoring, and alerting with broader reliability goals.
- Champion the use of AI tools such as Claude and internal tooling to speed up code review, root-cause analysis, and documentation without compromising compliance.
- Track key service metrics for the serving layer, identify trends, and propose data-driven improvements to capacity and scaling strategies.
Requirements
- U.S. citizenship is required for this role due to government cloud data access requirements.
- Hands-on experience operating distributed OLAP or serving-layer systems such as StarRocks, Trino, or ClickHouse in production environments.
- Demonstrated ability to tune complex queries and optimize performance at scale, including execution plans, indexing, and resource utilization.
- Experience owning data pipeline reliability, including monitoring, alerting, and incident response in regulated, high-compliance settings.
- Comfort using AI-assisted tools such as Claude, Cursor, or similar platforms to accelerate debugging, code review, and documentation.
- Independent ownership mindset with the ability to ramp on unfamiliar production infrastructure using research, code exploration, and AI assistance.
- Willingness to assume on-call responsibility and make time-bound decisions with minimal oversight in critical situations.
- Strong written and verbal communication skills to articulate technical tradeoffs to both engineering and non-technical stakeholders.
- Experience working in agile, distributed team environments with a preference for asynchronous communication and evidence-based discussions.
- Familiarity with compliance-driven workflows and the impact of audit requirements on data infrastructure operations.
- Ability to manage multiple priorities in a fast-paced setting while maintaining attention to detail and operational rigor.
- Understanding of secure configuration practices for government cloud deployments and data access controls.
- Commitment to following runbooks, contributing to post-incident reviews, and improving operational processes over time.
Nice to have
- Experience with StarRocks internals, query optimization, and execution planning in large-scale deployments.
- Prior work delivering infrastructure changes under strict compliance or audit deadlines in regulated environments.
- Contributions to open source OLAP projects or data platform tooling that demonstrates deep systems knowledge.
- Familiarity with backup, restore, and disaster recovery patterns for analytical databases in cloud environments.
- Background working alongside Forward Deployed Engineering teams to align platform capabilities with customer needs.
Practical notes
This is a full-time position based in the United States. The role may require periodic availability to support production incidents and compliance milestones. Travel is not currently required, and visa sponsorship is not available for this role.