Senior Data Platform Engineer
Job description
About the role
The Senior Data Platform Engineer role at TRM Labs centers on owning the performance, reliability, and operational resilience of TRM's government cloud data serving infrastructure. This position places you at the center of a high-stakes environment where AI-powered investigative platforms must remain fast, accurate, and compliant at all times. You will become the second expert on the StarRocks-backed serving layer, reducing single points of failure and enabling rapid incident response. Your work will directly support public sector investigators by ensuring that distributed analytical systems scale securely and predictably. You will leverage AI-assisted tooling to profile queries, debug production issues, and streamline code reviews under strict governance. The position demands hands-on systems ownership, from tuning query plans to hardening data pipelines against failure. You will play a key role in bridging the gap between commercial platform velocity and government cloud compliance requirements.
What you'll do
- Drive performance tuning on the StarRocks serving layer, using AI-assisted query profiling (Claude, internal tooling) to identify and resolve slow query patterns before they affect customers.
- Build and harden data pipelines that feed government cloud investigations, applying AI code review workflows to ship reliable changes faster while adhering to high-compliance standards.
- Reduce single-point-of-failure risk by becoming the second engineer capable of independently operating and troubleshooting the serving layer, ensuring continuity during critical incidents.
- Use AI-assisted debugging and log analysis to triage production issues in regulated environments, converting multi-hour investigations into rapid root-cause fixes.
- Collaborate closely with Forward Deployed Engineering and Product teams to maintain feature and parity between the government cloud and commercial platforms.
- Participate in weekly incident reviews, in-flight infrastructure work, and upcoming compliance milestones to align technical decisions with regulatory obligations.
- Own the reliability and scalability of distributed OLAP workloads, including query tuning, resource optimization, and capacity planning at scale.
- Contribute to sprint-based planning cycles, taking clear ownership for workstreams and ensuring timely delivery against compliance and operational goals.
- Document runbooks and operational procedures using AI-assisted workflows to accelerate onboarding and maintain consistent responses during outages.
- Engage in async communication across Slack and documentation platforms, supporting evidence-based technical discussions that drive rapid decisions.
Requirements
- U.S. citizenship is required for this role due to government cloud data access requirements.
- Hands-on experience operating distributed OLAP or serving-layer systems such as StarRocks, Trino, or ClickHouse, including query tuning and performance optimization at scale.
- Demonstrated ability to own data pipeline reliability and lead incident response in production environments.
- Comfort using AI tools such as Claude, Cursor, or similar platforms to accelerate debugging, code review, and documentation.
- Experience working in regulated, high-availability environments where mistakes have significant operational and compliance consequences.
- Strong understanding of data ingestion, transformation, and serving patterns that support investigative workflows.
- Proven capacity to ramp on unfamiliar production infrastructure quickly using AI-assisted research and code exploration.
- Willingness to take on-call responsibility with minimal oversight and to respond to incidents in a timely manner.
Nice to have
- Experience contributing to open source database or distributed query engines.
- Background implementing backup, restore, and retention pipelines for compliance-driven workloads.
- Familiarity with audit log retention changes and automated policy enforcement using configuration tooling.
About the Team
The Data Platform Serving team owns the layer that turns TRM's data into fast, reliable answers for the teams and systems built on top of it. We're distributed, not distant: the team communicates constantly across Slack and async docs, with a bias toward direct, evidence-based technical discussion. Decisions are made close to the data: engineers who operate the systems make the calls on architecture and tradeoffs, with input from the broader Data Platform org. We work closely with Forward Deployed Engineering and Product teams to keep the government cloud environment at parity with our commercial platform.
Team Operating Rhythms
- Weekly team sync to review open incidents, in-flight infrastructure work, and upcoming compliance milestones.
- Async daily updates in Slack on pipeline health, ongoing tickets, and blockers.
- Sprint-based planning cycles with clear ownership assigned per workstream.
- Retro after any production incident to capture learnings and adjust runbooks.
Learn about TRM Speed in this position
- A production database migration broke search-attribute registration across two environments right before a critical audit deadline. The on-call engineer traced the root cause, shipped a fix, and had both environments passing smoke tests again within the same day.
- With a compliance deadline days away, the team needed off-cluster backups for the government cloud database with zero prior tooling in place. An engineer designed and shipped the full backup and restore pipeline - three stacked PRs - in under two weeks.
- Audit log retention needed to jump to meet a new compliance floor with no advance notice. The engineer used AI-assisted config generation to update retention policies across the environment and verify compliance the same week.
Life at TRM
We are building a safer world. That promise shows up in how we work every day. TRM moves quickly. We are a high velocity, high ownership team that expects clarity, follow-through, and impact. People who thrive here are energized by hard problems, experimentation, and continuous feedback. If something takes months elsewhere, it will ship here in days.
Our work sits at the intersection of AI, national security, and fighting crime. The problems are complex, the stakes are real, and the environment evolves quickly. The pace and intensity of the work reflect the importance of the mission.