Data Platform Engineer
Job description
Data Platform Engineer at Trm Labs.
About the role
You will own the performance tuning and reliability of the StarRocks serving layer that powers government cloud investigations at TRM Labs. In this role, you will build and harden data pipelines to ensure they meet the strict standards of a high-compliance environment. You will become the second engineer capable of independently operating this critical infrastructure, reducing single points of failure and incident response times. The position requires using AI-assisted tools like Claude for query profiling, debugging, and code review to accelerate delivery and maintain system integrity. You will triage production issues in real time, turning complex multi-hour investigations into rapid root-cause fixes. This role places you at the center of a distributed, evidence-based team that makes decisions close to the data. Your work will directly support the mission of building a safer world by ensuring our analytical platforms are fast, reliable, and always available for public sector customers.
Key facts
What you'll do
- You will monitor and optimize query performance on the StarRocks serving layer using AI-assisted query profiling to resolve slow patterns before they impact customers.
- You will design, implement, and validate data pipelines that feed government cloud investigations, ensuring reliability and compliance through AI code review workflows.
- You will reduce operational risk by becoming the second engineer who can independently troubleshoot and support the serving layer, enabling faster incident response.
- You will leverage AI-assisted debugging and log analysis tools to triage production issues, converting lengthy investigations into quick root-cause resolutions.
- You will maintain high-availability standards in a regulated government cloud environment, ensuring system changes never compromise compliance or uptime.
- You will participate in on-call rotations to address serving layer incidents, applying AI tools to diagnose and remediate issues under tight timelines.
- You will collaborate with Forward Deployed Engineering and Product teams to align the government cloud platform with commercial capabilities and performance goals.
- You will contribute to incident retrospectives, documenting learnings and updating runbooks to prevent recurrence and improve system resilience.
- You will engage in weekly infrastructure reviews to track in-flight work, compliance milestones, and upcoming migration activities.
- You will help build and harden automation around backups, retention policies, and audit readiness using AI-assisted configuration workflows.
- You will ensure data pipeline reliability by identifying bottlenecks, implementing fixes, and validating changes through rigorous testing in production-like environments.
- You will support audit and compliance initiatives by rapidly adapting configurations and verifying policy enforcement across the serving and pipeline layers.
- You will mentor and enable other engineers by sharing AI-driven debugging techniques and operational best practices for distributed OLAP systems.
- You will contribute to the design of future serving layer enhancements, balancing performance, scalability, and regulatory requirements.
Requirements
- U.S. citizenship is required for this role due to government cloud data access requirements.
- Hands-on experience operating distributed OLAP or serving-layer systems such as StarRocks, Trino, or ClickHouse.
- Demonstrated ability to tune queries and optimize performance at scale in production environments.
- Experience owning data pipeline reliability and leading incident response efforts in high-stakes systems.
- Comfort using AI tools such as Claude, Cursor, or similar platforms to accelerate debugging, code review, and documentation.
- Independent ownership mindset with the ability to ramp quickly on unfamiliar infrastructure using AI-assisted research and code exploration.
- Willingness to take on-call responsibilities and make decisions with minimal oversight in a regulated environment.
- Strong understanding of compliance and operational risk in government cloud deployments.
- Ability to work effectively in a distributed team with a bias toward direct, evidence-based technical communication.
- Willingness to follow and contribute to runbooks, retrospectives, and sprint-based planning cycles.
- Experience working in regulated environments where auditability and reliability are critical.
- Capacity to manage multiple priorities during high-velocity production incidents and compliance deadlines.
- Commitment to maintaining system availability and performance during critical audit and investigation windows.
Nice to have
- Experience working with regulated, high-availability government cloud deployments.
- Background contributing to open source or distributed database projects like StarRocks.
- Familiarity with audit log retention and compliance-driven configuration management.
- Track record of building automated backup and restore workflows under tight timelines.
- Previous work with AI-assisted development workflows in fast-paced engineering environments.
Practical notes
This role is full-time and based in the United States. It requires availability to participate in on-call rotations and respond to production incidents as needed. No visa sponsorship is available for this position. The role is part of a sprint-based delivery cadence with weekly syncs, async daily updates, and post-incident retrospectives.