Senior Software Engineer, Data Infrastructure
Job description
Senior Software Engineer, Data Infrastructure at Trm Labs.
About the role
You will own the performance tuning and reliability of the StarRocks serving layer that powers government cloud investigations at TRM Labs. You will build and harden data pipelines that feed critical investigative workflows, ensuring they are resilient and fast in a high-compliance environment. In this role, you will become the second engineer who can independently operate and troubleshoot the serving layer, reducing single-point-of-failure risk and cutting incident response time. You will leverage AI-assisted debugging and log analysis to turn multi-hour production investigations into rapid root-cause fixes. You will use AI code review workflows to ship reliable changes faster while maintaining strict compliance standards. You will work closely with the engineer who currently owns this domain solo to ensure continuity and shared ownership. Your work will directly support public and private sector agencies that rely on TRM to trace illicit activity and build cases in real time.
Key facts
What you'll do
- You will own performance tuning on the StarRocks serving layer, using AI-assisted query profiling (Claude, internal tooling) to find and fix slow query patterns before they become customer-facing incidents.
- You will build and harden data pipelines feeding government cloud investigations, using AI code review workflows to ship reliable changes faster in a high-compliance environment where mistakes are costly.
- You will reduce single-point-of-failure risk on GovCloud data infrastructure by becoming the second engineer who can independently operate and troubleshoot the serving layer, cutting incident response time when the primary owner is unavailable.
- You will use AI-assisted debugging and log analysis to triage production issues in a regulated environment, turning multi-hour investigations into rapid root-cause fixes.
- You will participate in on-call rotation for the serving layer, responding to alerts and incidents with clear runbooks and communication.
- You will collaborate with Forward Deployed Engineering and Product teams to align the government cloud environment with the commercial platform's capabilities and performance.
- You will contribute to weekly incident reviews and retros, capturing learnings and updating runbooks to prevent recurrence.
- You will assist in designing and implementing data pipeline reliability improvements that meet compliance requirements without sacrificing velocity.
- You will evaluate and adopt new AI tools to accelerate debugging, code review, and documentation for complex production systems.
- You will ensure that all changes to the StarRocks-backed serving layer are backed by evidence-based analysis and testing in a regulated, high-availability environment.
Requirements
- U.S. citizenship is required for this role due to government cloud data access requirements.
- Hands-on experience operating distributed OLAP or serving-layer systems (StarRocks, Trino, ClickHouse, or similar), including query tuning and performance optimization at scale.
- Experience owning data pipeline reliability and incident response, and comfort using AI tools (Claude, Cursor, or similar) to accelerate debugging, code review, and documentation.
- Independent ownership mindset: you can pick up an unfamiliar piece of production infrastructure, use AI-assisted research and code exploration to ramp quickly, and take on-call responsibility with minimal oversight.
- Strong understanding of production observability, including logs, metrics, and traces, and how they inform rapid root-cause analysis.
- Experience working in a regulated, high-availability environment where compliance and auditability are non-negotiable.
- Demonstrated ability to work asynchronously and make decisions close to the data with minimal process overhead.
- Willingness to follow runbooks, contribute to their improvement, and adhere to strict operational standards.
Nice to have
- Experience with StarRocks internals or contributing to open-source database projects.
- Prior work in government cloud environments or compliance-driven industries.
- Familiarity with audit log retention and retention policy automation.
Practical notes
This role is full-time and based in the United States. It requires participation in on-call rotations and adherence to operational runbooks. Travel is not required for this position. No specific visa sponsorship details are provided. The engagement is contingent upon U.S. citizenship clearance for government cloud access.