Software Engineer, Data Infrastructure
Job description
About the role
Position: Data Infrastructure Engineer, Government Cloud
Introduction
TRM Labs builds technology that helps public and private sector teams investigate and stop crime. Our platforms enable the tracing of illicit networks and the construction of investigative pictures for decision-makers. The work intersects AI, national security, and public safety. The environment is high-velocity and high-stakes, demanding speed without sacrificing compliance.
We are expanding the Data Platform Serving team. This team owns the layer that converts raw data into fast, reliable answers for internal and external customers. You will join this team as a second engineer on a critical government cloud domain currently managed by a single owner. The goal is to harden infrastructure, reduce risk, and ensure performance under strict regulatory requirements. This role is based in the United States.
The Opportunity
You will join the Data Platform Serving team, which supports the StarRocks-backed layer used for government cloud investigations. You will work directly with the current domain owner, learning the system deeply while preparing it for scale and resilience. This is a hands-on role focused on distributed systems in a regulated, high-availability environment.
Your contributions will focus on four core areas: performance, reliability, automation, and compliance. You will use AI-assisted tools to accelerate delivery and ensure stability.
Responsibilities
- Performance Optimization: You will manage query tuning and configuration on the StarRocks serving layer. Using AI-assisted query profiling tools, you will identify slow patterns and implement fixes to prevent incidents.
- Pipeline Integrity: You will build and harden data pipelines for government investigations. AI code review workflows will be used to validate changes, ensuring reliability and speed in a high-compliance setting.
- Resilience Engineering: You will reduce single points of failure by becoming the second expert on the serving layer. This decreases incident response time and ensures continuity.
- Production Support: You will use AI-assisted log analysis and debugging to resolve production issues. The aim is to convert lengthy investigations into swift, root-cause resolutions.
- Compliance Readiness: You will coordinate activities such as backup strategies and retention policy updates to meet audit and compliance deadlines.
Requirements
- Clearance and
Location: USA
- Technical Expertise: Hands-on experience with distributed OLAP systems like StarRocks, Trino, or ClickHouse. You must have tuned queries in production environments.
- Reliability Focus: A record of owning data pipeline reliability and incident response. You must be comfortable using AI tools such as Claude and Cursor to accelerate work.
- Ownership Mindset: The ability to learn unfamiliar production systems quickly using self-directed research. You must be prepared to take-on call duties with minimal supervision.
Team and Collaboration
The Data Platform Serving team operates as a distributed unit. Communication is constant and occurs primarily through Slack and async documentation. The team operates on a bias for action, favoring direct, evidence-based technical discussions. Decisions are made by engineers who run the systems, informed by the broader Data Platform organization.
You will work closely with Forward Deployed Engineering and Product teams. This ensures the government cloud environment remains aligned with the commercial platform in both capability and timeline.
Operating Rhythms
- Weekly Sync: The team meets to review incidents, infrastructure work, and compliance milestones.
- Async Updates: Daily Slack messages cover pipeline health, tickets, and blockers.
- Sprint Planning: Work is organized in sprints with clear ownership for each workstream.
- Retrospectives: After any production incident, the team holds a retro to capture lessons and update runbooks.
Why This Role Matters
The problems are complex and the environment evolves rapidly. The role provides direct exposure to systems where failure has real-world consequences. The pace reflects the importance of the mission.
Life at TRM
TRM moves quickly. The team is high-velocity and high-ownership, requiring clarity, follow-through, and impact. People who succeed here enjoy hard problems, experimentation, and continuous feedback. If something takes months elsewhere, it will ship here in days.
Application
If you are motivated by complex infrastructure challenges and compliance-driven environments, this role is a fit. You will build scalable systems while contributing to a safer world.
What you'll do
- Manage query tuning and configuration on the StarRocks serving layer using AI-assisted query profiling tools to identify slow patterns and implement fixes that prevent incidents. This includes analyzing execution plans, optimizing join strategies, and adjusting resource allocation to meet investigative workload demands.
- Build and harden data pipelines for government investigations while employing AI code review workflows to validate changes, ensuring reliability and speed in a high-compliance setting. This involves creating robust ETL processes that maintain data integrity and support audit trails.
- Reduce single points of failure by becoming the second expert on the serving layer, decreasing incident response time and ensuring continuity for critical systems. You will document architectures, create runbooks, and implement monitoring to enable rapid detection and response.
- Utilize AI-assisted log analysis and debugging to resolve production issues, converting lengthy investigations into swift, root-cause resolutions. This includes correlating events across distributed systems and applying fixes based on empirical evidence.
- Coordinate activities such as backup strategies and retention policy updates to meet audit and compliance deadlines for government cloud workloads. You will ensure all data management practices align with regulatory frameworks and internal governance standards.
- Leverage distributed systems expertise in OLAP technologies like StarRocks, Trino, or ClickHouse to optimize performance and reliability. This includes capacity planning, scaling strategies, and technology evaluation for government-specific requirements.
- Maintain pipeline integrity through hands-on management of data flows that support government investigations and commercial platform requirements. You will implement validation checks, error handling, and monitoring to ensure continuous operation.
- Operate within a regulated, high-availability environment where speed without sacrificing compliance is essential. This involves balancing rapid iteration with thorough testing and documentation to meet governmental standards.
- Collaborate closely with the current domain owner to ensure a smooth transition and deep system understanding. You will participate in knowledge transfer sessions and contribute to architectural decisions.
- Drive ownership of infrastructure components to ensure resilience, scalability, and operational excellence. This includes proactive monitoring, capacity management, and leading incident post-mortems.
Requirements
- U.S. citizenship is mandatory due to the sensitivity of government cloud data.
- Hands-on experience with distributed OLAP systems like StarRocks, Trino, or ClickHouse is required. You must have tuned queries in production environments to ensure optimal performance and reliability.
- A record of owning data pipeline reliability and incident response is essential for success in this role. You must demonstrate the ability to manage complex data flows and respond effectively to outages.
- You must be comfortable using AI tools such as Claude and Cursor to accelerate work and improve efficiency. This includes leveraging these tools for code review, debugging, and documentation.
- The ability to learn unfamiliar production systems quickly using self-directed research is required. You should demonstrate resourcefulness in navigating new technologies and environments.
- You must be prepared to take-on call duties with minimal supervision and respond to production issues promptly. This includes participating in on-call rotations and maintaining situational awareness.
- A strong commitment to compliance and operational rigor is necessary to meet regulatory standards. You must understand the implications of government cloud requirements and implement controls accordingly.
- Clearance and location requirements include U.S. citizenship and the ability to work in the United States. Travel is not applicable, and visa requirements are not applicable due to citizenship requirements.
- Experience working in regulated environments is preferred, with demonstrated ability to balance speed with compliance.
- Strong written and verbal communication skills are required to collaborate effectively with distributed teams and stakeholders.
Practical notes
- Hours: Full-time during standard business hours in the United States.
- Travel: Not applicable.
- Visa: Not applicable due to United States location and citizenship requirements.
- Deadlines: Compliance-driven milestones may require timely delivery against audit and regulatory schedules.