Senior Platform Engineer
Job description
About the role
Solace is reimagining how patients navigate the American healthcare system by combining technology with human advocacy. As a Senior Platform Engineer, you will own the design and evolution of the core infrastructure that powers our patient-facing applications and internal tooling. You will partner closely with product and data teams to translate business requirements into robust, scalable platform capabilities. This role demands a high level of ownership where you will identify risks before they impact users and resolve complex outages at the intersection of technology and process. You will champion automation to reduce manual toil and build self-healing systems that improve reliability over time. Ultimately, you will be responsible for ensuring that our platform is secure, performant, and resilient enough to support rapid growth and mission-critical healthcare workflows.
Key facts
What you'll do
- Architect and implement cloud infrastructure components that serve as the foundation for Solace's healthcare platform.
- Evaluate the risk and downstream impact of proposed changes to production systems while balancing speed and stability.
- Participate in our global on-call rotation to provide timely responses to platform incidents and service disruptions.
- Diagnose and resolve intricate issues that arise from the interaction between technology, operational policies, and human workflows.
- Design and implement autonomous self-healing mechanisms that detect, isolate, and recover from failures with minimal human intervention.
- Act as a subject matter expert and resource for product and data engineers, collaborating across teams to unblock development and improve platform usability.
- Translate technical constraints and trade-offs into clear communication for both technical and non-technical stakeholders.
- Contribute to architectural decision records and long-term platform roadmaps that align with product strategy and compliance requirements.
- Mentor junior engineers by providing code reviews, guidance on best practices, and constructive feedback on system designs.
- Explore and prototype new tools and technologies that can enhance observability, deployment velocity, and infrastructure efficiency.
Requirements
- Have professional experience working in a start-up environment where ambiguity and rapid change are the norm.
- Have hands-on experience building scalable infrastructure to host web applications or data-intensive workloads in production.
- Demonstrate deep troubleshooting abilities by relentlessly digging through multiple abstraction layers to identify root causes under time pressure.
- Have a proven track record of independently diagnosing issues and implementing effective solutions while collaborating with cross-functional teams.
- Be fluent with command-line operations and highly comfortable working in Linux-based environments on a daily basis.
- Possess strong written and verbal communication skills to articulate complex technical concepts to diverse audiences.
- Bring cloud infrastructure expertise with experience in GCP, AWS, or Azure, with a preference for GCP background and knowledge of cloud landing zone patterns.
- Be familiar with GitOps toolchains and infrastructure-as-code practices using technologies such as Terraform.
- Have a solid understanding of networking fundamentals including VPCs, subnets, routing, peering, DNS, load balancers, L4/L7 routing, NAT gateways, CDNs, and TLS.
- Have direct experience with observability by integrating and using tools like Datadog and OpenTelemetry for infrastructure and application performance monitoring.
- Be comfortable writing queries, building dashboards, and collaborating with application engineers to improve service health.
Nice to have
- Have a demonstrated history of improving system performance, reliability, and cost across the entire technology stack from front-end to back-end as a site reliability engineer.
- Have hands-on experience with Kubernetes, including operation, optimization, and self-healing of containerized workloads at scale.
- Be polyglot and comfortable working with multiple programming languages while understanding the operational behavior of services written in different languages.
- Have a background in building and maintaining DevOps toolchains, including CI/CD pipelines and automated testing frameworks.
- Have experience building developer product experiences and designing low-friction tools that enable product and data engineers to be more effective.
- Have worked with complex data infrastructure, including moving databases, building data warehouses, and supporting analytics pipelines.
Practical notes
This role is full-time based in the United States.
Solace is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.