Software Engineer, Infrastructure
Job description
About the role
You will define and shape the technical direction of infrastructure and platform capabilities that directly support Rad AI's rapidly expanding AI product suite across reporting, impressions, and continuity. You will architect and evolve our cloud infrastructure on AWS, managing container orchestration with Kubernetes and Elastic Container Service, serverless patterns using Lambda, virtual machines on EC2, and scalable data stores to serve current and future product needs. You will work hand in hand with Platform leadership, product engineering, data, and ML teams to design systems that are robust, observable, and compliant within a regulated healthcare environment. In this role, you will define and drive infrastructure strategy, aligning roadmaps and setting standards in partnership with engineering leadership to prioritize work for maximum business impact. You will secure networking, identity, and access patterns across environments while improving reliability and operational excellence by defining and tracking SLOs, SLIs, and error budgets. You will lead and participate in blameless post-incident reviews, translating learnings into systemic improvements that strengthen our platforms. You will own observability and monitoring strategy across logging, metrics, and tracing, ensuring we can detect, debug, and prevent issues efficiently at scale. You will mentor and level up engineers across Platform and product teams by reviewing design documents, guiding architecture decisions, and modeling high standards for reliability, security, and maintainability. You will partner with security and compliance stakeholders to ensure our infrastructure and operational practices meet HIPAA and other healthcare requirements. You will advocate for and implement developer experience improvements, such as better CI/CD workflows, faster feedback loops, and tooling that reduces cognitive load for product teams.
Key facts
What you'll do
- Influence the technical direction for infrastructure and platform capabilities that support our rapidly growing AI product suite.
- Architect and evolve our cloud infrastructure (primarily on AWS) across container orchestration (Kubernetes, Elastic Container Service), serverless (e.g., Lambda), virtual machines (e.g., EC2), and data stores to support current and future products.
- Work closely with Platform leadership, product engineering, data, and ML teams to design systems that are robust, observable, and compliant in a healthcare environment.
- Define and drive infrastructure strategy for the Platform org - partnering with engineering leadership to align roadmaps, set standards, and sequence work for maximum business impact.
- Secure networking, identity, and access patterns across environments.
- Improve reliability and operational excellence by defining SLOs, SLIs, and error budgets for core platform services.
- Leading and participating in blameless post-incident reviews and translating learnings into systemic improvements.
- Own observability and monitoring strategy across logging, metrics, and tracing, ensuring we can detect, debug, and prevent issues efficiently.
- Mentor and level up engineers across Platform and product teams - reviewing design docs, guiding architecture decisions, and modeling high standards for reliability, security, and maintainability.
- Partner with security and compliance stakeholders to ensure our infrastructure and operational practices meet HIPAA and other healthcare requirements.
- Advocate for and implement developer experience improvements, such as better CI/CD workflows, faster feedback loops, and tooling that reduces cognitive load for product teams.
Requirements
- Bring 4+ years of hands-on infrastructure / platform development experience (or equivalent practical experience) in modern, cloud-native environments, with a track record of owning critical systems in production.
- Have deep expertise with AWS (preferred) and/or GCP, including core networking, compute, storage, and managed services.
- Are highly proficient in at least one programming/scripting language commonly used for infrastructure automation, such as Python, Go, or Bash.
- Have experience designing, building, and operating reliable distributed systems in the cloud, with familiarity with container orchestration platforms like Kubernetes and container runtimes.
- Are comfortable working with relational and NoSQL databases, as well as data stores common in analytics and machine learning pipelines.
- Have strong experience with Infrastructure as Code tools such as Terraform or CloudFormation, and configuration management tools.
- Have a strong understanding of networking fundamentals, including TCP/IP, load balancing, firewalls, and VPNs, as they apply to cloud environments.
- Are experienced with monitoring, logging, and observability tools and practices, and can use these signals to drive reliability improvements.