Platform Engineer
Job description
PLATFORM ENGINEER
The Challenge
Thirty-six million businesses in the United States require insurance; it is a non-negotiable obligation. A significant portion of that market is inadequately protected, with 77% underinsured and 40% possessing zero coverage. The existing distribution model has failed these businesses; it is sluggish, opaque, and difficult to navigate.
The majority of commercial insurance remains manually driven. Our objective is the complete inversion: a system driven by artificial intelligence exceeding 90%, pushing toward the upper 90s. We do not intend to repair outdated processes-we are constructing AI that enhances human effectiveness, optimizes the customer journey, and eradicates friction at every interaction point.
Our membership is expanding rapidly, adding approximately 1,000 clients monthly. The company has increased in size 100-fold year-over-year. This momentum will continue, creating the urgency behind our current hiring initiative.
Your responsibility is to construct the foundational systems upon which the engineering organization relies. When these systems perform optimally, the entire company accelerates.
Our Philosophy
Harper operates over 200 services across Railway and Amazon Web Services. We manage numerous parallel agentic workflows. Our artificial intelligence platforms make thousands of decisions daily, and each decision must be traceable, assessable, and more cost-efficient to execute tomorrow than it is today. The underlying infrastructure determines whether a company achieves scalable growth or encounters stagnation.
Exceptional platform engineering is not invisible; it is the catalyst for rapid product delivery. When observability identifies a silent failure before a client does. When a pooling mechanism eliminates connection exhaustion as a concern. When developer productivity doubles due to a tool that was specifically engineered for necessity. This use benefits every engineer on the team.
The Position
You are an infrastructure-focused engineer possessing expertise in databases, Amazon Web Services, networking, observability, site reliability engineering, continuous integration and delivery, and artificial intelligence tooling. You are responsible for building and maintaining the systems and tools that the engineering department depends on, with a focus on performance, reliability, and development speed.
You work directly with company founders and collaborate with product engineers who develop features atop your infrastructure. You identify points of failure at scale, design the solution, and deploy the fix before the issue escalates. When an incident occurs at 2 AM, you are the individual who receives the alert-and you are committed to constructing a system that prevents the alert from happening again.
We are recruiting platform engineers at all levels, which will be defined during the interview process.
What you'll do
- Manage Core Infrastructure: Oversee databases, AWS services, network architecture, CI/CD pipelines, and the systems that the entire engineering organization utilizes.
- Engineer for Scale: Support thousands of concurrent AI operations, ensuring sub-second response times while maintaining cost-effectiveness as load increases.
- Create Effective Observability: Develop system-level error detection and business-critical silent-failure identification; we aim to discover issues through alerts rather than human reports hours later.
- Orchestrate Workflows: Manage N parallel agent instances performing non-deterministic tasks, ensuring reliability and cost-efficiency as the volume of agents scales.
- Amplify Developer Productivity: Create the tools and abstractions that enable product engineers to deploy features in days rather than weeks.
- Ensure End-to-End Reliability: Manage service level objectives, on-call schedules, incident response, and post-incident reviews to minimize future occurrences.
- Develop Evaluation Systems: Construct infrastructure that assesses whether AI outputs are improving over time, transforming vague experimentation into concrete product decisions.
Requirements
- You have managed production systems at scale-not merely contributed to them. You have experienced the urgency of being paged during an outage.
- You write infrastructure code utilizing AI assistants such as Cursor and Claude Code, applying these tools to expedite platform work without sacrificing judgment.
- You possess depth across a minimum of two areas: databases, Amazon Web Services, networking, observability, or continuous integration and delivery.
- You have designed systems capable of handling real-world load, understanding retries, dead-letter queues, back pressure, and idempotency within distributed architectures.
- You prioritize developer experience, recognizing that internal tools are products that dictate how other engineers ship software.
- You prefer to construct systems that prevent fires rather than participating in perpetual firefighting.
- You are based in San Francisco or are willing to relocate to our primary operational location.
Compensation and Logistics
-
Salary: Ranging from $140,000 to $280,000 based on experience, complemented by performance bonuses and equity participation.
-
Location: USA
- Schedule: Monday through Friday, with very early morning hours, requiring attendance five days per week in the office.
- Benefits: Includes Uber commuter benefits; breakfast, lunch, and dinner are provided daily; snacks, drinks, and coffee are available continuously; free gym membership; and comprehensive health, dental, and vision insurance.
The Distinction
This is not a position focused on routine feature development. You will architect and maintain the core infrastructure that powers a rapidly scaling AI insurance enterprise. You will collaborate with the founders to establish technical direction and translate platform strength into market velocity. If you are passionate about developing infrastructure that empowers exceptional people to build exceptional products, we invite you to apply.