Senior Software Engineer
Job description
About the role
You will design and ship new Command skills that teach the model how to handle workflows such as sending money, managing invoices, and understanding cash flow in a reliable and secure manner. You will architect and own agentic workflows in Command, defining how multi-step agent interactions should behave as we extend what the product can do on a customer's behalf. You will collaborate with backend teams to define tool schemas for new capabilities, shaping the data contracts between Mercury's business logic and the model. You will own new capabilities end to end, from the system prompt to the frontend component that renders the response in a coherent and trustworthy way. You will maintain and evolve Command's prompt architecture, including the system prompt, skill loading system, session context, and the policy and compliance layers underneath. You will tune model behavior, reasoning effort, prompt caching strategy, fallback chains, and the streaming patterns that make the product feel fast and responsive. You will stay current with how models are evolving and bring that knowledge back to influence how Command is built and improved over time.
Key facts
What you'll do
- Design and ship new Command skills that teach the model how to handle workflows like sending money, managing invoices, and understanding cash flow.
- Architect and own agentic workflows in Command, defining the structure of multi-step agent interactions as we extend what the product can do on a customer's behalf.
- Collaborate with backend teams to define tool schemas for new capabilities, shaping the contracts between business logic and the model.
- Own new capabilities end to end, from the system prompt through to the frontend component that renders the model's response.
- Maintain and evolve Command's prompt architecture, including the system prompt, skill loading system, session context, and policy and compliance layers.
- Tune model behavior, including reasoning effort, prompt caching strategy, fallback chains, and streaming patterns to ensure the product feels fast and reliable.
- Stay current with advances in how models are evolving and translate that knowledge into concrete improvements for Command.
- Write and expand Command's eval harness, adding cases that cover new capabilities and scoring rubrics that detect regressions before users do.
- Partner with product and compliance teams to define what "working correctly" means for each new capability, then build the tests that prove it.
- Own the reliability and quality of what you ship, from initial design through post-launch monitoring and iteration.
- Prioritize work as part of a product in motion, choosing the next highest-leverage tasks as Command evolves.
- Contribute to the technical leadership of the LLM layer, guiding decisions on system prompts, loading mechanisms, and safety constraints.
- Collaborate across teams to ensure that agentic workflows are reliable, explainable, and safe to run on behalf of real users.
- Implement streaming interfaces and frontend components that deliver a responsive and trustworthy user experience.
Requirements
- Has 7 or more years of software engineering experience, with deep technical expertise building and scaling LLM-powered applications in production.
- Has gone beyond shipping a first version: you have scaled an LLM-powered product, dealt with the reliability and performance problems that come with real usage, and made it better over time.
- Has experience designing agentic systems and has opinions about how to architect multi-step workflows that are reliable, explainable, and safe to run on behalf of real users.
- Has built eval infrastructure and can write cases that actually measure whether the product works, not just whether the model outputs something plausible.
- Understands the real tradeoffs in LLM deployments: latency, cost, compliance, and what breaks in production that does not show up in demos.
- Has opinions about what makes an AI product trustworthy, not just impressive, and can build toward that standard.
- Is comfortable with TypeScript and willing to learn Haskell for backend tool work, or already comfortable with both.
- Can work across the full stack of an AI product, from the system prompt to the streaming frontend.
- Has a track record of mentoring engineers and raising the technical bar of their team.
- Must be eligible to work in the country of assignment without sponsorship.
- Must be available to work the hours defined in the source engagement details.
Nice to have
Only if the source specifies preferred items.
Practical notes
- This role is full time.
- Work location options include San Francisco, CA, New York, NY, Portland, OR, or Remote within Canada or United States.
- Engagement details and compensation specifics are to be found in the source information.
- Visa sponsorship is not available for this role.
- Candidates must be able to start within the timeframe outlined in the source engagement details.
- Hours and schedule are defined by the engagement details in the source.
- Travel is not expected for this role.