Manager, Engineering
Job description
About the role
This position is responsible for leading the development and delivery of tier 0 grid services and foundational data infrastructure that underpins millions of daily user requests across the organization. You will own the architecture, roadmap, and operational health of critical platform services that demand the highest levels of reliability and performance. The role involves managing a high-performing engineering team of 6 to 10 plus engineers, fostering a culture of technical excellence and continuous improvement. You will ensure 99.999 percent service availability by implementing robust incident management, reliability engineering, and on-call practices. A core part of this role is defining and maintaining low-latency systems capable of processing millions of requests daily while mentoring staff on distributed systems principles. You will partner with cross-functional stakeholders to manage dependencies and drive technical projects that align with business objectives. The position requires integrating AI into engineering lifecycles to enhance operational efficiency and automate routine tasks. Finally, you will define SLOs, SLAs, capacity plans, and observability strategies, reporting metrics and progress to executive stakeholders on a regular basis.
Key facts
What you'll do
- Lead and grow a team of 6 to 10 plus software engineers focused on mission-critical platform infrastructure, providing clear direction, feedback, and career development.
- Maintain 99.999 percent availability through rigorous incident management, reliability engineering practices, and structured on-call rotations to minimize downtime.
- Architect low-latency, scalable systems designed to process millions of requests daily while optimizing resource utilization and cost efficiency.
- Mentor staff and engineers on distributed systems concepts, scalability patterns, and platform design to elevate the overall technical capability of the team.
- Direct cross-functional technical projects, coordinating with product, infrastructure, and security teams to manage dependencies and delivery timelines.
- Define and enforce service-level objectives (SLOs), service-level agreements (SLAs), capacity planning strategies, and comprehensive observability frameworks.
- Partner with product managers, security specialists, and infrastructure leaders to collaboratively define and execute technical roadmaps that support organizational goals.
- Integrate artificial intelligence and machine learning into engineering workflows and tooling to improve development efficiency and system performance.
- Oversee quarterly and annual goal setting, tracking key performance indicators, and communicating metrics, outcomes, and trends to executive stakeholders.
- Design and evolve foundational data infrastructure that ensures consistency, reliability, and high throughput across all grid services.
- Implement best practices for concurrency control and data consistency to maintain integrity in high-volume transactional systems.
- Drive performance optimization initiatives, identifying bottlenecks and applying advanced techniques to improve response times and throughput.
- Build and maintain robust cloud-native backend services on AWS, leveraging managed services and infrastructure as code for scalability and resilience.
- Establish and refine observability standards, including metrics collection, distributed tracing, and alerting to enable proactive issue detection.
- Lead efforts in automating operational tasks, reducing manual toil, and improving the reliability of deployment and release processes.
- Collaborate with security teams to ensure that all platform services adhere to the highest standards of compliance and data protection.
- Evaluate and adopt new technologies and tools that enhance the efficiency and scalability of grid services over time.
- Facilitate regular design reviews, post-incident analyses, and knowledge-sharing sessions to promote a culture of learning and accountability.
- Coordinate with customer support and operations teams to ensure alignment on service expectations and incident response procedures.
- Develop and maintain runbooks, playbooks, and documentation to support operational excellence and enable team scalability.
- Guide technical decision-making through trade-off analysis, balancing innovation with the need for stability and operational simplicity.
Requirements
- Bring 8 plus years of professional software development experience with deep expertise in algorithms, data structures, and distributed systems.
- Demonstrate 3 plus years of hands-on engineering management experience leading teams responsible for high-availability infrastructure.
- Show a proven track record of operating tier 0 systems that maintain five-nines availability in production environments.
- Possess hands-on experience building and operating low-latency, high-throughput cloud-native services, particularly on AWS or equivalent platforms.
- Exhibit strong proficiency in observability practices, including metrics instrumentation, distributed tracing, and alerting strategy design.
- Have experience managing complex technical dependencies and scaling backend architectures to meet growing demands.
- Display the ability to balance technical excellence with operational pragmatism, making decisions that support both innovation and reliability.
- Hold a Bachelor of Science or higher degree in Computer Science, Engineering, or a related field, or demonstrate equivalent practical experience through a strong portfolio and professional background.
- Be legally eligible to work in the United States on an ongoing basis without requiring future sponsorship or visa changes.
Nice to have
- Demonstrated experience managing full-stack engineering teams, including frontend, backend, and DevOps collaboration.
- Prior exposure to AI or ML platform integration, including MLOps practices and model deployment strategies.
Practical notes
- This is a full-time position eligible for remote work within the United States where Smartsheet is a registered employer, with the primary office location in Bellevue, Washington.
- Benefits include subsidized medical, vision, and dental coverage, a 401k match offering 50 percent of employee contributions up to 6 percent of base salary, and 12 paid holidays annually.
- The role includes up to 24 weeks of paid parental leave, a monthly productivity stipend to support remote work, and one personal volunteer day per year.
- Employees have access to professional development resources such as Udemy learning subscriptions to support continuous skill growth.
- Smartsheet is committed to being an Equal Opportunity Employer, welcoming diverse talent and fostering an inclusive workplace for all employees.