Software Engineer II, Developer Tooling
Job description
About the role
You will own the design and delivery of incident management tooling that keeps Auth0 and its customers resilient under pressure. You will build and extend Slack-based automation that streamlines on-call workflows and accelerates response times. You will take ownership of the status page, ensuring it provides clear, timely communication during critical events. You will contribute to the reliability and roadmap of internal platforms such as incident.io and Vivaldi. You will implement features that help engineers understand and act on system behavior in real time. You will collaborate closely with platform and product teams to turn operational needs into robust tooling. You will help meet FedRAMP requirements by extending incident capabilities for regulated workloads. You will partner with both engineering and operations teams to ensure solutions serve both builders and customers.
Key facts
What you'll do
Investigate and define solutions for complex incident management challenges across distributed services.
Design and implement features for no-code automation platforms such as Tines to reduce manual toil and improve response coverage.
Build and maintain Slack bots and slash commands that automate incident response and integrate with internal workflows.
Own and evolve the customer-facing status page, improving clarity, accuracy, and speed of incident communication.
Develop and support local development environments and tooling such as Tilt to improve developer velocity and reliability.
Extend and operate services on AWS with Kubernetes, ensuring high availability and resilience for critical operations.
Consume and integrate with third-party SaaS platforms while respecting rate limits and operational boundaries.
Collaborate with cross-functional teams to deliver solutions that support both internal engineers and external customers.
Write and maintain automated tests, including end-to-end tests, to increase confidence in changes and reduce regressions.
Improve observability and resilience of internal tools by addressing technical debt and strengthening test coverage.
Contribute to architectural decisions for high-availability services and help evolve best practices across the platform.
Operate and monitor production services, responding to incidents and improving runbooks for long-term stability.
Support initiatives to align tooling with regulatory requirements such as FedRAMP.
Continuously learn new technologies and frameworks to solve problems and make day-to-day work easier for engineers.
Requirements
You must have at least 3 years of professional software engineering experience in relevant roles.
You must be proficient with TypeScript and Node.js, including writing clean and maintainable server-side code.
You must have experience with at least one additional programming language such as Golang or Python.
You must have experience building and operating production web applications using React and a framework such as Next.js.
You must understand server-side rendering and API routes within modern web frameworks.
You must have hands-on experience with cloud infrastructure from providers such as AWS, Google Cloud, or Azure.
You must have experience building and consuming REST APIs, including integration with third-party SaaS platforms.
You must be comfortable working with both relational and NoSQL databases such as Postgres and DynamoDB.
You must have experience managing infrastructure as code using tools such as Terraform.
You must have experience with Statuspage or similar customer-facing status and incident communication tools.
You must have experience operating and supporting services in production environments with strong reliability focus.
You must have contributed to the design of high-availability services and understand operational trade-offs.
You must be able to work in Toronto, Ontario, Canada, or be authorized to work in the country where the role is based.
You must be eligible to work in Canada without sponsorship for this position at this time.
Nice to have
You have experience with Redis and designing caching layers, including TTL strategy and serving fallback data.
You have experience with no-code automation platforms such as Tines or Zapier, or you can quickly learn new platforms and become effective.
You have built Slack applications, bots, or slash commands and understand their operational value.
You have experience with incident management processes or tooling such as incident.io or Datadog Incident Management.
You have written automated tests, including end-to-end test coverage, to ensure quality and reliability.
You have experience with data warehouses such as Snowflake and understand their role in analytics.
You have worked with ETL tools such as Apache Airflow to move and transform operational data.
Practical notes
This role is based in Toronto, Ontario, Canada, and requires the ability to work in that location or be authorized to work in Canada.
The engagement type and specific hours are outlined in the source details and may vary based on team needs.
Visa sponsorship is not available for this role at this time.
Travel is not required for this position.
Application deadlines are determined by the source and should be followed if mentioned there.