Software Engineer II
Job description
About the role
Abnormal is seeking a Platform and Infrastructure engineer to join our PI team and support our rapid growth by scaling and enhancing our core systems. In this role, you will be responsible for building, maintaining, and evolving the observability, monitoring, and alerting infrastructure that enables our engineering teams to deliver high-performance, reliable services. You will work closely with cross-functional teams to ensure our systems are resilient, scalable, and efficient, contributing to the overall stability and operational excellence of our platform. This position offers an to influence the infrastructure that underpins our innovative products and to work with tools and technologies in a fast-paced environment.
Key facts
What you'll do
- Design, develop, and maintain the observability stack, including tools such as Prometheus, Chronosphere, Grafana, and PagerDuty pipelines, to ensure comprehensive monitoring and alerting capabilities.
- Build and improve internal developer tools and platforms aimed at streamlining deployment processes, reducing operational friction, and enhancing developer productivity across product teams.
- Define, implement, and manage Service Level Agreements (SLAs) and Service Level Objectives (SLOs) to guarantee system resilience, performance, and cost-efficiency across multiple cloud environments, including US, EU, and GovCloud.
- Take ownership of services from initial technical scoping and design to deployment, ongoing monitoring, and iterative improvements, ensuring high availability and reliability.
- Participate actively in on-call rotations, diagnosing and resolving production incidents promptly to minimize downtime and impact on customers.
- Automate manual runbooks and operational procedures to improve efficiency and reduce human error, while identifying performance bottlenecks and implementing solutions to optimize system throughput and reliability.
- Mentor junior engineers, providing guidance through code reviews, technical discussions, and documentation to promote best practices and engineering standards within the team.
- Collaborate with product managers, software engineers, and other stakeholders to translate business requirements into scalable, reliable platform capabilities.
- Contribute to the development and refinement of system architecture, ensuring it adheres to best practices for fault tolerance, scalability, and maintainability.
- Implement and maintain monitoring dashboards, alerting rules, and performance metrics to ensure visibility into system health and operational metrics.
- Work with cross-functional teams to troubleshoot complex issues, perform root cause analysis, and implement long-term fixes to prevent recurrence.
- Support the continuous improvement of deployment pipelines and infrastructure automation to enable faster, safer releases.
Requirements
- 4+ years of professional experience in backend engineering, with a focus on operating and maintaining production-grade distributed systems.
- Strong proficiency in Python, including automation scripts, platform services, and Airflow DAGs for workflow orchestration.
- Working knowledge of Golang for developing high-performance infrastructure components and services.
- Experience managing data at scale, including stream processing, batch processing, and building high-throughput APIs.
- Proven ability to own a service from initial design and development through deployment, monitoring, and iteration, demonstrating end-to-end ownership.
- Deep understanding of fault tolerance patterns such as circuit breakers, retries, backpressure, and graceful degradation.
- Experience creating and maintaining technical design documents that clearly articulate architectural trade-offs and decisions.
- Ability to work effectively in an asynchronous, distributed environment, managing multiple priorities and collaborating across time zones.
- Strong knowledge of observability principles, including instrumenting services, defining SLIs/SLOs, and setting up alerting and monitoring pipelines.
- Familiarity with cloud environments, particularly AWS, and experience with cloud services such as EC2, EKS, S3, RDS, Lambda, SQS, and SNS.
- Experience with container orchestration tools like Kubernetes, including Helm charts and debugging at the cluster level.
- Knowledge of Infrastructure-as-Code tools such as Terraform, Pulumi, or CloudFormation for managing infrastructure deployments.
- Experience with CI/CD pipelines and tools like GitHub Actions or Jenkins to automate build, test, and deployment processes.
- Familiarity with web frameworks and protocols such as Django, gRPC, and Protobuf.
- Ability to communicate complex technical concepts clearly and effectively, both in writing and verbally.
- Demonstrated experience working in a collaborative, team-oriented environment with a focus on quality and continuous improvement.
Nice to have
- Hands-on experience with Prometheus, including writing PromQL queries, creating recording and alerting rules, and managing metric cardinality.
- Proficiency with Grafana for building dashboards, managing data sources, and visualizing system metrics.
- Familiarity with commercial observability platforms such as Chronosphere, Datadog, New Relic, or Honeycomb, and experience integrating these tools into existing systems.
- Experience managing alerting and incident response pipelines using PagerDuty, OpsGenie, or similar tools.
- Deep AWS expertise, including managing EC2 instances, EKS clusters, S3 storage, RDS databases, Lambda functions, and messaging services like SQS and SNS.
- Knowledge of Kubernetes, including cluster debugging, Helm chart management, and troubleshooting at the cluster level.
- Experience with Infrastructure-as-Code tools such as Terraform, Pulumi, or CloudFormation for automating infrastructure deployment and management.
- Familiarity with CI/CD tools like GitHub Actions, Jenkins, or similar platforms to enable continuous integration and deployment workflows.
- Experience working with Django web framework, gRPC, or Protocol Buffers for service development.
- Background in leading small engineering teams or building internal developer platforms, CLI tools, or automation frameworks.
Skills & tools
- Python, Golang, Prometheus, Chronosphere, Grafana, PagerDuty, Airflow, Spark, AWS, Kubernetes, Terraform, gRPC, Protobuf.
Practical notes
- Abnormal employs AI-assisted tools to help recruiters prepare interview questions based on resumes; however, all hiring decisions are made by human recruiters and hiring managers.
- The company is committed to providing an inclusive environment and is an equal opportunity employer, welcoming applicants regardless of background or identity.
- This role requires on-site presence at our Bangalore office, as the position is based in that location.
- Candidates should be prepared for technical interviews that assess their experience with distributed systems, observability, and infrastructure automation.
- The role offers an opportunity to work on impactful projects within a fast-growing company, with a focus on building scalable, reliable infrastructure that supports our innovative products.