Software Engineer II
AbnormalHybrid - Bangalore4d ago
EngineeringInfrastructurePlatformremotecurated-jd
Job description
Software Engineer II at Abnormal.
About the role
Abnormal is seeking a Platform and Infrastructure engineer to join our PI team and scale the systems supporting our growth. You will focus on building and evolving the observability, monitoring, and alerting infrastructure that empowers our engineering organization to maintain high-performance, reliable services.
Key facts
What you'll do
- Design and maintain the observability stack, including Prometheus, Chronosphere, Grafana, and PagerDuty pipelines.
- Build internal developer tools and platforms to streamline deployments and reduce operational friction for product teams.
- Define and manage SLAs and SLOs to ensure system resilience and cost-efficiency across US, EU, and GovCloud environments.
- Own services end-to-end, from technical scoping and implementation to production rollout and ongoing monitoring.
- Participate in on-call rotations to diagnose and resolve production incidents.
- Automate manual runbooks and identify performance bottlenecks to improve overall system reliability.
- Mentor junior team members and contribute to engineering standards through code reviews and design documentation.
- Collaborate with product and engineering teams to translate requirements into scalable platform capabilities.
Requirements
- 4+ years of professional experience in backend engineering and operating production-grade distributed systems.
- Proficiency in Python for automation, platform services, and Airflow DAGs.
- Working knowledge of Golang for high-performance infrastructure components.
- Experience managing data at scale, including stream or batch processing and high-throughput APIs.
- Proven ability to own a service from initial design through deployment and iteration.
- Strong grasp of fault tolerance patterns such as circuit breakers, retries, and backpressure.
- Experience writing technical design documents that clearly articulate trade-offs and architectural decisions.
- Ability to work effectively in an async-first, distributed environment.
- Solid understanding of observability principles, including instrumenting services and defining SLIs/SLOs.
Nice to have
- Hands-on experience with Prometheus (PromQL, recording/alerting rules, metric cardinality).
- Proficiency with Grafana for building dashboards and managing data sources.
- Familiarity with commercial observability platforms like Chronosphere, Datadog, New Relic, or Honeycomb.
- Experience managing alerting pipelines via PagerDuty or OpsGenie.
- AWS experience (EC2, EKS, S3, RDS, Lambda, SQS/SNS).
- Knowledge of Kubernetes, including Helm charts and cluster-level debugging.
- Experience with Infrastructure-as-Code (Terraform, Pulumi, or CloudFormation).
- Familiarity with CI/CD tools like GitHub Actions or Jenkins.
- Experience with Django, gRPC, or Protobuf.
- Background in leading small teams or building internal developer platforms and CLIs.
Skills & tools
- Python, Golang, Prometheus, Chronosphere, Grafana, PagerDuty, Airflow, Spark, AWS, Kubernetes, Terraform, gRPC.
Practical notes
- Abnormal utilizes AI-assisted tools to help recruiters prepare for interviews by suggesting questions based on resume content; however, all hiring decisions are made by humans.
- The company is an equal opportunity employer.