
Staff Software Engineer, Observability
Job description
About the role
You will architect and implement observability platforms that support billions of active time series and process petabytes of logs daily. You will manage infrastructure across nearly a hundred cloud regions, enabling all Databricks engineers and customers to monitor the reliability of our product. You will develop advanced workflows that accelerate incident diagnosis for Bricksters, allowing engineers to quickly derive insights from logs and metrics. You will leverage powerful capabilities of Databricks' own data intelligence platform to push the boundaries of troubleshooting practices in the industry. You will uplevel monitoring and reliability practices across Databricks engineering, developing opinionated tools that set common standards for managing structured logs, metrics, alerts, dashboards, and oncall rotations. You will mentor and uplevel engineers, fostering a culture of technical excellence within the team and broader observability community.
Key facts
What you'll do
Architect and scale observability infrastructure that ingests, processes, and queries petabytes of logs and metrics each day with high reliability and low latency.
Design and implement storage and indexing strategies for billions of active time series, ensuring efficient data retention, compression, and query performance across global regions.
Develop and operate alerting and notification systems that provide actionable insights while minimizing noise and alert fatigue for oncall engineers.
Build dashboards and visualization tools that surface service health, capacity trends, and anomaly signals to both technical and non-technical stakeholders.
Collaborate closely with SRE, platform, and product teams to define reliability goals, error budgets, and service level objectives.
Implement incident response tooling and runbooks that accelerate root cause analysis and improve mean time to resolution.
Partner with open source communities and internal teams to integrate observability standards and best practices across Databricks engineering.
Optimize cost and performance of observability pipelines by tuning data sampling, aggregation, and archival strategies without losing critical signals.
Lead design reviews and code reviews to ensure observability services are maintainable, observable, and secure.
Mentor engineers at all levels, sharing knowledge about monitoring patterns, debugging techniques, and reliability engineering.
Champion testing and validation practices for observability systems, including failure injection and synthetic monitoring.
Drive adoption of new observability features through documentation, training, and cross-team collaboration.
Contribute to the reliability and scalability of Databricks platform itself by instrumenting and troubleshooting critical services.
Explore and evaluate emerging observability technologies, translating findings into actionable platform improvements.
Own end-to-end delivery of observability initiatives from requirements gathering through production rollout and postmortem analysis.
Requirements
BS (or higher) in Computer Science, or a related field.
7+ years of production-level experience in one of: Go, Python, Java, Scala, Rust, C++, or similar languages.
Experience in software development, in large-scale distributed systems.
Experience driving large projects involving multiple teams.
Experience with cloud technologies, e.g. AWS, Azure, GCP, Docker, or Kubernetes.
Familiarity with observability infrastructure, monitoring patterns, and reliability practices.
Ability to work effectively in a fast-paced, ambiguous environment while maintaining a high standard of quality.
Strong written and verbal communication skills for collaborating with cross-functional teams and presenting technical designs.
Nice to have
Experience with open source observability projects such as Prometheus, Grafana, Loki, or Tempo.
Contributions to observability or reliability tooling at scale.
Deep understanding of time series databases, metrics pipelines, and log aggregation systems.
Experience with SRE methodologies, incident management, and postmortem practices.
Knowledge of Databricks products and data intelligence platforms.
Practical notes
This is a full-time position based in Mountain View, California.
Employment authorization is required for this role.
The compensation range provided reflects the expected base salary for this position in the designated work location and may be adjusted based on candidate qualifications and other factors as described in the pay transparency materials.