
Staff Software Engineer
Job description
About the role
Snowflake is actively building the foundational systems that enable the agentic enterprise to operate at global scale, and this role sits at the heart of that mission. You will own the design and evolution of the External Observability Platform, translating abstract reliability requirements into concrete, hyper-scalable backend distributed systems. In this capacity, you will define how thousands of enterprise customers perceive the reliability and performance of the Data Cloud through deep telemetry and actionable insights. You will partner closely with application engineers, security teams, and customer support to ensure telemetry is comprehensive, consistent, and trustworthy. Your work will directly influence the strategic direction of Snowflake's operational infrastructure and set long-term standards for observability. You are expected to move with an experimental mindset, rapidly validating new architectural patterns to achieve simpler and more powerful outcomes. Ultimately, your role is not only to build critical infrastructure but to help redefine how work gets done across the Data Cloud.
Key facts
What you'll do
- Architect and scale the External Observability Platform to handle trillions of events daily across multi-cloud deployments on AWS, Azure, and GCP.
- Define and enforce observability standards for metrics, logs, traces, and events to ensure schema consistency, zero data loss, and optimized storage access.
- Build high-throughput, low-latency backend services capable of ingesting, processing, and serving petabytes of telemetry data while meeting strict SLA guarantees.
- Implement infrastructure-as-code solutions using tools like Terraform or Pulumi to create self-healing, automated telemetry pipelines and dynamic monitoring topologies.
- Drive platform adoption by collaborating with Application Engineering, Security, and Customer Support teams to make systems measurable and translate telemetry into customer-facing insights.
- Lead technical design reviews and roadmaps, mentoring senior engineers and fostering engineering excellence across the broader organization.
- Serve as an expert incident commander and troubleshooter for complex system failures, conducting deep-dive post-mortems and automating fixes to prevent recurrence.
- Evaluate and integrate emerging capabilities to continuously improve the scalability, resilience, and usability of the observability platform.
- Partner with cross-functional stakeholders to prioritize features and technical debt in alignment with customer needs and business objectives.
- Contribute to the broader Snowflake engineering community by sharing insights, best practices, and architectural patterns that elevate the entire platform.
Requirements
- Bring a minimum of 10 years of professional experience building infrastructure and backend distributed systems at scale using languages such as Go, C++, Java, or Rust.
- Demonstrate a proven track record of architecting, deploying, and maintaining hyper-scale distributed platforms in public cloud environments including AWS, Azure, or GCP.
- Possess deep theoretical and practical computer science fundamentals covering data structures, algorithms, concurrency patterns, storage engines, distributed consensus, and networking protocols.
- Show strong experience with infrastructure-as-code tools such as Terraform or Pulumi to automate and manage cloud resources reliably.
- Exhibit hands-on expertise with time-series databases, high-cardinality metric stores, distributed tracing frameworks like OpenTelemetry or Jaeger, and log streaming systems such as Kafka, Flink, ClickHouse, or Prometheus/Thanos.
- Communicate with superior clarity and empathy, enabling alignment across cross-functional teams and facilitating technical decision-making with diplomacy.
- Collaborate effectively in a fast-paced, dynamic environment, balancing multiple priorities while maintaining rigorous standards for reliability and performance.
- Engage actively in the Snowflake culture of innovation, learning, and experimentation, contributing to the continuous improvement of platform capabilities.
Nice to have
- Massive Scale Experience: Demonstrated background in high-performance computing or managing global installations that process petabytes of telemetry per day.
- Customer-Facing Observability: Experience designing and building external or customer-exposed observability tools, APIs, and analytics dashboards with strict latency and access-control requirements.
- Network & Systems Deep Dive: Deep operational understanding of high-performance network systems and storage infrastructure, including load balancing, TCP/IP, TLS, and kernel-level performance tuning.
- Observability Ecosystem Leadership: Prior involvement in open-source observability projects or contributions to communities around OpenTelemetry, Jaeger, Prometheus, or similar frameworks.
- Data Platform Familiarity: Exposure to large-scale data platforms and query engines, including Snowflake, Kafka, ClickHouse, or similar systems that bridge streaming and analytics.
- Security and Compliance Experience: Knowledge of security controls, audit frameworks, and compliance regimes relevant to enterprise telemetry and data governance.
Practical notes
Location for this role is Bellevue, Washington (US-WA-Bellevue), with a hybrid work arrangement requiring 3 days per week in the office. This position is offered as full-time employment. Only candidates who meet the minimum qualifications listed above will be considered for further review.