Senior Observability Infrastructure Engineer
Job description
About the role
You will design and implement the next generation of logging, metrics, and tracing platforms to support global expansion and regulatory compliance. You will take full ownership of our hybrid infrastructure lifecycle, managing over 1,500 servers across bare-metal and Kubernetes environments in Amsterdam. You will automate operational toil by writing code in Go or Python to build self-healing systems and improve CI pipelines for safety and predictability. You will dive deep into performance optimization for distributed tracing and logging pipelines handling petabytes of high-volume data streams. You will participate in on-call rotations while focusing primarily on engineering solutions that prevent alerts before they fire. You will lead reliability initiatives by upgrading our stack, implementing automated guardrails, and designing safer API access patterns. You will partner with hundreds of product teams to ensure our observability platform enables real-time monitoring without compromising stability.
Key facts
What you'll do
- Build the next generation of our platform by designing and implementing the future architecture of our logging and metrics systems to support new global regions and ensure data isolation and regulatory compliance.
- Own infrastructure operations by managing the lifecycle of over 1,500 servers across both bare-metal and Kubernetes environments in a large-scale production setup.
- Automate to reduce toil by writing code in Go or Python to eliminate manual operational tasks and build self-healing systems that do not require intervention during the night.
- Optimize for scale and performance by tuning Elasticsearch clusters, optimizing Prometheus and VictoriaMetrics storage, and ensuring OpenTelemetry implementations handle peak traffic without data loss.
- Improve reliability and engineering practices by participating in on-call rotations focused on engineering solutions that stop alerts from firing through automation and proactive fixes.
- Upgrade our stack to the latest stable versions while ensuring the platform remains secure, performant, and available for hundreds of product teams.
- Enhance the self-service experience by implementing automated guardrails and quota management to prevent noisy tenants from destabilizing the shared observability platform.
- Design safer API access patterns and operational workflows to enable scalable and secure monitoring across distributed services.
Requirements
- Bring 10+ years of experience in the observability domain or in a relevant platform/infrastructure domain with a strong track record of delivery.
- Demonstrate observability stack expertise by operating core telemetry data stores at scale such as Elasticsearch, OpenSearch, VictoriaLogs, or ClickHouse for logging, Prometheus or VictoriaMetrics for metrics, and Grafana Tempo for distributed tracing.
- Show deep Linux experience with kernel-level understanding to debug complex networking, file system, and performance issues on both bare metal and virtualized hardware.
- Prove production Kubernetes experience by operating and troubleshooting production workloads on-premise and/or in the cloud, with advanced use of kubectl and Kubernetes primitives.
- Exhibit a software engineering mindset with proficiency in Go or Python, focusing on writing clean, maintainable, and scalable code rather than only configuring tools.
- Display strong ownership and problem-solving skills to turn recurring operational issues into automated, repeatable, and robust solutions.
- Communicate effectively with both technical and non-technical stakeholders while collaborating closely with product teams to align platform roadmaps with business needs.
- Adhere to strict reliability and security standards ensuring that platform changes do not compromise compliance, performance, or availability.
Nice to have
- Experience with large-scale observability platforms supporting thousands of tenants and petabytes of data.
- Deep knowledge of Elasticsearch performance tuning, shard allocation, and cluster resilience strategies.
- Familiarity with Prometheus remote storage options and VictoriaMetrics scaling patterns.
- Hands-on work with OpenTelemetry collectors, processors, and exporters in high-throughput environments.
- Background in building self-service platform engineering tools and automated quota management systems.
- Contributions to open source observability projects or published blogs demonstrating technical depth in the field.
Practical notes
This is a full-time position based in Amsterdam. No visa or travel details are provided in this listing.