Staff Software Engineer
Job description
About the role
Datadog is seeking a Staff Software Engineer to help shape the future of our Bring Your Own Cloud (BYOC) Logs offering by unifying observability pipelines with log management software that customers deploy and manage in their own infrastructure. This role will focus on building and scaling systems that process, route, and store high-volume observability data within customer-managed infrastructure. You will operate as a hands-on technical leader, driving architecture, cross-team delivery, and product direction across a complex and evolving space. This is a high-impact opportunity to influence product strategy, mentor engineers, and solve deeply technical challenges at scale. You will work closely with product, platform, and operations teams to ensure that customer deployments are reliable, observable, and secure while meeting stringent performance and compliance requirements in diverse cloud environments.
Key facts
What you'll do
Design, build, and operate customer-controlled log processing, routing, and storage systems that feel like a managed Datadog product when deployed in customer infrastructure.
Lead the architecture and implementation of high-throughput data pipelines capable of ingesting, transforming, and storing massive volumes of observability logs across heterogeneous environments.
Collaborate closely with product managers, SREs, and cross-functional engineers to define and deliver complex, multi-team initiatives that span backend, infrastructure, and customer-facing features.
Ensure deployed software is reliable, observable, and maintainable, with a focus on diagnostics, debugging, and rapid resolution of issues in production customer environments.
Drive improvements in system performance, scalability, and cost efficiency through careful trade-off analysis, capacity planning, and infrastructure optimization.
Partner with security and compliance teams to design and enforce secure-by-default configurations, access controls, and data protection mechanisms for customer deployments.
Mentor engineers across multiple teams, elevate technical standards, and influence long-term product direction for the BYOC Logs portfolio.
Contribute directly to critical code paths, including agents, collectors, and data plane components, to resolve challenging deployment and runtime issues.
Define and implement operational tooling that simplifies deployment, upgrades, configuration, and lifecycle management of log pipelines in customer-controlled clusters.
Champion best practices for reliability, testing, and observability within the team and across the organization to ensure consistent quality and user trust.
Work closely with customer success and technical support teams to translate real-world deployment challenges into actionable product improvements and technical requirements.
Evaluate and integrate emerging technologies and open source projects to enhance the capabilities and resilience of the log processing infrastructure.
Participate in on-call rotations to support production incidents and ensure rapid response to critical issues affecting deployed customer environments.
Drive documentation, runbooks, and knowledge-sharing initiatives that enable both internal teams and customers to operate software effectively and independently.
Continuously explore opportunities to reduce operational overhead and improve the end-to-end customer experience through automation and platform thinking.
Requirements
You have significant experience building software that is installed, deployed, and operated in customer environments rather than only as a fully managed SaaS service.
You possess strong expertise in distributed systems, including scalability, performance optimization, and high-throughput data processing across large-scale infrastructures.
You are proficient in systems-level programming (e.g., Go, Rust, or C++) and understand how software runs across diverse environments, including Linux-based systems and container runtimes.
You have deep experience with cloud platforms such as AWS, Azure, or GCP, and are comfortable troubleshooting infrastructure, networking, and security-related issues in production.
You have extensive hands-on experience with containerization and orchestration technologies such as Kubernetes in production environments, including cluster lifecycle management and networking.
You have a proven track record of leading large, cross-functional engineering efforts and influencing technical direction across multiple teams and stakeholders.
You demonstrate strong ownership of complex systems, including the ability to debug issues in distributed, multi-tenant, and air-gapped customer environments.
You have experience designing and operating systems that meet stringent reliability, security, and compliance requirements in regulated industries.
Nice to have
Experience with log management, observability pipelines, and telemetry data formats such as logs, metrics, and traces is preferred.
Familiarity with open source projects related to log collection, processing, and forwarding, such as Fluentd, Vector, or similar platforms.
Understanding of data retention, governance, and compliance frameworks relevant to log data in multi-tenant environments.
Experience with infrastructure-as-code tools and automation frameworks for deployment and configuration at scale.
Knowledge of monitoring and alerting practices for customer-managed deployments, including integration with observability platforms.
Practical notes
This is a full-time position based in New York, New York, USA.
Employment is at-will and may be subject to background checks.
Candidates must be eligible to work in the United States without sponsorship for this role.
Relocation support is not available for this position.
Hybrid work model is in effect, with a combination of in-office and remote work as defined by team and business needs.