
Senior Software Engineer, Observability
Job description
About the Role
Together AI is constructing the AI Acceleration Cloud, an end-to-end platform designed for the complete generative AI lifecycle. This platform integrates rapid inference capabilities with resilient cloud infrastructure to support demanding workloads. The role owns the observability platform for AI workloads, ensuring GPU clusters remain visible, stable, and responsive for critical inference paths. You will solve hard problems with code, not theory, by implementing robust solutions that directly impact platform reliability. The position requires a hands-on approach to building systems that provide deep insights into complex distributed environments. You will translate ambiguous operational challenges into concrete technical implementations using modern monitoring stacks. Your work will directly influence the performance and stability of production AI infrastructure serving high-scale workloads.
Key facts
What you'll do
- Design end-to-end observability pipelines using Prometheus, Grafana, ClickHouse, ClickStack, and OpenTelemetry to capture comprehensive system metrics.
- Build telemetry data pipelines and log aggregation workflows capable of handling metrics, logs, and traces at scale across distributed AI workloads.
- Develop automated monitoring, alerting, and anomaly detection systems that provide early warnings for potential service disruptions.
- Define SLIs and SLOs, create runbooks, and enable predictive analytics for critical services to improve operational maturity.
- Build and deploy custom observability tools and infrastructure-as-code using Go, Python, Terraform, Ansible, and Helm to ensure repeatable and reliable deployments.
- Collaborate with engineering teams to enhance distributed tracing and application monitoring, providing deeper visibility into service interactions.
- Lead incident response efforts and conduct post-mortem analysis to drive continuous improvement and prevent future occurrences.
- Define and evangelize observability best practices, establishing operational clarity and standardization across all data paths and teams.
- Operate storage and observability systems that sustain generative AI platforms, ensuring seamless data access and management under heavy load conditions.
- Partner with infrastructure teams to scale pipelines for high-volume ingestion and real-time querying across distributed nodes without degradation.
- Champion reliability by monitoring AI infrastructure and GPU clusters, tracking custom metrics for model performance and training pipelines.
- Implement monitoring for specific AI workload patterns, ensuring observability solutions are tailored to the unique needs of generative AI systems.
Requirements
Expertise in observability platforms is required, with a deep understanding of metrics, logs, and traces in distributed systems. You must have deep experience with Prometheus, Grafana, ClickStack, and OpenTelemetry to build comprehensive monitoring solutions. Experience with cloud-native monitoring services on AWS, GCP, and Azure is mandatory for operating in multi-cloud environments.
You need strong programming skills in Go and Python to develop efficient and reliable observability tools and agents. Fluency in infrastructure-as-code tools such as Terraform, Ansible, and Helm is required for automating deployment and configuration management.
You should have experience designing, operating, and scaling large-scale distributed systems, including the ability to handle complex failure modes and edge cases. This includes building pipelines for high-volume data ingestion and real-time querying across distributed nodes with strict performance requirements.
You must understand containerization with Docker and orchestration with Kubernetes deeply, including custom resource definitions and advanced scheduling scenarios.
You are expected to be knowledgeable about microservices architecture, service mesh technologies, CI/CD pipelines, and GitOps workflows to ensure reliable and auditable deployments.
You have expertise in managing databases such as PostgreSQL, MongoDB, Redis, and time-series databases handling high-cardinality data for efficient storage and retrieval of observability data.
You must be able to work effectively in a fast-paced environment, balancing multiple priorities while maintaining a strong focus on system reliability and performance.
You should demonstrate strong written and verbal communication skills to collaborate effectively with cross-functional teams and articulate complex technical concepts to diverse stakeholders.
You are required to have a proactive approach to problem-solving, with the ability to diagnose issues quickly and implement long-term solutions that improve system resilience.
You must be comfortable working with open-source technologies and contributing back to the community where appropriate to advance the state of observability practice.
Nice to have
Experience monitoring AI or ML infrastructure and GPU clusters is preferred, providing domain-specific insights into the challenges of training and inference workloads. Background in high-frequency, low-latency systems monitoring, chaos engineering, and reliability testing is a strong advantage for building robust observability solutions.
Contributions to open-source observability projects are valued as they demonstrate practical experience with real-world monitoring challenges and community collaboration. Familiarity with security monitoring and compliance frameworks is beneficial for ensuring observability practices meet organizational and regulatory requirements.
Practical Notes
Please apply with a summary of your experience using the specified Skills & Tools: Prometheus, Grafana, ClickHouse, ClickStack, OpenTelemetry, Go, Python, Terraform, Ansible, Helm.