Senior DevOps / Infrastructure Engineer
Job description
Senior DevOps / Infrastructure Engineer at Category Labs.
About the role
You will operate the infrastructure that powers a live Layer 1 blockchain, ensuring the global node fleet remains healthy and performant across mainnet and testnet environments. You will own the design and implementation of infrastructure-as-code, observability, and alerting systems that provide deep visibility while reducing manual intervention. In this role, you will build agentic tooling and deterministic guardrails that enable a small team to safely operate a large distributed fleet. You will collaborate closely with engineers who are pushing AI into the development and operations workflow, helping to shape how autonomous agents interact with production infrastructure. You will also stand up and operate the infrastructure required for internal model workloads as the organization increases its AI focus. This position is central to scaling secure, reliable, and automated operations in a fast-growing, open-source-first environment.
Key facts
What you'll do
- Operate the Monad node fleet, including health checks, synchronization management, upgrades, and recovery procedures for validators, full nodes, archive/historical nodes, and indexer services across mainnet and testnet.
- Own the full lifecycle of infrastructure-as-code, writing Ansible playbooks for fleet configuration, managing cloud resources with Terraform and Atlantis, and maintaining Kubernetes clusters using GitOps patterns with Flux.
- Design, build, and operate observability and alerting solutions using Prometheus, Grafana, and Loki, creating dashboards that provide early warning and reducing noise through careful threshold and alert design.
- Automate the release pipeline for node software, including staged canary rollouts, snapshot creation and restore procedures, and implementing safeguards that protect validator operations from risky automated changes.
- Design and implement agentic operations frameworks, developing AI agents, MCP servers, and runbooks-as-code that allow autonomous tools to investigate, diagnose, and execute routine operational tasks with built-in guardrails and human oversight.
- Codify institutional operational knowledge into reusable tools and automation that can be used by both human engineers and their agentic counterparts across the organization.
- Harden node and service configurations, manage secrets at scale, and continuously identify and eliminate sources of manual toil in operational workflows.
- Collaborate with engineering teams to support blockchain client operations, contributing to reliability and performance improvements where client behavior interacts with infrastructure.
- Partner with product and research teams to align infrastructure roadmaps with upcoming protocol changes, model training workloads, and new feature deployments.
- Mentor junior engineers and operators by documenting practices, improving runbooks, and promoting a culture of calm, methodical incident response.
Requirements
- You have 5 or more years of professional experience in DevOps, SRE, or infrastructure engineering, operating production systems at internet scale.
- You have strong fundamentals in Linux systems, systemd, networking, and security, and you are comfortable debugging live systems over SSH in high-pressure situations.
- You have deep, hands-on experience with infrastructure-as-code tools, specifically Ansible for configuration management and Terraform for provisioning cloud resources.
- You have extensive experience designing and operating observability stacks based on Prometheus for metrics, Grafana for visualization, and Loki or similar for log aggregation.
- You have hands-on fluency with AI-assisted engineering tools, including coding agents and LLM-based assistants, and you apply sound judgment in deciding where these tools provide value and where they introduce risk.
- You have a proven track record of designing automation that incorporates safe guardrails, and you conduct operations with a calm, methodical approach during incidents.
- You have practical programming and scripting experience, writing robust code in languages such as Python and Bash to automate complex tasks and integrate systems.
- You have direct experience with Kubernetes orchestration and GitOps workflows, having deployed and managed services using Flux or ArgoCD in production environments.
- You have prior experience building or operating AI agent tooling, including MCP servers, agent orchestration frameworks, or similar systems that enable automated operations.
- You have experience serving machine learning inference workloads, whether locally on-premises or through cloud-based endpoints, and understand the associated operational concerns.
- You have previously worked with blockchain clients, node software, or peer-to-peer network protocols, understanding the operational nuances of decentralized systems.
- You hold a Bachelor of Science degree in Computer Science, Engineering, or a closely related technical field.
Nice to have
Only items explicitly stated as preferred in the source appear here.
Practical notes
- Hours: Full-time.
LENGTH: 782 words
Output the page only.