Senior Infrastructure Engineer
Job description
About the role
The role is SRE-minded, focusing on making reliability measurable, rollouts safe, and infrastructure repeatable.
Platform and DevOps engineers build the systems that run everything else. They manage infrastructure, CI/CD pipelines, observability, and reliability. The work is about automation, scaling, and removing friction for product teams. Systems thinking is the core skill. Platform teams are measured by developer velocity and system reliability. Most companies run on-call rotations, and understanding incident response is part of the role.
Key facts
What you'll do
Define and maintain service-level objectives, indicators, and error budgets alongside observability data so that regressions are caught before users.
Measure reliability continuously to ensure infrastructure is dependable for every team.
Monitor distributed systems to manage capacity and tune performance as load and requirements evolve.
Requirements
Strong experience operating production infrastructure at scale with deep Linux fundamentals is required.
Experience with infrastructure-as-code tools such as Terraform or Pulumi alongside configuration management is required.
Experience running containers and orchestration platforms such as Docker and Kubernetes in production is required.
Strong programming skills in Go and/or TypeScript are required for building automation and internal tooling.
Experience with observability stacks including Prometheus, Grafana, and OpenTelemetry or equivalent is required.
Experience operating and monitoring distributed systems, including capacity planning and performance tuning, is required.
Comfort operating in high-stakes production environments and responding to incidents is required.
A genuine interest in crypto and on-chain systems is required.
Nice to have
Experience operating blockchain node infrastructure such as validators, RPC, and archive nodes for an L1 or L2.
Experience with high-performance networking, low-latency systems, or load balancing at scale.
Multi-region and geo-distributed deployments with robust failover strategies.
Security and key management practices involving HSMs, secrets management, and hardening.
EVM tooling and the wider Web3 infrastructure ecosystem.
What success looks like
Engineers deploy safely and frequently with confidence across environments.
Platform reliability is measurable through well-defined SLOs and continuously improving service health.
Infrastructure is automated, repeatable, and increasingly self-service for teams.
Incidents become less frequent, easier to diagnose, and faster to resolve.
Product teams spend more time shipping features and less time managing infrastructure.
Practical notes
You will rotate between team embedding and platform ownership to maintain consistency.
Typical interview steps
Platform interviews usually include an infrastructure scenario, a scripting or coding exercise, and operational questions. Candidates may be asked to design a deployment pipeline or debug an outage. Incident experience and an automation mindset are tested. Interviewers often ask about a past outage and how you handled it. Structured post-incident thinking, not heroics, is what they look for.
Career growth
Platform careers grow from engineer to senior, staff, and platform lead roles. Some people move into SRE leadership or cloud architecture. Breadth across networking, storage, and reliability becomes more important at senior levels. Platform careers reward breadth and calm under pressure. Experience automating your own work is the strongest signal for senior roles.
Questions to ask
Useful questions for the interview: what a typical week looks like, how work is assigned, what tools the team uses, and how feedback works. Asking how the role has changed recently and what the team wishes it had known when joining is also reasonable. Questions about the manager's priorities are especially valued.
About the company
If you're an tech enabled, data astute problem solver who wants to enhance our marketing team's performance at the cutting edge of Web3 through your analytical capabilities, this is the role you've been waiting for. SOMNIA is the Agentic L1 - a hyper-performance Blockchain operating at an order of magnitude faster than existing networks.