Senior SRE - Networks
Job description
About the role
The Senior SRE
Networks role at Fastly centers on owning the reliability, performance, and automation of the global edge network that powers the Fastly Edge Cloud Platform. You will be responsible for designing, operating, and iteratively improving the network infrastructure that spans six continents and over 100 points of presence. This position requires you to respond to significant traffic incidents, lead complex network outages, and resolve edge cases and failure scenarios using deep expertise in IP routing, particularly BGP. You will directly contribute to enhancing the robustness and performance of the internet for a diverse set of customers, including major online media streaming services, e-commerce entities, open source projects, and nonprofits. The role demands engineering best practices, operational rigor, and a strong commitment to automation and continuous improvement. You will act as a key technical leader, driving initiatives that strengthen the operational stability of the network and advance the company mission of building a more trustworthy internet.
Key facts
What you'll do
- Execute the operational and engineering responsibilities for the Fastly global network as a Senior SRE
- Networks, ensuring high availability and performance.
- Respond to significant traffic incidents and lead network incidents, resolving edge cases and failure scenarios using expertise in IP routing, particularly BGP.
- Write code that is performant, maintainable, clear, and concise, and actively contribute to code reviews to elevate the quality of the codebase and team processes.
- Partner in the development and iteration of tools and automation systems that enhance network operations, construction, and reliability.
- Innovate new methods for monitoring network performance with a focus on the end-user experience, and proactively identify and address potential issues before they impact customers.
- Conduct continual deep-dives of performance-based analytics to uncover trends, anomalies, and optimization opportunities across the network.
- Maintain close involvement with partner teams to ensure alignment and sustain a performant, scalable global network that meets business objectives.
- Advocate for the operational stability of the network by identifying opportunities and collaborating with engineering teams to shape roadmaps and software solutions.
- Mentor team members on the complexities of global routing, especially in an anycast-heavy environment, to elevate the overall technical capability of the organization.
- Leverage deep knowledge of internet protocols and practices, including TCP/IP, BGP, Anycast, and DNS, to inform decision-making and troubleshooting.
- Apply proficiency in automation and coding using languages such as Python, or similar, to build reliable and efficient operational tools.
- Demonstrate expertise in Tier 1 Internet service providers, Internet exchanges, and cloud providers to design resilient and high-performance network architectures.
- Analyze internet traffic patterns across multiple dimensions using flow-based tools to drive insights and improve network behavior.
- Utilize monitoring, alerting, and visibility tools such as Graphite, Grafana, Prometheus, or Splunk to maintain situational awareness and drive data-informed actions.
- Apply knowledge of cloud hosting solutions, including GCP, AWS, and Azure, within network design and operations to optimize infrastructure and services.
Requirements
- You possess experience in the protocols and practices that form the fabric of the global internet, including TCP/IP, BGP, Anycast, and DNS.
- You have experience in automation and coding using languages such as Python, or similar, to build operational tools and automation frameworks.
- You demonstrate proficiency in Tier 1 Internet service providers, Internet exchanges, and cloud providers, understanding their roles in global infrastructure.
- You can analyze internet traffic patterns across multiple dimensions using flow-based tools to diagnose issues and guide optimization efforts.
- You show proficiency with alerting, monitoring, and visibility tools such as Graphite, Grafana, Prometheus, or Splunk to maintain high levels of network observability.
- You have knowledge across cloud hosting solutions, including GCP, AWS, and Azure, and can apply this knowledge to network-related implementations.
- You follow DevOps practices and CI/CD pipelines using Git, Jenkins, and Ansible to ensure reliable and repeatable deployments.
- You understand Linux/Unix systems administration and network stack optimization to maintain high-performance and stable environments.
- You are adept at knowledge sharing and creating comprehensive documentation to empower teams and ensure continuity.
- You can collaborate with cross-functional teams to shape technical roadmaps, prioritizing initiatives that optimize automation tooling and the network.
- You strictly adhere to the requirement that only qualifications explicitly listed in the requirements are considered essential skills for this role.
- You apply a methodical and disciplined approach to operations, aligning with Fastly's commitment to operational rigor and reliability.
- You contribute to a culture of continuous learning and improvement, helping the team evolve alongside emerging network technologies and practices.
Nice to have
- We will be super impressed if you have experience writing production-level code using VCL or GO.
- Experience in network telemetry including sFlow and IPFIX analysis is also highly valued and would strengthen your contribution to network visibility and optimization.
Practical notes
This role is fully remote and based in the United Kingdom. The engagement is listed as Remote, and the compensation provided is $160,000 USD base for the 2024 timeframe. Please ensure you align with the expectations outlined in the requirements section, as only those qualifications explicitly listed are to be considered essential for this position. There are no additional hours, travel requirements, visa conditions, or application deadlines specified beyond the information provided in this description.