Senior Systems Engineer
EpirusUSA1w ago
Job description
About the role
Epirus is seeking a Senior Systems Engineer to join their growing team in Torrance, California. This role involves designing, building, and maintaining the core infrastructure that supports the company's daily operations and product development efforts. The Senior Systems Engineer will work closely with cross-functional teams to ensure all systems remain reliable, scalable, and secure across production environments. This position carries significant responsibility for architectural decisions and technical leadership within the broader engineering organization.
Key facts
What you'll do
- Design and implement scalable system architectures that meet both current and future business needs across the organization
- Monitor system performance and proactively troubleshoot issues across distributed infrastructure environments to maintain high availability
- Collaborate closely with development teams to integrate systems, resolve technical blockers, and improve deployment workflows
- Lead the evaluation and adoption of new tools and technologies for continuous infrastructure improvement
- Document system configurations, operational processes, and standard procedures to support team knowledge sharing and onboarding
- Mentor junior engineers and contribute meaningfully to their professional growth through code reviews and guidance
- Participate in on-call rotations to ensure timely and effective response to production incidents and outages
- Develop automation scripts and workflows to reduce manual operational tasks and minimize the risk of human error
- Conduct capacity planning and resource allocation to support growing workloads and evolving business demands over time
- Partner with security teams to enforce compliance standards and hardening practices across all production systems
- Evaluate existing system dependencies and identify areas for improvement in reliability and efficiency
- Coordinate with vendor partners and internal stakeholders to plan infrastructure upgrades and migrations
- Review and improve existing runbooks and operational documentation to reduce mean time to resolution
- Drive standardization of infrastructure patterns and best practices across the engineering organization
Requirements
- Demonstrated experience in systems engineering or a closely related technical discipline is required for this position
- Strong understanding of operating systems, computer networking, and distributed systems principles is essential for success
- Ability to work independently and make sound technical judgments in complex, fast-moving situations
- Excellent written and verbal communication skills for effective cross-team collaboration, coordination, and knowledge transfer
- Experience managing and maintaining production infrastructure in a fast-paced, high-availability, and demanding environment
- Bachelor's degree in computer science, engineering, or a related technical field is required
- Familiarity with infrastructure-as-code practices and configuration management approaches is strongly preferred by the team
- Proven track record of diagnosing and resolving complex system issues in production environments
- Experience with version control systems and collaborative development workflows is important
Nice to have
- Experience with containerization and orchestration platforms in production settings is highly valued
- Knowledge of major cloud service providers and their native infrastructure offerings is beneficial
- Exposure to high-availability and disaster recovery planning methodologies is a meaningful plus
- Background in working within regulated or compliance-driven industry environments is particularly helpful
- Prior experience contributing to open-source projects or internal tooling initiatives is encouraged
- Familiarity with incident management frameworks and post-mortem analysis practices is advantageous
Skills & tools
- Linux and Unix system administration, configuration, and troubleshooting in production environments
- Networking protocols, traffic analysis, firewall configuration, and network security fundamentals
- Scripting languages such as Python or Bash for automation and operational tasks
- Version control systems for managing and reviewing infrastructure code changes
- Monitoring and observability platforms for tracking system health and performance metrics
- Configuration management and infrastructure provisioning tools for consistent deployments
- Incident response practices and post-mortem analysis for continuous improvement
- Troubleshooting and debugging techniques for complex distributed systems
Practical notes
- This role is based in Torrance, California and requires regular on-site presence at the company office
- Standard full-time work hours with occasional flexibility for after-hours incident response and production support duties
- The engineering team values written documentation, clear technical communication, and consistent knowledge sharing practices
- Regular collaboration with product and development teams is expected throughout each working week
- Employees have access to professional development resources, training budgets, and conference attendance opportunities
- The company fosters a collaborative culture where engineers are encouraged to share ideas and take ownership of their work
- The position offers the chance to work on meaningful infrastructure challenges that directly impact product delivery