Site Reliability Engineer
vyncaRemote (USA)Full Time3d ago
PythonGoAWSKubernetesTerraformCI/CDPostgreSQLMySQLLinuxETLSecuritySOC
Job description
Site Reliability Engineer at vynca
About the role
We are seeking a Site Reliability Engineer to help construct and maintain the foundational technology that supports Vynca's healthcare services. This position involves a blend of software engineering, cloud operations, and system management to ensure our platforms are dependable, scalable, and efficient. You will be instrumental in keeping our production systems healthy and contributing to the evolution of our system architecture.
Key facts
What you'll do
- Build and manage cloud infrastructure in AWS using Terraform.
- Operate, maintain, and scale applications running on Kubernetes.
- Automate application packaging and deployment processes with Helm.
- Develop and enhance distributed, event-driven systems, addressing aspects like event sourcing, partitioning, and failure recovery.
- Establish and track Service Level Indicators (SLIs) and Service Level Objectives (SLOs) to balance system reliability with development speed.
- Create automation for deployments, scaling, monitoring, and incident response to minimize manual work and boost system resilience.
- Enhance platform observability through metrics, logging, tracing, monitoring, and alerting.
- Lead incident response activities, conduct blameless postmortems, and implement improvements to system reliability.
- Collaborate with Product and Engineering teams on capacity planning, performance tuning, and designing resilient systems.
- Implement and maintain security measures to comply with HIPAA and SOC 2 requirements.
- Participate in an on-call rotation to provide support for production systems.
Requirements
- Three to five years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or similar infrastructure roles.
- A Bachelor's degree in Computer Science, Information Systems, Software Engineering, or equivalent practical experience.
- Significant experience managing production workloads in AWS.
- Proven ability to manage infrastructure as code with Terraform.
- Experience operating and supporting production Kubernetes environments.
- Hands-on experience deploying and managing applications using Helm.
- Experience with distributed systems and event-driven architectures.
- Experience establishing and managing observability practices, including monitoring, logging, tracing, and alerting.
- Solid understanding of Linux administration, networking, cloud architecture, and distributed systems.
- Experience designing and implementing CI/CD pipelines and deployment automation.
- Strong analytical and problem-solving skills for complex infrastructure issues.
- Effective written and verbal communication abilities for cross-functional collaboration.
- A high degree of ownership, accountability, and initiative in pursuing operational excellence.
- Willingness to participate in an on-call rotation.
Nice to have
- Proficiency in programming or scripting with Python or Go.
- Experience with observability tools such as Prometheus, Grafana, Datadog, or OpenTelemetry.
- Familiarity with GitOps tools like ArgoCD or Flux.
- Experience managing databases such as PostgreSQL, MySQL, Redshift, or ClickHouse.
- Experience with secrets management solutions like AWS Secrets Manager or HashiCorp Vault.
- Experience supporting healthcare technology platforms or regulated environments.
- Familiarity with data infrastructure technologies including Snowflake, Redshift, and ETL/ELT pipelines.
- Experience with database performance tuning.
Skills & tools
- AWS
- Terraform
- Kubernetes
- Helm
- Linux
- Networking
- CI/CD
- Python (preferred)
- Go (preferred)
- Prometheus (preferred)
- Grafana (preferred)
- Datadog (preferred)
- CloudWatch (preferred)
- SigNoz (preferred)
- OpenTelemetry (preferred)
- ArgoCD (preferred)
- Flux (preferred)
- PostgreSQL (preferred)
- MySQL (preferred)
- Redshift (preferred)
- ClickHouse (preferred)
- AWS Secrets Manager (preferred)
- HashiCorp Vault (preferred)
- Snowflake (preferred)
Practical notes
- Applicants must reside in Arizona, California, Colorado, Florida, Georgia, Illinois, Nevada, North Carolina, Oregon, Texas, Utah, or Washington.
- Employment eligibility verification via E-Verify is required.
- A background check will be conducted prior to employment.
- Influenza vaccination is required for patient, client, or customer-facing roles, with accommodations considered.