Senior Devops Engineer
Job description
Senior Devops Engineer at Nightfall Ai.
About the role
In this role, you will be responsible for developing and maintaining the essential tools that empower our engineering team to function effectively. You will oversee AWS environments while balancing both engineering and operational responsibilities. This position calls for a proactive individual who is capable of guiding fellow team members and enhancing overall productivity. You will partner closely with security and product teams to embed data protection principles into the infrastructure that supports Nightfall's AI-driven DLP platform. The role requires a deep understanding of reliability and performance in the context of agentic workflows and sensitive data handling. You will contribute to the design of systems that uphold the integrity and confidentiality of customer data at every layer. Your work will directly influence the speed and safety with which Nightfall delivers cutting-edge data security solutions.
Key facts
What you'll do
- Architect and maintain the core infrastructure foundations that allow Nightfall's platform to process data securely and efficiently across multi-cloud environments.
- Define and enforce operational standards for AWS services, ensuring that networking, compute, and storage resources align with security and compliance objectives.
- Design and implement robust CI/CD pipelines that automate testing and deployment while preserving the ability to roll back changes safely and rapidly.
- Develop reusable modules and internal tools using high-level programming languages to reduce manual effort and accelerate feature delivery for data loss prevention workflows.
- Collaborate with engineering teams to containerize applications, standardize runtime environments, and optimize resource utilization for cost-effective scaling.
- Establish monitoring, alerting, and logging strategies that provide clear visibility into system health, data flows, and potential security anomalies in near real time.
- Lead incident response activities during outages or security events, coordinating with cross-functional teams to restore service and prevent recurrence.
- Mentor junior engineers on best practices for infrastructure as code, secure configuration, and automated operations within regulated environments.
- Evaluate and integrate emerging technologies such as service meshes to improve connectivity, resilience, and observability between microservices.
- Partner with database administrators to ensure that data stores like PostgreSQL and Cassandra are deployed, monitored, and tuned for high availability and durability.
- Implement and manage backup, retention, and disaster recovery procedures that meet enterprise standards for data protection and business continuity.
- Work with product and security stakeholders to translate compliance requirements such as SOC 2 and IAM governance into technical controls and automated checks.
- Optimize cloud spending by analyzing usage patterns, identifying idle resources, and recommending rightsizing actions without compromising performance.
- Support the rollout of new features by maintaining the reliability of underlying platforms and ensuring that deployments proceed smoothly in hybrid production environments.
Requirements
- Hold a Bachelor's or Master's degree in Computer Science or a related discipline from an accredited institution.
- Bring a minimum of 5 years of hands-on experience in infrastructure engineering covering provisioning, automation, and maintenance.
- Demonstrate at least 5 years of professional work with public cloud platforms, with a strong focus on AWS services and architectural patterns.
- Show at least 2 years of practical experience using Infrastructure as Code tools such as Terraform, Chef, Puppet, or Ansible in production settings.
- Provide evidence of at least 2 years of work with containerization technologies including Docker, Podman, or LXC for packaging and isolating applications.
- Present a track record of developing high availability and scalable SaaS or consumer-facing technologies that serve real-world users.
- Illustrate experience in deploying secure, reliable, and scalable services with automatic failover using containers and orchestration tools like Kubernetes.
- Exhibit proficiency in writing and maintaining infrastructure scripts using command-line shells, with a strong background in Bash or equivalent shell scripting.
- Display familiarity with virtualization technologies such as KVM, Hyper-V, or VMware and their integration with cloud environments.
- Show competence in designing, implementing, and managing complex AWS environments that span networking, compute, and storage components.
- Provide examples of in-depth Terraform usage, including module design, state management, and integration with remote backends.
- Demonstrate experience with Google Cloud Platform or Microsoft Azure, including core services and identity management concepts.
- Highlight practical exposure to AWS services such as EC2 Load Balancing, VPC, Route 53, Direct Connect, NAT Gateway, VPN, EC2 Networking, and Lambda.
- Show history of deploying container applications using Helm charts and managing the full lifecycle of Helm releases.
- Present experience managing database infrastructure, particularly with PostgreSQL deployments on RDS or Aurora, or Cassandra on AWS Keyspaces or Astra DB.
- Provide evidence of using monitoring and observability tools such as Datadog, ELK stack, OpenTelemetry, and on-call management systems like PagerDuty.
- Illustrate understanding of security best practices, including dependency vulnerability management, KMS secrets handling, and IAM/IRSA governance aligned with SOC 2 frameworks.
Nice to have
- A solid understanding of service meshes, particularly Linkerd, and their role in securing service-to-service communication.
- More than 5 years of experience in programming languages such as Go, Python, or similar high-level languages for automation and tooling.
- Over 5 years of hands-on work with Bash or other shell environments to build scripts that simplify operations.
- Familiarity with virtualization technologies like KVM, Hyper-V, or VMware in large scale deployments.
- Demonstrated success in architecting and running AWS environments with well-architected reviews and cost optimization strategies.
- In-depth knowledge of Terraform, including advanced use cases such as policy as code and state backends.
- More than 5 years of experience with Google Cloud Platform or Microsoft Azure, including multi-cloud strategies.
- Practical experience implementing AWS services related to networking, security, and serverless compute.
- Experience deploying container applications using Helm charts in varied environments and namespaces.
- Experience managing complex database infrastructure, including PostgreSQL on RDS or Aurora, and Cassandra on AWS Keyspaces or Astra DB.
- Hands-on familiarity with monitoring and observability stacks including Datadog, ELK, OpenTelemetry, and incident response tools like PagerDuty.
Practical notes
This position offers a hybrid working environment. Nightfall AI is committed to being an equal-opportunity employer, promoting a diverse and inclusive workforce. Hiring decisions are made based on merit, qualifications, and the specific needs of the business.