Staff Backend Engineer, AI Platform
Job description
Staff Backend Engineer, AI Platform at Nightfall Ai.
About the role
You will architect and own the core backend systems that power Nightfall's AI-native data loss prevention platform, ensuring reliability, security, and scalability at global scale. In this role, you will design low latency, real-time microservices that process and detect sensitive data across SaaS applications, GenAI tools, email, and endpoint environments. You will take ownership of mission-critical internal databases and services, optimizing them for performance, resilience, and maintainability. You will lead the development of high volume auto-scaling infrastructure that supports instant data processing and intelligent threat detection. You will play a key role in optimizing AI services to scale models for low latency, real-time inferencing across distributed environments. You will establish robust observability at scale, delivering detailed insights into customer usage, system performance, and security signals. You will contribute to clear, high quality documentation for internal and public services, enabling collaboration across engineering and product teams.
Key facts
What you'll do
- Architect and maintain highly available authentication and API services that serve a global customer base with strict security requirements.
- Build and operate mission-critical internal databases and services that store, retrieve, and protect sensitive metadata at scale.
- Design, optimize, and operate high volume, auto-scaling, reliable, real-time data infrastructure and services that support continuous data monitoring.
- Optimize AI services to scale models for low latency, real-time inferencing while balancing cost, throughput, and accuracy.
- Establish observability at scale, creating detailed insights into customer usage patterns, system performance, and anomalies.
- Write and maintain comprehensive documentation for internal and public services to support onboarding, operations, and long-term maintenance.
- Collaborate closely with data scientists and product teams to translate AI capabilities into reliable backend services and APIs.
- Lead the decomposition of complex business problems and guide cross-functional teams in implementing scalable solutions.
- Implement data processing pipelines that handle large scale and real-time complex data using technologies such as Kafka, Flink, Snowflake, and Databricks.
- Ensure that systems meet stringent reliability targets, supporting three nines availability and rapid response to incidents.
- Drive best practices in security, compliance, and data protection across backend components and integrations.
- Partner with platform and infrastructure teams to align services with operational standards and long-term architectural vision.
Requirements
- Expertise in one or more systems or high-level programming languages such as Go, Java, Python, or C++, with a demonstrated eagerness to learn additional languages and paradigms.
- Experience running scalable systems that sustain thousands of requests per second while maintaining three nines of reliability.
- Experience developing complex software systems that scale to substantial data volumes or millions of users, with proven production quality deployment, monitoring, and reliability practices.
- Experience with real-time ML inferencing or building Agentic AI systems at scale is viewed as a valuable asset.
- Experience with large-scale distributed storage and database systems, including relational options like Postgres and NoSQL solutions such as Cassandra.
- Ability to decompose intricate business challenges and lead technical efforts to solve them in a collaborative environment.
- Demonstrated experience with data processing, including the construction and maintenance of large scale and real-time complex data pipelines using tools such as Kafka, Flink, Snowflake, and Databricks.
- A minimum of 8+ years of professional experience with strong technical leadership in designing and operating scalable, distributed architectures.
Nice to have
- Prior exposure to AI platform infrastructure, model serving frameworks, or MLOps tooling in production environments.
- Hands on experience with data loss prevention, insider risk, or content classification technologies in previous roles.
- Familiarity with security and compliance frameworks relevant to highly regulated industries such as financial services.
- Contributions to open source projects or technical talks that demonstrate thought leadership in backend or AI infrastructure.
Practical notes
Nightfall AI takes pride in being an equal opportunity employer. We value a diverse and global talent pool and the collaboration that results from having a diverse and inclusive team. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status. Our hiring decisions are based exclusively on merit, qualifications, and business needs.
The role is based in Bengaluru and is offered as a full time engagement. Compensation will be determined based on interview performance, level of experience, specialization of skills, and market rate. During the offer discussion, your recruiter will review the finalized base salary, bonus for applicable roles, benefits and perks, and stock options as they will appear in the offer letter. Candidates must be eligible to work in India for this position. Travel requirements are not applicable for this role.