
Infrastructure Software Engineer, Enterprise GenAI
Job description
About the role
You will architect multi-cloud abstractions that enable the Scale Generative AI Platform to operate seamlessly across GCP, Azure, and AWS for customers in highly-regulated sectors. You will implement custom integrations that connect Scale AI's platform with diverse customer data environments, internal APIs, and data warehouses while working directly with platform, product teams, and enterprise clients. You will deliver experiments at high velocity with stringent quality standards to engage customers and support rapid iteration across the product lifecycle. You will own the full stack from conceptualization through production, ensuring infrastructure components are reliable, scalable, and observable. You will multi-task across competing priorities and learn new technologies quickly to adapt to evolving enterprise requirements. You will uphold software engineering best practices to build durable systems that meet strict compliance and security standards.
Key facts
What you'll do
- Architect multi-cloud systems and abstractions to allow the SGP platform to run on top of existing Cloud providers.
- Implement custom integrations between Scale AI's platform and customer data environments, including cloud platforms, data warehouses, and internal APIs.
- Collaborate with platform, product teams, and customers directly to develop and implement innovative infrastructure that scales to meet evolving needs.
- Deliver experiments at a high velocity and level of quality to engage our customers and validate hypotheses in production conditions.
- Work across the entire product lifecycle from conceptualization through production, ensuring alignment with product goals and operational requirements.
- Be able, and willing, to multi-task and learn new technologies quickly to address shifting priorities and emerging technical challenges.
- Design infrastructure components with observability, reliability, and performance in mind for enterprise workloads.
- Partner with cross-functional stakeholders to translate business requirements into technical specifications and implementation plans.
- Maintain and enhance existing codebases to support scalability, maintainability, and security across distributed environments.
- Participate in code reviews, technical design discussions, and documentation to uphold engineering standards and knowledge sharing.
- Implement automated testing and deployment pipelines to improve delivery speed and system resilience.
- Monitor system health and troubleshoot complex issues in distributed cloud environments to minimize downtime and impact.
- Support the refinement of infrastructure patterns and best practices to enable consistent developer experiences.
- Assist in evaluating new technologies and tools that can enhance the platform's capabilities and efficiency.
- Contribute to long-term roadmap planning for infrastructure capabilities in support of enterprise AI adoption.
Requirements
- 4+ years of full-time engineering experience, post-graduation, with a strong track record of delivering software systems in production.
- Experience scaling products at hyper growth startups where rapid iteration and operational resilience are essential.
- Experience tinkering with or productizing LLMs, vector databases, and other emerging AI technologies in real-world applications.
- Proficient in Python or Javascript/Typescript, and SQL with the ability to write complex queries and optimize data access patterns.
- Hands-on experience with Kubernetes for container orchestration, deployment, and management in production environments.
- Demonstrated experience with major cloud providers including AWS, Azure, and GCP across compute, storage, and networking services.
- Excellent communication skills with the ability to explain technical concepts clearly to both technical and non-technical audiences.
- Strong problem-solving skills and comfort working in ambiguous environments where requirements evolve quickly.
- Ability to work independently and as part of a distributed team across multiple time zones.
- Commitment to writing clean, maintainable, and well-documented code that adheres to industry best practices.
- Willingness to follow security and compliance guidelines relevant to enterprise AI and regulated data workloads.
- Proven experience with debugging complex distributed systems and resolving high-severity incidents.
- Understanding of infrastructure as code principles and experience with relevant tooling.
- Demonstrated ownership and accountability for system reliability, performance, and on-call responsibilities.
Nice to have
Only items indicated as preferred in the source are included; no additional preferences are added.
Practical notes
- Engagement details are to be found in the source material under "
"
- Compensation information is available in the source under pay transparency details for locations including San Francisco, New York, and Seattle.
- This role may be eligible for additional benefits such as a commuter stipend as noted in the source compensation section.
- Candidates must note the 90-day waiting period before reconsidering the same role as required by company policy.
- No specific hours, travel requirements, visa conditions, or deadlines are stated beyond those found in the source documentation.