Frontier Agents Engineer
Job description
Frontier Agents Engineer at Scale AI.
About the role
This role focuses on integrating advanced AI systems into complex enterprise environments. You will bridge the gap between frontier AI research and reliable production software, working directly with strategic customers. The position involves designing, deploying, and operating AI agents that automate workflows and reason over enterprise knowledge.
Key facts
What you'll do
- Design and implement production AI systems that integrate with customer cloud platforms, data warehouses, APIs, and proprietary software.
- Develop scalable agent architectures combining LLMs, retrieval, memory, tools, and structured data into dependable workflows.
- Build secure integrations for AI agents to interact with customer systems, adhering to enterprise security and compliance standards.
- Rapidly prototype new AI capabilities and transition successful prototypes into production-ready solutions.
- Construct the infrastructure necessary for frontier AI research to become stable enterprise software.
- Engineer agent runtimes, orchestration frameworks, context pipelines, and execution services for production AI systems.
- Design systems for high reliability, observability, low latency, scalability, and graceful degradation.
- Create human-in-the-loop workflows that blend AI automation with expert oversight.
- Implement deployment strategies for safe AI system evolution through continuous delivery and experimentation.
- Operationalize AI quality systems to ensure agents remain reliable as models, prompts, and data change.
- Deploy evaluation harnesses using benchmarks, experiments, and LLM-as-a-Judge to prevent quality regressions.
- Implement tracing, monitoring, guardrails, and safety mechanisms for confident AI operation.
- Collaborate with Applied AI engineers to productionize new evaluation methods, retrieval strategies, and reasoning architectures.
- Evaluate new models, agent frameworks, and tooling for safe adoption into production.
- Partner with enterprise customers to understand their technical infrastructure and workflows.
- Translate ambiguous customer needs into scalable production AI architectures.
- Work with customer engineering and product teams to deploy AI into critical workflows.
- Identify reusable engineering patterns for core capabilities across multiple enterprise deployments.
- Serve as the primary technical advisor for key enterprise accounts.
- Lead architecture discussions covering distributed systems, AI infrastructure, and enterprise integration.
- Document reusable architecture patterns, deployment strategies, and operational best practices.
- Collaborate with Scale's product, infrastructure, and Applied AI teams to enhance the platform.
Requirements
- 4+ years of software engineering experience with strong understanding of distributed systems, data structures, algorithms, and system design.
- Proficient in Python programming for building production software.
- Experience building or deploying AI applications using LLM APIs, agent frameworks, retrieval systems, or vector databases.
- Experience with cloud platforms (AWS, Azure, or GCP) and modern production infrastructure.
- Strong problem-solving skills to navigate ambiguous requirements and deliver production solutions.
- Excellent communication skills for direct interaction with enterprise engineering teams.
Nice to have
- Experience deploying production AI agents or autonomous systems.
- Experience designing distributed systems, APIs, or large-scale backend systems.
- Experience with cloud-native infrastructure, Docker, Kubernetes, Infrastructure as Code, and CI/CD.
- Experience integrating AI systems into enterprise software environments.
- Familiarity with modern agent architectures, retrieval systems, tool use, memory, and context engineering.
- Experience with evaluation frameworks, LLM observability, regression testing, tracing, and AI monitoring.
- Experience implementing guardrails, grounding, and safety mechanisms for production AI systems.
- Experience operationalizing new foundation models, agent frameworks, or AI infrastructure.
- Experience working directly with enterprise customers.
- Ability to translate complex technical requirements into scalable production systems.
- Strong written and verbal communication skills.
- Experience leading architecture reviews, technical workshops, or customer design sessions.
Skills & tools
- Python
- Distributed Systems
- Data Structures
- Algorithms
- System Design
- LLM APIs
- Agent Frameworks
- Retrieval Systems
- Vector Databases
- AWS
- Azure
- GCP
- Docker
- Kubernetes
- Infrastructure as Code
- CI/CD
Practical notes
Compensation packages include base salary, equity, and benefits. The salary range reflects the minimum and maximum target for new hire salaries and may include several career levels. It will be determined during the interview process based on location and other factors. Benefits include health, dental, vision, retirement, learning stipend, and PTO. A 90-day waiting period is required before reconsidering candidates for the same role.