Software Engineer, Cloud Infrastructure
Job description
About the role
We are seeking skilled Cloud and ML Infrastructure Engineers to lead the buildout of our AWS foundation and our LLM platform. You will design, implement, and operate services that are scalable, reliable, and secure. The role requires deep ownership of both cloud architecture and machine learning infrastructure, ensuring that our systems support real-world aviation workloads. You will work across the full stack from low-level networking to high-level application frameworks. This position is critical in turning our AI platform into mission-critical capabilities for our airline partners. You will be expected to operate in a fast-paced environment with minimal supervision. Your work will directly impact the safety and efficiency of aviation operations.
Key facts
What you'll do
- Design, provision, and maintain AWS infrastructure using infrastructure as code tools such as AWS CDK or Terraform.
- Build and enhance CI/CD and testing strategies for applications, infrastructure code, and machine learning pipelines using GitHub Actions, CodeBuild, and CodePipeline.
- Operate secure networking architectures including VPCs, PrivateLink, and VPC endpoints while managing IAM, KMS, Secrets Manager, and audit logging.
- Stand up and operate model endpoints using AWS Bedrock and SageMaker, and evaluate appropriate use of ECS, EKS, Lambda, or Batch for inference workloads.
- Build and maintain application services that call large language models through robust APIs, implementing streaming, batching, and backoff strategies.
- Implement prompt engineering and tool execution flows using LangChain or similar frameworks, including agent tools and function calling patterns.
- Design chunking and embedding pipelines for documents, time series data, and multimedia content, orchestrating workflows with Step Functions or Airflow.
- Operate vector search solutions using OpenSearch Serverless, Aurora PostgreSQL with pgvector, or Pinecone, while tuning for recall, latency, and cost efficiency.
- Build and maintain knowledge bases and data synchronization pipelines from S3, Aurora, DynamoDB, and external data sources.
- Create offline and online evaluation harnesses for prompts, retrievers, and chains, tracking quality, latency, and regression risk metrics.
- Instrument model and application telemetry using CloudWatch and OpenTelemetry, building dashboards for token usage and cost governance.
- Add guardrails, rate limits, fallbacks, and provider routing mechanisms to ensure resilience and reliability of AI services.
- Build ingestion and processing pipelines for structured, unstructured, and multimedia data, ensuring data integrity, lineage, and cataloging with Glue and Lake Formation.
- Optimize bulk data movement and storage in S3, Glacier, and tiered storage solutions, using Athena for ad-hoc analysis and querying.
Requirements
- Bachelor's degree in Computer Science, Engineering, or a related technical field or equivalent practical experience.
- Demonstrated experience designing and operating cloud infrastructure, with a strong background in AWS services.
- Solid understanding of machine learning infrastructure, including model deployment, inference optimization, and MLOps practices.
- Experience with infrastructure as code tools such as AWS CDK or Terraform for provisioning and managing cloud resources.
- Proficiency in at least one high-level programming language, with strong skills in Python for data engineering and automation.
- Familiarity with containerization and orchestration technologies such as Docker and Kubernetes, including ECS and EKS.
- Understanding of security and compliance principles, including data privacy, access control, and auditability in cloud environments.
Nice to have
- Experience with IoT AWS services for edge device management and secure communication.
- Hands-on work with retrieval-augmented generation and LangChain or similar agent frameworks.
- Knowledge of vector databases and optimization techniques for semantic search and retrieval.
- Experience building and operating CI/CD pipelines for machine learning and data engineering workflows.
- Familiarity with observability tools and practices for ML and cloud systems.
Practical notes
This role is full-time and based in San Carlos with a hybrid work arrangement. Travel is not required as part of the standard work arrangement. Employment is contingent on eligibility to work in the United States. The position does not require visa sponsorship at this time.