Software Engineer, Artificial Intelligence/LLM
Job description
About the role
You architect and ship LLM features, owning the full journey from retrieval design to production safety. You write services, build evals, and guard rails while watching cost and latency. You collaborate closely with ML infra and product teams in this safety-critical aviation domain.
Key facts
What you'll do
Architect user-facing LLM workflows using LangChain or lean primitives for simplicity.
Implement schema-bound JSON outputs with validation, retries, and graceful fallbacks.
Integrate function calling to connect internal tools, search, and data services.
Deliver API services and workers in Python or TypeScript with streaming and backoff logic.
Optimize latency and cost using caching, prompt templates, and intelligent context packing.
Connect with AWS Bedrock, OpenAI, Anthropic, or self-hosted endpoints as required.
Partner with infrastructure teammates on chunking, embeddings, and indexing for documents and multimedia.
Select and tune vector backends such as OpenSearch, pgvector, or Pinecone for performance.
Maintain fresh knowledge bases through data syncs from S3, Aurora, DynamoDB, and external sources.
Create offline evals and golden sets for prompts, retrievers, and tool behaviors.
Establish online metrics tracking task success, hallucination rate, retrieval precision, p95 latency, and cost.
Run A/B tests and prompt rollouts with guardrails and canary releases to limit risk.
Implement content checks, PII detection, access controls, and auditing for compliance.
Design human-in-the-loop paths for sensitive actions in aviation contexts.
Add tracing, logs, and dashboards for model calls, token usage, errors, and saturation.
Debug complex failures spanning retrieval, prompts, tools, and external providers.
Requirements
You have shipped LLM applications to users and iterated using real data feedback.
You are a strong builder who writes production code, tests, and documentation that is simple and observable.
You understand RAG and tool calling deeply, including embeddings, chunking tradeoffs, and vector search.
You have a quality mindset, designing evals, defining success metrics, and iterating on evidence.
You are cost and latency aware, tracking p95, meeting SLAs, and reducing spend without quality loss.
You communicate clearly, aligning partners across product, infrastructure, and security teams.
Nice to have
Experience with Bedrock, OpenSearch Serverless, pgvector, Pinecone, or Weaviate.
Prompt versioning, guardrails, and provider routing in production environments.
Multimodal work with time series or video data.
Familiarity with GPU inference, Triton, or TensorRT-LLM.
Exposure to aviation or other safety-critical domains.
DevOps basics for CI/CD, infrastructure as code, and secure secrets management.
Practical notes
This role requires U.S. citizenship, Green Card status, or lawful permanent residency due to export controls.
Hours are full-time during standard business hours with limited travel.
About the role
You architect and ship LLM features, owning the full journey from retrieval design to production safety. You write services, build evals, and guard rails while watching cost and latency. You collaborate closely with ML infra and product teams in this safety-critical aviation domain.
Key facts
What you'll do
Architect user-facing LLM workflows using LangChain or lean primitives for simplicity.
Implement schema-bound JSON outputs with validation, retries, and graceful fallbacks.
Integrate function calling to connect internal tools, search, and data services.
Deliver API services and workers in Python or TypeScript with streaming and backoff logic.
Optimize latency and cost using caching, prompt templates, and intelligent context packing.
Connect with AWS Bedrock, OpenAI, Anthropic, or self-hosted endpoints as required.
Partner with infrastructure teammates on chunking, embeddings, and indexing for documents and multimedia.
Select and tune vector backends such as OpenSearch, pgvector, or Pinecone for performance.
Maintain fresh knowledge bases through data syncs from S3, Aurora, DynamoDB, and external sources.
Create offline evals and golden sets for prompts, retrievers, and tool behaviors.
Establish online metrics tracking task success, hallucination rate, retrieval precision, p95 latency, and cost.
Run A/B tests and prompt rollouts with guardrails and canary releases to limit risk.
Implement content checks, PII detection, access controls, and auditing for compliance.
Design human-in-the-loop paths for sensitive actions in aviation contexts.
Add tracing, logs, and dashboards for model calls, token usage, errors, and saturation.
Debug complex failures spanning retrieval, prompts, tools, and external providers.
Requirements
You have shipped LLM applications to users and iterated using real data feedback.
You are a strong builder who writes production code, tests, and documentation that is simple and observable.
You understand RAG and tool calling deeply, including embeddings, chunking tradeoffs, and vector search.
You have a quality mindset, designing evals, defining success metrics, and iterating on evidence.
You are cost and latency aware, tracking p95, meeting SLAs, and reducing spend without quality loss.
You communicate clearly, aligning partners across product, infrastructure, and security teams.
Nice to have
Experience with Bedrock, OpenSearch Serverless, pgvector, Pinecone, or Weaviate.
Prompt versioning, guardrails, and provider routing in production environments.
Multimodal work with time series or video data.
Familiarity with GPU inference, Triton, or TensorRT-LLM.
Exposure to aviation or other safety-critical domains.
DevOps basics for CI/CD, infrastructure as code, and secure secrets management.
Practical notes
This role requires U.S. citizenship, Green Card status, or lawful permanent residency due to export controls.
Hours are full-time during standard business hours with limited travel.