Software Engineer, Artificial Intelligence/LLM
Job description
Software Engineer, Artificial Intelligence/LLM at Beacon AI.
About the role
Beacon AI is a dynamic team of aviation professionals and engineers focused on developing an AI platform aimed at enhancing the safety and efficiency of flight operations. Supported by prominent investors, we have secured numerous contracts with the Department of Defense and have collaborated with major airlines to provide essential systems. Our approach emphasizes agility, allowing small teams to take ownership of their projects, deliver quickly, and innovate at the intersection of human and AI collaboration in aviation.
In this position, you will be responsible for developing features powered by large language models (LLMs). Your work will involve designing retrieval and tool-calling processes, creating the services that support them, and ensuring quality, cost, and latency standards are met in production. You will collaborate closely with machine learning and infrastructure teams on aspects like embeddings and model hosting, as well as with product teams to enhance user experience and outcomes. We prioritize reliability in our safety-critical environment and are looking for engineers at various experience levels.
Key facts
What you'll do
- Develop user-facing features utilizing LLM technology, focusing on retrieval-augmented generation and tool-calling processes using frameworks such as LangChain or similar alternatives.
- Produce structured JSON outputs with validation mechanisms, retries, and fallbacks.
- Implement function calling to connect with internal tools, search functionalities, routing, and data services.
- Manage the service layer by deploying APIs and workers using Python or TypeScript, ensuring clear contracts and effective handling of requests.
- Optimize latency and cost through caching, request shaping, and context packing.
- Integrate with platforms like AWS Bedrock, OpenAI, and Anthropic, or utilize self-hosted solutions as necessary.
- Work with infrastructure colleagues to enhance capabilities for document chunking, embeddings, and indexing across various data types.
- Select and optimize vector backends such as OpenSearch, pgvector, or Pinecone.
- Maintain updated knowledge bases by synchronizing data from sources like S3, Aurora, and DynamoDB.
- Develop offline evaluations and benchmark sets for prompts, retrievers, and tools.
- Establish online metrics to track task success rates, hallucination occurrences, retrieval accuracy, latency, and cost per request.
- Conduct A/B testing and manage prompt/version rollouts with appropriate safeguards.
- Ensure safety and compliance by implementing content checks, PII detection, and access controls.
- Design processes for human oversight in sensitive operations.
- Handle aviation-related data with strict adherence to security protocols.
- Monitor the systems you create by adding tracing, logs, and dashboards to track model performance and error rates.
- Troubleshoot complex issues across retrieval, prompts, tools, and service providers.
Requirements
- Proven experience shipping LLM applications, with a track record of enhancing features based on user data.
- Strong programming skills, capable of writing production-level code, tests, and documentation while keeping systems simple and observable.
- In-depth understanding of retrieval-augmented generation, embeddings, chunking, and function calling.
- Ability to design evaluations, define success metrics, and iterate based on data-driven insights.
- Awareness of cost and latency implications, with a focus on meeting service level agreements and optimizing expenses without compromising quality.
- Excellent communication skills to articulate trade-offs and align with product, infrastructure, and security teams.
Nice to have
- Familiarity with AWS Bedrock, OpenSearch Serverless, pgvector, Pinecone, or Weaviate.
- Experience with prompt versioning, guardrails, and provider routing in live environments.
- Background in multimodal projects involving time series or video data.
- Knowledge of GPU inference technologies such as Triton or TensorRT-LLM.
- Experience in aviation or other safety-critical fields.
- Basic understanding of DevOps practices related to CI/CD, Infrastructure as Code, and secure handling of secrets.
Example problems you might tackle in month one
- Convert an internal knowledge base into a low-latency retrieval-augmented generation service, ensuring clear schemas and evaluations are in place.
- Implement tool-calling to automate repetitive tasks within cockpit or operations workflows, incorporating necessary safeguards and audit trails.
- Optimize request costs through improved chunking strategies, caching mechanisms, and prompt adjustments, while maintaining high success rates.
Work Location
This position is hybrid, based in San Carlos, CA, requiring at least three days per week onsite, with the flexibility to work remotely on other days.
Perks & Benefits (Full-Time Employees)
- Healthcare: Full coverage of employee medical premiums; partial coverage for dependents.
- Time Off: Three weeks of paid time off plus over 13 paid company holidays.
- Stipends: Monthly allowances for phone and wellness expenses.
- 401(k): Available (currently without employer matching, but future enhancements are planned).
Due to U.S. export control regulations, we can only consider U.S. Persons (citizens, Green Card holders, lawful permanent residents, or individuals granted asylum or refugee status). We are unable to offer visa sponsorship or transfer support. All work must be conducted within the United States.
Beacon AI is an equal opportunity employer, committed to creating a diverse and inclusive workplace. We do not discriminate based on race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected characteristic. We enforce a zero-tolerance policy for harassment or discrimination in the workplace and comply with all applicable employment laws.