Software Engineer, Applied AI
Job description
About the role
You will architect and deploy AI systems that turn real-world audio into structured, searchable intelligence. You will design and own the core voice-first interface that allows users to command Rilla through natural speech in the same way they would talk to a human collaborator. You will build the search engine that surfaces business-critical insights from conversations that were previously impossible to access or analyze. You will own the end-to-end audio intelligence pipeline, handling the messy, noisy, and unstructured nature of offline conversations across physical environments. You will work across the entire AI lifecycle, from data acquisition and model training to real-time inference and customer-facing chat interfaces. You will partner closely with customers in the field to translate their raw workflow challenges into scalable technical solutions. You will act like a founder, taking extreme ownership of problems and driving systems to production with high velocity and reliability. You will help define the foundational models and tooling that set the standard for voice and offline AI for years to come.
Key facts
What you'll do
Architect and deploy production-grade AI systems that enable voice and audio as first-class interfaces in the Rilla product stack.
Design and implement a voice-first interface that allows users to command Rilla directly through natural speech, translating intent into actions across the platform.
Build a search engine capable of uncovering business-critical insights from voice data that has never been searchable using traditional methods.
Own the end-to-end audio intelligence pipeline, processing messy, noisy, and unstructured conversations that occur in real-world offline settings.
Develop and integrate models using PyTorch, OpenAI APIs, Baseten, and LiteLLM to power real-time inference and high-quality transcription.
Collaborate with cross-functional teams to define requirements, prioritize features, and translate customer needs into technical specifications.
Implement data ingestion and storage strategies using AWS, PostgreSQL, Redis, and S3 to handle large-scale audio and metadata efficiently.
Partner directly with customers in the field, including visits and constant communication, to understand pain points and refine product solutions.
Write high-quality, maintainable code and contribute to the broader engineering culture of ownership, speed, and continuous learning.
Experiment with and evaluate new eval frameworks, agent tooling, and prompt engineering techniques to improve system performance and reliability.
Contribute to the design of real-time communication systems using LiveKit to support low-latency voice interactions.
Help establish best practices for AI model deployment, monitoring, and lifecycle management in production environments.
Drive innovation in audio processing and offline AI, exploring novel techniques that turn unstructured speech into structured, actionable insights.
Take responsibility for debugging complex issues across the stack, from edge inference to backend services and data pipelines.
Requirements
Experience building and deploying AI/LLM systems in production, with a proven track record of delivering reliable, scalable solutions.
Familiarity with modern AI tooling including eval frameworks, agent frameworks, and prompt engineering techniques to measure and improve system behavior.
Comfort working directly with customers to understand their needs, gather feedback, and iterate on solutions that solve real-world problems.
Strong proficiency in Python and TypeScript, with the ability to write clean, performant, and well-tested code.
Experience building and serving machine learning models, including integration with cloud platforms and inference optimization.
Deep understanding of cloud infrastructure, particularly AWS services, and experience managing stateful systems with PostgreSQL, Redis, and S3.
Ability to work effectively in an in-office environment in New York City, with a commitment to high-velocity collaboration and execution.
Willingness to work approximately 60 hours per week and embrace the intensity required to build a generational company from the ground up.
Practical notes
This role requires in-office work in New York City with a schedule of approximately 60 hours per week.
Travel is occasionally required for customer visits and field research.
No visa sponsorship is mentioned in SOURCE.
There is no application deadline specified in SOURCE.