Machine Learning Engineering Intern
Job description
About the role
You will architect and implement core components of Cresta's Knowledge Assist (KA) stack, owning the end to end lifecycle of AI features from research prototyping to scalable production deployment. You will partner intimately with product managers and domain experts to translate ambiguous customer conversation challenges into concrete technical solutions and high impact experiments. This role requires you to design and iterate on retrieval augmented generation pipelines that power real time agent assistance while rigorously measuring quality and latency. You will own the evaluation framework for language models, establishing benchmarks and automated tests that ensure reliability and safety in production. You will contribute directly to the research agenda of the KA team, exploring novel techniques to improve reasoning, efficiency, and robustness of AI systems under real world constraints. You will collaborate closely with front end and back end engineers to integrate your models into Cresta's platform and ensure seamless operator experiences for agents. Finally, you will document your work and share insights across the team, helping to define best practices and the long term technical roadmap for AI driven customer experiences at Cresta.
Key facts
What you'll do
Design and develop scalable machine learning pipelines that power real time knowledge retrieval and generation for contact center workloads.
Implement and iterate on Retrieval Augmented Generation (RAG) systems to ensure accurate, context aware responses grounded in enterprise knowledge bases.
Build and maintain evaluation frameworks and data driven experiments to measure model performance, quality, and cost efficiency.
Collaborate with software engineers to integrate AI solutions into Cresta's platform, optimizing for latency, reliability, and security in production environments.
Conduct research on model reasoning, alignment, and safety to address practical challenges observed in live customer conversations.
Work with cross functional stakeholders to translate business requirements into technical specifications and prototype solutions rapidly.
Lead investigations into performance bottlenecks and infrastructure constraints, proposing architectural improvements for scalability.
Contribute to the design of enterprise search capabilities that extend beyond contact centers, leveraging the GenKA and KS technology stacks.
Develop tools and abstractions that enable other engineers and researchers to prototype, test, and deploy AI features efficiently.
Drive best practices for data handling, model versioning, and experiment tracking to maintain reproducibility and auditability.
Engage with the AI research community, bringing cutting edge academic techniques into production ready systems at scale.
Support the deployment of generative AI features such as GenKA and Knowledge Search, ensuring they meet operational standards and user needs.
Requirements
Currently enrolled in a Bachelor's or Master's degree program in Computer Science, Artificial Intelligence, Machine Learning, or a related technical field.
Demonstrate strong proficiency in Python and experience with machine learning frameworks commonly used in research and production.
Possess a solid understanding of information retrieval, natural language processing, and modern retrieval augmented generation techniques.
Show experience with software engineering best practices, including version control, testing, and collaborative development workflows.
Have the ability to design and implement systems that balance accuracy, latency, and cost in real world applications.
Exhibit strong problem solving skills and the capacity to translate ambiguous requirements into concrete technical approaches.
Be comfortable working in a fast paced, dynamic environment where priorities evolve based on business and research needs.
Commit to maintaining high standards for code quality, model evaluation, and documentation throughout the development lifecycle.
Nice to have
Experience with large language models, including training, fine tuning, or inference optimization.
Knowledge of distributed systems and cloud infrastructure, and familiarity with deployment pipelines for AI services.
Background in conversational AI, customer experience, or contact center technologies.
Contributions to open source machine learning projects or published research in relevant domains.
Practical notes
This internship is based in Toronto, Canada, with a hybrid work model that balances in office collaboration and remote focus time. The role is time limited, aligned with academic calendars, and suitable for students who can commit for the duration of the internship. Applicants must be authorized to work in Canada without sponsorship for the duration of the engagement. No relocation support or visa sponsorship is provided for this position. Deadlines for application review are tied to cohort start dates, and early submission is encouraged to ensure full consideration. Working hours are constrained to ensure alignment with team availability and operational requirements.