Applied AI Engineer
Job description
About the role
You will be part of a team whose job is to make every team at HackerRank - from Go-To-Market to Finance to Product - dramatically more effective, by building AI agents, automations, and platform capabilities that eliminate the work that should not be manual in the first place. Most engineering roles put you deep inside one product surface. This one puts you across the whole company. In a given month you might be building a data analytics agent that answers ad-hoc product questions in Slack, writing an MCP connector that gives AI agents access to internal tooling, or designing the eval harness that lets the team ship agent changes without regressions. You will embed with internal teams to map workflows and identify the real bottlenecks which can be impacted by AI, and you will own agents end-to-end after launch: monitor quality, triage failures, review outputs, and iterate based on real usage data.
Key facts
What you'll do
- Embed with internal teams - from Marketing to Product to Finance - to map workflows and identify the real bottlenecks which can be impacted by AI.
- Scope, prototype, and ship AI agents and automations that eliminate high-leverage manual work across the company.
- Own agents end-to-end after launch: monitor quality, triage failures, review outputs, and iterate based on real usage data.
- Design and build the shared AI Agents platform: orchestration, routing, reusable components, access controls, logging, rate limiting.
- Build evaluation harnesses that measure agent quality programmatically, catch regressions before they reach users, and give the team confidence to ship changes.
- Author and maintain MCP (Model Context Protocol) connectors that give AI agents access to HackerRank's internal tools, APIs, and data systems.
- Own the outputs you ship. Write and review code with the same rigor you would apply to any production system - secure, well-tested, and maintainable.
- Partner with data and product teams to define metrics that show agent impact on speed, accuracy, and cost across the business.
- Implement retrieval and synthesis patterns that let agents work with large, structured datasets without exceeding cost or latency targets.
- Translate ambiguous business problems into agentic workflows, run experiments, and document learnings for the broader organization.
- Instrument agents for observability, debugging, and safety so that operations teams can understand and trust automated behavior.
- Continuously optimize prompts, tools, and handoffs based on telemetry and user feedback to improve reliability and adoption.
- Collaborate with security and compliance to ensure agent behavior respects data governance and privacy constraints.
- Mentor junior engineers on best practices for building and operating AI systems in production.
Requirements
- 1-4 years of software engineering experience with a track record of shipping production systems, not just demos or prototypes.
- Strong Python skills and solid fundamentals across the stack - you can build and deploy a service, wire it to APIs and databases, and keep it running.
- Hands-on experience building with LLMs (Anthropic, OpenAI, or similar) in production, including prompt engineering, context management, tool use, and debugging failure modes at real scale.
- You have built or operated structured evaluation pipelines for AI systems with metrics and regression detection.
- You diagnose before you build. You discover what is actually hampering productivity before jumping to a solution, and you are as likely to recommend a simple script or process change as a full agent.
- Comfort able with ambiguity and able to drive loosely defined problems to working solutions.
- Fluency with AI tools and agents - not just as a user, but as someone who builds production systems on top of them and understands their failure modes deeply enough to debug and improve them.
- Strong ownership and communication skills, with the ability to translate technical tradeoffs to non-technical stakeholders.
Nice to have
- Experience with agentic frameworks (LangChain, CrewAI, Claude Agent SDK or similar).
- Hands-on experience authoring MCP servers or building tool connectors for AI agents.
- Experience with RAG pipelines grounded in structured data models - vector databases, hybrid search, re-ranking - and AI observability tooling (Langfuse, OpenTelemetry).
Practical notes
- Hybrid in Bangalore, India.
- Engagement details are to be confirmed from source information.