Member of Technical Staff (MTS) -Senior Core Full-Stack
Job description
Member of Technical Staff (MTS) -Senior Core Full-Stack at Epifi.
About the role
You will architect and own the core production systems that transform how engineering organizations become AI-native, owning the full lifecycle of high-concurrency web flows and autonomous AI agent platforms from design through global scale. You will translate ambiguous product outcomes into resilient technical strategies, defining the stack boundaries and service contracts that let teams move fast without breaking critical production workflows. You will partner with AI product and platform teams to instrument deep usage telemetry, turning raw tool interactions into board-ready insights that prove the ROI of AI coding tools like Cursor, Claude Code, and GitHub Copilot. You will lead performance investigations into backend bottlenecks, database deadlocks, and latency spikes under extreme concurrency, designing caching and validation guardrails that keep non-deterministic AI outputs deterministic and safe. You will build and maintain the foundational libraries, internal frameworks, and developer workflows that raise the bar for reliability, maintainability, and operational excellence across engineering teams. You will mentor engineers on production-grade practices around concurrency, memory management, and asynchronous processing, elevating the whole organization's ability to ship complex features safely. You will collaborate closely with infrastructure and SRE partners to ensure deployments are secure, cost-efficient, and observable from development through production. You will be the technical authority who decides when to build bespoke solutions versus integrate specialized tools, always balancing innovation with the stability required for enterprise-grade AI products.
Key facts
What you'll do
Architect and deliver core services, high-concurrency web architectures, and transaction flows that power AI-native engineering platforms at scale.
Design and implement production-grade AI agent loops, validation guardrails, and caching strategies that move agents out of toy projects and into reliable, board-impacting workflows.
Diagnose and resolve complex backend bottlenecks, database deadlocks, and latency spikes under heavy load using deep performance engineering and profiling techniques.
Build and optimize data pipelines and telemetry systems that capture how AI coding tools are used, enabling precise measurement of developer effectiveness and tool ROI.
Define and enforce service contracts and API standards that allow multiple engineering teams to build independently while maintaining system-wide integrity.
Develop and maintain internal frameworks and libraries that abstract distributed complexity, making it easy for teams to build reliable, high-throughput features.
Implement robust queueing and caching layers using Redis and equivalent technologies to decouple workloads, smooth traffic spikes, and increase system resilience.
Partner with infrastructure and SRE teams to codify deployment, observability, and security standards that protect production while enabling rapid experimentation.
Lead post-incident reviews and reliability initiatives, turning operational failures into long-term architectural improvements and process changes.
Champion cost-aware engineering practices, optimizing resource usage and infrastructure spend without compromising performance or developer experience.
Mentor engineers on concurrency patterns, memory management, and asynchronous processing to raise the technical bar across the organization.
Act as a technical product partner, translating ambiguous market and customer problems into clear system designs that balance speed, safety, and scalability.
Requirements
Bring 6 to 8 years of intensive full-stack or backend engineering experience, with a strong track record of shipping complex, high-concurrency systems.
Have startup DNA, thriving in high-velocity environments where execution intensity is high and the destination is clear but the map is still being drawn.
Demonstrate deep understanding of software that scales, including memory management, concurrency primitives, asynchronous task processing, and how web protocols break under load.
Be proficient in Go or high-performance backend languages, with proven experience building services that handle demanding throughput and low-latency requirements.
Show mastery of relational and non-relational databases, designing schemas and queries that remain performant and maintainable as data volume and query complexity grow.
Have hands-on experience with Redis for cache and queuing, using it to build resilient, responsive systems that survive traffic bursts and backend failures.
Bring experience with modern, highly optimized web engineering frameworks and a commitment to clean architecture, testability, and maintainable codebases.
Possess strong fundamentals in distributed systems, including an understanding of consistency patterns, failure modes, and tradeoffs in large-scale architectures.
Nice to have
Experience building and operating AI tooling or platforms that instrument agent behavior, prompt usage, and workflow performance.
Background in AI coding tools such as Cursor, Claude Code, or GitHub Copilot, and familiarity with the developer experience problems they solve.
Practical notes
This is a full-time role based in Bangalore.
The work may require occasional travel within India for stakeholder engagements or on-site collaboration.
Candidates must be eligible to work in India without visa sponsorship for this position.
The role is expected to be demanding, with high execution standards and frequent context switching between product, platform, and reliability initiatives.
Success in this role requires comfort with ambiguity, a bias for action, and the ability to make technical decisions that balance long-term architecture with short-term business needs.