Engineering Manager - AI Engineering
MeeshoIndiaFull Time Employee3d ago
LLMAIEngineeringInfrastructurePlatformReliabilityremotecurated-jd
Job description
Engineering Manager - AI Engineering at Meesho.
About the role
Meesho is expanding its AI infrastructure team to support high-scale e-commerce operations. You will manage a technical group focused on building and optimizing the platforms that power real-time deep learning inferences for millions of users.
Key facts
What you'll do
- Direct the technical strategy and execution for a team of AI engineers.
- Design and scale infrastructure for multi-region model inference, distributed training, and feature engineering.
- Improve inference performance through GPU kernel tuning, quantization, and memory optimization.
- Refine open-weight models using techniques like speculative decoding, distillation, and serving-engine adjustments.
- Develop automated, agent-driven workflows to increase productivity in model training and rollout.
- Collaborate with Data Science and Product teams to deploy AI solutions that impact the consumer experience.
- Manage team operations including hiring, performance reviews, and goal setting.
Requirements
- Bachelor or Master degree in Computer Science or a related discipline.
- 9+ years of professional software engineering experience.
- 2+ years of experience in a management role.
- Proficiency with LLM inference tools such as TensorRT-LLM, vLLM, and SGLang.
- Experience with low-latency model serving in production environments.
- Knowledge of distributed training frameworks like PyTorch FSDP, DeepSpeed, Megatron, or Ray.
- Experience managing GPU fleets via Kubernetes, including scheduling and multi-region deployment.
- Familiarity with big-data and streaming technologies like Spark or Flink.
- Strong coding skills in Python and systems-level languages such as C++, Go, or Rust.
Nice to have
- Public contributions to open-source ML infrastructure or inference engines.
- Experience managing GPU costs and efficiency on cloud platforms.
- Background in building platforms for large-scale consumer applications.
- Expertise in ML observability, reliability, and incident management.
Skills & tools
- Python, C++, Go, Rust
- TensorRT-LLM, vLLM, SGLang
- PyTorch FSDP, DeepSpeed, Megatron, Ray
- Kubernetes, GKE
- Spark, Flink
- CUDA, GPU programming