Senior Software Engineer, Machine Learning Infrastructure
handshakeUSAFull Time3w ago
TypeScriptPythonGoAWSGCPDockerKubernetesTerraformCI/CDRedisMachine LearningPyTorch
Job description
Senior Software Engineer, Machine Learning Infrastructure at handshake.
About the role
Join our ML Infrastructure and Platform team to build the foundational systems supporting both our career marketplace and Handshake AI. You will focus on high-scale engineering, creating the reliable architecture required for training, evaluating, and serving generative AI models in production.
Key facts
What you'll do
- Design and maintain production-grade ML infrastructure, covering data pipelines, feature stores, and model training/serving.
- Scale our LLM platform by managing provider integrations, orchestration, and observability while controlling for latency, reliability, and costs.
- Create evaluation frameworks, including benchmarks, quality measurement pipelines, and LLM eval harnesses.
- Manage post-training workflows such as fine-tuning and reinforcement learning pipelines.
- Optimize inference systems through GPU serving, autoscaling, and batching.
- Collaborate with Data Science and Product teams to deploy new models and define engineering standards.
- Enhance the developer experience and system scalability across the platform.
Requirements
- 5+ years of professional software engineering experience using Go, Python, TypeScript, or related languages.
- Proven track record of operating cloud infrastructure on platforms like GCP or AWS.
- Proficiency with Docker, Kubernetes, CI/CD, and Terraform for production services.
- Direct experience building ML infrastructure, specifically model serving, training pipelines, embeddings, or observability.
- Familiarity with data platforms including BigQuery, Airflow, Spark, Dataflow/Beam, or streaming architectures.
- Practical experience deploying production systems utilizing generative AI or LLMs, including performance tuning and API orchestration.
- Strong systems design background with the ability to manage ambiguity.
Nice to have
- Experience with Ray, KubeRay, Ray Serve, Anyscale, vLLM, Triton, or PyTorch.
- Background in building LLM evaluation frameworks or quality regression testing.
- Familiarity with Redis, Bigtable, or Vertex AI.
- Experience with post-training methods like RLHF, reward modeling, or fine-tuning.
- Experience developing agentic systems, memory systems, tool use, or voice AI.
Skills & tools
- Python, Go, TypeScript
- AWS, GCP
- Kubernetes, Docker, Terraform
- BigQuery, Airflow, Spark, Dataflow
- LLM orchestration, GPU inference, ML observability
Practical notes
Benefits include equity, 401(k) matching, medical/dental/vision insurance, mental health support, and a $500 wellness stipend. We provide a $2,000 learning stipend, flexible PTO, 15 holidays plus 2 flex days, paid parental leave, and fertility benefits. Our San Francisco office offers free lunch and gym access.