Research Scientist / ML Engineer - Performance Engineering
The Biological Computing Co. (TBC)USA3w ago
MLEngineeringremotecurated-jd
Job description
Research Scientist / ML Engineer - Performance Engineering at The Biological Computing Co. (TBC).
About the role
The Biological Computing Co. is developing AI systems that combine generative models with biological computing and large-scale infrastructure. We need an optimization expert to refine our world-model and neural-optimizer projects by increasing the speed, throughput, and reliability of our production systems.
Key facts
What you'll do
- Refine inference performance for LLMs, diffusion models, video generation, and world-model systems.
- Manage serving efficiency using quantization, distillation, speculative decoding, KV caching, and memory optimization.
- Develop high-throughput inference pipelines for GPU clusters.
- Analyze model performance metrics including latency, memory usage, cost, and GPU utilization.
- Write custom kernels and low-level optimizations using Triton, CUDA, or PyTorch.
- Enhance distributed training and fine-tuning efficiency through parallelism, checkpointing, and data loading strategies.
- Collaborate with research teams to remove bottlenecks in model architecture and deployment workflows.
- Document performance gains as technical benchmarks for customer-facing applications.
- Balance the trade-offs between model quality, deployment costs, and inference speed.
Requirements
- Proven background in ML systems, high-performance infrastructure, or model optimization.
- Practical experience improving LLMs, diffusion models, or video generation systems.
- Proficiency in PyTorch and working at the intersection of model code and runtime environments.
- Ability to manage the balance between system performance and model output quality.
- Experience navigating ambiguous, early-stage project environments.
Nice to have
- PhD or MS in Computer Science, Machine Learning, Systems, Robotics, or a related field.
- History of optimizing large-scale generative models in research or production.
- Familiarity with stacks such as Triton, CUDA, vLLM, TensorRT, DeepSpeed, FSDP, or Ray.
- Experience specifically with world models.
Skills & tools
- PyTorch
- Triton
- CUDA
- Inference optimization (KV caching, quantization, pruning, distillation)
- Distributed training
- GPU profiling and debugging
Practical notes
- This role requires the ability to transition research prototypes into stable systems capable of supporting partner use cases and product demos.