
Forward Deployed Engineer (Inference & Post-Training)
Job description
About the role
This role collaborates with Solutions Architects as a deep-domain specialist to optimize engine performance and deployment pipelines. Through CX, Engineering, and Sales, the role helps complex proofs-of-concept succeed and guides platform adoption.
Software engineers turn product ideas into working code. Engineers work in small teams, review each other's work, and ship in small batches. Most teams follow agile practices such as sprints and daily standups. Engineers also write tests, fix bugs, and improve performance. The field values clear communication as much as technical skill. Engineers spend part of every week on planning, code review, and debugging, not just writing new code. The ability to explain a technical decision in plain words separates strong engineers from the rest.
Key facts
What you'll do
System components are selected, configured, and tuned to align with hardware, model architecture, and workload profiles. This work targets optimized inference engine behavior that meets throughput and latency goals for demanding deployments.
Critical proofs-of-concept and benchmarks are supported through configuration and performance tuning. KV cache settings are refined, speculative decoding is applied, tensor parallelism is determined, and quantization strategies are chosen to hit targets.
Experimentation flows into production as system design and execution are shaped with customer needs in mind.
Strategic accounts receive direct technical contact and monitoring. Endpoint configurations are observed to help customers extract maximum platform value and hit critical milestones.
Accelerated time-to-value results when the right setups are established early.
Field insights surface into software and model roadmap decisions. Product development is influenced through customer feedback, and early adoption is driven with strategic logos.
Requirements
Experience in technical roles focused on inference systems, open-source LLM deployment, or post-training workflows must total 5+ years. Demonstrated work in these areas is required.
Inference engine expertise must be hands-on and at an expert level. Familiarity with vLLM, TensorRT-LLM, and SGLang is expected, along with the ability to diagnose and resolve engine-level performance issues.
Deep knowledge of KV cache tuning, speculative decoding, tensor and pipeline parallelism, and quantization techniques must be applied in real scenarios.
Hands-on experience with fine-tuning and post-training pipelines, including LoRA, SFT, DPO, RLHF, and GRPO, must inform system design advice.
Awareness of state-of-the-art open-source models must be broad. Strong judgment guides model selection for specific customer use cases, hardware profiles, and performance targets.
Python coding must be strong, and operation in production environments must be comfortable and confident.
Practical notes
Typical interview steps
Hiring for engineering roles usually starts with a recruiter screen, followed by one or two technical rounds. Candidates often solve a coding problem, discuss past projects, and answer system design questions. Some loops include a take-home task. Final rounds typically cover team fit and give candidates a chance to ask questions. Interviewers look for how you break down an unfamiliar problem, not just whether you reach the answer. Practicing a few problems aloud and reviewing your own past projects are the best preparation.
Good to know
The role centers on inference optimization and post-training workflows in a research-driven AI environment. Common tools in this domain include vLLM, TensorRT-LLM, and SGLang. Open-source model ecosystems and hardware-aware optimization are central to the work. The position involves close collaboration with sales, customer success, and engineering teams.
Career growth
Engineering careers usually progress from individual contributor to senior, staff, and principal levels. Some engineers move into management and lead teams of five to twenty people. Others stay on the technical track. Growth follows demonstrated impact, not tenure alone. A typical engineering ladder has clear levels with defined expectations for scope, quality, and mentorship. Moving up usually requires owning outcomes end to end rather than completing assigned tickets.