Applied Scientist (LLM)
Job description
About the role
The owns the full lifecycle of bringing advanced LLM capabilities from experimental research into robust, production-grade reality. You will design and implement high-performance Generative AI features that operate seamlessly across both Cloud and Edge environments, balancing innovation with strict operational constraints. A core part of this role involves driving model compression and quantization strategies to enable private, real-time inference on edge devices without sacrificing critical capabilities. You will architect scalable Retrieval-Augmented Generation (RAG) systems and sophisticated multi-agent workflows that power next-generation cloud applications. Each day requires you to formulate rigorous hypotheses, generate high-fidelity synthetic datasets, and meticulously fine-tune LLMs while rigorously validating safety and alignment guarantees. Your work will culminate in clear, actionable technical reports that translate complex research findings into concrete engineering guidance. Ultimately, you own the bridge between state-of-the-art LLM research and reliable, performant products that serve real user needs.
Key facts
What you'll do
- Design and implement advanced methods in prompt orchestration, fine-tuning (SFT/RLHF/DPO), and autonomous agentic workflows to solve complex real-world problems.
- Curate high-quality training data from large-scale text and multi-modal sources, ensuring diversity, accuracy, and relevance for demanding production scenarios.
- Identify patterns in model hallucinations and visualize evaluation metrics for clear interpretation to guide targeted mitigation strategies.
- Tune hyperparameters and improve inference speed and accuracy through PEFT techniques such as LoRA and QLoRA, alongside advanced prompt engineering.
- Collaborate with Product and Data Engineering teams to seamlessly integrate LLM features into the broader ecosystem, maintaining high standards of reliability and usability.
- Track and report progress using industry-standard benchmarks like MMLU and HumanEval, as well as custom internal KPIs that reflect true business impact.
- Stay at the forefront of the field, including emerging areas such as State Space Models and new Transformer variants, and evaluate cutting-edge techniques for production readiness.
- Engage in continuous technical growth and actively mentor junior colleagues to elevate the team's overall expertise and capability.
- Implement robust evaluation frameworks that combine automated benchmarks with human judgment to ensure quality and alignment with product goals.
- Optimize data pipelines to handle massive datasets that exceed 100GB, working efficiently with text, image, and audio modalities in parallel.
- Build and maintain high-fidelity datasets and scalable data pipelines that support rigorous experimentation and reproducible results.
- Operate deployment-oriented workflows using production serving stacks such as Triton Inference Server, vLLM, TGI, or ONNX for latency-sensitive scenarios.
- Conduct thorough error analysis and surface-level debugging to rapidly resolve issues in fine-tuning, inference, or runtime behavior.
- Translate ambiguous product requirements into concrete technical tasks and deliver measurable improvements in latency, quality, and stability.
- Document experiments, configurations, and outcomes meticulously to enable knowledge sharing and long-term maintainability.
Requirements
- Bring 3+ years of commercial experience in Machine Learning, with a specific focus on the NLP or LLM domain, demonstrating consistent delivery of impactful models.
- Show strong knowledge of Python3, NumPy, pandas, and modern text-processing libraries, along with hands-on proficiency in PyTorch and Hugging Face ecosystems including Transformers, PEFT, and Accelerate.
- Demonstrate proficiency in PEFT and LoRA techniques as well as Reinforcement Learning methods applied to language models.
- Exhibit a deep understanding of attention mechanisms, tokenization strategies, context window management, and embedding spaces that influence model behavior.
- Have practical experience in at least one of the following: Retrieval-Augmented Generation (RAG), Fine-tuning, or Agentic frameworks, with evidence of deployed solutions.
- Proven ability to manage and analyze massive datasets exceeding 100GB across text, image, and audio formats while maintaining data quality and integrity.
- Hands-on experience crafting high-fidelity datasets and building robust, scalable data pipelines that feed into demanding ML workflows.
- Expertise in prompt engineering, agentic framework design, and LLM pipeline orchestration to coordinate complex reasoning tasks.
- Experience deploying LLMs to production environments using industry-standard serving solutions such as Triton Inference Server, vLLM, TGI, or ONNX-based stacks.
- Communicate effectively in written and spoken English, ensuring clarity and precision in both technical documentation and cross-functional collaboration.
Nice to have
- Practical experience with vector databases and similarity search platforms such as Pinecone, Weaviate, Milvus, or Chroma for enhancing retrieval-based applications.
- Advanced quantization approaches including GGUF, AWQ, and EXL2, along with pruning and knowledge distillation methods to optimize model footprint.
- Familiarity with agent orchestration frameworks such as LangChain, LlamaIndex, or AutoGen to build complex multi-step reasoning systems.
- Basic understanding of web and client-server architecture, including streaming API responses using Asyncio and aiohttp for responsive applications.
- Hands-on familiarity with evaluation tools such as RAGAS, DeepEval, or G-Eval to systematically measure retrieval and generation quality.
- Experience with container orchestration using Docker and Kubernetes, and cloud GPU management platforms like Run:ai or Lambda Labs for efficient resource utilization.
- Knowledge of C++, Triton, or CUDA to develop custom kernels that push the boundaries of performance and efficiency.
Practical notes
- Remote working mode is available within Ukraine only, aligning with location options in Kyiv, Lviv, and remote arrangements inside the country.
- Employment details specify a gig-contract engagement model as outlined in the source information.
- The position offers 21 paid vacation days per year, complemented by paid public holidays in accordance with Ukrainian legislation.
- Development opportunities include corporate courses, knowledge hubs, free English classes, and educational leaves to support continuous professional growth.
- Medical insurance is provided from day one, with coverage extended to sick leaves and medical leaves as part of the comprehensive benefits package.
- Free meals, fruits, and snacks are provided when working in the office to support focus and well-being during collaborative work.
- Applicants should be prepared to work across multiple time zones and collaborate closely with globally distributed teams while maintaining reliable communication.
- The role may require occasional on-site presence in Kyiv or Lviv for critical planning sessions, infrastructure reviews, or team gatherings as determined by project needs.
- Candidates must be legally authorized to work in Ukraine and comply with all local employment regulations and visa requirements where applicable.
- Deadlines for application review may apply, and early submission is encouraged to ensure full consideration in a competitive talent pool.