Senior/Principal Local LLM & Generative AI Platform Engineer
ParallelwirelessKfar SabaFull Time4w ago
LLMAIMobileSecurityOperationsSupportEngineeringPlatformTestingAutomationremotecurated-jd
Job description
Senior/Principal Local LLM & Generative AI Platform Engineer at Parallelwireless.
About the role
Parallelwireless is seeking a technical leader to architect and manage a secure, internal platform for large language models. You will enable engineering and business teams to utilize generative AI on proprietary data, including source code and technical documentation, while ensuring all information remains within company-controlled environments.
Key facts
What you'll do
- Define the technical roadmap and architecture for a reliable, self-hosted LLM platform.
- Collaborate with IT, security, legal, and engineering departments to prioritize high-impact use cases.
- Develop a modular inference layer featuring API stability, model routing, and concurrency management.
- Select and validate open-weight models, documenting their security, licensing, and hardware requirements.
- Optimize model serving across CPU and GPU resources using quantization, caching, and batching techniques.
- Build and maintain RAG and enterprise search pipelines that respect existing data permissions and access controls.
- Implement automated evaluation frameworks to monitor retrieval quality, latency, and factual accuracy.
- Create CI/CD workflows and release gates for model updates, prompts, and infrastructure changes.
- Ensure production observability, including error tracking, throughput monitoring, and resource utilization.
- Design secure agent workflows with sandboxing, input validation, and human-in-the-loop approval for sensitive actions.
- Establish operational standards for backups, disaster recovery, and capacity planning.
- Protect data through encryption, network isolation, and defenses against prompt injection or supply-chain risks.
Requirements
- BSc or MSc in Computer Science, Electrical Engineering, Data Science, or equivalent practical experience.
- 7+ years of experience in infrastructure, data, or ML platform engineering, with recent work shipping LLM-based systems.
- Proficiency in Python and experience building maintainable APIs and data pipelines.
- Deep understanding of transformer models, including context management, tokenization, and inference failure modes.
- Experience building production-grade RAG systems, including embedding strategies and hybrid search.
- Ability to define and execute systematic evaluation datasets for model performance and safety.
- Experience with Linux, Docker, and Kubernetes for containerized service deployment.
- Practical knowledge of GPU-based serving, performance profiling, and reliability engineering.
- Familiarity with distributed systems, authentication, API security, and data lifecycle management.
- Experience with Git, CI/CD, and infrastructure as code.
- Strong communication skills to explain AI behavior and technical tradeoffs to diverse stakeholders.
Nice to have
- Experience with air-gapped or private-cloud LLM deployments.
- Proficiency with inference runtimes like vLLM, SGLang, TensorRT-LLM, llama.cpp, Ray Serve, or KServe.
- Experience optimizing performance on NVIDIA or AMD hardware using CUDA or ROCm.
- Background in vector databases, model registries, or LLM tracing platforms.
- Familiarity with fine-tuning methods like LoRA/QLoRA and synthetic data generation.
- Experience building developer tools or code-intelligence assistants for large C/C++ and Python codebases.
- Knowledge of AI governance, data loss prevention, and enterprise identity providers like Active Directory.
- Experience red-teaming AI systems against security vulnerabilities.
- Familiarity with telecommunications, 3GPP, or Open RAN standards.
- Contributions to open-source AI or infrastructure projects.
Skills & tools
- Python
- LLM Inference (vLLM, TensorRT-LLM, etc.)
- RAG & Vector Search
- Kubernetes & Docker
- GPU Optimization (CUDA/ROCm)
- CI/CD & Infrastructure as Code
- Distributed Systems & API Security