Senior Software Engineer, Machine Learning Services
Job description
About the role
Join the Machine Learning Services (MLS) team to construct the foundational platform for UiPath's large-scale AI and Document Understanding products. This role involves developing the core infrastructure that supports high-volume inference and automated model training. You will be building the essential systems that enable UiPath's AI capabilities. The position requires deep collaboration with product and data science teams to turn strategic AI goals into reliable, production-grade services. You will own the design decisions that impact scalability, performance, and maintainability across the entire stack. This is an opportunity to solve demanding distributed systems problems at the intersection of machine learning and enterprise automation. The work you do will directly influence how UiPath delivers AI features to millions of users worldwide.
Key facts
What you'll do
- Design, create, and operate the central MLS platform, including the Rust-based API gateway, Python ML compute workers, and the distributed job orchestration system.
- Address complex concurrency, performance, and distributed systems challenges to ensure platform stability for production workloads under heavy load.
- Collaborate with product and ML science teams to develop scalable infrastructure for various models, from large GenAI to specialized classifiers with differing resource profiles.
- Develop a custom content-addressable storage layer over cloud object stores (GCS, S3, Azure Blob), including garbage collection and sharding to optimize storage efficiency and access patterns.
- Improve the asynchronous job-queueing system, built on the storage layer using compare-and-swap for atomicity, to reduce latency and increase throughput.
- Work across the entire stack, from Kubernetes and container orchestration, through gRPC communication, to performance tuning of ONNX inference on GPUs for optimal utilization.
- Produce clear, efficient, and thoroughly tested code, emphasizing simplicity, correctness, and peer review to maintain high software quality standards.
- Implement observability and monitoring solutions to gain insight into platform behavior, detect issues early, and support data-driven improvements.
- Participate in on-call rotations to respond to production incidents, investigate root causes, and apply fixes to maintain service reliability.
- Contribute to architectural discussions, proposing and evaluating alternative designs to balance trade-offs between performance, cost, and complexity.
- Mentor junior engineers by providing guidance, conducting code reviews, and sharing best practices in software engineering and machine learning infrastructure.
- Engage with the broader UiPath engineering community by documenting decisions, sharing knowledge, and promoting consistent standards across teams.
- Evaluate emerging technologies and assess their applicability to the MLS platform, driving innovation in a responsible and measured manner.
- Ensure that all implementations adhere to security and compliance requirements relevant to enterprise AI and document processing.
Requirements
- Over 5 years of experience engineering and architecting large-scale, distributed commercial services in demanding production environments.
- Strong proficiency in a systems-level language such as Rust, C++, or Go, with a commitment to becoming a Rust expert and writing idiomatic, safe code.
- Significant Python skills are also required for scripting, data processing, and integration with ML frameworks used in the stack.
- Practical experience with cloud environments (Azure, AWS, or GCP) and containerization technologies (Docker, Kubernetes) to build and deploy resilient services.
- A solid understanding of concurrency, multithreading, and asynchronous programming concepts applied to high-throughput, low-latency systems.
- Practical knowledge of computer science fundamentals, applied to real-world problem-solving involving data structures, algorithms, and system design.
- The ability to articulate and advocate for good code and architectural principles, contributing to continuous improvement across the engineering organization.
- Experience working with distributed storage systems, object stores, and queueing mechanisms to build reliable and scalable backends.
- Familiarity with machine learning operations concepts, including model deployment, versioning, and monitoring in complex production environments.
- A proactive approach to debugging, performance profiling, and root cause analysis, using logs, metrics, and traces effectively.
- Strong written and verbal communication skills to collaborate effectively with cross-functional teams and stakeholders at all levels.
- The willingness to learn and adopt new tools, frameworks, and processes as the platform evolves alongside UiPath's strategic goals.
- Commitment to writing tests, performing code reviews, and following engineering best practices to ensure long-term maintainability.
Nice to have
- Prior production experience with Rust in projects that involve networking, concurrency, or high-performance computing.
- Experience with MLOps, particularly managing model lifecycles in multi-tenant, high-availability systems with strict SLAs.
- Familiarity with building ML inference services, model serialization (for example ONNX), and GPU programming (CUDA) to optimize inference throughput.
- Experience developing or working on custom storage or job-queueing systems that address challenges around consistency, scalability, and fault tolerance.
- Background in document understanding, automation, or robotic process automation domains to better align platform work with customer needs.
- Contributions to open source projects related to distributed systems, machine learning infrastructure, or cloud-native development.
Practical notes
Applications are reviewed on a rolling basis, without a fixed deadline. The application window may close if a suitable candidate is found or if a high volume of applications is received. UiPath supports diversity and provides equal opportunities, including reasonable accommodations for candidates upon request. Please ensure your application reflects your genuine experience and capabilities, as the selection process will focus on matching the outlined requirements and role responsibilities. If you are based in or near London and meet the outlined criteria, you are encouraged to submit your application for prompt consideration. The role is full-time and based in the London office, with expectations for in-person collaboration as needed for team alignment and planning sessions.