Infrastructure Engineer
UMATRUSA5d ago
KubernetesAIEngineeringInfrastructureremotecurated-jd
Job description
Infrastructure Engineer at UMATR.
About the role
UMATR is hiring an Infrastructure Engineer to manage the deployment and reliability of on-premise AI systems for industrial clients. You will lead the architecture and scaling of platforms that run real-time data processing directly on customer hardware.
Key facts
What you'll do
- Manage the full deployment lifecycle from physical hardware setup to operational status.
- Configure and support Kubernetes clusters in single-node and multi-node on-premise settings.
- Implement GitOps workflows to manage infrastructure through version control.
- Provision and maintain cloud-based development and staging environments using Infrastructure as Code.
- Oversee CI/CD pipelines, including testing, build processes, and release management.
- Optimize AI inference performance on customer-owned GPU hardware.
- Adjust model serving infrastructure to meet latency and hardware requirements.
- Establish observability practices for infrastructure, applications, and AI workloads.
- Design secure systems for permissions, secret management, and industrial network connectivity.
- Develop technical documentation and automation to standardize deployment procedures.
Requirements
- Proven experience troubleshooting and operating production Kubernetes clusters, including networking and storage.
- Practical experience deploying Kubernetes on non-managed platforms like k3s, RKE2, or MicroK8s.
- Professional experience with Infrastructure as Code tools such as Terraform.
- Proficiency with Docker, including image optimization and multi-service management.
- Experience managing and improving CI/CD delivery pipelines.
- Experience running AI/ML workloads on GPUs, including model serving and performance tuning.
- Knowledge of GPU resource management, including quantization, batching, and VRAM usage.
- Advanced Linux and Bash skills with the ability to support Python applications.
- Ability to debug complex systems using first-principles thinking.
- Strong sense of ownership and the ability to work independently.
Nice to have
- Experience with GitOps tools like Argo CD or Flux.
- Background in edge computing or air-gapped infrastructure.
- Familiarity with industrial protocols, manufacturing systems, or PLCs.
- Knowledge of GPU scheduling, multi-model inference, or GPU sharing.
- Experience managing PostgreSQL or time-series databases.
- Familiarity with observability platforms like OpenTelemetry.
Skills & tools
- Kubernetes (k3s, RKE2, MicroK8s)
- Terraform
- Docker
- Linux/Bash
- Python
- GPU Infrastructure
- CI/CD
- GitOps
Practical notes
This role offers the chance to join an early-stage company and shape engineering strategy from the ground up. You will work directly with leadership on problems that combine physical hardware, AI, and distributed systems. Apply through the provided portal to be considered.