Platform Engineer (AI
Job description
About the role
The Platform Engineer (AI) will own the architecture and day-to-day operations of a dedicated on-premise artificial intelligence and data platform serving enterprise needs. This role focuses on building a resilient and efficient foundation that directly enables data science and machine learning initiatives. You will ensure the platform remains stable, performant, and capable of handling demanding computational workloads across the organization. The position requires close collaboration with data science teams to understand their requirements and adapt the infrastructure accordingly. You will act as the custodian of platform reliability, implementing rigorous capacity planning to meet aggressive project deadlines. Success in this role is measured by the platform's ability to deliver consistent performance and enable efficient, scalable experimentation.
Key facts
What you'll do
Architect, deploy, and manage the lifecycle of on-premise infrastructure specifically tailored for enterprise AI and data workflows.
Translate high-level business requirements into resilient technical environments that align with data flow and organizational security policies.
Implement and configure compute, storage, and networking resources while proactively analyzing cluster metrics to identify optimization opportunities.
Partner with analytics groups to synchronize environment changes with their testing and release schedules, ensuring minimal disruption.
Automate repetitive platform tasks to minimize manual intervention and enhance deployment consistency across all environments.
Monitor platform behavior through observability tools to identify potential issues before they escalate into critical incidents.
Document operational procedures and configurations clearly to preserve knowledge and facilitate smooth transitions or troubleshooting.
Collaborate with network specialists to resolve performance constraints affecting throughput and to optimize overall connectivity.
Validate configurations through structured testing before changes reach production to guarantee stability and security across staging, pre-production, and production.
Maintain accurate and accessible documentation that reflects the current state of the platform and supports ongoing operations.
Work closely with data science teams to understand their needs and adapt the platform to support evolving model training and inference services.
Ensure the platform remains highly available and efficient, enabling the organization to run demanding computational workloads without disruption.
Requirements
Candidates must possess hands-on expertise with container orchestration platforms and infrastructure provisioning methodologies as a core competency.
Practical experience with monitoring frameworks, log analysis, and visualization dashboards is mandatory for effective operational oversight.
A strong foundation in Linux administration, scripting, and computer networking principles is essential for successful daily execution of platform tasks.
The ability to juggle multiple projects simultaneously without compromising the quality of your work is a strict requirement of the role.
You must demonstrate reliability and a methodical approach to problem-solving in an enterprise on-premise environment.
Experience with infrastructure as code principles and version-controlled configuration management is expected.
Strong attention to detail is required to ensure configurations are secure, performant, and compliant with organizational standards.
The capacity to communicate technical concepts clearly to both technical and non-technical stakeholders is necessary for effective collaboration.
Nice to have
While not mandatory, experience with the specific technologies used in this ecosystem is highly valued for faster onboarding and effective contribution.
Proficiency in container management, workflow automation, messaging systems, and distributed processing frameworks is advantageous for handling complex workloads.
Competence in time-series databases and monitoring solutions will allow you to integrate seamlessly with existing observability practices and derive actionable insights.
Practical notes
This is a contract-based position with a defined duration.
The role is based in Vienna and offers the flexibility of remote work, with the expectation of occasional in-person meetings as necessary.
Start Date is Immediate.
Duration extends to the End of 2026, with potential for extension based on performance and project needs.
The work model is Remote preferred, with full-time allocation at 100%.