Data Infrastructure Engineer
Job description
About the role
You will architect and maintain the AWS cloud infrastructure that underpins the PayPay data ecosystem using Terraform as the primary tool. You will own the end-to-end reliability of the Databricks Lakehouse, from workspace provisioning to access governance and incident response. You will design platform services that enable data engineers across more than 75 million users to run their workloads without interruption. You will apply SRE practices to monitor, alert, and dashboard critical infrastructure components while enforcing security and compliance controls. You will collaborate closely with cross-functional teams to translate data platform requirements into robust infrastructure solutions. You will carry out projects independently, taking ownership of design decisions and long-term system health. You will contribute to a globally diverse engineering culture that values risk-taking, professionalism, and continuous learning.
Key facts
What you'll do
Design, build, and manage AWS cloud infrastructure using Terraform, implementing Infrastructure as Code and GitOps workflows to ensure reproducibility and scalability.
Operate and improve the Databricks Lakehouse platform, including workspace provisioning, configuration management, and day-to-day reliability to support high-availability data workloads.
Build and maintain platform services and tooling that abstract complexity and empower data engineers across PayPay Group companies to deliver new features quickly.
Strengthen system reliability through SRE practices, defining and refining monitoring, alerting, and dashboarding strategies to detect issues before they impact users.
Implement and enforce data governance and security controls, ensuring alignment with internal policies and external compliance requirements for data access and protection.
Collaborate with data engineering teams to gather infrastructure requirements, evaluate trade-offs, and deliver solutions that meet performance, cost, and operational goals.
Participate in on-call rotation and operational support, responding to platform incidents, coordinating remediation, and driving post-incident reviews to prevent recurrence.
Promote platform thinking by identifying repetitive operational tasks and automating them to reduce manual effort and increase team efficiency.
Champion best practices in networking, security, and cost optimization when designing cloud resources to support sustainable growth.
Contribute to the evolution of the data infrastructure roadmap by sharing insights, proposing improvements, and mentoring junior engineers.
Requirements
Candidates must bring a minimum of 3 years of professional experience in infrastructure or platform engineering roles with a proven track record of delivering reliable systems.
Hands-on experience with Terraform is required, including writing, reviewing, and maintaining Infrastructure as Code modules in production environments.
Working knowledge of AWS services is essential, with specific familiarity with IAM, VPC, S3, EC2, and multi-account strategies for scalable cloud architectures.
Familiarity with CI/CD pipelines and automation tools such as GitHub Actions or Jenkins is required to integrate infrastructure changes safely and efficiently.
Experience with monitoring and observability tools like Prometheus and Grafana is required to implement actionable insights and maintain platform health.
Scripting ability in Python and/or Shell is required to automate operational tasks, parse logs, and build custom tooling for infrastructure management.
Solid fundamentals in Linux/Unix systems, networking concepts such as DNS and HTTP, and the OSI model are required to troubleshoot complex infrastructure issues.
A strong understanding of security best practices and compliance principles is required to design controls that protect data and meet regulatory obligations.
Nice to have
Previous experience with Databricks is preferred, including workspace administration, cluster configuration, and job orchestration.
Familiarity with FinTech environments, high-availability systems, and payment-related data platforms is preferred.
Experience contributing to open source or building internal platform products is preferred.
Practical notes
This role operates on a hybrid work schedule with a mix of remote and in-office workdays.
The position requires participation in on-call rotation and occasional operational support outside regular business hours.
Employment is contingent upon successful completion of any required background checks and compliance verification.