Data / ML Automation Intern
Job description
About the role
Xendit is seeking an intern to develop automated systems and data workflows that support our operational teams. You will bridge the gap between data engineering, analytics, and machine learning to replace manual tasks with scalable, software-driven solutions. In this role, you will own the design and execution of data pipelines that ensure the reliability and accuracy of critical business information. You will be responsible for building and maintaining automation scripts that reduce manual effort and improve process efficiency across teams. The position requires a proactive approach to identifying bottlenecks in data workflows and implementing scalable fixes using modern tooling. You will collaborate closely with engineers and analysts to translate business requirements into technical automation strategies. This internship provides an opportunity to contribute directly to production-grade data and machine learning systems. You will gain hands-on experience in deploying solutions that monitor, validate, and optimize data processes in a real-world fintech environment.
Key facts
What you'll do
- Design and implement data pipelines for the extraction, transformation, and validation of information from multiple sources.
- Author and optimize SQL queries to perform data quality verification, reporting, and in-depth analysis of operational metrics.
- Develop internal tools and automation scripts using TypeScript or Python to streamline repetitive tasks and improve team productivity.
- Manage dataset lifecycles and storage on Google Cloud Platform and BigQuery for efficient and secure data handling.
- Configure and orchestrate workflows using schedulers like Airflow to ensure timely and reliable execution of data processes.
- Construct, evaluate, and monitor the performance of machine learning and LLM-based automation systems in production settings.
- Establish and maintain connections between third-party APIs and internal systems to enable seamless data flow and automation.
- Operate and refine operational dashboards while actively monitoring data quality, integrity, and pipeline health.
- Record technical decisions, data models, and architectural choices to ensure clarity and maintainability for future development.
- Investigate and debug complex workflow issues while actively participating in structured code review processes to uphold quality standards.
Requirements
- You are currently enrolled as a student or have recently graduated in Statistics, Computer Science, Data Science, Information Systems, Engineering, or a closely related discipline.
- You demonstrate competency in TypeScript or Python with the ability to write clean, maintainable, and efficient code.
- You possess fundamental SQL skills including filtering, aggregations, joins, subqueries, and the use of window functions for analytical tasks.
- You understand core software development basics such as APIs, Git, version control, and common data structures like arrays, hashes, and trees.
- You show a strong interest in analytics, automation, machine learning, or data engineering and are motivated to apply theoretical knowledge in practical scenarios.
- You are capable of navigating and learning unfamiliar systems while working with messy, incomplete, or poorly documented data.
- You communicate clearly in English to collaborate effectively with both technical and non-technical stakeholders across teams and time zones.
- You are able to work independently and take ownership of tasks with minimal supervision while maintaining attention to detail and accuracy.
- You are comfortable working in a fast-paced environment where priorities may shift and multiple responsibilities need to be managed simultaneously.
Nice to have
- Direct experience with GCP and BigQuery including hands-on implementation of datasets, queries, and access controls.
- Familiarity with ETL and ELT pipelines, including practical work with dbt, Databricks, Apache Kafka, or Airflow for data orchestration.
- Exposure to machine learning libraries such as TensorFlow, PyTorch, scikit-learn, or pandas for data manipulation and model development.
- Knowledge of agent workflows, prompt engineering, LLM APIs, and evaluation pipelines for building reliable AI-driven applications.
- Experience with containerization and orchestration tools such as Kubernetes, Docker, and infrastructure as code tools like Terraform.
- Familiarity with CI/CD practices using Git-based workflows and platforms such as ArgoCD for automated deployments.
- Experience with cloud data platforms including AWS and related services such as ArgoCD for scalable and resilient architectures.
- Ability to build and maintain dashboards using tools like Power BI, Tableau, Metabase, or Looker Studio to visualize key metrics.
- Demonstrated initiative through personal projects focused on machine learning, analytics, automation, or data pipelines that showcase technical depth and curiosity.
Practical notes
The internship is structured as a 6-month engagement based in Jakarta, Indonesia. Consistent on-site presence is expected throughout the duration of the internship to facilitate collaboration and learning. This role does not involve international travel as part of the standard work arrangement. Candidates must be authorized to work in Indonesia without sponsorship requirements for the duration of the internship. The position is not eligible for remote work arrangements outside of the Jakarta location. Applicants should be available to start at the earliest feasible date aligned with team requirements and project timelines. Clear adherence to company policies and timelines is essential for success in this role.