Software Engineer
Job description
About the role
You will architect and own the end-to-end design of large-scale data pipelines that ingest raw event streams and transform them into reliable production-ready datasets within the DataOS platform. You will build and operate high-throughput Spark and Kafka processing workflows that handle more than 150 billion events daily while ensuring strict data quality and performance standards. You will collaborate closely with engineers, product owners, and stakeholders to translate business requirements into robust, scalable data solutions that power critical analytics and product capabilities. You will champion data platform reliability, monitoring, and optimization to meet demanding production SLAs. You will help define and evolve engineering standards and operational best practices as the DataOS team continues to scale. You will take ownership of schema management, ETL logic, and the interfaces that connect storage layers with consuming applications. You will contribute to a polyglot engineering environment while focusing heavily on Scala and Python implementations within the data platform.
Key facts
What you'll do
Design and build large-scale data pipelines from raw event ingestion to production-ready datasets that serve analytics and product needs.
Develop and maintain Spark and Scala-based data processing workflows for both batch and streaming use cases with an emphasis on correctness and efficiency.
Own data platform reliability, quality, and performance in production environments through monitoring, alerting, and iterative improvements.
Collaborate with engineers, product owners, and stakeholders to define and deliver robust data solutions that align with business objectives.
Help shape engineering standards and operational practices as the DataOS team grows and matures its capabilities.
Implement and evolve schema management strategies to ensure data consistency, governance, and usability across the platform.
Integrate and optimize data flows across S3, Hadoop, Hive, and cloud-native storage systems to support scalable analytics.
Work with AWS services including EMR, EMR Serverless, and Athena to build cost-effective and resilient data processing solutions.
Support the delivery and maintenance of data interfaces for BigQuery, BigLake, Aerospike, and AWS Glue to enable downstream consumption.
Contribute to the reliability and performance of Delta Lake and Iceberg table formats within the data lake architecture.
Implement secure and efficient data pipelines while applying cloud security fundamentals and AWS authentication/authorization best practices.
Partner with data consumers to understand requirements and translate them into durable, production-grade data products.
Drive operational excellence by implementing CI/CD pipelines and infrastructure-as-code tooling using platforms such as Terraform.
Continuously evaluate new tools and technologies in the Spark, Kafka, and broader Hadoop ecosystem to improve platform capabilities.
Requirements
Bring 3+ years of software engineering experience with strong hands-on Scala and/or Java development in production environments.
Demonstrate strong Java core fundamentals and an understanding of the underlying runtime and concurrency models.
Show proven experience with Spark and Kafka in production, including tuning, monitoring, and operational troubleshooting.
Have hands-on experience with the Hadoop ecosystem and distributed data processing platforms such as AWS, GCP, Azure, CDP, or MapR.
Possess practical AWS experience, especially with S3 and EMR/EMR Serverless, where Athena knowledge is a strong advantage.
Have experience building and operating cloud-based, high-scale distributed systems that handle massive data volumes.
Communicate effectively across engineering and business teams with strong analytical and interpersonal skills.
Use professional spoken and written English to collaborate seamlessly within a global team.
Meet the hard bar of owning complex data workflows that impact platform reliability and data quality at scale.
Be comfortable working in a fast-paced environment where data platform demands grow rapidly across the organization.
Adhere to strict standards for code quality, testing, and documentation in production data pipelines.
Engage in cross-functional discussions where data architecture decisions affect multiple products and teams.
Maintain compliance with security and access control policies when designing data access layers.
Demonstrate ownership mindset by proactively identifying bottlenecks and driving improvements in pipeline performance and stability.
Nice to have
Knowledge of Scala is a plus that can deepen your impact on data processing logic.
Experience with AWS authentication/authorization and cloud security fundamentals is valued.
Hands-on background with CI/CD pipelines and tooling helps streamline platform operations.
Familiarity with DevOps and infrastructure-as-code tools such as Terraform is beneficial for infrastructure management.
Experience in SaaS product companies provides context for multi-tenant and global data platform requirements.
Referral by an AppsFlyer employee may accelerate your application review.
Practical notes
The role is based in Kyiv.
As a global company operating from 25 offices across 19 countries, we reflect the human mosaic of the diverse and multicultural world in which we live. We ensure equal opportunities for all of our employees and promote the recruitment of diverse talents to our global teams without consideration of race, gender, culture, or sexual orientation. We value and encourage curiosity, diversity, and innovation from all our employees, customers, and partners.