Data Engineer
Job description
About the role
This position involves developing and enhancing a sophisticated data platform for the automotive industry. You will be responsible for architecting and maintaining the data infrastructure that powers critical business insights. A core part of your ownership will be the creation of highly reliable data pipelines designed to handle millions of vehicle records with precision. You will drive improvements in data quality, ensuring that the information flowing through the system is accurate, consistent, and trustworthy. This role offers the chance to contribute directly to a modern, remote-first data ecosystem that values engineering excellence. You will work closely with cross-functional teams to translate complex business requirements into scalable technical solutions. The successful candidate will leverage cutting-edge tools to optimize how data is processed and utilized across the organization. This is a hands-on role where your technical decisions will have a direct impact on the company's analytical capabilities.
Key facts
What you'll do
- Architect and deploy resilient data pipelines that handle both high-volume batch processing and low-latency streaming ingestion.
- Engineer robust ETL and ELT workflows using Python and Spark, integrating them with advanced orchestration frameworks to automate data movement.
- Implement stringent data quality controls, validation logic, and monitoring dashboards to ensure the integrity and observability of the entire data platform.
- Design scalable data models that serve the dual purposes of supporting complex analytical queries and accelerating product development cycles.
- Develop high-throughput ingestion pipelines that securely connect to a diverse array of external APIs and internal databases.
- Optimize data processing jobs to achieve significant gains in speed and scalability, thereby increasing the reliability of core data services.
- Partner with product managers and software engineers to dissect ambiguous business problems and convert them into well-defined technical specifications.
- Incorporate the use of modern AI-assisted development tools into your daily workflow to write code, debug issues, and boost overall engineering efficiency.
- Establish comprehensive monitoring and alerting systems to proactively identify and resolve data pipeline failures before they impact downstream users.
- Document data architectures and processes meticulously to ensure that knowledge is shared and maintainable by the entire engineering team.
Requirements
- You must possess a minimum of 3 years of professional experience as a Data Engineer in a production environment.
- You must demonstrate strong proficiency in writing complex SQL queries and Python scripts for data manipulation and automation.
- You must have hands-on practical experience building and maintaining large-scale data processing jobs using Spark or PySpark.
- You must have a proven background in designing, building, and operating production-grade ETL or ELT pipelines.
- You must have a solid conceptual understanding of data modeling techniques, including dimensional modeling and entity relationship design.
- You must be experienced with workflow orchestration tools, specifically Apache Airflow or similar scheduling and monitoring systems.
- You must have familiarity with major cloud data platforms, holding experience with at least one of AWS, Azure, or GCP.
- You must exhibit a strong sense of ownership and the ability to work independently with minimal supervision.
- You must possess professional fluency in English, enabling clear communication in writing, speaking, and technical documentation.
Nice to have
- You have prior experience working with Databricks and its ecosystem.
- You have worked with Delta Lake or have a deep understanding of Lakehouse architecture patterns.
- You have used MongoDB or other NoSQL databases for handling unstructured or semi-structured data.
- You have hands-on experience with Change Data Capture (CDC) and streaming technologies such as Kafka or Debezium.
- You have implemented data quality frameworks like Great Expectations to automate compliance checks.
- You have used metadata management, data lineage visualization, or advanced observability tools.
- You have practical experience with AI-assisted development tools such as GitHub Copilot, Cursor, or Claude Code.
Practical notes
This is a remote position located within the Latin America region. The company is headquartered in Berlin and operates a long-term product strategy that provides ample opportunities for professional growth and skill development. The engagement is full-time, requiring dedication during standard business hours to ensure seamless collaboration with the global team. There is no specified deadline for applications, but early submission is encouraged to secure consideration. Travel is not required for this role, as the entire position is conducted remotely. Candidates must ensure they meet all the listed requirements before applying.