Data Engineer III
Job description
About the role
You will architect and own the core data ingestion and transformation pipelines that power Emburse's AI-driven expense intelligence platform. You will implement and maintain critical tenant data isolation mechanisms that ensure secure and compliant access across our global customer base. You will collaborate closely with data scientists to deliver the feature stores and datasets that drive advanced spend forecasting and anomaly detection models. You will integrate new AI data APIs and services directly into our SaaS product workflows to enhance real-time decision making. You will design and optimize analytical data structures within our Snowflake and Databricks environments to support high-performance business intelligence. You will translate ambiguous product questions into robust data operations that deliver actionable insights at scale. You will mentor junior engineers on best practices for data quality, pipeline reliability, and secure coding standards. You will continuously evaluate and introduce modern data technologies to keep our platform at the forefront of the industry.
Key facts
What you'll do
Design and develop scalable data ingestion pipelines using Python, Java, or .NET to extract, transform, and load data from diverse internal and external sources.
Architect and maintain the central data warehouse and data lake architectures on Snowflake and Databricks to ensure high availability and optimal query performance.
Implement and refine tenant data security models and isolation strategies to meet strict compliance requirements for multinational organizations.
Build and support data science platforms and feature stores that enable sophisticated AI-driven analytics and machine learning workflows.
Integrate AI data APIs and third-party services seamlessly into Emburse's SaaS products to enrich user experiences and automate decisions.
Create and maintain analytical data models and semantic layers that empower business users to generate ad-hoc insights without engineering support.
Develop automated scripts to monitor data quality, detect anomalies, and alert stakeholders on pipeline health and operational issues.
Optimize complex data workflows for scalability, reliability, and maintainability within our AWS cloud infrastructure.
Collaborate with product managers and software engineers to translate business requirements into robust data transformation and visualization logic.
Utilize advanced SQL techniques and modern ETL/ELT tools to build efficient data flows that support real-time and batch processing needs.
Construct data visualization datasets and support Looker implementations to deliver intuitive dashboards for enterprise stakeholders.
Perform in-depth code reviews and provide constructive feedback to peers to uphold high standards of code quality and documentation.
Champion the adoption of secure software development lifecycle practices with a strong focus on OWASP principles and data privacy.
Lead the investigation and resolution of medium-complexity production issues to minimize downtime and ensure data integrity.
Document all processes, pipelines, and architectural decisions to create a clear knowledge base for the analytics team.
Requirements
Bachelor's degree in Computer Science or a related field, or equivalent years of professional experience.
Possess advanced working knowledge of SQL and demonstrated experience with relational or columnar database systems.
Have hands-on experience with modern scalable data lakes and enterprise-grade data warehouse platforms.
Show proficiency with at least one major data acquisition, pipeline, or analytics codebase and a deep understanding of key sub-systems interoperability.
Demonstrate the ability to interpret ambiguous data requests and convert them into precise technical operations and queries.
Have substantial experience developing, testing, and deploying data pipelines within a product-oriented SDLC environment.
Apply strong understanding of data modeling techniques, including semantic data modeling for AI optimization and analytical use cases.
Exhibit solid debugging skills and the ability to optimize processes while writing clean, maintainable code.
Nice to have
Experience with AWS services such as S3, Glue, and Athena for building resilient data ecosystems.
Hands-on experience with Snowflake for advanced data warehousing and performance tuning.
Proven background using Looker or alternative Business Intelligence suites for dashboarding and data storytelling.
Experience with Fivetran or similar ETL/ELT platforms for automating data replication.
Familiarity with Databricks or other Spark-based frameworks for large-scale data processing.
Previous work in the financial services industry where data security and regulatory compliance are paramount.
Practical notes
This role is based in Barcelona and requires local presence due to operational constraints.
Full-time engagement is required with standard working hours aligned with the team schedule.
Travel is generally not required for this position as the role is fully remote-capable within the designated location.
Candidates must meet the stated experience and technical requirements without exception.
The employment opportunity is open to experienced professionals who can start promptly in the near term.