Data Engineer, Data Core
vibeFranceFull Time1w ago
AIMLAirflowSparkBigQueryUISalesMarketingFinanceOperationsSupportRecruiter
Job description
Data Engineer, Data Core at vibe.
About the role
Join the Data Platform team, which manages vibe's entire data infrastructure, including storage, batch processing, and real-time reporting. This role addresses the challenges of petabyte-scale data and high message throughput, offering ownership of systems from design to incident response.
Key facts
What you'll do
- Develop tools and services for storing and accessing petabyte-scale event data
- Build efficient compute solutions using Spark, DuckDB, and Trino for large datasets
- Create in-app reporting features on an OLAP database like ClickHouse
- Orchestrate data movement and jobs across the stack using both batch and real-time methods
- Identify and resolve performance and cost issues in data pipelines and queries
- Implement system instrumentation to prevent cost and latency problems
- Anticipate and address scaling limitations before they cause outages
- Participate in on-call rotations and lead incident response for owned systems
- Create and maintain alerts and runbooks to proactively manage system health
- Develop tooling, conventions, and data contracts to improve team efficiency
- Translate complex requests from Product, ML, and Finance into actionable deliverables
- Provide constructive feedback to stakeholders regarding project feasibility and timing
- Define technical specifications for the reporting UI built by the Front team
Requirements
- 5 or more years building and operating production data platforms at scale
- Extensive experience with lakehouse architectures, covering storage (Iceberg, Delta, or Hudi) and compute (Spark, DuckDB, or Trino)
- Production experience with a modern orchestrator like Dagster, Airflow, or Prefect
- A demonstrated example of improving pipeline or query speed or cost
- Experience with on-call duties and incident response, including creating alerts or runbooks
- Ability to adjust communication style for diverse audiences
Nice to have
- Production experience with a column-store analytics engine (ClickHouse, Druid, Pinot, or BigQuery at significant scale)
- Experience with streaming data technologies
- Background in AdTech or programmatic advertising
- Experience with Dagster, Spark, DuckDB, Iceberg, or ClickHouse Cloud
Skills & tools
- Spark
- DuckDB
- Trino
- ClickHouse
- Dagster
- Airflow
- Prefect
- Iceberg
- Delta
- Hudi
- Druid
- Pinot
- BigQuery
Practical notes
Full health insurance coverage via Alan. Meal vouchers provided by Swile. Annual company offsite and quarterly in-person Engineering and Product syncs. The interview process includes a recruiter screen, manager interview, technical interview (practical exercise, no LeetCode), and a final bar raiser with the CTO.