Sr. Data Engineer
futurefitaiRemote (North America)Full Time2d ago
PythonGoAWSPostgreSQLMongoDBMachine LearningNLPAIMLData EngineeringETLAirflow
Job description
Sr. Data Engineer at futurefitai
About the role
Join a fast-paced, high-trust team dedicated to making a significant impact. We are looking for a motivated individual who is driven by purpose and embraces ambitious goals. This role offers a chance to build and own the core data infrastructure for a product that helps people find better jobs more efficiently.
Key facts
What you'll do
- Develop and maintain the systems that ingest and transform diverse data sources, ensuring reliable and scalable data flow into our data warehouse and product.
- Define and manage the structure, quality, and documentation of our central data models, including taxonomies for skills, occupations, and careers.
- Create the data transformation layers and datasets necessary for internal analytics, reporting tools like Looker and Quicksight, and customer-facing insights.
- Construct and manage pipelines that supply data for our matching and recommendation machine learning models, collaborating with engineering and data scientists on production deployment and monitoring.
Requirements
- Approximately 4+ years of experience in data engineering, with a proven track record of building and managing production ETL/ELT pipelines.
- Proficiency in Python and deep expertise in SQL, including experience with data modeling in data warehouses or data lakes, not just querying.
- Hands-on experience with modern data orchestration and transformation tools such as Airflow and dbt, and familiarity with cloud data warehouses.
- Experience integrating data from various external sources, including APIs and flat files, while managing schema changes, delivery inconsistencies, and data quality issues.
- Ability to work effectively with large, unstructured, and inconsistent datasets, making sound decisions on data cleaning, modeling, or source engagement.
- A proactive approach to ensuring pipeline reliability through testing, monitoring, and debugging.
- Strong communication skills, capable of explaining technical data concepts and their implications to non-technical stakeholders.
Nice to have
- Prior experience with labor market data, HR systems, or job-related data.
- Familiarity with skills or occupation frameworks like O*NET or ESCO.
- Experience with hierarchical data structures, ontologies, classification systems, or entity resolution across disparate data sources.
- Background in building data infrastructure for machine learning, including feature pipelines, model deployment, and monitoring tools like SageMaker.
- Contributions to open-source projects, publications, or presentations demonstrating data engineering expertise.
Skills & tools
- Languages: Python, SQL
- Orchestration & Transformation: Airflow, dbt
- Storage & Warehousing: PostgreSQL, Redshift, MongoDB
- Cloud Platform: AWS
- Visualization & Reporting: Looker, Quicksight
- Machine Learning & NLP: scikit-learn, AWS SageMaker, modern NLP and embedding tools
Practical notes
- This role is open to candidates located in Canada or the United States.
- Occasional travel, up to once per quarter, may be required for team events and off-sites.
- The hiring process typically takes around 6 weeks and includes an online application, initial screen, interviews, and a performance challenge.
- FutureFit AI utilizes AI tools to enhance hiring efficiency and equity in screening and interview support, with all final decisions made by humans.
- Reasonable accommodations will be provided for individuals with disabilities during the application and interview process.