Data Engineer
Job description
Data Engineer at Stuut Ai.
About the role
Stuut is actively engaged in the modernization of accounts receivable for B2B enterprises through the automation of manual financial workflows. In this capacity, you will serve as the first data hire, taking full ownership of architecting our entire data foundation from the ground up. You will be responsible for defining the methodologies by which we transform raw, unstructured information into actionable intelligence for clients spanning manufacturing, chemicals, and industrials. The role requires you to act as the central technical authority on data strategy, ensuring that every pipeline and model aligns with long-term business objectives. You will own the end-to-end lifecycle of data, from initial ingestion to advanced transformation and final delivery to end-users. This position demands a high degree of initiative, as you will set the standards for data quality, reliability, and scalability without precedent within the organization. You will work closely with executive leadership to ensure that data capabilities directly support product development and commercial growth.
Key facts
What you'll do
- Design and implement a robust data infrastructure capable of ingesting and modeling information from diverse sources such as ERP systems and payment processors.
- Construct a semantic layer that enforces metric consistency across internal reporting, product analytics, and AI-driven applications.
- Build canonical data models that normalize heterogeneous source data while embedding rigorous quality testing and system observability.
- Develop event and signal pipelines that generate labeled datasets essential for machine learning workflows and new product features.
- Partner with engineering and product teams to integrate data lineage tracking and quality standards directly into the software development lifecycle.
- Create and maintain DataOps workflows to ensure ongoing high data accuracy, system reliability, and efficient issue resolution.
- Define critical key performance indicators and construct interactive reporting dashboards that provide strategic insight for senior business decisions.
- Optimize the data platform architecture to support seamless scaling from an initial set of dozens of customers to a base of hundreds.
- Establish data governance frameworks that clarify ownership, definitions, and access controls across all data assets.
- Conduct in-depth data investigations to troubleshoot complex issues and provide clear narratives that explain variances or anomalies.
- Implement monitoring and alerting systems to detect data drift, schema changes, and pipeline failures before they impact stakeholders.
- Facilitate documentation and knowledge transfer sessions to ensure continuity and enable other team members to leverage the data platform effectively.
Requirements
- Possess a minimum of 3 years of professional experience developing production-grade data pipelines using Python in a live environment.
- Demonstrate proficiency in SQL and navigating modern cloud data warehouse environments such as Snowflake or BigQuery.
- Have practical experience designing and maintaining ETL/ELT workflows using orchestration tools like Airflow and transformation tools such as dbt.
- Show a background in building semantic or metrics layers to enforce data consistency across reporting and analytics.
- Exhibit the ability to design canonical schemas that impose structure on messy, heterogeneous, and unstructured data sources.
- Bring experience managing real-world data originating from SaaS APIs, enterprise resource planning systems, and third-party integrations.
- Show a steadfast commitment to data quality, including the implementation of automated testing, proactive anomaly detection, and comprehensive lineage tracking.
- Prove the capability to function effectively in ambiguous and fast-paced environments while building critical systems from the ground up with minimal guidance.
- Hold the right to work in the United States and be eligible for full-time employment in the state of California without sponsorship at this time.
Nice to have
- Have direct experience collaborating with machine learning teams on the development of feature pipelines or related infrastructure.
- Possess a background or demonstrated interest in fintech, B2B SaaS, or familiarity with accounts receivable and accounts payable workflows.
Practical notes
- This is an on-site role based in San Francisco, requiring consistent presence in the office.
- Benefits for U.S. employees include medical, dental, and vision insurance, a 401(k) plan with company match, flexible paid time off, and parental leave.