Member of Technical Staff, Data Engineering
Job description
About the role
Join our team to construct the core data infrastructure for model training, product development, and company-wide analytics. You will shape the systems that power our AI agents, ensuring data integrity and accessibility. This role involves designing and implementing scalable data solutions from ingestion to serving. You will be responsible for translating ambiguous product requirements into robust data structures that enable experimentation and insight. The position requires close collaboration with machine learning engineers and analysts to align data strategy with business objectives. You will own the reliability and performance of critical data services that directly impact our product roadmap. Your work will establish the foundational layer that allows our AI agents to interact with the web at scale. This is an opportunity to define data standards and best practices for the entire engineering organization.
Key facts
What you'll do
- Design and implement data pipelines for large-scale ingestion and transformation across diverse web data sources.
- Develop high-throughput storage and serving layers that guarantee fast and reliable data querying for AI model workloads.
- Establish end-to-end systems for data quality, lineage, and observability to maintain trust in our analytical outputs.
- Proactively identify and resolve scaling bottlenecks in the data platform before they impact product performance.
- Make pragmatic architectural decisions to ensure the platform meets current demands and future growth trajectories.
- Collaborate with cross-functional partners to define data contracts and schemas that enable efficient product development.
- Implement monitoring and alerting frameworks to detect anomalies and failures in real-time data flows.
- Optimize data processing workflows to reduce latency and improve resource utilization across our infrastructure stack.
- Evaluate emerging data technologies and assess their fit within our existing architecture for long-term viability.
- Lead initiatives to standardize data practices and documentation across engineering and analytics teams.
- Partner with machine learning teams to ensure data pipelines support advanced model training and inference requirements.
- Drive improvements in data reliability by implementing robust testing and validation frameworks for all data products.
- Analyze complex data problems and propose solutions that balance performance, cost, and maintainability.
- Act as a technical leader in data engineering, mentoring junior engineers and elevating the overall technical bar.
Requirements
- Strong understanding of distributed data processing systems and the trade-offs involved in scaling them horizontally.
- Expertise in data modeling principles for both transactional and analytical workloads across relational and NoSQL stores.
- Deep intuition for system reliability, including fault tolerance, failure modes, and recovery strategies in distributed environments.
- Ability to analyze trade-offs between batch and streaming data approaches to select the optimal architecture for each use case.
- Experience evaluating storage cost versus query performance to design solutions that meet business and technical constraints.
- Skill in deciding between custom-built and off-the-shelf infrastructure solutions based on long-term maintainability and flexibility.
- Commitment to data accuracy, implementing rigorous validation checks and data quality frameworks to ensure dependable systems.
- Strong proficiency in SQL and at least one modern programming language for building data processing applications and infrastructure.
- Demonstrated experience working with large datasets in cloud environments, managing scalability and performance challenges.
- Understanding of data security and compliance best practices for handling sensitive information in web-scale applications.
- Proven ability to debug complex data issues in production environments using logs, metrics, and tracing tools.
- Excellent communication skills to articulate technical concepts to both technical and non-technical stakeholders.
- Willingness to participate in on-call rotations to support critical data infrastructure and respond to incidents promptly.
- Passion for building and maintaining high-quality data pipelines that serve as the backbone of machine learning and analytics.
Skills & tools
- Distributed data processing
- Data modeling
- System reliability
- Data pipeline design
- Data storage and serving
- Data quality systems
- Data lineage
- Data observability
- Cloud platforms and infrastructure as code
- Containerization and orchestration technologies
- Version control and collaborative development practices
Practical notes
Visa sponsorships are available. Benefits include 401K plans, daily lunch and office snacks, dinner at the office, unlimited vacation, and Caltrain pass reimbursement. Applicants must be authorized to work in the United States. The position is based in our San Francisco or Palo Alto office and requires consistent on-site presence. This role is not eligible for remote work arrangements outside the designated locations. Travel is not required for this position. Employment is at-will and contingent upon standard background verification processes. The start date is flexible but aligned with quarterly team intake cycles.