Senior Staff Engineer, Data
Job description
About the role
Staff Data Engineer, Lakehouse Platforms
The role centers on architecting and advancing FloQast's foundational data infrastructure. You will establish the patterns that govern how data is ingested, secured, stored, processed, and accessed across every product and analytics function. Apache Spark serves as the primary engine driving this ecosystem. Your responsibilities will ensure that data platforms deliver reliability, scalability, and security for every team that builds on them.
Success requires deep operational expertise. You think in terms of shuffle mechanics and partition methodologies long before reviewing any reference material. You understand the inner transformations of the query planner and you write code that works efficiently with that design. You have made the strategic choice between PySpark and Scala Spark based on concrete production requirements. You have operated real-time data flows that move events through distributed messaging systems. You have investigated performance bottlenecks using detailed execution traces and determined why specific stages run slowly. You have constructed large-scale data operations that manage complex table formats and their associated metadata. You treat data correctness and infrastructure health as core responsibilities, not optional considerations. At this level of impact, your technical direction dictates whether future growth is achieved through evolutionary changes or complete rewrites.
In this position, you will apply that expertise across the entire lakehouse environment at FloQast. FloLake provides a foundation built on open table formats, storing information on object storage. Orchestration occurs through managed workflow systems, with metadata registered in central catalog services. Analytical engines interact with this layer using standardized query interfaces. Event streams transport timely information through high-throughput channels. You will be responsible for the compute strategies that support thousands of independent operations. You will define how data processing jobs are structured, governed, and monitored. You will establish the expectations for how the engineering organization approaches distributed computing challenges. Professionals at all levels will rely on your judgment for guidance and validation.
Location and Compensation
This position is based in San Jose, California. It is a full-time role with a total annual compensation package ranging from $190,000 to $230,000 USD.
Core Responsibilities
Designing robust pathways for data ingestion that feed structured lakehouse environments.
Establishing access controls and governance policies across multiple query execution engines.
Implementing processing workflows that preserve data correctness during structural changes.
Analyzing execution plans to verify efficient data scanning and minimize unnecessary data movement.
Deploying enhancements to table formats while automating routine maintenance activities.
Collaborating with product teams to integrate event-driven architectures into larger platforms.
Determining compute frameworks that influence stability and performance for a large user base.
Defining validation procedures that verify correctness before updates reach production systems.
Qualifications and Experience
You possess extensive hands-on experience with distributed processing frameworks in production settings. You understand the behavior of memory systems under load without needing to check documentation. You can analyze how a query is transformed step by step and adjust your implementations accordingly. You have practical experience selecting execution languages based on workload characteristics. You have managed streaming use cases involving high-volume ingestion and processing. You have used diagnostic tools to isolate and resolve performance issues. You have created data processing jobs that interact with modern table formats at significant scale. You treat metadata management and snapshot isolation as standard parts of the job. You determine architectural patterns that dictate how easily the system can scale.
Additional Assets
Contributions to public repositories involving data processing frameworks are valuable. Familiarity with scheduling and coordination tools used in data workflows is beneficial.
Technical Stack
Core technologies in this role include Apache Spark, Apache Iceberg, and event streaming platforms. You will work within environments supported by AWS managed services and query engines.
Practical Information
All details regarding qualifications and application procedures are confirmed on the official career site.