Staff Data Engineer
Job description
About the role
Cribl is a data observability and pipeline company that helps organizations manage, route, and process their machine-generated data at scale. The Staff Data Engineer will play a critical role in designing and building the data infrastructure that powers Cribl's platform and enables customers to gain visibility into their data flows. This individual will work closely with engineering teams across the organization to solve complex data challenges and drive architectural decisions that shape the product roadmap. As a senior technical leader, this engineer will mentor other data engineers, contribute to hiring efforts, and help evolve Cribl's data ecosystem to meet growing customer demands. The Staff Data Engineer will partner with product managers and site reliability engineers to ensure that data infrastructure scales alongside Cribl's growing customer base and feature set.
Key facts
What you'll do
Design and implement scalable data pipelines that handle high-volume log and telemetry data across distributed systems and production environments, ensuring low latency and high throughput. Architect data processing workflows using modern streaming and batch processing frameworks to meet evolving product requirements and customer expectations. Collaborate with product and engineering teams to define data models, schemas, and storage strategies that support Cribl's platform growth. Optimize existing data infrastructure for performance, reliability, and cost efficiency while maintaining high availability in production deployments and minimizing operational overhead. Lead the evaluation and adoption of new data technologies that improve system capabilities and developer productivity across the organization. Build and maintain data quality monitoring systems to ensure accuracy, consistency, and completeness across all data flows and processing stages. Contribute to the open-source Cribl project by authoring and reviewing code for data processing components and pipeline features. Develop comprehensive documentation and runbooks that enable other engineers to understand, operate, and troubleshoot data systems effectively. Participate in on-call rotations and incident response procedures to troubleshoot data pipeline failures, diagnose root causes, and restore service in production environments. Mentor junior and mid-level data engineers through code reviews, design discussions, and structured knowledge-sharing sessions within the team.
Requirements
Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience and professional background. Seven or more years of professional experience in data engineering, data infrastructure, or a closely related technical discipline. Strong proficiency in Python and SQL for building data pipelines, transformations, and analytical workflows in production settings, with the ability to write clean, maintainable code. Experience with distributed systems concepts including message queues, stream processing, and event-driven architectures at scale. Familiarity with cloud platforms such as AWS, GCP, or Azure for deploying and managing data services and infrastructure. Demonstrated ability to design and implement data models that support high-throughput ingestion, processing, and querying workloads. Experience working with containerization technologies like Docker and orchestration tools such as Kubernetes for deploying data applications. Proven track record of mentoring other engineers and contributing to team-wide technical standards and engineering best practices.
Nice to have
Experience with data streaming platforms such as Apache Kafka, Apache Pulsar, or similar distributed messaging technologies. Familiarity with data observability tools and practices for monitoring data freshness, volume, and schema changes in real time. Contributions to open-source data projects or a strong personal portfolio of data engineering work and technical writing. Knowledge of data governance frameworks and security best practices for handling sensitive machine-generated data responsibly. Experience with data warehousing solutions such as Snowflake, BigQuery, or Redshift for analytical workloads.
Skills & tools
Python for data pipeline development, scripting automation tasks, and building reusable data utilities
SQL for querying and transforming structured and semi-structured data sets
Apache Kafka or equivalent distributed messaging and streaming systems
Docker and Kubernetes for containerized deployment and orchestration
AWS or another major cloud provider for data infrastructure services
Git for version control, collaborative software development workflows, and code review practices
Practical notes
This role is fully remote and open to candidates located anywhere within the United States. Cribl operates on a flexible work schedule, but expects regular availability for team collaboration, synchronous meetings, and cross-functional discussions across multiple time zones within the United States. The hiring process includes technical interviews focused on system design and data architecture problem-solving exercises. Candidates should expect a multi-stage interview process that includes a technical design exercise, a coding assessment, and behavioral conversations with senior leadership. Cribl is an equal opportunity employer and welcomes applicants from all backgrounds and experiences.