Analytics Engineer in Data Collection
Job description
About the role
At Similarweb, we are not just another data company; we are the leading digital intelligence platform used by global giants like Google, eBay, and Adidas. 🌍 Our insights power the digital strategies of over 4,300 companies worldwide, and we are growing fast. After going public on the New York Stock Exchange in 2021, we continue to break new ground, and now we are expanding our dynamic team in Prague! 🇨🇿 We are looking for an Analytics Engineer to join the Data Collection Operations team and contribute to the configuration writing, monitoring and big data analysis of Similarweb's data collection services. This role involves working as part of a team responsible for ensuring the accuracy and reliability of our data collection processes across mobile and desktop devices. This position is crucial because as the most trusted platform for measuring online behavior, millions of people rely on Similarweb's insights daily as the ground truth for their knowledge of the digital world. Producing these insights requires large scale raw data to be collected in high quality and with minimal outages. As the Analytics Engineer, you will have the opportunity to perform hands-on work and take ownership of Similarweb's raw data collection, monitoring and quality assurance processes. Your work will have a direct impact on the insights our products deliver to our customers.
Key facts
What you'll do
Write and maintain configuration files that precisely define the data elements to be captured from web pages and APIs.
Monitor data pipelines in real time using Grafana and other observability tools to detect anomalies and ensure continuity of collection.
Utilize AI tools to automate configuration autohealing and reduce alert fatigue across monitoring dashboards.
Continuously learn and evaluate emerging technologies to enhance the efficiency and reliability of data collection workflows.
Collaborate with product managers and data team leaders to align collection strategies with evolving product requirements.
Perform systematic code reviews to validate logic, improve monitoring systems, and answer data-related queries from stakeholders.
Investigate and resolve data quality issues by tracing discrepancies from source to dashboard with methodical debugging.
Document collection procedures and configurations to ensure clarity, consistency, and ease of onboarding for future team members.
Participate in on-call rotations to provide rapid response and support for critical data collection incidents around the clock.
Champion best practices in data governance, ensuring that all collection activities comply with internal standards and external regulations.
Drive iterative improvements by analyzing historical collection metrics and proposing data-driven enhancements to the pipeline.
Mentor junior analysts by sharing knowledge on SQL, Pyspark, and Python to strengthen the overall capability of the data collection function.
Champion the use of version control and testing frameworks to increase reliability and reproducibility of configuration changes.
Promote a culture of curiosity and innovation by encouraging experimentation with new tools that optimize data collection performance.
Requirements
Bring 3+ years of professional experience in data analysis or a closely related technical role.
Demonstrate advanced proficiency with core data stack technologies such as Databricks, Spark, and Airflow.
Show proven ability to perform complex data analysis using SQL, Pyspark, or Python in production environments.
Possess solid understanding of HTML, JavaScript, and web data collection concepts to effectively interpret source structures.
Act as a trusted data partner by delivering fast, reliable responses to data-related questions while identifying and resolving issues promptly.
Comfortably accept challenging assignments and show eagerness to learn new technologies and adapt to changing requirements.
Communicate clearly and professionally when collaborating with cross-functional teams including product managers and engineering colleagues.
Adhere to strict attention to detail to ensure high data quality and consistency across all collection processes.
Nice to have
Knowledge of AI tools and their application in automating data collection and monitoring tasks.
Experience with version control systems such as Git for managing configuration files.
Familiarity with data quality frameworks and automated testing for analytics pipelines.
Understanding of privacy regulations and data compliance relevant to digital analytics.
Practical notes
This role operates under a hybrid work model with 3 days in our office and 2 days working from home. Team-building events, parties, weekly lunches, and happy hours are held throughout the year to strengthen team spirit. You will work with top hardware including a MacBook Pro M3, dual monitors, and an electric standing desk. Access to equity programs is available, allowing you to become a shareholder and join in the success of Similarweb. The office is located at DOCK IN (Palmovka) in Prague, stocked with snacks, drinks, and space to unwind. Enjoy 5 weeks of vacation, an extra day off during your birthday month, 3 sick days, and a multisport card.