Software Development Engineer II, Search Data
Job description
About the role
Mapbox is seeking a Software Development Engineer II for the Search Data team to own the design and execution of data pipelines that power global search capabilities. You will own the end to end lifecycle of address, place, and points of interest data from ingestion through serving. This role requires deep ownership of complex geospatial data transformations, quality assessment, and optimization of large scale systems. You will partner closely with Maps and Navigation to ensure search data meets strict latency, accuracy, and reliability requirements. You will shape the technical direction of search data infrastructure while mentoring peers and driving operational excellence. This is an opportunity to work on foundational data that enables location aware applications across web, mobile, and vehicle platforms. You will thoughtfully integrate AI into engineering workflows to improve how search data is processed, validated, and utilized.
Key facts
What you'll do
Design and build scalable batch and streaming ingestion systems that process terabytes of data daily from thousands of sources.
Implement robust ETL data pipelines using AWS technologies such as Lambda, S3, Athena, Glue, and EMR to prepare data for search engines.
Collaborate with engineers across Maps, Navigation, and other teams to understand geospatial data requirements and deliver tailored solutions.
Simplify and strengthen Mapbox's processes for designing, deploying, and monitoring data processing and querying workloads in cloud environments.
Document technical decisions, data schemas, and pipeline behaviors to ensure clarity and long term maintainability of search data systems.
Lead presentations and discussions that translate complex geospatial concepts into accessible narratives for both technical and non technical stakeholders.
Mentor other software engineers through design and code reviews, fostering skill development and a culture of continuous improvement.
Champion operational excellence by implementing rigorous testing, monitoring, and alerting for data pipelines and search services.
Reduce technical debt by investing in reusable components, improving tooling, and sharing knowledge across the team.
Promote a culture of collaboration, transparency, creativity, inclusion, and data driven decision making in all aspects of the search data lifecycle.
Requirements
5+ years of experience building scalable backend systems and data pipelines in production environments.
Hands on experience with AWS technologies including Lambda, S3, Athena, Glue, and EMR for building and operating data platforms.
Strong proficiency in SQL for complex querying, optimization, and data transformations across large datasets.
Strong proficiency in Python for scripting, pipeline development, and automation of data workflows.
Proficiency in at least one modern programming language such as NodeJS, Scala, or Java to build backend services and data processing applications.
Demonstrated history of designing batch and real time data processing systems with the judgment to implement new data pipelines and best practices.
Familiarity with Apache Spark or other Hadoop based technologies for large scale data processing and analytics.
Familiarity with CI/CD processes to automate testing, deployment, and monitoring of data pipelines and services.
Experience with introducing quality and operational metrics into data ETL pipelines to monitor performance, reliability, and correctness.
Experience integrating data with APIs and querying data through APIs to support search services and downstream applications.
Experience with AI tools in the software development lifecycle to improve development efficiency, testing, and decision making.
Nice to have
Experience with geospatial data analysis and processing using specialized libraries and data formats.
Experience with Docker for containerizing services and ensuring consistent environments across development and production.
Experience with machine learning infrastructure to support data driven features and predictive capabilities within search pipelines.
Practical notes
Annual base compensation for this role ranges from $160,650 to $217,350 for most US locations, with a 5% to 10% increase for locations with a higher cost of labor.
Job level and actual compensation are determined based on factors including, but not limited to, individual skills and experience.
This position is full time and based in Mapbox US.
This role involves on call responsibilities to support the health of search data services.
Candidates should be legally authorized to work in the country where the position is located without requiring sponsorship for employment authorization at the time of hire.