Data Engineer
Job description
About the role
Newsela operates a digital platform that delivers curated news and educational content to classrooms across the United States. The Data Engineer position at Newsela focuses on constructing and sustaining the data infrastructure that powers analytics and reporting throughout the organization. This role requires designing reliable data pipelines, maintaining data quality standards, and enabling teams to make informed decisions based on accurate information. The Data Engineer collaborates with product managers, analysts, and fellow engineers to translate business needs into effective data solutions that drive product improvements. This position is fully remote and open to candidates based anywhere within the United States, offering flexibility while contributing to a mission-driven company.
Key facts
What you'll do
Architect and build data pipelines that extract, transform, and load information across multiple storage systems reliably
Develop automated workflows that ensure data freshness and consistency across all downstream consumers and applications
Partner with multidisciplinary teams to understand their data requirements and deliver scalable solutions that meet those needs
Monitor pipeline performance and respond to failures or degradation to maintain high availability of data services
Write clear documentation describing data schemas, transformation logic, and infrastructure configurations for team reference
Implement data validation and quality checks that identify anomalies, duplicates, or inconsistencies in incoming datasets
Optimize existing data infrastructure for improved throughput, reduced latency, and lower operational costs over time
Support business intelligence and analytics initiatives by providing well-structured, accessible datasets to stakeholders
Participate in code reviews and contribute to engineering standards and best practices within the data group
Evaluate emerging tools and technologies that could enhance the company's data processing and storage capabilities
Troubleshoot and resolve data quality issues by tracing problems back through source systems and transformation logic
Contribute to the evolution of the company's data strategy by proposing improvements and sharing knowledge with peers
Requirements
Demonstrated experience designing and operating data pipelines in a production or production-like environment
Strong proficiency in at least one general-purpose programming language used for data processing and engineering
Familiarity with cloud-based data services and managed infrastructure offerings for building scalable systems
Solid understanding of relational databases and non-relational data stores and their appropriate application scenarios
Ability to diagnose and resolve complex issues in data flows, transformations, and integrations effectively
Experience with version control systems and collaborative software development practices in a team setting
Excellent written and verbal communication skills for explaining technical concepts to diverse audiences
Bachelor's degree in computer science, software engineering, mathematics, or a closely related discipline
Comfortable working in an agile or iterative development environment with regular releases and feedback cycles
Aptitude for learning new tools and frameworks quickly and adapting to evolving business and technical requirements
Nice to have
Exposure to real-time or streaming data sources and event-driven processing patterns in production environments
Experience with containerization technologies and cluster management platforms for deploying and scaling data services
Background in education technology, media, or content-driven companies and understanding their unique data challenges
Knowledge of data governance practices, cataloging, metadata management, and maintaining organized data ecosystems
Familiarity with data visualization platforms and the ability to create dashboards for stakeholder consumption
Skills & tools
SQL and Python for querying, transforming, and moving data between systems efficiently
Cloud platform services for data storage, compute, and managed analytics offerings
Data modeling techniques for structuring analytical datasets and warehouse schemas
Workflow scheduling and orchestration frameworks for managing pipeline execution and dependencies
Git and collaborative development workflows for versioning code and infrastructure definitions
Logging and monitoring solutions for tracking pipeline health, latency, and error rates
Data integration and API tools for connecting external systems and internal platforms together
Practical notes
This role operates in a fully remote setting, and candidates must be based within the United States
Newsela is committed to building a diverse and inclusive workplace and encourages all qualified individuals to apply
The interview process typically includes technical discussions, coding exercises, and conversations with prospective teammates
The Data Engineer should expect to work across time zones and communicate progress asynchronously with colleagues
Newsela values transparency and open communication, and expects all team members to share progress and blockers regularly