Data Engineer
Job description
About the role
You will design and implement scalable data architectures that power our mission to transform healthcare with artificial intelligence. You will own the end to end lifecycle of data pipelines, ensuring they are reliable, efficient, and secure. You will define and enforce internal standards for style, maintenance, and best practices for a high-scale data platform. You will collaborate closely with researchers and stakeholders to understand data needs for model training and production monitoring systems. You will take ownership of key data engineering projects and deliver high quality solutions independently. You will evaluate and recommend data technologies to improve the efficiency of our data engineering processes. You will continuously monitor, maintain, and optimize the performance and stability of our data infrastructure.
Key facts
What you'll do
- Design and implement data architecture that is scalable, flexible, and efficient using pipeline authoring tools like Metaflow and large-scale data processing technologies like Spark.
- Define and extend internal standards for style, maintenance, and best practices for a high-scale data platform.
- Collaborate with researchers and stakeholders to uncover data requirements for model training and production monitoring systems.
- Take ownership of key data engineering projects and execute them independently from design through deployment.
- Ensure data quality, integrity, and security by implementing robust data validation, monitoring, and access controls.
- Evaluate and recommend data technologies and tools that enhance the efficiency and effectiveness of the data engineering process.
- Continuously monitor, maintain, and improve the performance and stability of the data infrastructure.
- Partner with product and engineering teams to align data strategies with business objectives and regulatory requirements.
- Build and optimize ETL workflows that handle large volumes of healthcare data with precision and reliability.
- Instrument data pipelines for observability, enabling proactive identification of issues and performance bottlenecks.
- Support the deployment of data solutions into production environments with an emphasis on reliability and scalability.
- Document data architectures, pipelines, and processes to ensure clarity and maintainability for cross-functional teams.
- Explore emerging tools and frameworks to future proof our data platform and drive innovation.
- Mentor junior engineers and contribute to technical discussions that shape the direction of our data ecosystem.
Requirements
- 5+ years relevant experience in data engineering with a proven track record of delivering robust data solutions.
- Expertise in designing and developing distributed data pipelines using big data technologies on large scale data sets.
- Deep and hands-on experience designing, planning, productionizing, maintaining and documenting reliable and scalable data infrastructure and data products in complex environments.
- Solid experience with big data processing and analytics on AWS, using services such as Amazon EMR and AWS Batch.
- Experience in large scale data processing technologies such as Spark.
- Expertise in orchestrating workflows using tools like Metaflow.
- Experience with various database technologies including SQL, NoSQL databases (e.g., AWS DynamoDB, ElasticSearch, Postgresql).
- Hands-on experience with containerization technologies, such as Docker and Kubernetes.
- Prior Software Engineering experience is a big plus.
- Demonstrated ability to work effectively in a fast paced, dynamic environment.
- Strong problem solving skills and a meticulous approach to debugging and optimization.
- Excellent written and verbal communication skills for collaborating with technical and non technical stakeholders.
- Commitment to maintaining high standards of data quality, security, and compliance in all data operations.
- Willingness to learn and adapt to new technologies in the rapidly evolving healthcare AI landscape.
Nice to have
- Experience working at an early stage startup.
- Experience in a HIPAA compliant environment.
- Experience working on machine learning or healthcare related projects.
Practical notes
- Employment is full time based in the United States.
- This role requires collaboration with cross functional teams and may involve travel to customer sites or industry events as needed.
- Candidates must be authorized to work in the United States without sponsorship for this position.
- The position is eligible for participation in Rad AI benefits programs upon hire, subject to plan eligibility requirements.
- Rad AI is an equal opportunity employer and welcomes applicants from diverse backgrounds to join our mission of transforming healthcare through AI.