Software Engineer, Data Platform
Job description
About the role
This role centers on the design and construction of scalable data platforms that power healthcare analytics and decision support. The position operates within a hybrid work model and requires U.S. citizenship or permanent resident status. You will own the full lifecycle of data platform components, from initial design through deployment and ongoing optimization in production environments. The role demands close collaboration with cross-functional partners to translate analytical needs into robust technical solutions. You will be responsible for ensuring that data pipelines are reliable, performant, and secure by design. Success in this position requires balancing innovation with the practical constraints of operating in regulated environments. You will mentor junior engineers and contribute to architectural discussions that shape the long-term vision for the data platform.
Key facts
What you'll do
Bring 3+ years of experience building production software, with strong command of SOLID principles, design patterns, and clean architecture to ensure maintainable and extensible systems.
Demonstrate deep understanding of algorithms and data structures, with a focus on distributed computing principles such as concurrency, partitioning, and shuffling for large-scale datasets to optimize performance and resource utilization.
Apply fluency across unit, integration, and end-to-end testing, treating automated tests as core to shipping to maintain high confidence in releases and reduce regression risk.
Show strong proficiency in Python and SQL, with the ability to quickly learn new languages and frameworks, and familiarity with JavaScript as a plus to broaden your contribution across the stack.
Apply deep practical knowledge of open table formats such as Delta Lake and big data file formats including Apache Parquet and Apache Avro to optimize storage, schema evolution, and query performance.
Implement Infrastructure as Code principles and tools for automated deployment and management of data pipelines to enable reproducibility, version control, and consistent environments.
Design robust data ingestion frameworks via RESTful APIs and build event-driven architectures for real-time data flow to support timely insights and operational decisions.
Design and scale cloud-native data platforms and orchestrate data workloads using AWS and Kubernetes to ensure elasticity, resilience, and efficient cost management.
Work comfortably in an agile startup environment, balancing quality with rapid iteration to deliver value quickly while maintaining technical excellence and operational stability.
Gain experience with agentic systems, Claude or LLM integration, or AI-assisted development workflows to explore how emerging technologies can enhance data platform capabilities and developer productivity.
Implement observability, monitoring, and tracing for production systems to enable proactive issue detection, faster debugging, and clear insights into system health and performance.
Work with open table formats (e.g., Delta Lake), big data file formats (Parquet, Avro), or stream processors (Kafka, Flink) to build flexible, high-throughput data processing pipelines.
Apply prior experience in regulated, high-security industries with knowledge of data security principles, especially around Protected Health Information (PHI), to design compliant and auditable data controls.
Draw on experience with cloud platforms, containerization, and orchestration tools to automate operations, enforce policies, and simplify complex workflows for data teams.
Requirements
Bring 3+ years of experience building production software, with strong command of SOLID principles, design patterns, and clean architecture.
Demonstrate deep understanding of algorithms and data structures, with a focus on distributed computing principles such as concurrency, partitioning, and shuffling for large-scale datasets.
Apply fluency across unit, integration, and end-to-end testing, treating automated tests as core to shipping.
Show strong proficiency in Python and SQL, with the ability to quickly learn new languages and frameworks, and familiarity with JavaScript as a plus.
Apply deep practical knowledge of open table formats such as Delta Lake and big data file formats including Apache Parquet and Apache Avro.
Implement Infrastructure as Code principles and tools for automated deployment and management of data pipelines.
Design robust data ingestion frameworks via RESTful APIs and build event-driven architectures for real-time data flow.
Design and scale cloud-native data platforms and orchestrate data workloads using AWS and Kubernetes.
Work comfortably in an agile startup environment, balancing quality with rapid iteration.
U.S. Citizenship or Permanent Residency is required.
Candidates must hold a Bachelor's degree.
Individuals must have a minimum of 3 years of relevant professional experience.
The role requires the ability to quickly learn new technologies and apply them in production settings.
Strong written and verbal communication skills are essential for effective collaboration with technical and non-technical stakeholders.
Nice to have
Gain experience with agentic systems, Claude or LLM integration, or AI-assisted development workflows.
Implement observability, monitoring, and tracing for production systems.
Work with open table formats (e.g., Delta Lake), big data file formats (Parquet, Avro), or stream processors (Kafka, Flink).
Apply prior experience in regulated, high-security industries with knowledge of data security principles, especially around Protected Health Information (PHI).
Practical notes
U.S. Citizenship or Permanent Residency is required.
Typical interview steps
Hiring for engineering roles usually starts with a recruiter screen, followed by one or two technical rounds. Candidates often solve a coding problem, discuss past projects, and answer system design questions. Some loops include a take-home task. Final rounds typically cover team fit and give candidates a chance to ask questions. Interviewers look for how you break down an unfamiliar problem, not just whether you reach the answer. Practicing a few problems aloud and reviewing your own past projects are the best preparation.
Good to know
Data engineers build pipelines that move and transform critical healthcare information.
The stack emphasizes modern data platforms, cloud services, and automated testing.
Security and compliance are important when handling sensitive health data.
Teams work in agile environments with rapid iterations and shared ownership.
Career growth
Engineering careers usually progress from individual contributor to senior, staff, and principal levels. Some engineers move into management and lead teams of five to twenty people. Others stay on the technical track. Growth follows demonstrated impact, not tenure alone. A typical engineering ladder has clear levels with defined expectations for scope, quality, and mentorship. Moving up usually requires owning outcomes end to end rather than completing assigned tickets.