Senior Software Engineer, Data
Job description
About the role
You will architect and maintain the data pipelines that power Abridge's healthcare AI, ensuring the reliable flow of information from patient encounters to actionable insights. This role centers on building the robust infrastructure that allows our platform to process medical conversations at scale. You will own the design of systems that capture, transform, and serve complex healthcare data with precision and care. Your work will directly influence how clinicians interact with technology and how we validate the accuracy of our AI outputs. You will collaborate with cross-functional partners to translate analytical needs into durable data solutions. This position demands a high level of ownership and the ability to solve intricate problems in a regulated environment. Ultimately, you will ensure that the data underpinning our mission is trustworthy, accessible, and performant.
Key facts
What you'll do
- Construct and evolve scalable data services and pipelines that capture unstructured application feedback for machine learning training and evaluation.
- Architect and manage OLAP databases, ELT processes, and advanced data tooling to power analytics, business intelligence, and product feature development.
- Partner closely with frontend and backend engineers, product managers, and analysts to align data strategy with product goals.
- Optimize data infrastructure to improve system throughput, reduce latency, and strengthen overall reliability.
- Investigate anomalies flagged by data operations monitors, leveraging specialized tools and reports to resolve issues swiftly.
- Design cohesive data integrations and establish a rigorous data quality framework to support healthcare datasets.
- Implement infrastructure that supports the efficient ingest, transformation, and governance of structured and unstructured data types.
- Utilize modern data stack best practices to build solutions that are maintainable and scalable in a production environment.
- Integrate feedback mechanisms that allow data scientists and analysts to assess model performance against ground truth.
- Leverage containerization and orchestration technologies to deploy resilient data services in cloud environments.
- Document data architectures and processes to ensure clarity and maintainability for current and future engineers.
- Contribute to the development of data products that are intuitive, well-modeled, and easy to maintain across the organization.
- Balance multiple competing priorities while maintaining a high standard of work in a fast-paced startup setting.
- Collaborate with security and compliance teams to ensure data handling meets industry standards and regulations.
Requirements
- Bring 8+ years of hands-on experience in Data Engineering or Backend Engineering with a strong focus on data systems.
- Demonstrate proficiency in at least one general purpose programming language such as Python, Java, or Scala, alongside SQL proficiency in any variant.
- Show mastery of at least one major cloud provider platform, including GCP, AWS, or Azure, and their associated data services.
- Exhibit experience building systems that handle the full lifecycle of data, including ingest, transformation, and management of structured and unstructured formats.
- Apply deep knowledge of contemporary data infrastructure principles, including storage, processing, and orchestration patterns.
- Have practical experience with distributed systems and familiarity with frameworks like Spark, Flink, or similar processing engines.
- Utilize infrastructure as code tools such as Terraform, and orchestration platforms like Kubernetes, along with containerization methods.
- Leverage experience deploying machine learning models into production at scale where applicable.
- Focus on building data products that are well-documented, clearly modeled, and straightforward to maintain.
- Thrive in a dynamic environment where shifting priorities require quick adaptation and continuous learning.
- Communicate technical concepts effectively to both technical and non-technical stakeholders.
- Write clean, testable code and participate in code reviews to uphold engineering excellence.
- Take ownership of on-call responsibilities to support critical data infrastructure operations.
- Mentor junior engineers by sharing knowledge and providing guidance on best practices.
Practical notes
LENGTH: 700-900 words. No HTML, no markdown, no em dashes.
Output the page only.