Software Engineer - Data Platform
Job description
About the role
SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.
The Data Platform team builds and operates the infrastructure responsible for all large-scale data transport and processing across the company. We own and manage core systems including Apache Kafka, HDFS, Spark, Flink, and Trino, enabling real-time ML pipelines, feed ranking, experimentation, analytics, and observability at petabyte scale. Our team deals with latency-critical workloads, high-throughput streaming, and distributed compute systems that require fault tolerance, performance, and absolute reliability.
As a software engineer on the Data Platform team, you will design, build, and operate the distributed systems powering SpaceXAI's data movement and compute. You will take ownership of infrastructure components that process trillions of events daily, driving the scalability, performance, and reliability of the systems that power product and ML workloads across the company.
Responsibilities
- Design and implement high-throughput, low-latency data ingestion and transport systems.
- Scale and optimize multi-tenant Kafka infrastructure supporting real-time workloads.
- Extend and tune Spark, Flink, and Trino for demanding production pipelines.
- Build interfaces, APIs, and pipelines enabling teams to query, process, and move data at petabyte scale.
- Debug and optimize distributed systems, with a focus on reliability and performance under load.
- Collaborate with ML, product, and infrastructure teams to unblock critical data workflows.
Basic Qualifications
- Proven expertise in distributed systems, stream processing, or large-scale data platforms.
- Proficiency in Rust, Go, Scala or similar systems languages.
- Hands-on experience with Kafka, Flink, Spark, Trino, or Hadoop in production.
- Strong debugging, profiling, and performance optimization skills.
- Track record of shipping and maintaining critical infrastructure.
- Comfortable working in fast-moving, high-stakes environments with minimal guardrails.
Compensation and Benefits
The role offers a base salary within the range of $180,000 - $440,000 USD. Base salary is just one part of our total rewards package at SpaceXAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.
About the role
You will architect and maintain the data transport backbone that moves information across the company at every scale. You will own the critical path for high-throughput streaming and batch processing, ensuring no data loss and strict reliability. You will take direct responsibility for the performance and fault tolerance of core distributed systems. You will partner with product and ML teams to turn complex data requirements into robust production pipelines. You will lead debugging efforts for latency-critical workloads under heavy load conditions. You will contribute directly to the company mission by building infrastructure that enables scientific discovery. You will exercise ownership by proposing improvements that span design, implementation, and on-call support. You will communicate clearly with both technical and non-technical stakeholders to align priorities and share knowledge.
What you'll do
- Design and implement high-throughput, low-latency data ingestion and transport systems that handle petabyte-scale flows.
- Scale and optimize multi-tenant Kafka infrastructure to support real-time workloads with strict reliability requirements.
- Extend and tune Spark, Flink, and Trino for demanding production pipelines that drive product and ML outcomes.
- Build interfaces, APIs, and pipelines enabling teams to query, process, and move data efficiently at massive scale.
- Debug and optimize distributed systems, focusing on performance, reliability, and behavior under sustained heavy load.
- Collaborate with ML, product, and infrastructure teams to unblock critical data workflows and align on long-term architecture.
- Implement observability and monitoring strategies that provide deep insight into system health and data quality.
- Evaluate emerging technologies and integrate components that improve throughput, resilience, and operational simplicity.
- Own the full lifecycle of data platform services from design and deployment through iteration and incident response.
- Champion best practices for security, governance, and compliance across data movement and storage systems.
- Drive automation to reduce manual effort, improve reliability, and accelerate delivery of platform capabilities.
- Mentor engineers by sharing deep expertise in distributed systems, stream processing, and large-scale data platforms.
Requirements
- You possess proven expertise in distributed systems, stream processing, or large-scale data platforms with a track record of handling high-volume environments.
- You demonstrate proficiency in Rust, Go, Scala, or similar systems languages, writing robust and performant code.
- You have hands-on experience with Kafka, Flink, Spark, Trino, or Hadoop in production, managing complex deployment and tuning scenarios.
- You show strong debugging, profiling, and performance optimization skills, able to isolate issues in complex distributed flows.
- You maintain a track record of shipping and maintaining critical infrastructure that serves high-stakes production workloads.
- You are comfortable working in fast-moving, high-stakes environments with minimal guardrails, making sound decisions with incomplete information.
- You exhibit excellent communication skills, concisely and accurately sharing knowledge with teammates across disciplines.
- You understand the importance of reliability, fault tolerance, and performance in systems that process trillions of events daily.
Nice to have
- Experience contributing to open-source data platforms or distributed systems projects that demonstrate engineering excellence.
- Familiarity with machine learning data pipelines and the specific requirements of training and inference workflows.
- Background in building or operating observability stacks for large-scale distributed systems.
Practical notes
-
Location: USA
-
Engagement: Full-time
- Travel: Not required
- Employment eligibility to work in the United States is required.
- No current visa sponsorship information is provided.
- Applications will be reviewed on a rolling basis until the role is filled.