Software Engineering Intern - Distributed Simulation Systems
Job description
Software Engineering Intern - Distributed Simulation Systems at Astera.
About the role
You will architect and implement core components of large-scale distributed simulation infrastructure that powers biological and physics-based research workloads. This role owns the networking and orchestration layers that allow simulation workloads to scale across heterogeneous compute clusters in production. You will build high-performance C++ services and tooling that directly impact how researchers model complex biological systems and 3D environments. You will collaborate closely with distributed systems engineers to design reliable communication protocols and performance optimization strategies. A significant part of your contribution will involve debugging intricate performance bottlenecks and improving system throughput under demanding simulation loads. You will also develop internal developer tools for visualization, workflow orchestration, and operational monitoring of simulation pipelines. This position offers deep exposure to scientific infrastructure where your work will directly support cutting edge research into biological and physical simulation at scale.
Key facts
What you'll do
- Architect and extend distributed simulation infrastructure that orchestrates complex multi-node workloads across on-premise and cloud environments.
- Design and implement robust networking and communication systems that enable low-latency data exchange for multi-node simulations.
- Develop and integrate high-performance C++ services that serve as the backbone for MuJoCo-based and other 3D simulation environments.
- Build and maintain systems that support large-scale biological simulations, focusing on correctness, scalability, and reproducibility.
- Profile and debug performance bottlenecks across the full software stack to improve simulation throughput, reliability, and resource utilization.
- Construct internal tooling for simulation orchestration, visualization, and end-to-end developer workflows to streamline researcher productivity.
- Partner with computational biologists and physics researchers to translate scientific requirements into resilient software solutions.
- Optimize system pipelines to handle demanding computational workloads while maintaining strict reliability and observability standards.
- Implement monitoring and diagnostic capabilities that provide deep insight into simulation execution across distributed clusters.
- Mentor and collaborate with cross-functional teams to ensure that infrastructure components meet evolving research and operational needs.
Requirements
- Demonstrate strong programming fundamentals with a focus on systems-level thinking and problem decomposition.
- Bring hands-on experience with C++, Python, or similar systems languages to build efficient and reliable software components.
- Show comfort working within Linux development environments, using standard toolchains, debuggers, and profiling utilities.
- Exhibit a solid grasp of core data structures, algorithms, and concurrency fundamentals that underpin high-performance software.
- Express a deep interest in distributed systems, simulation, or systems engineering with a track record of personal projects or contributions.
- Prove your ability to learn quickly and work independently on complex technical problems with minimal supervision.
- Maintain strong written and verbal communication skills to collaborate effectively with researchers and engineers across disciplines.
- Commit to upholding the highest standards of reliability, security, and maintainability in the systems you build and operate.
Nice to have
- Hands-on experience with C++ projects that demonstrate mastery of modern C++ features and software design patterns.
- Familiarity with networking and distributed systems concepts such as consensus, replication, and fault tolerance in practical deployments.
- Direct experience with physics simulators such as MuJoCo, Unity, or Unreal, including integration and performance tuning.
- Demonstrated background in Python, PyTorch, or scientific computing libraries for data processing and analysis workflows.
- Exposure to GPU programming or CUDA for accelerating compute-intensive simulation kernels and data-parallel workloads.
- Active involvement in open-source, research, robotics, or simulation projects that showcase engineering rigor and collaboration skills.
- Strong interest in biology, neuroscience, or scientific infrastructure with examples of prior work or study in these domains.
Practical notes
This is a hybrid role at our office in Emeryville, CA. Some travel may occasionally be required for collaboration and team events.
Applicants must be currently authorized to work in the United States without the need for employer sponsorship, now or in the future.