Staff Software Engineer, Agentic Infrastructure
Job description
Space.
About the role
You will own the end to end reliability and evolution of a distributed stack of agentic systems that power our engineering and manufacturing workflows. You will solve problems that have not been solved before by building infrastructure where good off the shelf solutions do not yet exist. You will collaborate closely with forward deployed engineers to package and deliver portions of the stack as services to internal and external customers. Your work will directly shape how agents, sessions, model routing, and context pipelines behave in production. You will ensure that local LLM inference across Mac Studio and GPU hardware remains performant, stable, and observable. The role demands comfort with ambiguity and the ability to drive technical decisions across distributed systems and cross functional teams.
Key facts
What you'll do
Design, implement, and maintain the infrastructure that powers our agentic systems, including agents, sessions, model routing, context pipelines, and inter agent communication.
Build and maintain foundational capabilities where existing off the shelf solutions are insufficient for our aerospace workloads and reliability requirements.
Deploy and operate local LLM inference clusters across Mac Studio and GPU hardware, tuning performance and stability for production usage.
Manage Linux and macOS systems, Proxmox virtual machines, containers, storage, backups, and recovery processes to ensure high availability.
Maintain robust multi site networking including L3 routing, VLANs, DNS, firewalls, and connectivity for distributed infrastructure components.
Implement and operate service discovery, load balancing, and secure internal communication between services and agentic workloads.
Own the reliability of inter agent message buses and parametric CAD pipelines that connect design, simulation, and manufacturing workflows.
Operate and optimize a print farm and related scheduling systems that align production capacity with engineering and manufacturing demand.
Collaborate with forward deployed engineers to turn prototypes into hardened production infrastructure that can be consumed internally and externally.
Instrument observability, monitoring, and alerting across the full stack to enable rapid troubleshooting and data driven decision making.
Define and enforce reliability standards, runbooks, and automation to reduce manual toil and improve incident response.
Partner closely with software, avionics, and manufacturing teams to align infrastructure capabilities with evolving product and business needs.
Contribute to technical design reviews, architecture decisions, and long term roadmaps for the agentic infrastructure platform.
Mentor engineers on best practices for building, operating, and scaling distributed, agent centric systems in a fast moving environment.
Requirements
You have deployed an agentic harness such as OpenClaw, Hermes, SybilClaw, or a comparable system in a real work environment.
You understand what it takes to make agentic systems reliable when engineers depend on them every day, and you have the operational experience to prove it.
You are comfortable owning and evolving complex software systems that move from prototype to production at production scale.
You have strong experience with Linux and macOS system administration, including performance tuning, security hardening, and troubleshooting.
You are proficient in managing containerized workloads, orchestration where appropriate, and infrastructure as code practices.
You have hands on experience with local LLM inference, model optimization, and hardware considerations that affect throughput and reliability.
You understand networking fundamentals including L3 routing, VLANs, DNS, firewalls, and secure connectivity between sites and services.
You are capable of designing and operating service discovery, load balancing, and resilient communication patterns for distributed services.
You have experience with parametric CAD pipelines and understand the challenges of integrating design, simulation, and manufacturing data.
You have operated print farm or scheduling systems or are willing to deeply learn and own this workload in production.
You communicate clearly and collaborate effectively with cross functional teams in a fast paced, mission critical environment.
You are comfortable working with minimal process and making decisions through technical discussion and consensus.
You are passionate about building infrastructure that empowers other engineers to deliver reliable and innovative solutions.