Applied Machine Learning Engineer
Job description
About the role
Vulcan Elements is committed to producing American rare-earth permanent magnets to support a secure and resilient future. This role involves developing and enhancing the company's data architecture, which is fundamental to operations and business analytics. You will be responsible for designing, implementing, and maintaining systems that facilitate the collection, storage, and utilization of data across various departments. Working closely with engineering, operations, and IT teams, you will ensure that the data infrastructure is robust, scalable, and compliant with relevant standards. Your contributions will enable the company to leverage data effectively for decision-making, process optimization, and future growth initiatives. This position offers an exciting opportunity to work on mission-critical systems in a fast-growing organization focused on advanced manufacturing technologies.
Key facts
What you'll do
- Develop, implement, and manage the overall data architecture, including operational data stores, data lakes, Lakehouses, ETL pipelines, and analytics layers, to support current operations and future scalability.
- Evaluate, select, and justify platforms for data Lakehouse, ETL tools, and operational databases, ensuring they meet scalability, compliance, and cost requirements.
- Create comprehensive documentation of data architecture, including data flow diagrams, system specifications, and standards, to facilitate team understanding and compliance with data security standards such as CUI and ITAR.
- Record architectural decisions and design choices to support ongoing development, troubleshooting, and future upgrades by team members and stakeholders.
- Design data systems capable of scaling from pilot projects to full-scale production environments, ensuring reliability, performance, and security.
- Implement best engineering practices, including version control, automated testing, and documentation, to support a growing data engineering team.
- Develop and maintain ETL (Extract, Transform, Load) pipelines that process, enrich, and prepare operational data for analytics, reporting, and AI applications.
- Establish reliable data ingestion pathways from manufacturing systems, laboratory systems, and other data sources, ensuring data quality and integrity.
- Collaborate with cross-functional teams to understand their data needs and translate these requirements into scalable architectural solutions.
- Automate manual data workflows to improve efficiency, reduce errors, and enable real-time or near-real-time data processing.
- Monitor data pipelines for performance issues, data quality problems, and security compliance, implementing alerts and corrective actions as needed.
- Address data quality issues proactively, establishing data validation and monitoring systems to detect anomalies early and ensure high data integrity.
- Support compliance with data handling standards, including CUI and ITAR, by implementing appropriate security measures and documentation practices.
- Contribute to the continuous improvement of data infrastructure, staying current with industry best practices, emerging technologies, and regulatory requirements.
- Work closely with data scientists, analysts, and other stakeholders to ensure that data architecture supports advanced analytics, machine learning, and AI initiatives.
- Participate in planning and executing data migration, upgrades, and system integrations as the company's data landscape evolves.
- Assist in training and mentoring junior team members, fostering a culture of best practices in data engineering and architecture.
Requirements
- Over 8 years of experience in data engineering or a related technical field, with a proven track record of delivering production-grade data systems.
- Extensive experience designing and building data lakes, Lakehouses, or analytical data stores, with the ability to evaluate and justify platform choices based on project needs.
- Proficiency in creating and deploying ETL/ELT pipelines that process large volumes of data efficiently and reliably.
- Strong understanding of data modeling principles, including normalization, denormalization, and schema design for operational and analytical workloads.
- Familiarity with metadata standards and practices to ensure data context, usability, and compliance.
- Experience working with data security standards such as CUI and ITAR, including implementing appropriate controls and documentation.
- Excellent problem-solving skills related to data quality, pipeline performance, and system integration.
- Ability to work collaboratively with cross-disciplinary teams, translating business needs into technical solutions.
- Strong documentation skills, with the ability to clearly communicate architecture decisions and system specifications.
- Knowledge of version control systems, automated testing frameworks, and best practices in software engineering for data systems.
- Experience working in a regulated environment with strict compliance standards.
- Ability to adapt to evolving project requirements and rapidly changing technology landscapes.
Nice to have
- Experience working with machine learning frameworks and tools, supporting data science and AI initiatives.
- Knowledge of compliance requirements specific to defense and aerospace sectors, including handling of sensitive data.
- Familiarity with cloud-based data platforms such as AWS, Azure, or Google Cloud, and related data storage solutions.
- Experience with containerization and orchestration tools like Docker and Kubernetes.
- Understanding of data governance, privacy, and security best practices in enterprise environments.
Skills & tools
- Proficiency in SQL and Python for data processing and pipeline development.
- Experience with ETL frameworks and data integration tools.
- Familiarity with cloud platforms and data storage solutions, including data lakes and warehouses.
- Strong analytical skills, with a focus on data quality, validation, and troubleshooting.
- Ability to document complex architectures clearly and effectively.
- Knowledge of version control systems such as Git.
- Skills in automating workflows and implementing monitoring solutions for data pipelines.
Practical notes
This position is based in Research Triangle Park, NC. The role begins in Durham, NC and is expected to transition to Benson, NC upon the completion of a new facility. Candidates must be eligible to work in the United States. The company emphasizes a collaborative environment and values candidates who are proactive in maintaining high standards for data security and quality. The position offers an opportunity to work on mission-critical data systems in a fast-growing organization dedicated to advanced manufacturing and national security.