DevOps Engineer I (Remote)
Job description
About the role
You will perform daily system monitoring to verify the integrity and availability of all hardware, server resources, systems, and key processes while reviewing system and application logs. You will provision and configure hardware, peripherals, services, settings, directories, and storage in strict accordance with established standards and project or operational requirements. You will research, develop, and implement innovative and where possible automated approaches for system administration tasks to drive efficiency. You will provide Tier II support for incidents and requests from various constituencies, investigating and troubleshooting issues as they arise. You will maintain and actively contribute to our knowledge base and documentation to ensure continuity and clarity. You will influence broader technology groups in adopting Cloud technologies, processes, and best practices to align with enterprise standards. You will exercise creative thinking to propose solutions that grow our business by delighting our clients and enhancing their experience. You will comply with all policies and standards, ensuring adherence to governance and regulatory expectations.
Key facts
What you'll do
Perform daily system monitoring, verifying the integrity and availability of all hardware, server resources, systems, and key processes, reviewing system and application logs, and verifying completion of automated processes.
Perform ongoing performance tuning, infrastructure upgrades, and resource optimization as required to maintain efficiency and reliability.
Provision and configure hardware, peripherals, services, settings, directories, storage, etc. in accordance with standards and project/operational requirements to ensure consistency and compliance.
Research, develop, and implement innovative and where possible automated approaches for system administration tasks to reduce manual effort and increase scalability.
Provide Tier II support for incidents and requests from various constituencies, investigating and troubleshooting issues to resolution in a timely manner.
Maintain and contribute to our knowledge base and documentation to capture procedures, configurations, and lessons learned for future reference.
Influence broader technology groups in adopting Cloud technologies, processes, and best practices to promote standardization and continuous improvement.
Exercise creative thinking and propose solutions to grow our business by delighting our clients and enhancing their overall experience with our platform.
May perform other duties as assigned to support team objectives and organizational priorities as needed.
Complies with all policies and standards, ensuring that operational activities align with regulatory, security, and governance requirements.
Supports the implementation of infrastructure changes within defined processes to minimize risk and maintain stability.
Assists in the evaluation and integration of new tools and technologies that improve operational efficiency and service delivery.
Collaborates with cross-functional teams to identify bottlenecks and implement corrective actions that improve system performance.
Contributes to disaster recovery and business continuity initiatives by maintaining and testing backup procedures and configurations.
Participates in on-call rotations as required to provide timely response to production incidents and service disruptions.
Requirements
2+ years of professional experience with AWS or an AWS Certification to demonstrate foundational cloud knowledge and practical exposure.
Bachelor degree in Computer Science, MIS, or equivalent professional experience to provide a baseline understanding of technology and systems concepts.
2+ years of Linux systems administration and configuration experience to ensure proficiency in managing server environments.
Strong scripting skills - Python / Bash preferred - to automate tasks and manipulate system processes effectively.
Experience with best practices for Cloud Networking and Security to protect data and infrastructure integrity.
Experience with monitoring and alerting solutions to proactively identify and resolve potential issues before they impact users.
Experience with DNS and web server configurations to support application delivery and availability.
Ability to use AI to quickly come up to speed on large complex codebases and leverage intelligent tools for productivity enhancement.
Knowledge of automating infrastructure builds utilizing tools such as Gitlab, Terraform and CDK to enable reliable and repeatable deployments.
Knowledge of containers and orchestration tools such as ECS or Kubernetes to manage scalable and resilient applications.
Knowledge of CI / CD pipelines and Agile methodologies to support iterative development and continuous delivery.
Highly motivated, Innovative, self-directed thinker with an eagerness to stay up to date with current trends and a desire to impress stakeholders with results.
Excellent written and verbal communication skills to facilitate collaboration and convey technical concepts to diverse audiences.
Strong troubleshooting and problem-solving skills to diagnose issues and implement effective solutions under pressure.
Ability to thrive in a fast-paced, innovative environment and adapt to changing priorities and emerging business needs.
Practical notes
This role is eligible to participate in the Businessolver Shared Equity Program.
This position is remote within the United States.
This position may be filled within 30 days of posting.
This role is eligible for the Businessolver Shared Equity Program.