Senior Platform Engineer
Job description
About the role
This position is for a Senior Platform Engineer who will own the design and evolution of the core developer platform at Kojo. You will build and maintain the golden paths and self-service abstractions that allow product teams to move quickly without sacrificing reliability. The role focuses on infrastructure as code, cloud architecture, and platform automation rather than traditional application feature work. You will shape how internal tools are consumed and standardize best practices across engineering organizations. Collaboration with product and infrastructure teams is central to ensuring the platform scales with business needs. You will have direct influence on the long-term technical strategy for Kojo's cloud foundation. This is an opportunity to define platform standards from the ground up in a fast-growing environment.
Key facts
What you'll do
- Design and maintain typed infrastructure component libraries using Pulumi and Terraform to standardize multi-environment stack deployments.
- Build and evolve the EKS platform by configuring Karpenter autoscaling and GitOps delivery pipelines with ArgoCD, Kargo, and Helm.
- Create and optimize CI/CD workflows in GitHub Actions, including release automation and secure promotion across environments.
- Develop developer-experience tooling that integrates AI-driven workflows, including skills, hooks, MCP servers, and managed agents.
- Implement observability standards with OpenTelemetry and Datadog, ensuring vendor-agnostic telemetry across services.
- Strengthen network and edge security through CloudFront, WAF, ACM certificates, and VPC segmentation strategies.
- Automate SOC 1 and SOC 2 evidence collection and monitoring to support compliance and audit readiness.
- Partner with product teams to deliver reliable backend services using Aurora PostgreSQL, Prisma, and Slonik for optimized data access.
- Operate and scale DynamoDB and other NoSQL databases to meet performance and availability goals.
- Coordinate event-driven architectures using EventBridge and SQS to build resilient and idempotent systems.
- Lead production Kubernetes operations on EKS, including node management, VPC CNI with NetworkPolicy, and multi-cluster strategies.
- Guide platform adoption by creating documentation, samples, and internal enablement content for engineering teams.
- Evaluate and integrate emerging tools for AI-assisted development and infrastructure management.
- Continuously improve platform reliability, cost efficiency, and security through measurable outcomes and feedback loops.
Requirements
- Demonstrate strong AWS services knowledge across compute, networking, storage, and security services.
- Show advanced proficiency in JavaScript, TypeScript, and Node.js for building platform tools and integrations.
- Apply solid grasp of parallel programming concepts in real-world distributed systems scenarios.
- Execute deep PostgreSQL and ORM optimization techniques using Prisma, Slonik, and similar libraries.
- Use NoSQL databases such as Redis, DynamoDB, and Elasticsearch in production environments.
- Practice strong system design and microservice architecture with event-driven, idempotent consumers and fanout patterns.
- Work with message brokers and streaming platforms to enable reliable asynchronous communication.
- Leverage infrastructure as code tools such as Terraform, Pulumi, or AWS CDK to manage cloud resources.
- Build and maintain RESTful APIs that are scalable, well-documented, and secure.
- Support production Kubernetes workloads, especially on EKS, with focus on networking, node groups, and multi-cluster setups.
- Apply observability standards using OpenTelemetry and APM platforms in vendor-neutral implementations.
- Perform network segmentation using VPCs, subnets, security groups, firewall rules, and DNS configurations.
- Harden edge and network security with CloudFront, WAF, ACM certificates, and layered VPC protections.
- Configure pipeline infrastructure using GitHub Actions and related CI/CD tooling.
- Work effectively with AI-driven development tools like Cloude, Codex, Cursor, and related frameworks.
Nice to have
- Experience with GraphQL or gRPC for platform service design.
- Hands-on GitOps practice with ArgoCD, Kargo, and Helm in complex environments.
- Relevant certifications in AWS, CNCF, Linux Foundation, or similar programs.
- Background with security tooling such as SAST, Software Composition Analysis, IAST, SIEM, and threat management.
- Deep engagement with SRE and reliability practices including SLOs, error budgets, incident response, and runbooks.
- Knowledge of FinOps and cost management techniques for EKS, including right-sizing and Savings Plans.
- Experience automating compliance workflows for SOC 2 evidence, IAM Identity Center, and least-privilege access models.
Practical notes
This role is full-time and based in Mexico. Travel requirements and visa sponsorship details are not specified in the available information. Applicants should review deadlines and submission instructions through official Kojo hiring channels when applying.