OpenSearch Engineer
Job description
OpenSearch Engineer at Ruby Labs.
About the role
Join a team dedicated to building a high-volume payment processing platform that handles millions of transactions on a monthly basis. As an OpenSearch Engineer, you will be responsible for managing the entire lifecycle of our AWS OpenSearch infrastructure, including its design, data ingestion processes, performance optimization, and recovery procedures. This role requires a blend of search infrastructure management, data engineering expertise, and production operations skills to ensure the system's reliability, scalability, and efficiency. You will work closely with backend, database, infrastructure, and product teams to develop and maintain a robust search platform that supports critical payment services.
Key facts
What you'll do
- Manage and operate AWS OpenSearch Service in a production environment, ensuring high availability and performance.
- Design and implement indices, mappings, analyzers, templates, aliases, shard configurations, and retention policies tailored for high-volume payment data.
- Optimize search queries, aggregations, filtering, pagination, and bulk indexing operations to achieve low latency and high throughput.
- Troubleshoot and resolve issues such as slow queries, indexing bottlenecks, hot shards, memory pressure, and rejected requests to maintain system health.
- Plan and execute safe upgrades, migrations, reindexing, and capacity scaling activities without disrupting ongoing operations.
- Develop and maintain reliable data pipelines to synchronize PostgreSQL data with OpenSearch, ensuring data consistency and integrity.
- Support data operations including inserts, updates, deletes, backfills, and full reindexing, ensuring minimal downtime and data accuracy.
- Utilize AWS services such as DMS, OpenSearch Ingestion, Lambda, SQS, S3, ECS, or EKS to build, monitor, and optimize data pipelines.
- Implement robust error handling mechanisms, including retries, idempotency, deduplication, replay, and dead-letter queues, to ensure data reliability.
- Monitor ingestion lag, throughput, failed records, and data freshness metrics to proactively identify and resolve issues.
- Create reconciliation processes to detect and correct missing, stale, or inconsistent documents within the search index.
- Ensure that data ingestion processes do not negatively impact PostgreSQL transactional workloads, maintaining overall system stability.
- Monitor cluster health, search and indexing latency, storage utilization, CPU load, JVM pressure, and shard distribution to optimize performance.
- Develop CloudWatch dashboards, alerts, and operational runbooks to facilitate ongoing system monitoring and incident response.
- Design and implement snapshot, restore, backup, and disaster recovery procedures to safeguard data and ensure business continuity.
- Participate actively in production incident response, root cause analysis, and post-mortem reviews to improve system resilience.
- Implement secure access controls using IAM, role-based permissions, encryption, and least-privilege principles to protect sensitive data.
- Collaborate with cross-functional teams to improve system architecture, data models, and operational workflows.
- Assist in designing efficient OpenSearch document models for various payment-related entities such as payments, customers, subscriptions, and transactions.
- Define standards for searchable fields, schema changes, and index versioning to maintain consistency across the platform.
- Document architecture diagrams, data flows, recovery procedures, and operational best practices to support team knowledge sharing.
Requirements
- Proven experience managing OpenSearch or Elasticsearch in a production environment, demonstrating reliability and scalability.
- Practical experience with AWS OpenSearch Service, including deployment, configuration, and performance tuning.
- Experience building and maintaining data pipelines between PostgreSQL and OpenSearch, ensuring data consistency and timeliness.
- Strong understanding of index and shard design principles, including best practices for high-volume data.
- Deep knowledge of mappings, analyzers, and their impact on search performance and accuracy.
- Proficiency with OpenSearch Query DSL for crafting complex search queries and aggregations.
- Experience with aggregations, sorting, filtering, and pagination techniques to optimize search results.
- Skilled in bulk indexing operations and cluster performance optimization to handle large datasets efficiently.
- Understanding of data consistency mechanisms, retries, idempotency, and replay strategies to ensure data integrity.
- Experience managing large datasets with high indexing and query volumes, maintaining system responsiveness.
- Ability to perform safe backfills and reindexing operations with minimal impact on live systems.
- Strong troubleshooting skills, capable of diagnosing and resolving production issues related to search and indexing.
Nice to have
- Experience with AWS DMS or OpenSearch Ingestion for data migration and ingestion workflows.
- Familiarity with AWS Lambda, SQS, S3, ECS, EKS, or Terraform for building and managing infrastructure and data pipelines.
- Experience developing custom ingestion services using Go, Java, or Python to meet specific data processing needs.
- Background working within payments, billing, fintech, or subscription systems to understand domain-specific requirements.
- Experience working with multi-tenant or high-volume search platforms to support scalable architectures.
- Knowledge of analytical systems such as ClickHouse, Tinybird, or similar tools used for data analysis and reporting.
- AWS certifications that demonstrate cloud expertise and best practices.
Skills & tools
- OpenSearch
- Elasticsearch
- AWS OpenSearch Service
- PostgreSQL
- AWS DMS
- OpenSearch Ingestion
- AWS Lambda
- AWS SQS
- AWS S3
- AWS ECS
- AWS EKS
- Terraform
- Go
- Java
- Python
- CloudWatch
- IAM
Practical notes
This position is on-site in Ukraine, and applicants from any country are welcome to apply, provided they are located within approximately +/- 4 hours of Central European Time (CET). The company offers benefits including unlimited paid time off, paid national holidays, and a company-provided MacBook. Employment is structured through a flexible independent contractor agreement. The interview process involves a recruiter screening, a technical interview, and a final interview. Candidates should be prepared to discuss their experience with search infrastructure, data pipelines, and production operations. The role requires strong troubleshooting skills, the ability to work collaboratively across teams, and a focus on system reliability and performance.