Staff Machine Learning Engineer
OKXUSA4w ago
Job description
About the role
OKX is hiring a Staff Machine Learning Engineer to build and operate ML systems that detect fraud, account takeovers, scams, and other financial risks across a global crypto exchange. You will work with data spanning on-chain activity, fiat transactions, trading behavior, device signals, and identity information. The role covers the full ML lifecycle from problem definition through real-time deployment and continuous iteration. You will also help define how AI and LLM-based agents are used safely and effectively across the risk organization.
Key facts
What you'll do
- Create and deploy ML models for payment fraud, account takeover, scam detection, deposit and withdrawal risk, promotional abuse, customer risk scoring, and transaction monitoring
- Own production ML systems end to end, covering feature pipelines, training workflows, model serving, decision integrations, monitoring, alerting, drift detection, retraining, and incident response
- Collaborate with risk strategy and product teams to convert model outputs into production controls such as approvals, rejections, reviews, cooldowns, limit changes, and account restrictions
- Work with risk operations to understand investigation workflows, gather reviewer feedback, improve explainability, and refine labels and training data
- Use LLM coding tools throughout the engineering workflow to speed up implementation, testing, debugging, analysis, and documentation while following security and review standards
- Build AI-powered risk capabilities including investigation agents, case summarization, evidence collection, review recommendations, alert triage, suspicious-entity mining, and automated decision support
- Move research-stage models into reliable production by validating feature logic, reviewing data quality, addressing latency and scalability constraints, and ensuring consistency between offline training and online inference
- Ensure models and decision systems are explainable, traceable, and documented for risk operations, product stakeholders, governance teams, and regulators
- Design and deploy LLM-based agents for case triage, evidence retrieval, transaction analysis, alert summarization, review recommendations, and automated action orchestration
- Build production-grade agent architectures using tool calling, retrieval-augmented generation, workflow orchestration, structured outputs, memory, guardrails, and human-in-the-loop controls
- Develop evaluation frameworks for LLM agents measuring factual accuracy, task completion, decision consistency, latency, cost, reviewer acceptance, and operational impact
- Implement permission controls, audit logs, data privacy protections, prompt and tool security, fallback mechanisms, and escalation paths for safe LLM agent operation
Requirements
- Significant professional experience in ML engineering, applied data science, or a related field with a track record of taking models from prototype to production
- Strong Python skills and hands-on experience with ML frameworks such as PyTorch, TensorFlow, XGBoost, LightGBM, or scikit-learn
- Solid understanding of applied ML fundamentals including supervised learning, anomaly detection, representation learning, class-imbalanced modeling, model calibration, and evaluation under shifting data distributions
- Demonstrated fluency with AI-assisted engineering, regular use of LLM coding tools, and understanding of productivity benefits alongside security, reliability, and governance risks
- Familiarity with model explainability techniques such as SHAP, feature attribution, reason-code generation, and model scorecards
- Hands-on experience designing and deploying production LLM agents including agentic workflows, tool calling, retrieval-augmented generation, prompt and context management, structured output generation, and multi-step task orchestration
- Experience integrating LLM agents with internal systems, APIs, databases, search tools, case-management platforms, or decision engines
- Strong understanding of LLM-agent evaluation and reliability including hallucination control, grounding, observability, permissions, failure handling, human review, latency, and cost optimization
- Strong communication and collaboration skills for working with engineers, data scientists, risk specialists, product managers, operations teams, and legal or compliance stakeholders
Nice to have
- Experience building AI agents for fraud, risk, compliance, customer operations, cybersecurity, or other high-stakes domains
Skills & tools
- Python, PyTorch, TensorFlow, XGBoost, LightGBM, scikit-learn
- SHAP, feature attribution, model scorecards
- LLM agent frameworks, tool calling, retrieval-augmented generation, structured outputs
Practical notes
- Base salary range: $214,666 to $321,999, depending on skills, experience, and market location
- Compensation may include performance bonus and long-term incentives plus medical, financial, and other benefits
- Apply via Okcoin or OKX internal or external careers site
- Equal opportunity employer; considers qualified applicants with arrest and conviction records per San Francisco Fair Chance Ordinance