
Token Engineer
Job description
Token Engineer at Twelve Labs.
About the role
The Token Engineer at Twelve Labs owns the design and operation of the core infrastructure that measures, controls, and optimizes how large language models and AI coding tools consume tokens across the entire product lifecycle. This role bridges deep infrastructure work with product impact by building the measurement and control systems that make AI usage safe, predictable, and cost-effective. You will be responsible for turning token data from multiple AI providers into actionable insights and guardrails that protect the business while enabling experimentation. The position requires a strong partnership mindset to align with research, product, security, and finance teams to translate ambitious AI strategies into robust technical controls. You will define success metrics for AI usage and implement the automation that enforces those standards without human intervention. Ultimately, this role shapes the governance layer that allows the organization to scale AI adoption responsibly.
Key facts
What you'll do
- Design and implement an instrumentation pipeline that collects token usage, cost, model, latency, errors, retries, user, and team data from multiple AI coding tools and model APIs into a standard data model.
- Build near real-time dashboards for daily usage and cost visibility, automated alerting, and anomaly detection systems while continuously validating alignment between provider billing reports and internal measurement.
- Design and operate an AI gateway and control layer that centrally manages authentication, authorization, usage limits, budget criteria, request rate limiting, model and provider routing, caching, retries, and audit logging.
- Define rational thresholds and guardrails based on historical usage patterns and actual business outcomes to prevent excessive costs and operational risk without hindering productivity or experimentation.
- Design metrics that explain actual efficiency beyond raw token counts or invoice amounts, including success rate per task, latency, failure rates, and retry patterns.
- Build secure, scalable services on AWS and Kubernetes, applying Infrastructure as Code, secret management, role-based access control, and data retention policies.
- Define service level objectives, operational procedures, and failure response playbooks, and build tools to rapidly narrow down root causes and recover from cost spikes or provider outages.
- Provide self-service tools and documentation so that every team can understand their own usage and improvement opportunities, reducing manual verification work.
- Analyze quality, cost, latency, and stability across new models and providers, and propose and implement routing and optimization strategies aligned with specific use cases.
- Translate requirements from diverse stakeholders into technical designs, clarify trade-offs, and drive ambiguous problems to completion as production-ready systems.
Requirements
- Have worked in infrastructure, platform, backend, or engineering roles for five or more years, or have equivalent depth of production environment build and operations experience.
- Have built and operated real services using AWS and Kubernetes, and have created repeatable environments using Infrastructure as Code tools such as Terraform.
- Have developed long-running production services using Python or Go, rather than only internal operation tools.
- Have designed observability data including logs, metrics, and distributed tracing, and have experience finding root causes of outages or performance degradation using data.
- Have experience in one or more areas such as API gateways, proxies, networking, authentication, and authorization management.
- Have designed or operated usage metering, billing, quota allocation, cost allocation, or large-scale event processing systems.
- Understand the trade-offs between security, stability, cost, and developer experience and can make realistic design decisions in ambiguous situations.
- Have experience structuring ill-defined problems and driving initiatives from design through deployment and operations when requirements are not fully specified.
- Have the ability to explain technical concepts clearly to both engineers and non-technical stakeholders and to build consensus across multiple organizations.
- Prefer to automate short-term manual responses and to create system improvements that prevent the same problems from recurring.
Nice to have
- Have built internal platforms for large language model APIs, AI gateways, model serving, or AI coding tools.
- Have built and operated observability platforms using OpenTelemetry, Prometheus, Grafana, Loki, and Tempo.
- Have used FinOps techniques to technically measure and optimize cloud and AI costs, including chargeback and internal billing systems.
- Have implemented request-level usage metering, subscriptions and billing, quotas, rate limiting, and policy engines.
- Have enterprise security experience including OIDC, SSO, SCIM, role-based access control, secret management, and audit logs.
- Have applied anomaly detection, cost forecasting, or capacity planning to time-series data in production.
- Have designed request routing or fallback paths across multiple model providers based on quality, cost, and latency, and handled failure scenarios.
- Have reduced organizational overhead by building developer platform or self-service infrastructure.
Practical notes
- Hours: Full-time
- Travel: Not mentioned
- Visa: Not mentioned
- Deadline: Not mentioned
Token Engineer at Twelve Labs.
About the role
The Token Engineer at Twelve Labs owns the design and operation of the core infrastructure that measures, controls, and optimizes how large language models and AI coding tools consume tokens across the entire product lifecycle. This role bridges deep infrastructure work with product impact by building the measurement and control systems that make AI usage safe, predictable, and cost-effective. You will be responsible for turning token data from multiple AI providers into actionable insights and guardrails that protect the business while enabling experimentation. The position requires a strong partnership mindset to align with research, product, security, and finance teams to translate ambitious AI strategies into robust technical controls. You will define success metrics for AI usage and implement the automation that enforces those standards without human intervention. Ultimately, this role shapes the governance layer that allows the organization to scale AI adoption responsibly.
Key facts
What you'll do
- Design and implement an instrumentation pipeline that collects token usage, cost, model, latency, errors, retries, user, and team data from multiple AI coding tools and model APIs into a standard data model.
- Build near real-time dashboards for daily usage and cost visibility, automated alerting, and anomaly detection systems while continuously validating alignment between provider billing reports and internal measurement.
- Design and operate an AI gateway and control layer that centrally manages authentication, authorization, usage limits, budget criteria, request rate limiting, model and provider routing, caching, retries, and audit logging.
- Define rational thresholds and guardrails based on historical usage patterns and actual business outcomes to prevent excessive costs and operational risk without hindering productivity or experimentation.
- Design metrics that explain actual efficiency beyond raw token counts or invoice amounts, including success rate per task, latency, failure rates, and retry patterns.
- Build secure, scalable services on AWS and Kubernetes, applying Infrastructure as Code, secret management, role-based access control, and data retention policies.
- Define service level objectives, operational procedures, and failure response playbooks, and build tools to rapidly narrow down root causes and recover from cost spikes or provider outages.
- Provide self-service tools and documentation so that every team can understand their own usage and improvement opportunities, reducing manual verification work.
- Analyze quality, cost, latency, and stability across new models and providers, and propose and implement routing and optimization strategies aligned with specific use cases.
- Translate requirements from diverse stakeholders into technical designs, clarify trade-offs, and drive ambiguous problems to completion as production-ready systems.
Requirements
- Have worked in infrastructure, platform, backend, or engineering roles for five or more years, or have equivalent depth of production environment build and operations experience.
- Have built and operated real services using AWS and Kubernetes, and have created repeatable environments using Infrastructure as Code tools such as Terraform.
- Have developed long-running production services using Python or Go, rather than only internal operation tools.
- Have designed observability data including logs, metrics, and distributed tracing, and have experience finding root causes of outages or performance degradation using data.
- Have experience in one or more areas such as API gateways, proxies, networking, authentication, and authorization management.
- Have designed or operated usage metering, billing, quota allocation, cost allocation, or large-scale event processing systems.
- Understand the trade-offs between security, stability, cost, and developer experience and can make realistic design decisions in ambiguous situations.
- Have experience structuring ill-defined problems and driving initiatives from design through deployment and operations when requirements are not fully specified.
- Have the ability to explain technical concepts clearly to both engineers and non-technical stakeholders and to build consensus across multiple organizations.
- Prefer to automate short-term manual responses and to create system improvements that prevent the same problems from recurring.
Nice to have
- Have built internal platforms for large language model APIs, AI gateways, model serving, or AI coding tools.
- Have built and operated observability platforms using OpenTelemetry, Prometheus, Grafana, Loki, and Tempo.
- Have used FinOps techniques to technically measure and optimize cloud and AI costs, including chargeback and internal billing systems.
- Have implemented request-level usage metering, subscriptions and billing, quotas, rate limiting, and policy engines.
- Have enterprise security experience including OIDC, SSO, SCIM, role-based access control, secret management, and audit logs.
- Have applied anomaly detection, cost forecasting, or capacity planning to time-series data in production.
- Have designed request routing or fallback paths across multiple model providers based on quality, cost, and latency, and handled failure scenarios.
- Have reduced organizational overhead by building developer platform or self-service infrastructure.
Practical notes
- Hours: Full-time
- Travel: Not mentioned
- Visa: Not mentioned
- Deadline: Not mentioned