Data Engineer, Red Tape Index
Job description
Data Engineer, Red Tape Index at Infinity Constellation.
About the role
You will architect and maintain the data ingestion pipelines that feed our regulatory indices, ensuring the integrity and reproducibility of every public record used in our rankings. This role owns the transformation logic that converts messy government outputs into structured, auditable datasets ready for analysis. You will implement the storage strategies and processing workflows that allow our mathematical models to run efficiently and transparently. You will work directly with raw source files and API streams to validate data quality and document any anomalies for review. Collaboration with quantitative analysts and domain experts will be essential to align the data structures with the evolving needs of our index methodologies. You will leverage modern AI-assisted development practices to iterate quickly on complex data problems while maintaining strict version control. Your work will directly influence the transparency and credibility of the indices that stakeholders use for decision-making.
Key facts
What you'll do
Design and implement scalable data ingestion workflows that pull from diverse government and public sources using Python-driven automation.
Construct robust data validation frameworks that apply business rules and anomaly detection to ensure the reliability of imported datasets.
Develop modular ETL pipelines using Prefect 3 to orchestrate complex data transformations across distributed compute environments.
Maintain version-controlled data schemas and migration scripts with Alembic to guarantee reproducibility across deployments.
Integrate with object storage solutions such as S3 to persist large parquet datasets and intermediate processing artifacts efficiently.
Optimize query performance on Postgres instances by designing strategic indexing and partitioning strategies for time-series regulatory data.
Build interactive exploratory environments using marimo to enable rapid hypothesis testing and stakeholder feedback on data quality.
Apply polars and pyarrow to perform high-throughput data wrangling operations that minimize memory footprint and maximize execution speed.
Implement monitoring and logging mechanisms for data pipelines to detect failures and performance regressions in near real-time.
Collaborate with AI agent frameworks to generate specification documents and automated tests that keep the codebase self-documented.
Evaluate and prototype integrations with large language model tools such as pydantic-ai and AWS Bedrock to automate data documentation tasks.
Perform regular audits of public datasets to track changes in schema, naming conventions, and availability over time.
Contribute to open-source style internal libraries that standardize common tasks like HTTP requests, browser automation, and TLS-fingerprint evasion.
Support the deployment of data services to ECS Fargate clusters using infrastructure as code principles implemented in Terraform.
Champion data security and compliance practices by enforcing strict access controls and encryption standards for sensitive regulatory information.
Requirements
Demonstrate advanced proficiency in Python 3.14, leveraging its standard library and ecosystem packages for data manipulation and scripting.
Utilize uv as the primary package and dependency manager to ensure deterministic builds across development and production environments.
Adopt ruff as the standard linter and formatter, maintaining consistent code style and eliminating preventable syntax issues.
Employ ty for runtime type validation, strengthening the correctness of data transformations and interface contracts.
Showcase experience with polars, constructing efficient dataframe operations that handle large volumes of regulatory records without degradation.
Implement Prefect 3 flows that define clear task dependencies, retry policies, and state management for critical pipeline stages.
Leverage marimo notebooks to create reproducible analysis artifacts that blend code, visualizations, and narrative explanations.
Design database schemas in Postgres, using SQLAlchemy and Alembic to manage evolutionary changes to the data model over time.
Apply strong knowledge of HTTP/2 clients to interact with modern REST APIs while respecting rate limits and concurrency constraints.
Configure TLS-fingerprint evasion techniques where necessary to comply with access policies imposed by certain government endpoints.
Automate browser workflows using browser automation tools to scrape data from legacy portals that lack public APIs.
Deploy and manage containerized applications on ECS Fargate, ensuring high availability and efficient resource utilization.
Apply infrastructure as code practices with Terraform to provision and govern cloud resources in a secure and auditable manner.
Use pyarrow to interchange data between different processing frameworks and storage formats without unnecessary serialization overhead.
Integrate pydantic-ai models to validate and structure semi-structured data extracted from complex government documents.
Experiment with AWS Bedrock services to assess the viability of large language models for automating data quality checks.
Leverage Claude Code and Cursor as part of the development workflow to accelerate coding tasks and reduce manual intervention.
Demonstrate a solid understanding of data indexing strategies that improve the speed of regulatory lookups and aggregation queries.
Maintain strict adherence to security protocols, ensuring that all data transfers are encrypted and credentials are managed appropriately.
Embrace a mindset of continuous learning, adapting to new tools and methodologies introduced by the evolving AI-native development landscape.
Nice to have
Experience in fields such as actuarial science, quantitative research, or data science focused on ranking and index development.
Familiarity with government open data portals and their typical structures for permits, energy, environment, and economics.
Comfort working in an agent-forward development environment where AI tools are used for specification-driven documentation.
Practical notes
This position is fully remote, allowing you to work from anywhere. We offer an opportunity to engage in impactful work at the intersection of artificial intelligence and critical infrastructure regulation, giving you complete ownership of indices from their raw data origins to the published methodologies. You will be part of a small, dynamic team where your assessments of feasibility will directly influence our product offerings. Our development environment is modern and AI-native, designed to foster innovation and efficiency.