Head of Data Infrastructure
Job description
Head of Data Infrastructure at Beacon Software.
About the role
You will own the technical roadmap for Beacon's data infrastructure across a portfolio of more than twenty-five software companies, designing the foundational lakehouse, warehouse, and pipeline architecture that connects diverse source systems under multi-tenant constraints. You will define the ingestion operating model so that onboarding a new portco is fast, repeatable, and resilient to messy source data and varying maturity levels. You will set modeling and transformation standards, including dbt, Spark, and lineage practices, to ensure a canonical model that unifies analytics across verticals instead of fragmenting it. You will establish real-time and batch processing frameworks, choosing streaming or batch appropriately for operational analytics, AI feature engineering, and executive reporting. You will architect and own multi-tenant isolation, security boundaries, and regional residency for regulated verticals, including KMS scoping, IAM design, and auditability. You will manage infrastructure-as-code for the entire data stack, writing and maintaining Terraform or Pulumi modules that keep provisioning, cost, and observability reliable and repeatable. You will partner closely with AI and product teams to build feature stores and vector infrastructure, ensuring the data layer is ready for low-latency serving and embedded intelligence. You will lead vendor stack evaluation between Snowflake and Databricks, make buy vs build decisions, and act as the senior technical voice during acquisition diligence to assess data maturity and integration risk.
Key facts
What you'll do
Architect the central data platform for Beacon's portfolio, designing a lakehouse, warehouse, and pipeline architecture that survives contact with very different source systems, data quality issues, and portco maturity levels.
Define and enforce ingestion strategy that connects SaaS apps, transactional databases, Stripe, Salesforce, and custom APIs into a low-latency data layer with an operating model that makes onboarding portco 50 as fast as portco 5.
Establish data modeling and transformation standards using dbt, Spark, or equivalent tools, creating a canonical model that enables cross-portco queries such as "show me sales across all portcos" while preventing platform fragmentation.
Determine the appropriate mix of real-time and batch processing frameworks for operational analytics, AI feature engineering, and executive reporting, with the discipline to know which workloads belong to which paradigm.
Implement multi-tenant isolation and security posture, including per-portco data, compute, and credential boundaries, cross-cloud connectivity across AWS and Azure, regional residency for regulated verticals, KMS scoping, IAM design, and defensible audit surfaces.
Own infrastructure-as-code for the data stack, writing and maintaining Terraform or Pulumi modules that manage cloud provisioning, cost management, and observability for portability and repeatability.
Partner with AI and product teams to build feature stores and vector infrastructure, ensuring the underlying data layer is performant and reliable for low-latency serving and intelligence layers as products embed AI capabilities.
Evaluate and own the vendor stack for cloud warehouses, orchestration, cataloging, and BI tooling, guiding buy vs build decisions with long-term leverage in mind and leading the active evaluation between Snowflake and Databricks.
Serve as the senior technical voice in acquisition diligence, assessing target companies' data maturity, integration complexity, and risk to reduce surprises during post-close integration and maximize leverage in pre-close decisions.
Set patterns for operational excellence, reliability, cost control, and security across the portfolio, ensuring the platform compounds value instead of fragmenting as the portfolio grows.
Define and drive a roadmap for data platform scalability, balancing standardization with the flexibility required by very different portcos and regulatory constraints.
Establish observability, monitoring, and alerting for data pipelines and infrastructure to quickly surface issues and maintain trust across product and executive stakeholders.
Lead cross-functional collaboration with product, engineering, and compliance teams to align data capabilities with business outcomes and regulatory requirements in regulated verticals.
Mentor and shape the data infrastructure culture, defining practices and guardrails so that future hires can build on a coherent platform instead of repeating foundational work.
Requirements
Hands-on background in distributed systems, query engines, or cloud data platforms with the ability to write production Python, SQL, and infrastructure-as-code such as Terraform or Pulumi.
Deep familiarity with modern lakehouse and warehouse internals including Iceberg, Delta, Parquet, Snowflake, and Databricks, along with strong grasp of streaming architectures like Kafka, Kinesis, and Flink and CDC patterns.
Experience building for both analytical and operational workloads at scale, with an understanding of when to apply streaming versus batch to achieve the right balance of freshness, cost, and reliability.
Strong opinions about data architecture that are loosely held, capable of balancing idealism with the practical constraints of messy source systems and evolving portco requirements.
Comfort operating in small-team, high-autonomy environments without heavyweight process, while still being rigorous about correctness, repeatability, and long-term maintainability.
Clear written and verbal communication skills to translate complex technical designs into decisions that non-technical stakeholders can understand and support.
Proven ability to work independently and make high-impact technical decisions in ambiguous environments while owning outcomes across a multi-tenant portfolio.
Commitment to building for the long term, with the judgment to know when to standardize and when to allow controlled variation across portcos to maximize compounding value.
U.S. employment eligibility is required, and candidates must be able to work full-time in San Francisco, CA.
Nice to have
Experience with regulated verticals that have data residency requirements and familiarity with compliance considerations in those markets.
Background operating at scale with SaaS portfolios where multi-tenant isolation, uptime, and reliability are business-critical.
Practical notes
This role is full-time based in San Francisco, CA.
Candidates must be able to work onsite in San Francisco, CA.
No sponsorship is available for this role.