Data Platform Engineer
About the team & role
The Data Platform team builds and operates the shared infrastructure that powers how data is stored, moved, and consumed across Alpaca. Our platform supports financial transactions, customer data, API and system events, enriched datasets, and third-party sources that power critical decisions and products for internal teams and external stakeholders. We process hundreds of millions of events each day and that scale continues to grow as Alpaca serves more customers and launches new products.
As a Data Platform Engineer, you will build and operate reliable systems across Alpaca’s data platform. You’ll contribute to our distributed query, lakehouse, CDC, and real-time streaming infrastructure while collaborating with experienced engineers and partner teams. You’ll own projects from implementation through production operation and help improve the platform’s scalability, reliability, and developer experience.
Our team is 100% distributed and remote.
Responsibilities:
Design, develop, and evolve Alpaca’s core data platform, including distributed query engines, orchestration, lakehouse storage, metadata, and cataloging.
Own our lakehouse infrastructure as code, building reliable and repeatable deployment workflows with Terraform and Ansible on Kubernetes.
Build and operate scalable, low-latency streaming and CDC pipelines alongside dependable batch ingestion paths into Apache Iceberg.
Develop and scale our serving and BI infrastructure, giving downstream teams and agents performant, governed, self-service access to lakehouse data.
Take end-to-end ownership of platform reliability through observability, actionable alerting, on-call practices, incident response, maintenance procedures, runbooks, and service-level objectives.
Partner with DevOps, Analytics Engineering, and stakeholders across Alpaca to close infrastructure gaps, shape technical requirements, and deliver reusable platform capabilities.
Must-Haves:
4+ years of experience in software engineering, data infrastructure, or platform engineering, with demonstrated ownership of complex production systems and technical projects.
Strong understanding of distributed systems and production experience with query engines such as Trino or Presto.
Strong experience operating Kubernetes-based data infrastructure using infrastructure-as-code and deployment tools such as Terraform, Helm, Ansible, or Argo CD.
Experience designing and operating compute platforms across multiple cloud regions, balancing workload locality, peak and bursty demand, resource isolation, workload prioritization, and cost-performance trade-offs.
Hands-on experience designing and operating cloud lakehouse architectures on GCP, AWS, or Azure, including object storage, metastores, and open table formats, particularly Apache Iceberg.
Experience operating large-scale streaming and CDC systems using technologies such as Kafka, Redpanda, and Debezium.
Strong Python and SQL skills, with experience building reliable data pipelines and platform tooling using orchestration and ingestion frameworks such as Airflow and Airbyte.
Strong operational judgment and the ability to make sound technical decisions, communicate trade-offs, and work effectively through ambiguity.
Nice to Haves:
Experience with semantic or metrics layers such as Cube, dbt, or Looker.
Experience deploying and operating infrastructure for AI agents.
Experience building or integrating context layers that connect structured and unstructured organizational data.
Experience with data governance and access-control frameworks such as Apache Ranger.