Principal Software Architect - Data Platform

SecurityScorecard · Hybrid (Austin, TX) · Engineering

Posted 2026-09-11

Apply for this role →

About the Role:

SecurityScorecard is hiring a Principal Software Architect to lead the system design of our data platform. Rating 12 million companies continuously means ingesting internet-scale measurement data, processing it across streaming, microbatch, and batch paths, storing it so it stays queryable and affordable as it grows, and serving analytics fast enough that customers can explore their own risk in real time. The data is not a byproduct of our product. It is the product.

That also raises the stakes on correctness. We publish a number about other companies, they dispute it, and underwriters price against it. A quiet data quality regression here doesn't produce a stale dashboard, it moves someone's score. Quality, contracts, and lineage are therefore architecture problems on this platform, not administrative ones.

This is an individual contributor role reporting to the Chief Architect, alongside a Principal Architect focused on AI and agentic systems and a Principal Front-end Architect, and partnering closely with engineering leadership, Product, and Data Science.

We're looking for someone who is opinionated about data architecture and persuasive about it: an architect whose designs get adopted because the reasoning is visible, not because they carry a title.

Like the rest of our architecture function, you'll prototype to prove out decisions rather than implement full solutions, and you'll set direction through Technical Design Reviews (TDRs), and standards.  You will lead the data domain, and you'll bring enough general distributed systems judgment to review designs across the wider platform.

What You'll Do

Own the system design for our data platform end to end, from ingestion through to the serving layer

Define the service boundaries and data contracts between producers and consumers, including schema ownership, compatibility rules, and what happens when a producer needs to make a breaking change

Design the lakehouse: table format, partitioning strategy, schema evolution, compaction, and metadata growth at scale

Architect the analytical serving layer for three workload classes with conflicting demands, isolated so that one never degrades another: low-latency, high-concurrency queries from customers in the product, ad-hoc exploration from internal analytics and Data Science, and bulk delivery to external feeds and partners

Set direction on languages and frameworks in the data stack

Engineer data quality and observability into the platform rather than bolting them on: validation and quarantine paths, freshness and completeness SLOs, drift detection, and lineage and metadata generated by the pipeline itself instead of maintained by hand

Design for correctness and reproducibility in the ratings pipeline, including backfills and historical restatement when scoring logic changes

Write the TDRs, design docs, and standards that set data architecture direction across teams, and push that intent into the repos themselves so engineers and coding agents both have it in local context

Review TDRs from across engineering, giving teams substantive feedback on architecture and risk, not only on data work

Partner with the AI & Front End Architects on the data access patterns, mentor senior and staff engineers on data system design

Required Qualifications:

10+ years of software or data engineering experience, including significant time architecting large-scale data platforms

Deep expertise in stream and batch processing at scale with Kafka, Flink, and Spark or close equivalents, and clear judgment about which path a given workload belongs in

Strong Python and PySpark, solid Java for Flink stream processing, and enough Scala to read and reason about an existing Spark codebase

Hands-on experience designing lakehouse storage in production: columnar formats such as Parquet, open table formats such as Iceberg, and the partitioning, compaction, and schema evolution decisions that come with them

Experience architecting OLAP and analytical serving layers (ClickHouse, Druid, Pinot, BigQuery, Snowflake or similar)

Strong distributed systems fundamentals as they apply to data: exactly-once versus at-least-once semantics, ordering, backpressure, late and out-of-order data, and pipeline failure modes

A track record of building data quality, contracts, and observability as engineered system properties, meaning assertions, schema enforcement, and lineage that live in code

Experience owning a large-scale data migration, including preserving history and correctness through a cutover

A track record of influence without authority: presenting a technical direction to skeptical engineers and earning genuine buy-in, and giving rigorous design review feedback on systems you didn't build yourself

Strong technical writing and mentorship: TDRs, design docs, and decision records that teams can act on without hand-holding, plus a history of raising the technical bar around you

Comfort operating as a senior individual contributor, driving outcomes through prototyping and technical credibility

Preferred Qualifications:

Experience building the data layer underneath ML or LLM systems: feature stores, vector stores, or retrieval pipelines

Experience with internet-scale scan, telemetry, or observability data, or a cybersecurity industry background

Track record of bringing platform cost down at scale through storage tiering, query governance, or compute right-sizing

Our Tech Stack

Our platform runs on Node.js and TypeScript, with a React Microfrontend Architecture and PostgreSQL and ClickHouse for storage. We use Kafka for event streaming, and our infrastructure runs on AWS with Kubernetes, Terraform, Helm, and ArgoCD. We are actively expanding our AI/ML infrastructure.

On the data side, Kafka is the backbone for event flow. Stream processing runs on Flink in Java, while batch and microbatch run on Spark. Some high performance pipeline components are written in C++. ClickHouse serves our analytical workloads.

You do not need to have used every tool here, but you should be comfortable reasoning across a stack of this kind and making principled architectural trade-offs within it.

Benefits:

Specific to each country, we offer a competitive salary, stock options, Health benefits, and unlimited PTO, parental leave, tuition reimbursements, and much more!

The estimated total compensation range for this position is $240,000 - $300,000 (base plus bonus). Actual compensation for the position is based on a variety of factors, including, but not limited to affordability, skills, qualifications and experience, and may vary from the range. In addition to base salary, employees may also be eligible for annual performance-based incentive compensation awards and equity, among other company benefits.

Apply for this role →

← Back to all jobs