Staff Software Engineer - Reporting, Data Platform & Observability
The Challenge
OneTrust is seeking a Staff Software Engineer to join the Reporting and Data Platform team. This is a hands-on individual contributor role focused on designing, building, operating, and improving distributed backend services and data-processing platforms.
You will work across Java microservices and Python/PySpark data pipelines, with a strong focus on reliability, scalability, performance, and observability. You will take complex or ambiguous problems from investigation through production delivery and help improve the systems that power reporting and data-driven experiences.
Your Mission
Technical Ownership and Delivery
Own complex features and technical improvements from discovery through production rollout, making substantial hands-on contributions across backend services and data-processing pipelines.
Investigate ambiguous problems, identify root causes, evaluate trade-offs, and implement pragmatic solutions that improve code quality, maintainability, automated testing, and operational readiness.
Review code and technical designs, document important implementation decisions and system behavior, and partner with product managers, engineers, and other teams to clarify requirements and deliver outcomes.
Apply AI-assisted engineering tools such as Devin, Claude, or similar systems to accelerate delivery while maintaining production-quality design, code, tests, security, and operational readiness.
Backend and Distributed Systems
Design and implement production services using Java, Spring Boot, and Maven, including APIs, asynchronous workflows, report generation, aggregation, export, and scheduling capabilities.
Develop event-driven functionality using Kafka and related messaging patterns, and work with caching technologies, relational storage, and service-to-service integrations.
Improve service performance, scalability, fault tolerance, and resource efficiency through appropriate patterns for retries, idempotency, caching, backpressure, concurrency, and failure recovery.
Diagnose issues across services, queues, databases, and downstream dependencies, and modernize established capabilities incrementally while maintaining production stability.
Data Engineering
Build and maintain ingestion and transformation pipelines using Python, PySpark, Azure Databricks, and Delta Lake across batch and streaming workloads.
Implement schema evolution, checkpoint management, deduplication, replay, late-arriving-data handling, and standardized data-layer patterns.
Optimize Spark joins, partitioning, Delta operations, cluster utilization, and query performance while troubleshooting failed, delayed, or inefficient Databricks workloads.
Protect tenant boundaries across joins, aggregations, deduplication, and Delta operations; implement data-quality controls; and monitor data freshness, completeness, and correctness.
Work securely with Azure storage, identities, secrets, and encryption mechanisms.
Observability, On-Call, and Operational Excellence
Improve observability across backend services, event-driven workflows, and data pipelines using meaningful metrics, structured logs, traces, and business telemetry.
Build and maintain actionable dashboards, monitors, and alerts using Datadog and Grafana, applying OpenTelemetry, Prometheus, and Micrometer patterns where appropriate.
Participate in the on-call rotation and incident-response workflows, using PagerDuty, Datadog monitors, or equivalent platforms to diagnose production issues and drive sustainable resolution.
Reduce recurring alerts and operational toil by improving alert quality, eliminating noisy or non-actionable monitors, creating runbooks and diagnostic tools, and implementing corrective actions from blameless incident reviews.
Improve end-to-end correlation and monitor availability, error rates, latency, ingestion lag, data freshness, event throughput, consumer lag, job health, rejected records, checkpoint health, tenant-specific failures, data-quality violations, and Spark resource utilization.
What Success Looks Like
You require limited direction after understanding the desired outcome and relevant constraints, and you break ambiguous problems into concrete, deliverable work.
You own work through design, implementation, testing, deployment, production validation, and ongoing operation.
You use production evidence and telemetry to prioritize improvements and resolve root causes rather than repeatedly treating symptoms.
You reduce alert volume and operational toil over time without hiding genuine system risks, leaving systems easier to operate after each incident.
You make sound trade-offs among delivery speed, reliability, performance, security, cost, and maintainability while collaborating constructively without formal authority.
You Are
You are a self-directed, hands-on Staff Engineer who enjoys solving complex problems across distributed services and data platforms. You think in systems and trade-offs, take ownership of production behavior, and use clear design thinking to simplify solutions and reduce code-delivery cycles. You are motivated by building reliable, secure, and maintainable systems and by improving them over time.
Comfortable working across service, platform, data, and partner-team boundaries without needing formal authority.
Pragmatic about when to build, reuse, or modernize, with sound judgment around reliability, performance, security, cost, and maintainability.
Committed to test-driven development, early validation, and production-quality engineering practices.
Thoughtful about using AI-assisted development tools to accelerate implementation while preserving engineering judgment and accountability.
Focused on reducing recurring failure modes, alert noise, and operational burden through engineering improvements.
Your Experience Includes
Required
Strong professional experience building and operating production software systems as a highly autonomous individual contributor.
Strong proficiency in Java and Spring Boot, with experience designing and operating distributed systems and microservices.
Production experience with asynchronous or event-driven systems, preferably Apache Kafka.
Strong experience with Python, PySpark, Apache Spark, and Delta Lake, plus production experience with Azure Databricks or a comparable managed Spark platform.
Hands-on experience with test-driven development, automated testing strategies, and quality gates that support fast, reliable delivery.
Strong understanding of metrics, logs, distributed tracing, dashboards, monitoring, and alerting, including hands-on experience with Datadog and Grafana.
Experience creating or responding to PagerDuty incidents, Datadog alerts, or equivalent production alerting workflows, and willingness to participate in an on-call rotation.
Experience using AI engineering tools such as Devin, Claude, or similar systems to produce production-ready code, tests, documentation, and operational improvements.
Strong design-thinking skills and the ability to reduce delivery-cycle time through clear architecture, smaller increments, reusable patterns, and pragmatic technical trade-offs.
Ability to independently diagnose complex performance and reliability problems and communicate implementation decisions and technical trade-offs clearly.
Preferred
Experience with both batch and streaming data pipelines and with optimizing Spark or Databricks workloads for performance, reliability, and cost.
Experience with Databricks SQL, Databricks SDKs, Delta operations, and schema migrations.
Familiarity with Azure Blob Storage, Azure Identity, and Azure Key Vault.
Experience operating reporting, analytics, dashboard, or large-scale export systems.
Experience with Kubernetes, containers, CI/CD, and infrastructure as code.
Experience defining or applying service-level indicators, service-level objectives, and error budgets, and using incident and alert trends to prioritize engineering work.
Understanding of data governance, encryption, audit-ability, and tenant isolation.
Experience modernizing established production systems incrementally.
For California, Colorado, Connecticut, Nevada, New York, Rhode Island, and Washington-based candidates: the annual base pay range for this role is listed below. Within this range, individual pay is determined by several factors, including location, job-related skills, work experience, and relevant education and/or training. This role may also be eligible for discretionary bonuses, equity, and/or commissions, as well as benefits.
Salary Range
$139,725—$209,588 USD