The Challenge
OneTrust is seeking a Staff Software Engineer to join the Reporting and Data Platform team. This is a hands-on individual contributor role focused on designing, building, operating, and improving distributed backend services and data-processing platforms.
You will work across Java microservices and Python/PySpark data pipelines, with a strong focus on reliability, scalability, performance, and observability. You will take complex or ambiguous problems from investigation through production delivery and help improve the systems that power reporting and data-driven experiences.
Your Mission
Technical Ownership and Delivery
- Own complex features and technical improvements from discovery through production rollout, making substantial hands-on contributions across backend services and data-processing pipelines.
- Investigate ambiguous problems, identify root causes, evaluate trade-offs, and implement pragmatic solutions that improve code quality, maintainability, automated testing, and operational readiness.
- Review code and technical designs, document important implementation decisions and system behavior, and partner with product managers, engineers, and other teams to clarify requirements and deliver outcomes.
- Apply AI-assisted engineering tools such as Devin, Claude, or similar systems to accelerate delivery while maintaining production-quality design, code, tests, security, and operational readiness.
Backend and Distributed Systems
- Design and implement production services using Java, Spring Boot, and Maven, including APIs, asynchronous workflows, report generation, aggregation, export, and scheduling capabilities.
- Develop event-driven functionality using Kafka and related messaging patterns, and work with caching technologies, relational storage, and service-to-service integrations.
- Improve service performance, scalability, fault tolerance, and resource efficiency through appropriate patterns for retries, idempotency, caching, backpressure, concurrency, and failure recovery.
- Diagnose issues across services, queues, databases, and downstream dependencies, and modernize established capabilities incrementally while maintaining production stability.
Data Engineering
- Build and maintain ingestion and transformation pipelines using Python, PySpark, Azure Databricks, and Delta Lake across batch and streaming workloads.
- Implement schema evolution, checkpoint management, deduplication, replay, late-arriving-data handling, and standardized data-layer patterns.
- Optimize Spark joins, partitioning, Delta operations, cluster utilization, and query performance while troubleshooting failed, delayed, or inefficient Databricks workloads.
- Protect tenant boundaries across joins, aggregations, deduplication, and Delta operations; implement data-quality controls; and monitor data freshness, completeness, and correctness.
- Work securely with Azure storage, identities, secrets, and encryption mechanisms.
Observability, On-Call, and Operational Excellence
- Improve observability across backend services, event-driven workflows, and data pipelines using meaningful metrics, structured logs, traces, and business telemetry.
- Build and maintain actionable dashboards, monitors, and alerts using Datadog and Grafana, applying OpenTelemetry, Prometheus, and Micrometer patterns where appropriate.
- Participate in the on-call rotation and incident-response workflows, using PagerDuty, Datadog monitors, or equivalent platforms to diagnose production issues and drive sustainable resolution.
- Reduce recurring alerts and operational toil by improving alert quality, eliminating noisy or non-actionable monitors, creating runbooks and diagnostic tools, and implementing corrective actions from blameless incident reviews.
- Improve end-to-end correlation and monitor availability, error rates, latency, ingestion lag, data freshness, event throughput, consumer lag, job health, rejected records, checkpoint health, tenant-specific failures, data-quality violations, and Spark resource utilization.
What Success Looks Like
- You require limited direction after understanding the desired outcome and relevant constraints, and you break ambiguous problems into concrete, deliverable work.
- You own work through design, implementation, testing, deployment, production validation, and ongoing operation.
- You use production evidence and telemetry to prioritize improvements and resolve root causes rather than repeatedly treating symptoms.
- You reduce alert volume and operational toil over time without hiding genuine system risks, leaving systems easier to operate after each incident.
- You make sound trade-offs among delivery speed, reliability, performance, security, cost, and maintainability while collaborating constructively without formal authority.
You Are
You are a self-directed, hands-on Staff Engineer who enjoys solving complex problems across distributed services and data platforms. You think in systems and trade-offs, take ownership of production behavior, and use clear design thinking to simplify solutions and reduce code-delivery cycles. You are motivated by building reliable, secure, and maintainable systems and by improving them over time.
- Comfortable working across service, platform, data, and partner-team boundaries without needing formal authority.
- Pragmatic about when to build, reuse, or modernize, with sound judgment around reliability, performance, security, cost, and maintainability.
- Committed to test-driven development, early validation, and production-quality engineering practices.
- Thoughtful about using AI-assisted development tools to accelerate implementation while preserving engineering judgment and accountability.
- Focused on reducing recurring failure modes, alert noise, and operational burden through engineering improvements.
Your Experience Includes
Required
- Strong professional experience building and operating production software systems as a highly autonomous individual contributor.
- Strong proficiency in Java and Spring Boot, with experience designing and operating distributed systems and microservices.
- Production experience with asynchronous or event-driven systems, preferably Apache Kafka.
- Strong experience with Python, PySpark, Apache Spark, and Delta Lake, plus production experience with Azure Databricks or a comparable managed Spark platform.
- Hands-on experience with test-driven development, automated testing strategies, and quality gates that support fast, reliable delivery.
- Strong understanding of metrics, logs, distributed tracing, dashboards, monitoring, and alerting, including hands-on experience with Datadog and Grafana.
- Experience creating or responding to PagerDuty incidents, Datadog alerts, or equivalent production alerting workflows, and willingness to participate in an on-call rotation.
- Experience using AI engineering tools such as Devin, Claude, or similar systems to produce production-ready code, tests, documentation, and operational improvements.
- Strong design-thinking skills and the ability to reduce delivery-cycle time through clear architecture, smaller increments, reusable patterns, and pragmatic technical trade-offs.
- Ability to independently diagnose complex performance and reliability problems and communicate implementation decisions and technical trade-offs clearly.
Preferred
- Experience with both batch and streaming data pipelines and with optimizing Spark or Databricks workloads for performance, reliability, and cost.
- Experience with Databricks SQL, Databricks SDKs, Delta operations, and schema migrations.
- Familiarity with Azure Blob Storage, Azure Identity, and Azure Key Vault.
- Experience operating reporting, analytics, dashboard, or large-scale export systems.
- Experience with Kubernetes, containers, CI/CD, and infrastructure as code.
- Experience defining or applying service-level indicators, service-level objectives, and error budgets, and using incident and alert trends to prioritize engineering work.
- Understanding of data governance, encryption, audit-ability, and tenant isolation.
- Experience modernizing established production systems incrementally.