Observability Architect – Dynatrace
Observability Architect – Dynatrace
Role Overview
We are looking for a highly experienced Observability Architect with deep expertise in Dynatrace, application and infrastructure observability, and ServiceNow ITSM integration.
The architect will be responsible for defining and implementing enterprise observability gold standards, establishing best practices for Dynatrace deployment and configuration, and leading the rollout of observability capabilities across applications and infrastructure.
This is a hands-on architecture and implementation role. The individual is expected to architect solutions, perform proof-of-concepts, configure and implement Dynatrace, integrate with ServiceNow, troubleshoot complex observability challenges, and provide technical recommendations to engineering and operations teams.
Key Responsibilities
1. Observability Architecture & Gold Standards
Define and maintain enterprise Observability Gold Standards covering applications, infrastructure, cloud platforms, databases, APIs, Kubernetes, containers, and critical business services.
Establish Dynatrace architecture, deployment, configuration, and monitoring standards.
Define standards for:
Metrics, logs, and traces
Distributed tracing
Application Performance Monitoring (APM)
Infrastructure monitoring
Kubernetes/container monitoring
Real User Monitoring (RUM)
Synthetic monitoring
API monitoring
Service-level objectives (SLOs) and SLIs
Dashboards and management reporting
Alerting and event management
Tagging and entity management
Business application/service mapping
Develop reusable observability patterns and implementation templates.
2. Dynatrace Implementation & Rollout
Lead enterprise Dynatrace implementation and rollout across applications, infrastructure, and cloud environments.
Define deployment strategies for OneAgent, ActiveGate, Kubernetes monitoring, cloud integrations, and application instrumentation.
Establish Dynatrace configuration and operational best practices.
Conduct application assessments and identify observability gaps.
Build and validate dashboards, alerts, service flows, dependency maps, and application health views.
Establish standards for monitoring critical business applications and services.
Lead troubleshooting and optimization of Dynatrace implementations.
Ensure implementations are scalable, secure, and aligned with enterprise architecture standards.
3. ServiceNow ITSM Integration
Architect and implement Dynatrace–ServiceNow integration for enterprise incident and event management.
Define standards for:
Event ingestion
Alert filtering
Event correlation
Automatic incident creation
Incident enrichment
Assignment and routing
Deduplication
Priority/severity mapping
Business service mapping
Incident closure and feedback
Integrate Dynatrace with ServiceNow Incident, Event Management, CMDB, and related ITSM capabilities.
Optimize the integration to minimize alert noise and ensure only actionable events generate incidents.
Establish best practices for automated incident triage and escalation.
4. Observability Operations & Monitoring
Define operational processes for enterprise monitoring and observability.
Establish monitoring coverage and health-check standards for critical applications and infrastructure.
Develop proactive monitoring strategies to identify performance degradation and potential incidents before they impact users.
Establish standards for alert thresholds, anomaly detection, noise reduction, and event prioritization.
Review monitoring effectiveness and continuously improve observability coverage.
Support major incident investigations through Dynatrace-driven topology, dependency, and root-cause analysis.
5. Proof of Concepts & Technology Evaluation
Conduct hands-on POCs and technical evaluations of Dynatrace capabilities and complementary observability technologies.
Evaluate new Dynatrace features and determine applicability to enterprise use cases.
Develop technical recommendations, reference architectures, and implementation roadmaps.
Benchmark observability solutions and recommend improvements to the enterprise monitoring strategy.
6. Governance & Adoption
Establish an enterprise observability governance framework.
Create implementation checklists and certification standards for application onboarding.
Review and certify application observability implementations against established gold standards.
Partner with application, cloud, DevOps, SRE, infrastructure, and ITSM teams to drive adoption.
Conduct technical workshops and enablement sessions for engineering and operations teams.
Required Technical Skills
Dynatrace
Strong hands-on experience with Dynatrace SaaS.
Deep knowledge of Dynatrace Application Observability and Infrastructure Monitoring.
Experience with:
OneAgent
ActiveGate
Kubernetes monitoring
Application monitoring
Distributed tracing
OpenTelemetry
Logs and metrics
Service flow and topology
Davis AI / anomaly detection
Dashboards and reporting
SLOs/SLIs
Synthetic monitoring
RUM
APIs and automation
Experience designing enterprise-scale Dynatrace architectures.
ServiceNow
Strong experience integrating Dynatrace with ServiceNow ITSM/Event Management.
Knowledge of ServiceNow Incident Management, Event Management, CMDB, and Service Mapping.
Experience designing event-to-incident workflows.
Understanding of alert correlation, enrichment, routing, deduplication, and noise reduction.
Preferred Experience
8+ years of experience in infrastructure, application monitoring, SRE, DevOps, or observability.
4+ years of hands-on Dynatrace experience.
Experience leading enterprise observability implementations.
Experience working with large, distributed application environments.
Dynatrace certification preferred.
ServiceNow certification or implementation experience preferred.
Experience developing enterprise architecture standards and reference architectures.
Experience working directly with customers and senior technology stakeholders.
Key Success Measures
The success of this role will be measured by:
Enterprise Dynatrace Gold Standards established and adopted.
Successful rollout of Dynatrace across targeted applications and infrastructure.
Improved monitoring coverage and application visibility.
Reduction in alert noise and false-positive incidents.
Successful Dynatrace–ServiceNow ITSM integration.
Improved incident detection, triage, routing, and resolution.
Increased use of topology and dependency information for root-cause analysis.
Standardized observability architecture across application teams.
Successful POCs and recommendations for emerging observability technologies.