Purpose
The Contingency and Resiliency Pattern defines the approved reference architecture for recoverability, resiliency, operational continuity, restoration, and service sustainability within University-managed technology environments.
This pattern describes how workloads, platforms, identity services, monitoring services, data-protection services, recovery services, and operational teams interact to support restoration and continued operation following failures, disruptions, outages, or disaster events.
The pattern provides a common resiliency architecture supporting workload sustainability, restoration capabilities, and operational continuity across University-managed technology environments.
Applicable standards, workload classifications, data classifications, business requirements, and obligations determine the recovery capabilities implemented within this architecture.
Use Cases
This pattern applies when:
- Hosting applications
- Hosting databases
- Hosting platform services
- Hosting integrations
- Hosting data products
- Implementing workload recoverability
- Implementing operational continuity capabilities
- Implementing restoration capabilities
- Implementing resilience architectures
- Supporting critical business, academic, research, or administrative services
This pattern does not define:
- Recovery time objectives
- Recovery point objectives
- Backup schedules
- Operational procedures
- Testing schedules
- Business continuity procedures
- Disaster recovery procedures
- Retention requirements
- Specific recovery-control requirements
Pattern Application
The Contingency and Resiliency Pattern provides a common recovery architecture for University-managed workloads.
Innovation Workloads
Innovation workloads may implement simplified resiliency architectures appropriate for research, experimentation, pilot initiatives, and proof-of-concept activities.
Enterprise Workloads
Enterprise workloads implement the Contingency and Resiliency Pattern using the standard enterprise recovery architecture.
Regulated Workloads
Regulated workloads implement the Contingency and Resiliency Pattern using additional resiliency capabilities, operational continuity capabilities, recovery boundaries, and obligation-specific extensions.
Design Principles
Resiliency by Design
Recovery and resiliency capabilities are incorporated into architecture designs rather than introduced after deployment.
Failure Assumption
Architectures assume components, services, platforms, integrations, and dependencies may fail and provide recovery capabilities supporting restoration of service.
Recoverability
Architectures provide capabilities supporting workload, platform, service, and data restoration.
Operational Continuity
Architectures support continued operation or restoration of operation following disruptions affecting technology services.
Composable Recovery Architecture
Recovery capabilities are assembled from reusable architecture components supporting multiple workload classifications and hosting models.
Shared Recovery Services
Recovery capabilities may leverage approved enterprise recovery, monitoring, identity, governance, and data-protection services.
Logical Architecture
Workload
|
+-- Applications
+-- Databases
+-- Platform Services
+-- Integration Services
+-- Data Products
|
v
Resiliency Architecture
|
+-- Data Protection Services
+-- Recovery Services
+-- Monitoring Services
+-- Identity Services
+-- Shared Recovery Capabilities
|
v
Recovery Operations
|
v
Restored Service
Resiliency Architecture Components
Workload Resiliency Component
Provides the architecture supporting continued operation, service restoration, and workload sustainability.
Data Recovery Component
Provides the architecture supporting restoration of protected data required for workload recovery.
Recovery Services Component
Provides recovery capabilities supporting restoration of applications, services, platforms, and supporting resources.
Operational Continuity Component
Provides capabilities supporting service continuation, service restoration, and operational sustainability.
Recovery Coordination Component
Provides the architecture supporting coordination between workload, platform, monitoring, identity, and data-protection capabilities during restoration activities.
Recovery Visibility Component
Provides operational visibility supporting failure detection, recovery activities, restoration validation, and recovery operations.
Shared Recovery Component
Provides reusable recovery services supporting multiple workloads, platforms, and hosting environments.
Recovery Models
Application Recovery
Applications implement recovery architectures supporting restoration of application capabilities and services.
Data Recovery
Data platforms implement recovery architectures supporting restoration of required information assets and data services.
Platform Recovery
Platform services implement recovery architectures supporting restoration of shared platform capabilities.
Identity Recovery
Identity services provide recovery capabilities supporting continued authentication, authorization, and identity-management functions.
Monitoring Recovery
Monitoring services provide visibility supporting recovery activities, restoration activities, and operational investigations.
Integration Recovery
Integration services provide recovery capabilities supporting restoration of communications, data exchange, and dependent service interactions.
Recovery Boundaries
Workload Recovery Boundary
The workload recovery boundary contains recovery capabilities directly supporting a workload and its dependencies.
Platform Recovery Boundary
The platform recovery boundary contains shared recovery services supporting multiple workloads.
Data Recovery Boundary
The data recovery boundary contains the capabilities supporting restoration of protected information assets.
Enterprise Recovery Boundary
The enterprise recovery boundary contains centrally operated recovery capabilities supporting multiple technology environments.
External Dependency Boundary
The external dependency boundary contains recovery dependencies associated with systems, services, providers, and platforms outside direct University operational control.
Resiliency Integration
Identity Integration
The Contingency and Resiliency Pattern integrates with the Identity Pattern to support recovery of authentication, authorization, administrative access, and workload identities.
Monitoring Integration
The Contingency and Resiliency Pattern integrates with the Monitoring Pattern to support failure detection, recovery visibility, operational awareness, and restoration activities.
Network Integration
The Contingency and Resiliency Pattern integrates with the Network Pattern to support restoration of connectivity, communication paths, and service interactions.
Data Protection Integration
The Contingency and Resiliency Pattern integrates with the Data Protection Pattern to support restoration of protected information assets.
Shared-Service Integration
The Contingency and Resiliency Pattern integrates with approved enterprise shared services supporting recovery and operational continuity.
Landing Zone Integration
Landing Zones implement this pattern through approved recovery architectures and reusable resiliency components.
Operational Responsibilities
Enterprise Service Providers
Enterprise service providers own enterprise recovery services, recovery architecture, recovery platforms, and shared resiliency capabilities.
Platform Engineering
Platform Engineering owns reusable resiliency architecture components, onboarding automation, and Landing Zone resiliency integration.
Information Security
Information Security provides security requirements consumed by recovery architectures and participates in reviews and investigations.
Operations Teams
Operations teams use resiliency capabilities to support restoration activities, operational continuity activities, and service recovery.
Workload Teams
Workload teams own workload recovery architecture, workload restoration activities, workload dependencies, and workload-specific resiliency integration.
Automation Pattern
Resiliency architectures should be implemented through approved platform automation whenever practical.
Automation may support:
- Recovery onboarding
- Platform integration
- Data-protection integration
- Monitoring integration
- Identity integration
- Landing Zone integration
- Recovery-service integration
- Workload onboarding
Recovery Definition
|
+-- Recovery Services
+-- Data Recovery
+-- Monitoring Integration
+-- Identity Integration
+-- Operational Continuity
|
v
Reusable Recovery Components
|
v
Landing Zone or Workload Deployment
|
v
Integrated Resiliency Architecture
Recovery automation components should be reusable across supported hosting architectures, workload classifications, and platform products.
Reference Architecture Outcomes
A workload implementing this pattern should provide:
- Recoverability
- Service restoration capability
- Operational continuity support
- Data restoration capability
- Identity-service recovery integration
- Monitoring-service recovery integration
- Platform-recovery integration
- Shared recovery-service integration
- Operational supportability
- Reusable resiliency architecture
- Repeatable onboarding
- Support for workload and obligation-specific extensions
Exceptions
Exceptions to this pattern must follow approved architecture governance and information security exception processes.
Approved exceptions must be periodically reviewed and must not be treated as permanent architecture patterns.