{"description":"Standing operator workflow for the monthly IT availability and resilient-failure-operations cycle: batch and processing monitoring with incident and problem resolution, backup and restore verification, availability-to-SLA tracking, network capacity and DoS-mitigation posture, fail-secure and alternate-communications readiness, and the mean-time-to-failure replacement queue with verified fail-safe procedures, as a decision-aware flow that escalates an urgent reliability gap immediately and rolls routine findings into a single end-of-cycle corrective-action log. Each monthly instance runs against the EXISTING Control item for this IT-availability and resilient-failure-operations control (frequency: monthly; framework: nist-800-53 + sox; domains: business_continuity_disaster_recovery, network_communications_security, incident_management_response) — it enriches that standing control with the cycle's evidence rather than creating a new record: per-stream evidence attaches as documents on the workflow instance's steps, operational failures and gaps become Issue items linked to that Control, and the named deliverable is the signed monthly control record (the classify-cycle-disposition form plus its disposition summary), backed by the monitoring roll-up, incident-and-problem log, backup-and-restore summary, availability dashboard, capacity and fail-secure verification records, and the corrective-action register. In scope: the in-scope batch jobs and system processing, network services and communications paths, and components tracked for mean-time-to-failure named in the operating brief, measured against their defined availability and processing service expectations; out of scope: application change management, access provisioning, and physical-environment controls, which are operated by their own workflows. This is a standing monthly cycle with no upstream workflow dependency and no downstream handoff — its inputs are the organization's own scheduler and monitoring feeds, backup history, capacity telemetry, and asset/MTTF inventory; its archived record seeds the next monthly instance of the same control.","edges":[{"id":"e-maintain-reliability-queue-and-fail-safe-readiness-classify-cycle-disposition","label":"Within tolerance","source":"maintain-reliability-queue-and-fail-safe-readiness","target":"classify-cycle-disposition","whenValue":"within_tolerance"},{"id":"e-maintain-reliability-queue-and-fail-safe-readiness-escalate-urgent-reliability-gap","label":"Urgent gap","source":"maintain-reliability-queue-and-fail-safe-readiness","target":"escalate-urgent-reliability-gap","whenValue":"urgent_gap_identified"},{"id":"e-escalate-urgent-reliability-gap-classify-cycle-disposition","source":"escalate-urgent-reliability-gap","target":"classify-cycle-disposition"},{"id":"e-classify-cycle-disposition-close-and-archive","label":"Healthy","source":"classify-cycle-disposition","target":"close-and-archive","whenValue":"healthy"},{"id":"e-classify-cycle-disposition-log-corrective-actions","label":"Gaps","source":"classify-cycle-disposition","target":"log-corrective-actions","whenValue":"gaps_identified"},{"id":"e-log-corrective-actions-close-and-archive","source":"log-corrective-actions","target":"close-and-archive"}],"isPublic":true,"metadata":{"capabilities":[],"controlVerbs":{},"controls":["UC-ASSET-12","UC-NET-08","UC-VULN-08"],"department":"it","domains":["controls"],"library":{"aliases":[],"canonicalUrl":"https://workflow-library.com/all/?w=controls-it-availability-resilient-failure-operations","contentDigest":"sha256:4c158ff94db3daeacbef5d536ca11ed76c08fa3eb7514bebaf1b0847e4d071bc","prerequisites":{"status":"undeclared"},"provenance":[],"releaseId":"sha256:4c158ff94db3daeacbef5d536ca11ed76c08fa3eb7514bebaf1b0847e4d071bc","schemaVersion":1,"sourceTemplateId":"workflow-library:controls-it-availability-resilient-failure-operations"},"lineOfDefense":"operate","mappingStatus":"mapped","risks":[],"slug":"controls-it-availability-resilient-failure-operations","source":"coworkcanvas-gallery","standards":["nist-800-53","sox"],"teams":["it"]},"name":"IT Availability & Resilient Failure Operations","nodes":[{"data":{"decisionField":"reliability_readiness","description":"Agent tracks components against MTTF tolerance and verifies fail-safe procedures; human decides whether the cycle is within tolerance or carries an urgent reliability gap","formData":{"fields":[{"key":"reliability_readiness","label":"Reliability Readiness","options":[{"label":"Within tolerance","value":"within_tolerance"},{"label":"Urgent gap identified","value":"urgent_gap_identified"}],"required":true,"type":"select"}],"resultType":"form","submittedAt":null,"values":{}},"instructions":"**Objective** — Determine whether every component tracked for mean time to failure (MTTF) is inside its replacement window and every fail-safe procedure reaches a known safe state, so a degrading component or untested fail-safe logic is caught before it causes an outage. The control owner picks the branch.\n\nBefore deciding, the agent assembles the evidence:\n1. Query the asset inventory and MTTF model for every tracked component, computing remaining tolerance and flagging any component inside or past its replacement window.\n2. Create or update a reliability-queue Issue for each flagged component (predicted failure window, replacement owner, target date) and link it to the anchor Control — there is no native Asset/Component item to link to.\n3. Execute or review the fail-safe procedure test for in-scope systems: on a simulated failure condition the system must enter a known safe state, preserve security protections, alert designated personnel, and prevent unsafe continuation; record each result.\n4. Compile the reliability-queue and fail-safe-verification summary and attach it.\n\n**Decision criteria**\n- `within_tolerance` — every tracked component is inside its replacement window (none past MTTF tolerance unreplaced) AND every fail-safe test reached a known safe state that alerted designated personnel and prevented unsafe continuation. Any routine findings roll into the end-of-cycle corrective-action log rather than immediate escalation.\n- `urgent_gap_identified` — at least one component is past its MTTF tolerance still in service, OR a fail-safe test failed to reach a known safe state (failed open, lost security protection, did not alert designated personnel, or allowed unsafe continuation). Either condition can precipitate an imminent, unsafe outage and must be escalated now, not at cycle close.\n\n**Record in AssureSwarm**\n- Step form — submit the `reliability_readiness` SELECT with the step result and the step's approver record.\n- Item create — for each component inside or past its MTTF replacement window, create a reliability-queue Issue (issue_type: deficiency, source: management_identified, severity, description, issue_owner, target_remediation_date) — there is no native Asset/Component item, so the flagged component is tracked as an Issue (coach-item-create).\n- Item relationship — link each such Issue to the anchor Control (coach-items-link).\n- Step document — attach the reliability-queue and fail-safe-verification summary (coach-document-upload) built from the MTTF and inventory queries (coach-query-data); each fail-safe test result is recorded in it.\n- Record the decision rationale with evidence references in the step result and the decision owner in the step's approver record.\n\n**Exit criteria** — The `reliability_readiness` routing selector is submitted and the step result contains a rationale citing the queue and test evidence; the unused branch is prunable.","kind":"decision","label":"Maintain reliability queue and fail-safe readiness","performedBy":{"primitives":["coach-query-data","coach-item-create","coach-items-link","coach-document-upload"]}},"id":"maintain-reliability-queue-and-fail-safe-readiness"},{"data":{"description":"Agent packages the reliability or fail-safe deficiency and an interim mitigation for immediate escalation; human confirms leadership is notified and the interim mitigation is in force","instructions":"**Objective** — Escalate an urgent reliability gap immediately — a component past MTTF tolerance or a fail-safe test that failed to reach a known safe state — with an interim mitigation in force before the cycle disposition is classified.\n\n**Inputs**\n- The reliability-queue and fail-safe-verification summary (from the reliability-readiness decision's reviewed output), isolating the specific deficiency.\n- The accountable owner and IT operations leadership escalation contacts.\n- Available interim-mitigation options (expedited replacement, manual monitoring, temporary isolation, compensating alert).\n\n**Procedure**\n1. Parse the summary to isolate the exact deficiency: the component past its replacement window, or the fail-safe procedure that failed to reach a known safe state, alert designated personnel, or prevent unsafe continuation.\n2. Create an urgent corrective-action item capturing the deficiency, its outage or safety impact, the chosen interim mitigation, owner, and due date; link it to the reliability-queue or fail-safe-verification record.\n3. Draft an escalation notice to the accountable owner and IT operations leadership stating the deficiency, the interim mitigation now in force, and the residual risk until remediation completes; attach it.\n4. Flag the item for priority follow-up ahead of the standard monthly cadence rather than waiting for the next interval.\n\n**Record in AssureSwarm**\n- Item create — create the urgent corrective-action Issue (issue_type: deficiency, severity: high or critical, source: management_identified, description, remediation_plan = the interim mitigation, issue_owner, target_remediation_date) (coach-item-create).\n- Item relationship — link it to the reliability-queue Issue (or fail-safe finding) it escalates and to the anchor Control (coach-items-link).\n- Step document — attach the escalation notice (coach-document-upload) and issue the notification (coach-notify).\n\n**Exit criteria** — The escalation is received and acknowledged, the interim mitigation is actually in force, and the corrective action is owned and dated before the cycle disposition is classified.\n\n> **⚡ Audit Artist accelerator:** `/coach-notify` — sends the escalation notice to the accountable owner and IT operations leadership and records the acknowledgement.","label":"Escalate urgent reliability gap","performedBy":{"primitives":["coach-item-create","coach-items-link","coach-document-upload","coach-notify"]}},"id":"escalate-urgent-reliability-gap"},{"data":{"decisionField":"cycle_disposition","description":"Judge complete processing failure resolution, demonstrated recovery, live DoS protection, secure failure and working alternate communications against availability and cycle thresholds.","formData":{"fields":[{"key":"cycle_disposition","label":"Cycle Disposition","options":[{"label":"Healthy, within tolerance","value":"healthy"},{"label":"Gaps identified","value":"gaps_identified"}],"required":true,"type":"select"}],"resultType":"form","submittedAt":null,"values":{}},"instructions":"**Objective** — Judge complete processing failure resolution, demonstrated recovery, live DoS protection, secure failure and working alternate communications against availability and cycle thresholds.\n\n**Inputs**\n- The anchor Control item — the EXISTING monthly IT-availability control (frequency: monthly; framework: nist-800-53, sox) this cycle's workflow instance runs against; Control.framework and Control.family carry the governing standard references (SOX ITGC operations monitoring and incident/problem management, NIST 800-53 SI-13), so there is no separate storage for them.\n- The in-scope job/process catalog — Process items (process_type: it_general_control or operational, process_owner, frequency) linked to the anchor Control; each job's defined run window and availability/processing service-expectation thresholds have no native item field, so they live in the operating-brief document referenced on this step.\n- The job scheduler feed and system-processing monitoring feed for the cycle window (run status, start/end time, duration, exception codes) — external systems queried via coach-query-data; the AssureSwarm copy is the roll-up document attached here.\n- The prior-cycle monitoring roll-up document (on the previous archived instance's monitoring step), to carry forward any run that spans the cycle boundary.\n- Existing open incident and problem records carried from prior cycles, plus system logs, dependency/health status, and prior incident history for root-cause analysis.\n- This is a standing monthly cycle with no upstream workflow; inputs are the organization's own scheduler and monitoring feeds.\n- The anchor Control item this cycle runs against, and the in-scope system list with each system's backup schedule and recovery objective (RTO/RPO) — there is no native Asset/System item type, so this list and its RTO/RPO targets are carried as a document on this step, referenced against the anchor Control.\n- The backup job history for the cycle window (completion status, size, integrity-check result) — external backup platform, queried via coach-query-data.\n- The restore-test rotation schedule identifying which systems are due this cycle.\n- An isolated recovery target to restore into.\n- Control references: NIST 800-53 CP-9 (backup), CP-10 (recovery), SI-17.\n- Standing monthly cycle with no upstream workflow; inputs are the organization's own backup history and test rotation.\n- The in-scope network services and their peak historical and projected demand.\n- Current capacity/utilization telemetry and capacity-warning thresholds.\n- The active rate-limiting and filtering rule sets.\n- The upstream DoS-mitigation service status, coverage map, and recent activation/anomaly events.\n- Control references: NIST 800-53 SC-5 (denial-of-service protection), SC-6 (resource availability).\n- Standing monthly cycle with no upstream workflow; inputs are the organization's own capacity telemetry and mitigation-service status.\n- The in-scope network and security components and their documented fail-secure configuration.\n- The alternate-communications-path register (out-of-band channel, secondary provider, or failover route) and its provisioning status.\n- Control references: NIST 800-53 SC-24 (fail in known state), SC-47 (alternate communications paths).\n- Standing monthly cycle with no upstream workflow; inputs are the organization's own component configuration and alternate-comms register.\n- Each in-scope system's defined service expectation (uptime target, on-time completion rate, mean-time-to-resolve target), its availability/uptime telemetry and processing-completion data for the cycle window, and the prior three cycles' metrics for trend.\n- The incident-and-problem log (from the incident-resolution checkpoint's reviewed output), to attribute each deviation to its driving incident or problem.\n- The other streams' reviewed evidence: the monitoring roll-up, backup and restore summary, capacity and fail-secure verification records, and the reliability-queue and fail-safe-verification summary, including any urgent reliability gap escalated this cycle.\n\n**Procedure**\n_This checkpoint absorbs “Resolve processing incidents and problems”, “Verify backups and restore tests”, “Assess capacity and DoS-mitigation posture”, “Verify fail-secure state and alternate comms”. The agent runs the preparation, evidence assembly and record updates below; the named owners retain the substantive decisions and approvals stated in the procedure._\n1. Resolve processing incidents and problems: Query the job scheduler and system-processing monitoring feed for every scheduled run in the cycle window, capturing run status, duration, start/end timestamps, and exception codes.\n2. Reconcile the observed runs against the confirmed catalog of in-scope jobs and processes: flag any job that did not run, ran outside its defined window, ran with a nonzero exit or exception code, or is entirely absent from the feed. Treat a job missing from monitoring as a failure, not a pass — a silent gap is the most dangerous finding here.\n3. Classify each run as success, failure, late/out-of-window, or missing, and quantify totals: total runs, successes, failures, and missed/late runs.\n4. Create a monitoring roll-up record summarizing the totals and enumerating each failed, missed, or late run with its job identity, timestamp, and exception code; link each exception run to its job definition record.\n5. Attach the roll-up document.\n6. For each failed or missed run in the roll-up that lacks an incident, open an incident record capturing symptom, affected system, first-observed time, and business impact; link it to the monitoring roll-up.\n7. Investigate root cause per incident against system logs, upstream/downstream dependency status, and prior incident history; record the confirmed or hypothesized cause and findings on the record.\n8. For failures resolved within the cycle, document the fix applied and the verification that the next run succeeded, then close the incident.\n9. For recurring or unresolved failures, open a problem record capturing root-cause hypothesis, current workaround, permanent-fix owner, and target date; link it to all originating incidents so recurrence is visible over time.\n10. Compile the incident-and-problem log, attach it, and have the control owner verify the disposition of every exception run: nothing silently dropped from the roll-up, and each failure closed with verification or carried as an owned, dated problem.\n11. Verify backups and restore tests: Query the backup job history for every in-scope system, confirming completion status, backed-up size against expectation, and integrity-check result for each scheduled backup in the window; flag any missed, partial, or integrity-failed backup.\n12. Identify the systems due for a periodic restore test this cycle per the rotation schedule (a subset each cycle, so every system is exercised across the year).\n13. Restore each due system's backup to the isolated recovery target and validate the restored set: recovery time achieved vs. RTO, data currency vs. RPO, and integrity/consistency of the restored data.\n14. Record a restore-test result per system tested (recovery time, data integrity, any gap against objective) and link it to the source backup job.\n15. Compile the backup-completion and restore-test summary and attach it.\n16. Assess capacity and DoS-mitigation posture: Compute current capacity headroom per in-scope service against peak historical and projected demand; flag any service inside its capacity-warning threshold.\n17. Pull the active rate-limiting and filtering rule sets and confirm each in-scope service is covered by an enforced rule, not a disabled or draft policy.\n18. Retrieve the upstream DoS-mitigation service status and coverage map; confirm every in-scope service is within the mitigation subscription and that the mitigation is in an active (not bypassed) state.\n19. Cross-check recent traffic anomalies or mitigation-activation events against the coverage map to confirm the mitigation actually engaged when triggered — evidence of function, not just configuration.\n20. Compile the capacity-and-mitigation posture summary and attach it.\n21. Verify fail-secure state and alternate comms: Query the fail-secure configuration for each in-scope network and security component; confirm it denies or closes to a known secure state on failure (fails closed, not open) and that essential state information is preserved rather than lost.\n22. Retrieve the alternate-communications-path register and confirm the out-of-band channel, secondary provider, or failover route is currently provisioned, funded, and reachable.\n23. Coordinate a test activation of the alternate communications path and record whether operations could genuinely continue on it — reachability plus a functional exchange, not just a dial tone.\n24. Record any component that fails open, loses essential state, or any alternate path that could not be activated, as a finding for the reliability disposition.\n25. Compile the fail-secure and alternate-comms verification record and attach it.\n26. Classify cycle disposition: Compute availability and processing-completion metrics per in-scope system against its service expectation: uptime %, on-time completion rate, and mean time to resolve.\n27. Build an availability dashboard showing each system against its threshold with a trend line against the prior three cycles, so chronic degraders are visible even when a single cycle passes.\n28. For every system below its service expectation, create a follow-up item capturing the deviation, its magnitude vs. threshold, and the driving incident or problem record; link the follow-up to that record.\n29. Note any system trending toward its threshold (inside a defined warning band) as an early-warning follow-up even if not yet breached.\n30. Compile the availability-tracking summary and attach it.\n31. Compute the cycle metrics: unresolved processing incidents or open problem records; backup or restore-test failures; unresolved availability deviations; any capacity, DoS-mitigation, fail-secure, or alternate-comms gap; and any reliability-queue component past MTTF tolerance or fail-safe test that failed to reach a known safe state, including any urgent reliability gap escalated this cycle.\n32. Build a cycle-health dashboard showing each metric against its threshold with the trend against prior cycles.\n33. List every breach with its owner and the supporting evidence, drawing on the monitoring roll-up, incident-and-problem log, backup and restore summary, availability-tracking summary, capacity and fail-secure verification records, and the reliability-queue and fail-safe-verification summary already attached this cycle; confirm any escalated urgent reliability gap has its interim mitigation recorded and is not silently dropped.\n34. Draft the disposition summary and attach it.\n35. The control owner reads the dashboards and breach list against the criteria below — no deviation untracked, no gap unowned — and submits the disposition.\n\n**Decision criteria**\n- `healthy` — every metric is within tolerance: no unresolved processing incident, no open unowned problem, no backup or restore failure, no unresolved availability deviation, no capacity/DoS/fail-secure/alternate-comms gap, and no reliability or fail-safe gap. Any urgent gap escalated this cycle has been resolved or carries a recorded interim mitigation and owned corrective action.\n- `gaps_identified` — any unresolved incident, backup or restore failure, availability deviation, network-posture gap, or reliability or fail-safe gap exists. Route to corrective-action logging so each gap is owned and dated before close.\n\n**Record in AssureSwarm**\n- Step document — attach the monitoring roll-up (XLSX) and the incident-and-problem log to this step (coach-document-upload), built from the scheduler, system-processing, log, and dependency queries (coach-query-data); the roll-up itemizes every failed, missed, and late run with its job identity, timestamp, and exception code.\n- Item create — open an incident Issue (issue_type: exception, source: management_identified, severity, description, root_cause, issue_owner, identified_date = first-observed time) for each exception run needing follow-up, and a parent problem Issue (issue_type: exception, source: management_identified, root_cause, remediation_plan = permanent-fix plan, issue_owner, target_remediation_date) for each recurring or unresolved failure (coach-item-create).\n- Item update — on each incident resolved this cycle, record the fix in management_response and set actual_remediation_date, then close it (coach-item-update).\n- Item relationship — link each incident Issue to the anchor Control and to its Process (job/process catalog) item where one exists; link each problem Issue to all its originating incident Issues so recurrence is visible over time (coach-items-link).\n- Step document — attach the backup-completion and restore-test summary (coach-document-upload) built from backup-history queries (coach-query-data), recording per-system completion status and each restore test's recovery time vs. RTO, data currency vs. RPO, and restored-data integrity (there is no Asset/System item to hold per-test results).\n- Item create — for any missed, partial, or integrity-failed backup, or a restore test that missed its objective, create an Issue (issue_type: exception, source: management_identified, severity, description, identified_date) and link it to the anchor Control (coach-item-create, coach-items-link).\n- Step document — attach the capacity-and-DoS-mitigation posture summary to this step (coach-document-upload), built from capacity, rule-set, and mitigation-status queries (coach-query-data); capacity headroom, rule coverage, and mitigation state have no native item field, so this document (referenced against the anchor Control) is the record of the posture, and any flagged shortfall is carried into the cycle-disposition review.\n- Step document — attach the fail-secure and alternate-comms verification record (coach-document-upload) built from configuration queries (coach-query-data); the alternate-communications test result is recorded in this document, as there is no native Asset/Component item to hold it.\n- Item create — for any component that fails open or loses essential state, or an alternate path that could not be activated, create an Issue (issue_type: deficiency, source: management_identified, severity, description, identified_date) and link it to the anchor Control (coach-item-create, coach-items-link).\n- Step form — submit the `cycle_disposition` SELECT with the step result and the step's approver record; this signed form plus its summary is the month's control-execution record on the anchor Control.\n- Dashboard — build the availability dashboard (each system against its service expectation, trended against the prior three cycles) and the cycle-health dashboard (each metric against its threshold with the prior-cycle trend) (coach-dashboard-create).\n- Item create + relationship — for every system below (or trending toward) its service expectation, create a deviation follow-up Issue (issue_type: observation, source: management_identified, description, issue_owner, target_remediation_date) and link it to its driving incident or problem Issue and to the anchor Control (coach-item-create, coach-items-link).\n- Step document — attach the availability-tracking summary and the cycle disposition summary, both built from the availability, processing, and metric queries (coach-document-upload, coach-query-data).\n- Record the decision rationale with evidence references in the step result and the decision owner in the step's approver record.\n\n**Exit criteria**\n- Every in-scope job and process for the cycle window is accounted for in the roll-up, with each failure, late run, and missing run itemized and linked to its job definition; every failure has either a documented resolution with verification or an owned, dated problem record; no failure is left uninvestigated; the control owner has confirmed nothing was silently dropped and the closure or carry state of each.\n- Backup completion is confirmed for every in-scope system; the scheduled restore test(s) demonstrably recovered data within objective on an isolated target; the control owner has confirmed recovery was proven, not assumed.\n- Capacity headroom is adequate (or a flagged shortfall is documented), rate limiting and filtering are enforced, and DoS-mitigation coverage is current and proven-engaging for every in-scope network service; the control owner has confirmed the posture.\n- Every in-scope component is confirmed to fail to a known secure state with essential state preserved; the alternate communications path was tested and works; the control owner has confirmed both.\n- Every in-scope system has a computed metric against its service expectation and every deviation has an owned, dated follow-up linked to its driver; the `cycle_disposition` routing selector is submitted and the step result contains a rationale citing the consolidated evidence; the unused branch is prunable.","kind":"decision","label":"Classify cycle disposition","performedBy":{"primitives":["coach-query-data","coach-item-create","coach-item-update","coach-items-link","coach-document-upload","coach-dashboard-create"]}},"id":"classify-cycle-disposition"},{"data":{"description":"Agent converts each identified gap into an owned corrective action; human confirms every gap is owned, dated, and escalated where required","instructions":"**Objective** — Convert every gap from the end-of-cycle disposition review into an owned, dated corrective action so nothing degrades the standing availability control unaddressed.\n\n**Inputs**\n- The cycle disposition summary consolidating the processing, backup, availability, network-resilience, reliability-queue, and fail-safe findings (from the classify-cycle-disposition decision, gaps_identified branch).\n- Any urgent reliability gap already escalated this cycle, to roll forward without duplication.\n- The corrective-action register.\n\n**Procedure**\n1. Parse the disposition summary to list each gap with its root cause and the metric, component, incident, or test that surfaced it.\n2. Create a corrective-action item per gap capturing root cause, owner, due date, and interim mitigation; link it to the driving metric, component, or incident.\n3. Roll forward any already-escalated urgent reliability gap as the same tracked item, not a duplicate, preserving its existing owner and mitigation.\n4. Raise capability-level improvement items for systemic gaps — a chronically understaffed restore-test rotation, an unfunded component-replacement backlog, or a repeatedly failing fail-safe procedure — so recurring root causes are addressed, not just instances.\n5. Attach the corrective-action register.\n\n**Record in AssureSwarm**\n- Item create — create one corrective-action Issue per gap (issue_type: deficiency, source: management_identified, root_cause, issue_owner, target_remediation_date, remediation_plan = interim mitigation), and raise capability-level improvement Issues (issue_type: observation) for systemic root causes (coach-item-create).\n- Item relationship — link each corrective-action Issue to the anchor Control and to its driving incident/problem Issue, metric, or component Issue; roll an already-escalated urgent reliability gap forward as the same Issue, not a duplicate (coach-items-link).\n- Step document — attach the corrective-action register (XLSX) (coach-document-upload).\n\n**Exit criteria** — Every gap has a named owner and due date; any escalated urgent reliability gap is reflected once without duplication; the control owner has confirmed nothing is left untracked before closure.","label":"Log corrective actions","performedBy":{"primitives":["coach-item-create","coach-items-link","coach-document-upload"]}},"id":"log-corrective-actions"},{"data":{"description":"Automatically archive the authorized cycle record and carry open actions into the next cycle.","instructions":"**Objective** — Automatically preserve the authorized cycle record and its carry-forward actions after the preceding decision.\n\n**Inputs**\n- The full operating record for the cycle (monitoring roll-up, incident and problem log, backup and restore summary, availability summary, capacity and fail-secure records, reliability-queue and fail-safe-verification summary, decision forms, and corrective-action register).\n- Open corrective actions, upcoming restore tests, and components approaching their next MTTF review.\n- The designated evidence repository and its retention controls; the control execution log.\n\n**Procedure**\n1. Export the full operating record and archive it in the designated evidence repository under retention controls; record the archive location and reference.\n2. Create carry-forward items for open corrective actions, upcoming restore tests, and components approaching their next MTTF review; link each to its source so it arrives as an explicit input to the next cycle.\n3. Update the control execution log with the cycle result and key metrics, and confirm the next monthly cadence review is scheduled.\n4. Attach the closure record.\n\n**Record in AssureSwarm**\n- Workflow instance — export the full operating record as the month's audit trail (coach-workflow-export) and archive it under retention.\n- Item create — create carry-forward Issues (issue_type: observation or exception, source: management_identified, issue_owner, target_remediation_date) for open corrective actions, upcoming restore tests, and components approaching their next MTTF review, and link each to its source and to the anchor Control so it arrives as an explicit input to the next cycle (coach-item-create, coach-items-link).\n- Step document — attach the closure record (coach-document-upload).\n\n**Exit criteria** — The archived record is immutable and retrievable; the next review is scheduled; nothing remains open without a tracked owner; the authorized cycle record is complete.","label":"Close and archive","performedBy":{"primitives":["coach-workflow-export","coach-item-create","coach-items-link","coach-document-upload"]},"requiredApprovals":0},"id":"close-and-archive"}],"sourceTemplateId":"workflow-library:controls-it-availability-resilient-failure-operations"}
