{"description":"Standing operator workflow for the weekly IT operations and capacity management cycle. It runs on the EXISTING **Process** item \"IT Operations & Capacity Management\" (process_type: operational, process_owner: IT operations manager, frequency: weekly): each weekly cycle is a new workflow instance attached to that Process — the Process is enriched, never recreated — with the operated **Control** items (job-scheduling, infrastructure-monitoring, and capacity controls; frequency: weekly) linked to it via a Control ↔ Process relationship. The instance executes and reconciles daily job scheduling, processing, infrastructure monitoring, and facility management against documented procedure, then assesses capacity and utilization against forecast demand, triggers capacity additions and priority-allocation safeguards ahead of threshold breach, and produces the named evidence set: the weekly operations review minutes, the consolidated exception log, and the capacity-and-utilization report (plus a recurring capacity-and-utilization dashboard). In scope: daily job scheduling, processing, infrastructure monitoring, and facility management for the operator's named in-scope systems and facilities, plus the weekly capacity and utilization review of infrastructure, data, and software against forecast demand. Out of scope: incident response, change management, and disaster-recovery testing, which run as their own workflows; this cycle raises corrective actions and capacity-addition requests (tracked as Issue items) that feed change management in substance but models no explicit handoff node — it is a terminal, self-originating operate-line cycle that does not itself execute infrastructure changes. No upstream workflow feeds it: its inputs are the operator's own job schedules, monitoring and facility feeds, the governing operating-procedure and capacity-threshold and quota **Policy** items, and the prior cycle's carryover **Issue** items (created when the signed weekly operations review closes the cycle) that arrive as tracked inputs — and the cycle seeds its own next-cycle carryover the same way.","edges":[{"id":"e-reconcile-outcomes-and-log-exceptions-correct-operational-deviations","label":"Deviations","source":"reconcile-outcomes-and-log-exceptions","target":"correct-operational-deviations","whenValue":"deviations_require_correction"},{"id":"e-reconcile-outcomes-and-log-exceptions-assess-capacity-and-confirm-safeguards","label":"Within procedure","source":"reconcile-outcomes-and-log-exceptions","target":"assess-capacity-and-confirm-safeguards","whenValue":"within_procedure"},{"id":"e-correct-operational-deviations-assess-capacity-and-confirm-safeguards","source":"correct-operational-deviations","target":"assess-capacity-and-confirm-safeguards"},{"id":"e-assess-capacity-and-confirm-safeguards-trigger-capacity-additions-and-alert-personnel","label":"Action required","source":"assess-capacity-and-confirm-safeguards","target":"trigger-capacity-additions-and-alert-personnel","whenValue":"action_required"},{"id":"e-assess-capacity-and-confirm-safeguards-compile-weekly-operations-review","label":"Within thresholds","source":"assess-capacity-and-confirm-safeguards","target":"compile-weekly-operations-review","whenValue":"within_thresholds"},{"id":"e-trigger-capacity-additions-and-alert-personnel-compile-weekly-operations-review","source":"trigger-capacity-additions-and-alert-personnel","target":"compile-weekly-operations-review"}],"isPublic":true,"metadata":{"capabilities":[],"controlVerbs":{},"controls":["UC-BCDR-12","UC-BCDR-05"],"department":"it","domains":["controls"],"library":{"aliases":[],"canonicalUrl":"https://workflow-library.com/all/?w=controls-it-operations-capacity-management-cycle","contentDigest":"sha256:6c08d1ef2edeb59f2859ef25dc798d5e8c3d049174647623a14d5ae0d013b541","prerequisites":{"status":"undeclared"},"provenance":[],"releaseId":"sha256:6c08d1ef2edeb59f2859ef25dc798d5e8c3d049174647623a14d5ae0d013b541","schemaVersion":1,"sourceTemplateId":"workflow-library:controls-it-operations-capacity-management-cycle"},"lineOfDefense":"operate","mappingStatus":"mapped","risks":[],"slug":"controls-it-operations-capacity-management-cycle","source":"coworkcanvas-gallery","standards":["cobit-2019","nist-800-53","nist-csf-2","soc2"],"teams":["it"]},"name":"IT Operations & Capacity Management Cycle","nodes":[{"data":{"decisionField":"operations_disposition","description":"Agent runs and reconciles the cycle's job scheduling, processing, infrastructure monitoring, and facility management against documented procedure and compiles the consolidated exception log; human decides whether deviations require correction","formData":{"fields":[{"key":"operations_disposition","label":"Operations Disposition","options":[{"label":"Within procedure, no correction needed","value":"within_procedure"},{"label":"Deviations require correction","value":"deviations_require_correction"}],"required":true,"type":"select"}],"resultType":"form","submittedAt":null,"values":{}},"instructions":"**Objective** — Execute and reconcile this cycle's job scheduling, processing, infrastructure monitoring, and facility management against documented procedure, then resolve for the IT operations manager whether operations ran within procedure or contain deviations that must be corrected — decided off a single consolidated exception log so nothing is judged on an incomplete picture.\n\n**Inputs**\n- The job-scheduling calendar and actual run history for the in-scope systems — pulled at run time from the external job scheduler; the extract lands as a **step document** (CSV/XLSX) on this step (no System/Job item type exists to hold the schedule natively).\n- Infrastructure monitoring feeds for compute, network, and storage health (alert stream, uptime of each monitored agent, alert-resolution status) and the facility management logs — power, cooling, physical-access checks, and environmental sensor readings for the in-scope facilities — pulled from the external monitoring stack and BMS/facility logs; both extracts land as **step documents** on this step.\n- The documented operating procedure for each job stream (batch, interface transfer, processing run) and for required monitoring coverage and facility-check frequency — the governing **Policy** items (policy_type: procedure, policy_owner: IT operations manager), each with its procedure document attached; they define expected start time, duration, downstream dependency order, which systems must be monitored, and how often physical/environmental checks occur.\n- The capacity-threshold and priority-allocation **Policy** items in force (read-only context — a starved queue or throttled resource explains a late run).\n- Any carryover run or rerun still open, unresolved alert, or missed check from the prior cycle — **Issue** items (issue_type: exception, source: management_identified) created by the prior cycle's close and linked to the anchor Process.\n- The documented tolerance for reliable service delivery (COBIT DSS01; SOC 2 A1.1).\n\n**Procedure**\n_Items 1–11 are agent-run (items 1–5 folded from the former \"Execute scheduled jobs and processing\" step, items 6–9 from the former \"Monitor infrastructure and facility operations\" step); the human moment is item 12._\n1. Pull the job schedule and actual run history for the cycle across batch jobs, interface transfers, and processing runs on every in-scope system.\n2. For each run, compare actual start time, duration, and completion status against the documented procedure and its expected run window. Flag any late start, failure, abnormal termination, rerun, or manual intervention, and note whether a dependency chain was broken (a failed upstream job that left downstream jobs skipped).\n3. Confirm run-order dependencies held — no downstream job started before its prerequisite completed successfully — and that any manual intervention was authorized and logged.\n4. Compile the cycle run-log record summarizing completed, failed, and rerun jobs with counts and the on-time rate; name each flagged run's source system in the run log so it is visible for reconciliation.\n5. Attach the job-processing log and the flagged-run list.\n6. Pull infrastructure monitoring alerts (compute, network, storage) and facility management logs (power, cooling, physical access, environmental sensors) for the cycle.\n7. Verify monitoring coverage was continuous: check each monitored system's agent uptime and flag any system that went dark, any monitoring blind spot, and any alert left unresolved past its response target.\n8. Verify facility checks ran at their documented frequency: flag any missed physical-access or environmental check and any sensor reading outside its safe band (e.g., temperature/humidity outside the data-hall envelope).\n9. Compile the monitoring-and-facilities record for the cycle capturing coverage percentage, open alert count, and missed-check count, naming each flagged gap's affected system or facility; attach the infrastructure monitoring summary and the facility check log.\n10. Consolidate the flagged runs and the monitoring/facility gaps into one exception log; each entry tags the expected outcome, the actual outcome, and the variance, classified by severity and likely cause (transient, configuration, capacity-related, procedural).\n11. Cross-check for any exception already under an open corrective action carried from a prior cycle, and compute the cycle's exception rate against the documented tolerance for reliable service delivery (COBIT DSS01; SOC 2 A1.1).\n12. The IT operations manager reviews the run log, the monitoring-and-facilities record, and the consolidated exception log — confirming processing, monitoring, and facility operations are fully accounted for against procedure — then submits the disposition below.\n\n**Decision criteria**\n- **within_procedure** — Choose when every outcome matches its expected outcome or falls within documented tolerance: no failed job without a successful rerun, no dark monitored system, no missed facility check, no unresolved alert past target, and the exception rate at or below tolerance.\n- **deviations_require_correction** — Choose when one or more exceptions remain open — a failed/late run not yet re-run to success, a monitoring or facility gap, an unresolved alert, or an exception rate above tolerance — such that service delivery is not yet confirmed reliable.\n\n**Record in AssureSwarm**\n- Query the schedule/run-history, monitoring, and facility-log extracts and the open prior-cycle carryover Issues (coach-query-data, coach-workflow-scan).\n- Record the cycle run-log (completed/failed/rerun counts and the on-time rate) and the monitoring-and-facilities record (coverage percentage, open alert count, missed-check count) as **step documents** (XLSX/PDF) attached to this step (coach-document-upload); there is no native operational-record item type, so these step documents are the canonical home (an Operational Record type would be their natural home). Name each flagged run's source system and each flagged gap's affected system or facility inside the documents — no System/Job/Facility item type exists to link them to.\n- Attach the job-processing log, the flagged-run list, the infrastructure monitoring summary, and the facility check log.\n- Attach the consolidated exception log and reconciliation summary (XLSX exception log + summary) on this decision step.\n- Submit the `operations_disposition` SELECT (within_procedure | deviations_require_correction). Record the rationale and evidence references in the step result, and the decision owner in the step's approver record.\n\n**Exit criteria** — Every in-scope job stream is accounted for against its documented window and the run-log document exists with an on-time rate; monitoring coverage and facility-check frequency are evidenced against procedure, with every late, failed, or manually-intervened run, dark system, missed check, and unresolved alert flagged against its named system or facility; the consolidated exception log is complete and attached with the exception rate computed against tolerance; the `operations_disposition` field is submitted with a rationale and owner recorded; the unchosen branch is prunable.","kind":"decision","label":"Reconcile outcomes and log exceptions","performedBy":{"primitives":["coach-query-data","coach-workflow-scan","coach-document-upload","coach-item-create","coach-items-link"]}},"id":"reconcile-outcomes-and-log-exceptions"},{"data":{"description":"Agent drives root-cause correction and reruns for each logged exception; human verifies deviations are corrected and service delivery restored","instructions":"**Objective** — Correct every deviation on the exception log so day-to-day service delivery returns to reliable operation, producing a reviewed corrective-action log before capacity is assessed.\n\n**Inputs**\n- The consolidated exception log and reconciliation summary — the **step document** attached to the reconcile decision step, with each entry's severity and likely cause.\n- The affected system, monitoring feed, or facility named in each exception entry (there is no System/Facility item type to link to — the affected asset is named in the exception-log document).\n- Escalation contacts / accountable owners for systems that need vendor or owner engagement.\n\n**Procedure**\n1. For each exception, draft the corrective action fitted to its cause: rerun a failed batch job, reconfigure a monitoring agent or job dependency, restore a facility check, escalate to the system owner, or engage the vendor. Create a corrective-action **Issue** (issue_type: exception, source: management_identified, with severity, root_cause, remediation_plan, issue_owner, target_remediation_date) and link it to the anchor Process and the relevant Control.\n2. Execute or re-check each correction, then re-run or re-query the affected job, monitoring feed, or facility item and record the actual outcome against the original expected outcome — the exception only clears when the re-check passes.\n3. For any exception that cannot be corrected within the cycle, escalate to the accountable owner with a target date and flag it as explicit carryover so it arrives as an input to the next cycle rather than silently reopening.\n4. Re-tally the exception log: confirm each entry is either cleared by a passing re-check or carried with an owner and target date, and recompute the residual exception rate.\n5. Attach the corrective-action log.\n\n**Record in AssureSwarm**\n- Create one corrective-action **Issue** per exception (coach-item-create) — issue_type: exception, source: management_identified, severity, root_cause, remediation_plan, issue_owner, target_remediation_date; set actual_remediation_date when the re-check passes.\n- Link each corrective-action Issue to the anchor Process and the relevant Control (coach-items-link).\n- Re-run/re-query the affected feed to confirm the fix (coach-query-data).\n- Attach the corrective-action log as a **step document** (XLSX) on this step (coach-document-upload).\n\n**Exit criteria** — Every exception is cleared by a passing re-check (actual_remediation_date set) or carried as an Issue with a named owner and target date; the residual exception rate is recorded; the corrective-action log is attached; the IT operations manager has confirmed reliable service delivery is restored.","label":"Correct operational deviations","performedBy":{"primitives":["coach-query-data","coach-item-create","coach-items-link","coach-document-upload"]}},"id":"correct-operational-deviations"},{"data":{"decisionField":"capacity_disposition","description":"Agent determines capacity and utilization against current and forecast demand and verifies priority-allocation/quota safeguards; human decides whether capacity action is required","formData":{"fields":[{"key":"capacity_disposition","label":"Capacity Disposition","options":[{"label":"Within thresholds, safeguards confirmed","value":"within_thresholds"},{"label":"Capacity action required","value":"action_required"}],"required":true,"type":"select"}],"resultType":"form","submittedAt":null,"values":{}},"instructions":"**Objective** — Resolve, for the IT operations manager, whether processing capacity sits within thresholds with allocation safeguards confirmed, or whether a capacity action is required — determined from utilization versus forecast demand across infrastructure, data, and software.\n\n**Decision criteria**\n- Assess current utilization and forecast demand for infrastructure (compute, storage, network), data volumes, and software licensing/throughput, computing headroom against each resource's documented capacity threshold. Confirm the priority-based allocation rules and quota configurations for shared resources are still in force, correctly ranked, and that no resource has breached or is trending toward its quota. Reference COBIT DSS01 capacity management and NIST SC-6 resource-availability safeguards.\n- Pick **within_thresholds** when every resource sits within its threshold with adequate forecast headroom, and the priority-allocation and quota safeguards are confirmed present and effective for all shared resources.\n- Pick **action_required** when any resource is at or forecast to exceed its threshold within the planning horizon, OR a priority-allocation or quota gap is found (a missing, mis-ranked, or breached quota that leaves a shared resource unprotected).\n\n**Record in AssureSwarm**\n- Submit the `capacity_disposition` SELECT (within_thresholds | action_required).\n- Record the rationale and evidence references in the step result, and the decision owner in the step's approver record.\n- Build the capacity-and-utilization dashboard (coach-dashboard-create) showing current usage, forecast demand, threshold, and trend per resource; query utilization and the forecast model (coach-query-data). The documented capacity thresholds and quota/priority-allocation rules are read from the governing **Policy** items (policy_type: standard/policy). Attach the capacity-and-utilization report as a **step document** (XLSX/PDF, per-resource headroom vs threshold) on this decision step (coach-document-upload).\n\n**Exit criteria** — The capacity report and dashboard exist with per-resource headroom against thresholds, allocation/quota safeguards are evidenced, the `capacity_disposition` field is submitted with a rationale and owner, and the unchosen branch is prunable.","kind":"decision","label":"Assess capacity and confirm safeguards","performedBy":{"primitives":["coach-query-data","coach-dashboard-create","coach-document-upload"]}},"id":"assess-capacity-and-confirm-safeguards"},{"data":{"description":"Agent initiates capacity additions and routes threshold-breach alerts; human confirms capacity added and shared resources protected before demand exceeds thresholds","instructions":"**Objective** — Add capacity ahead of forecast threshold breach, correct any priority-allocation or quota gap, and alert responsible personnel for every actual breach, producing a reviewed capacity-addition and alert-delivery record.\n\n**Inputs**\n- The capacity-and-utilization report (**step document** on the capacity decision step) and the capacity-and-utilization **dashboard**, listing each resource at or forecast to exceed its threshold and any allocation/quota gap found.\n- The provisioning and procurement paths for each resource type (infrastructure scale-out, storage expansion, license/throughput increase) and their lead times.\n- The responsible-personnel routing (on-call / owner contacts) for threshold-breach alerts, and the acknowledgment convention.\n\n**Procedure**\n1. For each resource flagged at or approaching threshold, draft the capacity-addition request sized to close the forecast gap plus target headroom — scale compute, expand storage, raise license count or throughput, or reallocate quota — accounting for provisioning lead time so capacity lands before demand crosses the threshold. Create a capacity-addition **Issue** (issue_type: observation, source: management_identified, remediation_plan = the sized addition, issue_owner, target_remediation_date) — Issue is reused here as the request tracker (a dedicated Request/Action type would be cleaner) — and link it to the anchor Process.\n2. For every resource that has actually breached its defined threshold, route a threshold-breach alert to the responsible personnel and record delivery and acknowledgment; escalate if not acknowledged within the response target.\n3. Where a priority-allocation or quota gap was found, correct the configuration so shared resources remain protected (restore the missing quota, re-rank priorities, tighten the throttle) and re-query to verify the fix held.\n4. Confirm each capacity-addition request has an owner and target date and each breach alert is acknowledged or escalated.\n5. Attach the capacity-addition log and the alert-delivery record.\n\n**Record in AssureSwarm**\n- Create one capacity-addition **Issue** per at-threshold resource (coach-item-create) — issue_type: observation, source: management_identified, remediation_plan = the sized addition, issue_owner, target_remediation_date — and link each to the anchor Process (coach-items-link).\n- Re-verify the corrected allocation/quota configuration in the external infrastructure (coach-query-data).\n- Attach the capacity-addition log and the alert-delivery record (delivery + acknowledgment/escalation) as a **step document** on this step (coach-document-upload); alert delivery/acknowledgment has no native field, so it rides in this document.\n\n**Exit criteria** — Every at-threshold resource has an owned capacity-addition Issue in motion ahead of the forecast breach; every actual breach alert reached and was acknowledged by responsible personnel; priority allocation and quotas again protect shared resources; the IT operations manager has confirmed all three.","label":"Trigger capacity additions and alert personnel","performedBy":{"primitives":["coach-query-data","coach-item-create","coach-items-link","coach-document-upload"]}},"id":"trigger-capacity-additions-and-alert-personnel"},{"data":{"description":"Agent assembles the review minutes, exception log, and capacity report into the cycle's evidence set, then archives it under retention and seeds next cycle's carryover; human signs the weekly review, and that signature closes the cycle","instructions":"**Objective** — Bring daily operations and the capacity review into one weekly evidence set, obtain the IT operations manager's sign-off on it, and preserve that signed set under retention while seeding the next cycle with anything still open — the sign-off recorded here is the cycle's closure.\n\n**Inputs**\n- The job run-log and monitoring-and-facilities records for the cycle (**step documents** on the reconcile decision step).\n- The consolidated exception log and the corrective-action log (or the within_procedure disposition, when no corrections were needed).\n- The capacity report and dashboard, plus any capacity-addition items and breach alerts triggered (or the within_thresholds disposition, when no capacity action was needed).\n- The prior cycle's carryover items, to confirm each was addressed or re-carried, and the list of open corrective actions, in-flight capacity additions, and unacknowledged alerts to carry forward.\n- The designated evidence repository and its retention/immutability policy, the control execution log, and the review-cadence schedule.\n\n**Procedure**\n_Items 7–10 close the workflow (folded from the former \"Close and archive\" step); the sign-off recorded in item 6 is the closure — there is no separate confirmation._\n1. Assemble the review minutes covering the cycle trigger and scope, the job-scheduling/processing results, the infrastructure-monitoring and facility-management results, the exception log and its corrections, the capacity report, and the capacity additions or alerts triggered — querying the final state of each so the minutes reflect closed status, not mid-cycle status.\n2. Prepare the weekly review package and record overall disposition, open items with owners, exception rate, capacity headroom and next cadence date in the step result. Use native approval for the manager’s sign-off.\n3. Cross-check completeness: confirm every exception and every capacity action carries a closed status or an owned carryover date, and scan for any linked item left orphaned (no owner, no resolution).\n4. Reconcile the two decision dispositions into the overall statement (e.g., within_procedure + within_thresholds = clean cycle; any action = summarize what was corrected or added).\n5. Attach the review minutes, exception log, and capacity report as the consolidated evidence set.\n6. The IT operations manager reviews the evidence set for internal consistency and approves the weekly review natively; that signature closes the operating rhythm for this cycle and authorizes the archival and carry-forward below.\n7. Export the full operating record and archive the signed review minutes, exception log, and capacity report in the designated evidence repository under retention controls; record the archive location and reference on the workflow.\n8. Create carry-forward items for any open corrective action, in-flight capacity addition, or unacknowledged alert, and link each to its source so it arrives as an explicit tracked input to next cycle rather than being rediscovered.\n9. Record this cycle's result — exception rate, capacity headroom, and any deficiency — in the closure record (Control carries no native execution-log field, so the quantitative entry rides in the closure document while the linked Control items remain ACTIVE with control_owner confirmed), and confirm next week's review is scheduled on the cadence.\n10. Verify the archived set is retrievable and immutable (a read-back check) and that nothing remains open without a tracked owner, then attach the closure record.\n\n**Record in AssureSwarm**\n- Query the final state of each cycle artifact and the corrective-action and capacity-addition Issues (coach-query-data).\n- Record overall disposition, open items with owners, exception rate, capacity headroom and next cadence date in the step result. Record the IT operations manager’s sign-off through native approval.\n- Confirm no corrective-action or capacity-addition Issue is orphaned (coach-workflow-scan).\n- Attach the weekly operations review minutes (DOCX/PDF), the consolidated exception log, and the capacity report as the consolidated evidence set — a **step document** on this step (coach-document-upload).\n- Export the operating record (coach-workflow-export); the workflow instance itself, attached to the anchor Process item, is the durable audit trail.\n- Create carry-forward **Issue** items (coach-item-create) for every open corrective action, in-flight capacity addition, or unacknowledged alert — issue_type: exception, source: management_identified, issue_owner, target_remediation_date — and link each to its source Issue/step and to the anchor Process (coach-items-link); these become next instance's prior-cycle carryover input.\n- Attach the exported operating record and the closure record (retention/immutability evidence, archive reference, and the cycle-result metrics) as a **step document** on this step.\n\n**Exit criteria** — The evidence set is assembled, internally consistent, and signed by the IT operations manager; every exception and capacity action is closed or carried as an owned Issue with a target date; the archived set is immutable and retrievable under retention with its location referenced; the cycle-result metrics are recorded in the closure record and next week's review is scheduled on the cadence — the signed weekly review is the cycle's formal closure.","label":"Compile weekly operations review","performedBy":{"primitives":["coach-query-data","coach-form-create","coach-workflow-scan","coach-document-upload","coach-workflow-export","coach-item-create","coach-items-link"]}},"id":"compile-weekly-operations-review"}],"sourceTemplateId":"workflow-library:controls-it-operations-capacity-management-cycle"}
