{"description":"Each quarterly instance runs against the existing Control item for independent third-party AI evaluation (framework aiuc-1 + iso-42001 + eu-ai-act; domains ai_governance; frequency quarterly; control_owner AI Product Lead) — the run enriches that Control's execution history and is its evidence of operation, it never creates a duplicate Control; the one record it does create is a per-quarter Audit item (audit_type: it_audit) for the evaluator engagement. The decision-aware cycle confirms the quarter's in-scope systems and risk taxonomy, engages an independent evaluator with the test plan, categories, and pass thresholds fixed in advance, provisions contained test access, triages every finding by category and severity against the tested control, routes failed thresholds through remediation and evaluator retest, gates acceptance of the evaluator report, publishes the accepted evidence (trust-portal summary, customer-facing attestation, evidence register), and tunes guardrails from the findings. In scope: every in-scope AI system and agent under the AIUC-1 program and its six mandatory third-party test categories (B001 adversarial robustness, C010 harmful outputs, C011 out-of-scope outputs, C012 agent-specific risk, D002 hallucinations, D004 tool calls). Out of scope: internal red-teaming, pre-release model evaluation, and vendor AI due diligence, which run in their own workflows. It hands off only to the next quarterly run of itself, seeding it with the open carry-forward finding Issues that the publish-evidence-and-tune-guardrails step links to the anchor Control.","edges":[{"id":"e-engage-evaluator-and-fix-test-plan-provision-test-access-and-safeguards","source":"engage-evaluator-and-fix-test-plan","target":"provision-test-access-and-safeguards"},{"id":"e-provision-test-access-and-safeguards-assess-threshold-results","source":"provision-test-access-and-safeguards","target":"assess-threshold-results"},{"id":"e-assess-threshold-results-remediate-failed-categories-and-retest","label":"Thresholds failed","source":"assess-threshold-results","target":"remediate-failed-categories-and-retest","whenValue":"thresholds_failed"},{"id":"e-assess-threshold-results-accept-evaluator-report","label":"All thresholds met","source":"assess-threshold-results","target":"accept-evaluator-report","whenValue":"all_thresholds_met"},{"id":"e-remediate-failed-categories-and-retest-accept-evaluator-report","source":"remediate-failed-categories-and-retest","target":"accept-evaluator-report"},{"id":"e-accept-evaluator-report-return-report-for-rework","label":"Return for rework","source":"accept-evaluator-report","target":"return-report-for-rework","whenValue":"returned_for_rework"},{"id":"e-accept-evaluator-report-publish-evidence-and-tune-guardrails","label":"Accepted","source":"accept-evaluator-report","target":"publish-evidence-and-tune-guardrails","whenValue":"accepted"},{"id":"e-return-report-for-rework-publish-evidence-and-tune-guardrails","source":"return-report-for-rework","target":"publish-evidence-and-tune-guardrails"}],"isPublic":true,"metadata":{"capabilities":[],"controlVerbs":{"UC-AI-18":"tests","UC-AI-19":"tests","UC-AI-20":"tests"},"controls":["UC-AI-21","UC-AI-18","UC-AI-20","UC-AI-19"],"department":"ai-governance","domains":["controls"],"library":{"aliases":[],"canonicalUrl":"https://workflow-library.com/all/?w=controls-quarterly-third-party-ai-evaluation-cycle","contentDigest":"sha256:e542a1851c8c176ba2d527bcdd1e574e8c9b9b5ee51fbbb9ca5a86307c7b2d84","prerequisites":{"status":"undeclared"},"provenance":[],"releaseId":"sha256:e542a1851c8c176ba2d527bcdd1e574e8c9b9b5ee51fbbb9ca5a86307c7b2d84","schemaVersion":1,"sourceTemplateId":"workflow-library:controls-quarterly-third-party-ai-evaluation-cycle"},"lineOfDefense":"operate","mappingStatus":"mapped","risks":[],"slug":"controls-quarterly-third-party-ai-evaluation-cycle","source":"coworkcanvas-gallery","standards":["nist-ai-tevv-athlon","aiuc-1","iso-42001","eu-ai-act"],"teams":["ai-governance","procurement"]},"name":"Quarterly Third-Party AI Evaluation Cycle","nodes":[{"data":{"controls":["UC-AI-21"],"description":"Approve complete system/category scope, evaluator independence, methodology and locked thresholds before access is provisioned.","formData":{"fields":[{"key":"independence_status","label":"Independence from the AI product team","options":[{"label":"Fully independent — no role in building, tuning, or operating the in-scope systems","value":"fully_independent"},{"label":"Prior or current involvement exists — disclosed below","value":"involvement_disclosed"}],"required":false,"type":"select"},{"key":"prior_involvement_detail","label":"If involvement exists, describe it and any safeguard applied","required":false,"type":"textarea"},{"key":"named_testers_and_qualifications","label":"Named testers and their qualifications for each test category","required":false,"type":"textarea"},{"key":"methodology_and_attack_set_provenance","label":"Methodology per category and provenance of the attack sets and benchmarks used","required":false,"type":"textarea"},{"key":"testing_window_start","label":"Testing window start date","required":false,"type":"date"},{"key":"testing_window_end","label":"Testing window end date","required":false,"type":"date"}],"resultType":"form","submittedAt":null,"values":{}},"instructions":"**Objective** — Approve complete system/category scope, evaluator independence, methodology and locked thresholds before access is provisioned.\n\n**Inputs**\n- The anchor Control item (the quarterly independent third-party AI evaluation control, framework aiuc-1 + iso-42001 + eu-ai-act) this quarterly instance runs against — the run enriches its execution history, it never creates a duplicate Control.\n- The AI system inventory and agent register: every in-scope production system, including each agent's tool set and authorized scope — queried live from the external model registry (coach-query-data); no native AI System item type exists, so the query result is retained on this step as the point-in-time scope evidence.\n- The current risk taxonomy per system: adversarial robustness and jailbreak resistance, harmful outputs, out-of-scope outputs, agent-specific high-risk outputs, hallucination rates, and unsafe or unauthorized tool calls.\n- The prior quarter's evaluation record and any open carry-forward Issue that the prior quarter's publication and closure step linked to the anchor Control.\n- The confirmed quarter scope memo from the scope step.\n- The evaluator's independence attestation (no role in building, tuning, or operating the in-scope systems), tester qualifications, and prior methodology — solicited through this step's form.\n- The harmful-output taxonomy, each system's intended-use and scope statement, each agent's tool allow-list, and last quarter's threshold values.\n\n**Missing-input gate** — Before any form request below, inspect the existing source records, reports and correspondence. Reuse every established fact and record its source. Send a form only when a listed fact remains genuinely unresolved and the named respondent is outside the complete roster of people executing or approving any checkpoint in this workflow. If the respondent is on that roster, record their contribution in native results and approvals. Ask only the unresolved fields; leave known, unasked or inapplicable fields optional and blank. Skip the form entirely when no missing facts remain. Attach evidence documents and record sign-off through native approval. References below to form answers or completion also accept the existing authoritative record or native participant contribution.\n\n**Procedure**\n_This checkpoint absorbs “Confirm quarter scope and risk taxonomy”. The agent runs the preparation, evidence assembly and record updates below; the named owners retain the substantive decisions and approvals stated in the procedure._\n1. Confirm quarter scope and risk taxonomy: Fix which AI systems and agents this quarter's independent third-party evaluation covers and which risk categories each must be tested against, so the evaluator's scope is complete before any engagement terms are drafted (AIUC-1 B001, C010, C011, C012, D002, and D004 each require the test every 3 months).\n2. Query the inventory (coach-query-data) and list every system released, materially changed, or newly agent-enabled since the last evaluation; a system added mid-quarter enters scope now, not next quarter.\n3. Assign each system its applicable categories: every system carries adversarial-robustness, harmful-output, out-of-scope, and hallucination testing; every agent additionally carries agent-specific high-risk output and tool-call testing.\n4. Pull the open carry-forward Issues (coach-query-data) and mark each as a mandatory retest target for the category it failed.\n5. Draft the quarter scope memo — systems, categories, retest targets, and exclusions with rationale — and attach it (coach-document-upload).\n6. Engage evaluator and fix test plan and thresholds: Engage an independent third-party evaluator and fix the test scope, methodology, and pass thresholds in advance, so results are judged against criteria agreed before testing rather than negotiated after (AIUC-1 B001, C010, C011, C012, D002, D004).\n7. Send this step's form to the evaluator's engagement lead and confirm independence and qualification from the answers (coach-query-data): the evaluator and its named testers must be organizationally and commercially independent of the AI product team; record the attestation. An `involvement_disclosed` answer must be resolved — safeguard documented or evaluator replaced — before the plan is approved.\n8. Fix the test plan per system and category — adversarial robustness and jailbreak resistance (B001), harmful outputs (C010), out-of-scope outputs (C011), agent-specific high-risk outputs (C012), hallucination rates (D002), and unsafe or unauthorized tool calls (D004) — with attack-set sizes, sampling approach, grading rubric, and reproducibility requirements.\n9. Set the pass threshold per category in advance (for example maximum jailbreak success rate, maximum harmful-output rate, maximum hallucination rate on the grounded benchmark, zero tool calls outside authorized scope) and lock them so they cannot move after results arrive.\n10. Create the engagement record (coach-item-create): one Audit item, audit_type: it_audit, scope = the systems and categories, lead_auditor = the evaluator lead, period_start and period_end = the testing window from the form; link it to the anchor Control (coach-items-link).\n11. Attach the signed engagement letter, test plan, and locked threshold sheet (coach-document-upload).\n\n**Record in AssureSwarm**\n- Step document — the quarter scope memo (coach-document-upload) with the inventory query result embedded as evidence.\n- Item relationship — each carry-forward finding Issue linked to the anchor Control as this quarter's retest input (coach-items-link).\n- Step form — the form on this step, answered by the independent evaluator's engagement lead, uses the assigned engagement lead identity and captures independence status and any disclosed involvement, named testers and qualifications, per-category methodology and attack-set provenance, the testing window; obtain the signed independence attestation as an attached document.\n- Item create — the evaluator engagement Audit item (audit_type: it_audit) with scope, lead_auditor, period_start, period_end (coach-item-create), linked to the anchor Control (coach-items-link).\n- Step document — the signed engagement letter, test plan, and locked threshold sheet (coach-document-upload).\n\n**Exit criteria**\n- Every in-scope system and agent carries its risk categories; every exclusion has a written rationale; the AI Product Lead confirms the scope memo before the evaluator is engaged.\n- Independence is attested through the evaluator's form answers with any disclosure resolved; every system-category pair has a method and a locked pass threshold; the AI Product Lead approves the plan before access is provisioned.\n\n**Form recipient** — The independent evaluator engagement lead supplies only independence status and disclosed involvement; tester qualifications; methodology and attack-set provenance; proposed testing window only when this gate permits the assigned form. The workflow executor records their own analysis in the step result.","label":"Engage evaluator and fix test plan and thresholds","performedBy":{"primitives":["coach-query-data","coach-items-link","coach-document-upload","coach-item-create"]}},"id":"engage-evaluator-and-fix-test-plan"},{"data":{"controls":["UC-AI-21"],"description":"Agent provisions evaluator access, production-representative environments, and blast-radius safeguards for agent and tool-call testing; human verifies the setup is realistic yet contained","instructions":"**Objective** — Give the independent evaluator representative access to every in-scope system and agent while containing the blast radius of adversarial, harmful-output, and tool-call testing, so the evaluation is realistic without harming production users or data.\n\n**Inputs**\n- The approved test plan and testing window from the engagement step.\n- Each system's deployment topology: production endpoints, staging mirrors, model versions, and the guardrail configuration the evaluator must face (input screening, output filtering, endpoint rate limiting, tool allow-lists, approval gates, sandboxing).\n- Data-handling and confidentiality terms from the engagement letter.\n\n**Procedure**\n1. Provision evaluator accounts and API keys scoped to the testing window on environments that mirror production model versions and guardrail configuration (coach-query-data to confirm version parity); a test against a weaker staging configuration does not evidence the production control.\n2. Confirm endpoint rate limits and anti-scraping controls stay enabled unless the plan explicitly exempts a test, and record any exemption with its expiry.\n3. For agent and tool-call testing, stand up sandboxed tool backends with synthetic data so an unauthorized or irreversible action is observed rather than executed, and confirm approval gates remain in the path.\n4. Agree the stop conditions and the incident channel for any live harm the evaluator surfaces, and notify the on-call owner and security team of the window (coach-notify).\n5. Attach the access record, environment parity confirmation, exemption log, and safeguard checklist (coach-document-upload).\n\n**Record in AssureSwarm**\n- Step document — the access record, environment parity confirmation, exemption log, and safeguard checklist (coach-document-upload); credentials themselves are never stored in AssureSwarm.\n- Notification — the on-call owner and security team notified of the testing window and stop conditions (coach-notify).\n\n**Exit criteria** — Every in-scope system and agent is reachable by the evaluator on a production-representative configuration; tool-call testing is sandboxed; stop conditions are agreed; the AI Product Lead confirms the setup before testing begins.","label":"Provision test access, environments, and safeguards","performedBy":{"primitives":["coach-query-data","coach-notify","coach-document-upload"]}},"id":"provision-test-access-and-safeguards"},{"data":{"controls":["UC-AI-21","UC-AI-18","UC-AI-20","UC-AI-19"],"decisionField":"threshold_outcome","description":"Agent ingests the evaluator's results, books every reproducible failure as an Issue mapped to the control it tested, and computes each system's observed rate beside its locked threshold; human decides whether every in-scope system met its threshold in every category, with any failure routing to remediation and retest","formData":{"fields":[{"key":"threshold_outcome","label":"Threshold Outcome","options":[{"label":"All thresholds met","value":"all_thresholds_met"},{"label":"One or more thresholds failed","value":"thresholds_failed"}],"required":true,"type":"select"}],"resultType":"form","submittedAt":null,"values":{}},"instructions":"**Objective** — Turn the evaluator's raw results into a triaged finding register and then decide whether every in-scope system met its pre-agreed pass threshold in every tested category, so failed categories go to remediation and evaluator retest rather than being explained away after the fact. Owned by the AI Product Lead.\n\n**Inputs**\n- The evaluator's results package: per-category pass rates, reproduction transcripts, attack sets, and methodology notes.\n- The locked threshold sheet and test plan from the engagement step.\n- The engagement Audit item and the anchor Control.\n\n**Procedure**\n_Items 1–6 are agent-run (folded from the former \"Receive results and triage findings\" step); the human moment is the threshold decision in item 7._\n1. Ingest the results (coach-query-data) and compute each system's observed rate per category beside its locked threshold: jailbreak success (B001), harmful outputs (C010), out-of-scope outputs (C011), agent-specific high-risk outputs (C012), hallucination rate (D002), unsafe or unauthorized tool calls (D004).\n2. Create one Issue per reproducible failure (coach-item-create): issue_type: finding, source: external_audit, severity by exploitability and harm, identified_date = the report date, description naming the system, category, prompt or tool sequence, and observed behaviour.\n3. Link each Issue to the engagement Audit and to the tested Control (coach-items-link): injection and jailbreak successes to UC-AI-18; harmful, out-of-scope, and hallucinated outputs to UC-AI-20; tool calls outside authorized scope or past an approval gate to UC-AI-19.\n4. Verify a sample of reproductions independently; a finding that does not reproduce is disputed with the evaluator in writing, never silently dropped.\n5. Build the comparison table (coach-query-data): system by category, locked threshold, observed rate, pass or fail, and the linked finding Issues. Flag every critical-severity finding and every incomplete category as an automatic fail.\n6. Attach the triage register, reproduction sample log, threshold comparison, and decision brief (coach-document-upload).\n7. Read the comparison against the criteria below and submit the branch: every category on every system must carry an observed rate beside its threshold and every reproducible failure must be an owned Issue mapped to the control it tested before the decision is made.\n\n**Decision criteria**\n- Select `all_thresholds_met` when every system's observed rate is at or better than its locked threshold for adversarial robustness (B001), harmful outputs (C010), out-of-scope outputs (C011), agent-specific high-risk outputs (C012), hallucination rate (D002), and tool calls (D004), and no critical-severity finding exists regardless of rate.\n- Select `thresholds_failed` when any system misses any category threshold, when any single finding is critical (a working jailbreak producing harmful output, a tool call outside authorized scope that would have been irreversible, or exposure of the system prompt or internal data), or when the evaluator could not complete a category. Thresholds fixed at engagement are not renegotiated here.\n\n**Record in AssureSwarm**\n- Step form — submit `threshold_outcome` (SELECT) with the step result and the step's approver record (coach-form-fill).\n- Item create — one finding Issue per reproducible failure (coach-item-create): issue_type: finding, source: external_audit, severity, identified_date, description.\n- Item relationship — each finding linked to the engagement Audit and its tested Control (coach-items-link).\n- Step document — the triage register, reproduction sample log, threshold comparison, and decision brief (coach-document-upload).\n\n**Exit criteria** — Every category on every system has an observed rate beside its threshold; every reproducible failure is an owned Issue mapped to the control it tested; the form is submitted with rationale and owner and the unused branch is prunable. `all_thresholds_met` proceeds straight to report acceptance; `thresholds_failed` routes through remediation and retest first.","kind":"decision","label":"Assess results against pre-agreed thresholds","performedBy":{"primitives":["coach-query-data","coach-form-fill","coach-document-upload","coach-item-create","coach-items-link"]}},"id":"assess-threshold-results"},{"data":{"controls":["UC-AI-18","UC-AI-20","UC-AI-19","UC-AI-21"],"description":"Agent drives each failed finding through an owned fix at the control that failed and commissions the evaluator's retest of the category; human confirms every failure is verified closed or carried with a dated owner","instructions":"**Objective** — Fix every failed category at the control that failed and have the independent evaluator retest it, so closure is evidenced by a repeated third-party test rather than by an internal claim (AIUC-1 requires every finding tracked to remediation and retest).\n\n**Inputs**\n- The finding Issues routed here by the threshold decision, each linked to UC-AI-18, UC-AI-20, or UC-AI-19.\n- The evaluator's reproduction transcripts and attack sets for each failed category.\n- The retest terms in the engagement letter.\n\n**Procedure**\n1. For each finding, set the remediation on the Issue (coach-item-update): remediation_plan naming the concrete control change — a tightened injection or jailbreak screen or endpoint rate limit (UC-AI-18); an output filter, scope limit, grounding step, or withheld system prompt (UC-AI-20); a narrowed tool allow-list, task-scoped permission, approval gate, or sandbox (UC-AI-19) — plus issue_owner and target_remediation_date sized to severity.\n2. Commission the evaluator's retest of each failed category using the original attack set plus variants (coach-notify), so a fix that blocks only the exact reproduction is caught.\n3. On a passed retest, set actual_remediation_date and verified_date on the Issue (coach-item-update); on a failed retest, leave verified_date clear, raise severity if warranted, and record the next fix.\n4. Link the retest evidence to the engagement Audit (coach-items-link) and attach the evaluator retest report and remediation log (coach-document-upload).\n\n**Record in AssureSwarm**\n- Item field update — remediation_plan, issue_owner, target_remediation_date on every routed finding; actual_remediation_date and verified_date on retest-passed findings (coach-item-update).\n- Item relationship — retest evidence linked to the engagement Audit (coach-items-link).\n- Step document — the evaluator retest report and remediation log (coach-document-upload).\n\n**Exit criteria** — Every failed category has a retest result from the independent evaluator; every finding is verified closed or carries a dated owner; the AI Product Lead confirms before the report is judged.","label":"Remediate failed categories and commission retest","performedBy":{"primitives":["coach-item-update","coach-items-link","coach-notify","coach-document-upload"]}},"id":"remediate-failed-categories-and-retest"},{"data":{"controls":["UC-AI-21"],"decisionField":"report_acceptance","description":"Human decides whether the evaluator's final report is complete, reproducible, and consistent with the triage and retest record, or must go back for rework","formData":{"fields":[{"key":"report_acceptance","label":"Report Acceptance","options":[{"label":"Accept report","value":"accepted"},{"label":"Return for rework","value":"returned_for_rework"}],"required":true,"type":"select"}],"resultType":"form","submittedAt":null,"values":{}},"instructions":"**Objective** — Decide whether the independent evaluator's final report can be accepted as this quarter's evaluation evidence or must be returned for rework. Owned by the AI Product Lead; the report is what customers, the AIUC-1 certifier, and regulators will read.\n\n**Decision criteria**\n- Select `accepted` when the report covers every in-scope system and every category (B001, C010, C011, C012, D002, D004), states the locked thresholds and observed rates, documents methodology and attack sets well enough to reproduce, reflects every retest result, and carries the evaluator's independence statement and signature.\n- Select `returned_for_rework` when any system or category is missing, thresholds or rates differ from the triage register, retest outcomes are absent or misstated, methodology is insufficient to reproduce, or unsupported conclusions appear. Disagreement with a finding is not grounds for return; only accuracy and completeness are.\n\n**Agent procedure** (produces the acceptance brief this decision reads)\n1. Reconcile the draft report against the triage register, threshold comparison, and retest log (coach-query-data), listing every discrepancy by page and finding.\n2. Check that each finding Issue's status matches the status the report states for it.\n3. Draft the acceptance brief with the discrepancy list and attach it together with the draft report (coach-document-upload).\n\n**Record in AssureSwarm**\n- Step form — submit `report_acceptance` (SELECT) with the step result and the step's approver record (coach-form-fill).\n- Step document — the draft report and acceptance brief (coach-document-upload).\n\n**Exit criteria** — The form is submitted with rationale and owner and the unused branch is prunable. `accepted` proceeds to publication; `returned_for_rework` routes through the rework step first.","kind":"decision","label":"Accept the evaluator report","performedBy":{"primitives":["coach-query-data","coach-form-fill","coach-document-upload"]}},"id":"accept-evaluator-report"},{"data":{"controls":["UC-AI-21"],"description":"Agent returns the itemized rework list to the evaluator, tracks the corrected report to its deadline, and re-reconciles it; human confirms every rework point is resolved before publication","instructions":"**Objective** — Return the report to the independent evaluator with an itemized rework list, obtain a corrected report inside the quarter, and confirm every rework point is resolved, so the evidence published is accurate without slipping the quarterly cadence.\n\n**Inputs**\n- The acceptance brief and discrepancy list from the acceptance decision.\n- The engagement letter's rework and delivery terms.\n- The quarter-end date by which accepted evidence must exist (AIUC-1 requires the third-party evaluation every 3 months).\n\n**Procedure**\n1. Send the evaluator the rework list — each discrepancy with its evidence reference and the correction required — together with the corrected-report deadline (coach-notify).\n2. Track the rework to its deadline (coach-workflow-scan); if the corrected report will not land before quarter end, create an Issue (coach-item-create) — issue_type: exception, severity: high, source: management_identified, target_remediation_date = quarter end — linked to the anchor Control, and escalate it.\n3. Re-reconcile the corrected report against the triage register, threshold comparison, and retest log, confirming every rework point is closed and no new discrepancy was introduced.\n4. Attach the rework list, evaluator correspondence, corrected report, and re-reconciliation memo (coach-document-upload).\n\n**Record in AssureSwarm**\n- Item create — an at-risk Issue only when the rework threatens the quarterly cadence (coach-item-create): issue_type: exception, severity: high, source: management_identified, target_remediation_date = quarter end.\n- Step document — the rework list, evaluator correspondence, corrected report, and re-reconciliation memo (coach-document-upload).\n\n**Exit criteria** — Every rework point is resolved in the corrected report; the corrected report matches the triage and retest record; the AI Product Lead confirms the corrected report as accepted evidence before publication.","label":"Return report for rework and re-review","performedBy":{"primitives":["coach-notify","coach-workflow-scan","coach-item-create","coach-document-upload"]}},"id":"return-report-for-rework"},{"data":{"controls":["UC-AI-21"],"description":"Agent publishes the accepted evaluation evidence to the trust portal, customer attestation, and evidence register, turns recurring findings into guardrail tuning actions, and archives the quarterly record under retention while seeding next quarter's scope; human approves what is disclosed and what changes, and that approval closes the cycle","instructions":"**Objective** — Publish the accepted quarterly evaluation evidence in the forms customers and the AIUC-1 certifier rely on, feed what the evaluation found back into the interface, output, and agent guardrails, and close the quarter's record under retention — the AI Product Lead's approval of the disclosure is the cycle's closure.\n\n**Inputs**\n- The accepted evaluator report (directly, or the corrected report from the rework step).\n- The finding Issues with their remediation and retest status, including any not yet verified closed.\n- The trust-portal template, the customer-facing attestation template, and the evidence register.\n- The evidence repository's retention controls and the next quarter's evaluation date.\n\n**Procedure**\n_Items 1–5 publish and tune; items 6–9 close the workflow (folded from the former \"Close and archive\" step) — the disclosure approval recorded on this step is the closure, with no separate confirmation._\n1. Close the engagement record (coach-item-update): on the Audit item set rating (satisfactory when every threshold was met or verified closed on retest, otherwise needs_improvement), report_date, and fieldwork_end.\n2. Publish the trust-portal summary — systems and categories tested, thresholds, pass results, remediation status — and issue the customer-facing attestation naming the independent evaluator, the quarter, and the categories covered (B001, C010, C011, C012, D002, D004); notify customer-facing owners (coach-notify).\n3. File the report, methodology, attack-set references, and remediation evidence in the evidence register (coach-document-upload).\n4. Build the evaluation dashboard (coach-dashboard-create): pass rate per category per system across quarters, open findings by control, and retest cycle time.\n5. Convert recurring failure patterns into guardrail tuning requests for the responsible owners — injection and jailbreak screening and endpoint limits (UC-AI-18), output filtering and grounding (UC-AI-20), tool allow-lists and approval gates (UC-AI-19) — each as a dated action noted on the related finding Issue (coach-item-update).\n6. Export the full operating record (coach-workflow-export) — scope memo, engagement and locked thresholds, access and safeguard records, triage register, both decisions, retest report, accepted report, and published evidence — and archive it under retention controls, recording the archive reference.\n7. Confirm every open finding and tuning action is linked to the anchor Control (coach-items-link) so it arrives as a mandatory retest target in next quarter's scope step; do not re-create Issues already booked this cycle.\n8. Move the engagement Audit item's status to complete (coach-item-update) and confirm the next quarterly evaluation is scheduled within 3 months of this one, with the evaluator's independence to be re-attested at engagement.\n9. Attach the closure record (coach-document-upload) and record the AI Product Lead's approval of what is disclosed and what changes — that approval closes the cycle.\n\n**Record in AssureSwarm**\n- Item field update — engagement Audit rating, report_date, fieldwork_end, then status moved to complete; guardrail-tuning actions noted on the related finding Issues (coach-item-update).\n- Dashboard — the quarterly evaluation dashboard (coach-dashboard-create).\n- Item relationship — every open finding and tuning action linked to the anchor Control as next quarter's input (coach-items-link).\n- Workflow instance — the full operating record exported (coach-workflow-export) and archived under retention controls; the completed instance against the anchor Control IS the control execution record.\n- Step document — the trust-portal summary, customer-facing attestation, evidence-register filing, and closure record (coach-document-upload).\n\n**Exit criteria** — The attestation and trust-portal summary are published; the evidence register holds the report, methodology, and remediation evidence; every tuning action has an owner and date; the archived record is immutable and retrievable; the next quarterly evaluation is scheduled; nothing remains open without a tracked owner; the AI Product Lead's approval of the disclosure is recorded and closes the quarterly cycle.","label":"Publish evaluation evidence and tune guardrails","performedBy":{"primitives":["coach-item-update","coach-dashboard-create","coach-notify","coach-document-upload","coach-workflow-export","coach-items-link"]}},"id":"publish-evidence-and-tune-guardrails"}],"sourceTemplateId":"workflow-library:controls-quarterly-third-party-ai-evaluation-cycle"}
