NIST TEVV-Athlon (draft, Aug 2026)
50 records. Direct records match this source; context records explain their connections.
Read the first JSON page · Data retrieval guide
Mappings may provide partial coverage. Read mapping properties, residual requirements and source notes before relying on a connection.
control · Direct
NIST-TEVV-01 — Define evaluation objectives, context, and measurements
The framework starts with organizational goals and the system's operational context, then constructs measurement concepts, test events, and tools that address those objectives.
control · Direct
NIST-TEVV-02 — Run evaluations and examine results and limitations
The framework applies the selected tests and synthesizes their results into evidence about system performance, while interrogating the measurements and what the findings support.
control · Direct
NIST-TEVV-03 — Evaluate AI systems in realistic operating settings
The draft discusses testing in settings that better reflect real use, including interactions with users and the operating environment, to complement model tests and benchmarks.
control · Direct
NIST-TEVV-04 — Test for disclosure of confidential information
Appendix B includes confidentiality attacks that test whether an AI system reveals confidential information or internal functionality, including user information and system-prompt leakage.
control · Direct
NIST-TEVV-05 — Test direct and indirect prompt injection
Appendix B includes integrity tests for direct and indirect prompt injection, poisoning, obfuscated inputs, and retrieval weaknesses that could alter intended outputs or outcomes.
control · Direct
NIST-TEVV-06 — Test agent tool misuse and unauthorized external actions
Appendix B includes tests for misuse of connected tools and external actions, including unsafe tool selection, excessive agency, unauthorized action attempts, and harmful task execution.
risk · Context
AI accountability gaps and organizational liability
Ambiguous responsibility across developers, deployers, and operators means no party is clearly accountable when AI causes harm; organizations face reputational damage, regulatory sanctions, and civil liability from AI failures at scale, and operational disruption when AI is unavailable.
risk · Context
Adversarial attacks, data poisoning and prompt injection
Data-poisoning corrupts training data and embeds backdoors; adversarial evasion, prompt injection, and jailbreaks fool deployed models at inference; model extraction steals proprietary weights/logic — enabling harmful or policy-violating outputs.
risk · Context
Unauthorized or unsafe autonomous agent actions and tool calls
AI agents with excessive permissions or weak action controls execute tool calls, transactions, or code outside their authorized task scope — through prompt injection, misinterpretation, or emergent behaviour — causing data loss, financial loss, or irreversible changes in connected systems.
risk · Context
Harmful AI bias and discrimination against protected groups
Models encode/amplify historical bias, producing allocative harm (biased hiring/lending/housing/benefits), representational harm (stereotyping), and evaluation bias masking disparate subgroup performance — systematically disadvantaging protected groups.
risk · Context
Misuse of AI systems for offensive cyber operations or catastrophic harm
Users obtain meaningful uplift from an AI system for offensive cyber operations (malware development, vulnerability exploitation, autonomous intrusion) or for biological, chemical, nuclear, or radiological harm, exposing the deploying organization to severe legal, regulatory, and societal consequences.
risk · Context
Emergent behaviour and unsafe AI system integration
Large-scale or multi-model pipelines exhibit emergent capabilities/failures not present in any component and not predictable from component testing; integration with legacy systems introduces interface mismatches and configuration errors.
risk · Context
AI endpoint abuse, scraping and model extraction
Adversaries scrape inference endpoints at scale to extract model behaviour or proprietary data, exhaust compute budgets through unbounded consumption, or harvest system prompts and technical details disclosed in outputs or documentation, degrading service and eroding competitive and security posture.
risk · Context
Rights violations from AI in law enforcement
Because AI for individual risk assessment, evidence evaluation, profiling, or predictive policing (Annex III(6)) operates without strict accuracy, human oversight, logging, and fundamental-rights safeguards, it can drive wrongful enforcement action, resulting in unjust detention, profiling harm, and rights violations.
risk · Context
Inaccurate, unreliable or hallucinated AI outputs
AI outputs contain factual errors, hallucinations, or confidently wrong predictions; inappropriate proxy metrics, overfitting/underfitting, or insufficient pre-deployment testing undermine trust in decisions made on their basis.
risk · Context
Inappropriate human/AI task allocation and end-of-life risk
Tasks needing contextual judgment, ethics, or accountability are improperly delegated to AI; and decommissioning raises data-persistence in weights, loss of institutional knowledge, and service-continuity gaps.
risk · Context
Insufficient human oversight and automation complacency
Because consequential AI decisions run under full automation with reviewers who lack the authority, information, or AI literacy to intervene, meaningful human control is absent and automation complacency erodes vigilance, so erroneous or harmful automated decisions reach individuals unchecked.
risk · Context
Model/data drift and inadequate post-deployment monitoring
Distribution shift between training and deployment data silently degrades accuracy, and without ongoing monitoring, model decay and emerging failure modes go undetected with no trigger to retrain or decommission. Uncontrolled updates alter behaviour and invalidate prior assessments.
risk · Context
Insufficient AI resilience and fallback mechanisms
AI systems lacking fallback, redundancy, or graceful degradation fail catastrophically under adversarial conditions, infrastructure outages, or out-of-distribution inputs, disrupting dependent business processes.
risk · Context
AI privacy leakage and re-identification
Model inversion and membership-inference attacks reconstruct training data or reveal individuals in the training set; AI inference re-identifies anonymized data and infers sensitive attributes; training on data without consent/legal basis creates regulatory liability.
risk · Context
AI safety failures causing physical or psychological harm
AI errors in safety-critical systems (autonomous vehicles, medical devices, industrial controls) cause injury or death; safety-constraint violations by agentic AI, cascading failures across coupled systems, and AI-generated misinformation/deepfakes cause harm.
risk · Context
Credential and secret leakage through AI inputs, outputs, logs and generated code
API keys, tokens, private keys, and connection strings pasted into prompts, returned in outputs, hardcoded in generated code, or captured in conversation logs are exposed to unauthorized parties or persisted outside secret management, enabling account takeover and lateral movement.
risk · Context
User error and mishandling of sensitive information
Authorized users make mistakes — incorrect data entry, misconfiguration, improper procedures, incorrect privilege settings, or spilling/mishandling sensitive information — causing harm to information assets without malicious intent.
risk · Context
Internet-exposed or misconfigured systems
Adversary gains access through the Internet to systems not authorized for Internet connectivity or that do not meet configuration requirements, and exploits attacks over unauthorized ports, protocols, and services.
risk · Context
Unauthorized disclosure / breach of sensitive information
Unauthorized disclosure of information to parties not entitled to receive it, whether by insecure controls (insecurity), spillage, or authorized users induced to expose data — resulting in identity theft, economic loss, and loss of trust.
risk · Context
Data exfiltration and theft of information by attackers
Adversary (outsider, insider, nation-state, or competitor) installs malware or sniffers to locate and exfiltrate sensitive/proprietary information, or steals data by external actors — including systems-security losses from hacking.
risk · Context
Privacy-program non-compliance (GDPR, CCPA, state laws)
Failure to honour data-subject rights (access, deletion, portability, restriction) on time, missing lawful-basis/consent documentation, defective consent mechanisms, invalid cross-border transfer mechanisms, or inadequate notices — driving fines and private rights of action.
risk · Context
Cross-border personal-data transfer without safeguards
Transferring personal data to jurisdictions lacking equivalent protection without SCCs, BCRs, adequacy decisions, or other recognized mechanisms, exposing individuals and the organization to legal risk.
risk · Context
Vulnerabilities introduced during software development
Inherent weaknesses in programming languages and development environments introduce errors and exploitable vulnerabilities into software products, and software malfunctions cause incorrect outputs, crashes, or security weaknesses.
risk · Context
Inadequate vulnerability scanning and pre-release testing
Software released without adequate testing, and no regular vulnerability scanning or penetration testing, leaves exploitable defects undiscovered until they manifest — or are exploited — in production.
risk · Context
Exploitation of known, unpatched vulnerabilities
Use of software with publicly known, unpatched flaws (CVEs) that adversaries readily exploit — including recently discovered vulnerabilities exploited before mitigations are in place, and internal-system vulnerability exploitation.
standard · Direct
NIST TEVV-Athlon (draft, Aug 2026)
NIST AI 200-2: TEVV-Athlon Framework for Evaluating AI Systems
unified · Context
UC-AI-05 — Set responsible AI development objectives and requirements
Define objectives for responsible AI development, such as fairness, safety, security, transparency, and accountability, and embed them in a documented design and development process. Specify and document requirements for each AI system, including intended purpose, performance criteria, and constraints, before build begins. Evidence includes the development process definition, per-system requirement specifications, and design-stage approvals.
unified · Context
UC-AI-07 — Verify, validate, and control AI deployment and changes
Verify and validate each AI system against its requirements and responsible-AI objectives, documenting test plans, acceptance criteria, and results before release approval. Gate deployment on a documented deployment plan and sign-off confirming requirements are met. Route updates, retraining, and other changes through the same assessment and approval process, including impact reassessment where relevant, and retain verification records and deployment and change approvals.
unified · Context
UC-AI-08 — Log and monitor AI system behavior in operation
Ensure AI systems automatically record event logs that enable traceability of operation over the system's lifetime, including events relevant to identifying risk situations and substantial modification. Retain logs for at least the mandated regulatory minimum, or longer where required. Monitor deployed systems against defined performance and behavior metrics with alerting and escalation for anomalies and drift, and retain logs and monitoring reviews as evidence.
unified · Context
UC-AI-18 — Defend AI interfaces against adversarial input, injection, and endpoint abuse
Protect the inference and agent interfaces of AI systems with layered input defenses: screen prompts, uploaded content, retrieved data, and tool results for prompt-injection and jailbreak patterns before they reach the model or trigger actions; detect and alert on adversarial-input campaigns; and rate-limit, authenticate, and monitor endpoints to prevent scraping, model extraction, and resource-exhaustion abuse. Tune detections from evaluation findings and retain filter configurations and detection logs as evidence.
unified · Context
UC-AI-19 — Constrain agent actions and tool use to authorized scope
Bound what autonomous agents may do: allow-list the tools, connectors, and actions each agent may invoke; scope its permissions to the task, user, and context; require human approval for irreversible, high-value, or out-of-policy actions; execute agent-generated code only in isolated sandboxes; and scan agent configuration artifacts such as hooks, skills, and rules for injected instructions. Log every tool call with its authorization decision and review denied and escalated calls.
unified · Context
UC-AI-20 — Prevent harmful, out-of-scope, hallucinated, and over-exposed AI outputs
Filter and shape every AI output before release: block or transform content that matches the system's harmful-output taxonomy, keep responses within the declared scope and capabilities, detect agent-specific high-risk outputs and route them to defined responses by severity, ground factual claims in cited sources and verify them to limit hallucination, withhold system prompts, internal data, and other over-exposed information, and sanitize outputs consumed by downstream systems so they cannot carry executable or injected payloads. Measure filter effectiveness and retain configurations, block logs, and review samples as evidence.
unified · Context
UC-DATA-11 — Control data flows, leakage, and cross-border transfers
Enforce approved authorizations for information flows within and between systems using technical flow-control mechanisms, and deploy data-leakage-prevention measures on systems and channels that could exfiltrate sensitive data. Transfer personal data across borders only under a valid transfer mechanism (adequacy decision, standard contractual clauses, binding corporate rules, or a documented derogation), with the transfer risk assessed and the safeguard recorded.
unified · Context
UC-VULN-02 — Test security through independent penetration exercises
Commission penetration tests of systems, applications, and networks at least annually and after material changes, performed by qualified testers independent of the target's operation and governed by documented rules of engagement. Include both internal and external testing perspectives, validate the exploitability of identified weaknesses, and report results to accountable management. Track corrective actions from each exercise to verified closure, and use the results as a separate evaluation of whether security controls are present and functioning.
workflow · Context
Technical Security Testing & Pentest Engagement
Runs ON an existing Audit item (audit_type: it_audit) that represents the authorized penetration-test engagement — the workflow instance attaches to that record and enriches it (scope, ratings, dates, and the assurance conclusion write back to its fields); it never creates a duplicate engagement record. In scope: authorized technical testing (reconnaissance, discovery, exploitation validation, severity rating, reporting, and retest) of the defined system boundary against its control baseline and assessment objective, with each confirmed finding recorded as an Issue linked to the anchor Audit and its affected Controls. Out of scope: any testing beyond the agreed rules of engagement, and the downstream remediation program itself — confirmed control gaps and open POA&M findings are handed to the Security Control Assessment & POA&M Remediation workflow, and the assurance conclusion to the Cybersecurity Assurance Review workflow. No upstream workflow feeds this engagement; its inputs are the anchor Audit, the in-scope system boundary (Process items), the control baseline (Control items, framework nist-800-53), the assessment objective and signed authorization, and any open Issue items (source: penetration_test / vulnerability_scan) from prior engagements.
workflow · Context
Network Segmentation & Boundary Rule Management
Network Segmentation & Boundary Rule Management as a decision-aware operator workflow covering trust-zone maintenance, gated rule-change implementation, boundary threat monitoring, the quarterly segmentation and rule-set review, and cross-domain exchange-policy enforcement. This cycle runs on the EXISTING boundary-protection Control item (domains network_communications_security; framework NIST 800-53, NIST CSF 2.0, ISO 27001, SOC 2; frequency quarterly; the boundary control owner as control_owner) — each quarterly run is one workflow instance attached to and enriching that Control, never a new control record. In scope: the trust-zone model, the managed interfaces (firewalls, gateways, proxies) mediating traffic between zones, the external boundary and key internal boundaries, and the cross-domain interconnection points, all under a deny-by-default baseline (NIST 800-53 SC-7, SC-16, SC-46; NIST CSF 2.0 PR.IR-01; ISO 27001 A.8.22; SOC 2 CC6.6). No upstream workflow feeds this cycle; it self-seeds from its own prior-cycle carry-forward — open corrective-action Issue items linked to the Control plus the prior run's closure record and archived rule-change record log. Named deliverables: the reconciled trust-zone model, the verified rule-change record log, the boundary threat-monitoring summary, the cross-domain exchange-policy enforcement report, the quarterly segmentation & rule-set review report, the boundary-posture dashboard, and corrective-action Issues linked to the Control. Downstream, it hands off only to its own next cycle via the carry-forward closure record; a confirmed intrusion is escalated to the incident-response workflow (out of scope), as are host/endpoint controls.
workflow · Context
AI Operations Monitoring & Incident Response
Each monthly instance runs against the existing Control item for AI operations monitoring and incident response (framework eu-ai-act + iso-42001; domains ai_governance / logging_monitoring_detection / incident_management_response; frequency monthly; control_owner AI Operations Manager) — the run enriches that Control's execution history and is its evidence of operation, it never creates a duplicate. The decision-aware cycle covers event-logging and retention verification, alert resolution, human-oversight confirmation, AI-use screening against legal prohibitions, and AI concern-and-incident triage through required communications and regulator reporting, consolidating all five streams into a named cycle report and cycle dashboard while every residual gap is booked as an Issue (source: management_identified) linked to the anchor Control. In scope: every deployed production, limited-risk, and high-risk AI system under the EU AI Act, its event-log sources and monitoring dashboards, the human-oversight roster, the AI use inventory, and the AI concern-reporting channels. Out of scope: pre-deployment model development and conformity assessment, and third-party/vendor AI due diligence, which are governed by their own workflows. This cycle consumes no upstream workflow handoff; it hands off only to the next monthly run of itself, seeding it with the open carry-forward Issues (corrective actions, in-flight blocks, incident follow-ups) that the cycle-metrics-and-report step links to the anchor Control so next month can query them as its prior-cycle input.
workflow · Context
AI Guardrail Configuration & Agent Permission Review
Each monthly or release-triggered instance runs against the existing Control item for AI guardrail configuration and agent permission review (framework aiuc-1 + iso-42001 + eu-ai-act; frequency monthly and per release; control_owner AI Platform Security Lead) — the run enriches that Control's execution history and is its evidence of operation, never a duplicate. The decision-aware cycle confirms the in-scope agent population and baseline; reviews input defenses and endpoint limits, tool allow-lists, permissions and sandboxing, output filters and grounding, misuse refusals, secrets redaction, and secure-code-generation defaults; decides on agent permission scope and on guardrail drift with a remediation branch for each; and consolidates the results into a signed guardrail attestation with cycle metrics and owned actions, booking every residual gap as an Issue (source: management_identified) linked to the anchor Control. In scope: every production AI agent and inference endpoint, its guardrail configuration, tool-call and detection logs, and configuration artifacts. Out of scope: model development, pre-deployment evaluation, and vendor AI due diligence, which have their own workflows. The cycle hands off only to its next instance through the carry-forward Issues that the guardrail attestation step links to the anchor Control.
workflow · Context
Quarterly Third-Party AI Evaluation Cycle
Each quarterly instance runs against the existing Control item for independent third-party AI evaluation (framework aiuc-1 + iso-42001 + eu-ai-act; domains ai_governance; frequency quarterly; control_owner AI Product Lead) — the run enriches that Control's execution history and is its evidence of operation, it never creates a duplicate Control; the one record it does create is a per-quarter Audit item (audit_type: it_audit) for the evaluator engagement. The decision-aware cycle confirms the quarter's in-scope systems and risk taxonomy, engages an independent evaluator with the test plan, categories, and pass thresholds fixed in advance, provisions contained test access, triages every finding by category and severity against the tested control, routes failed thresholds through remediation and evaluator retest, gates acceptance of the evaluator report, publishes the accepted evidence (trust-portal summary, customer-facing attestation, evidence register), and tunes guardrails from the findings. In scope: every in-scope AI system and agent under the AIUC-1 program and its six mandatory third-party test categories (B001 adversarial robustness, C010 harmful outputs, C011 out-of-scope outputs, C012 agent-specific risk, D002 hallucinations, D004 tool calls). Out of scope: internal red-teaming, pre-release model evaluation, and vendor AI due diligence, which run in their own workflows. It hands off only to the next quarterly run of itself, seeding it with the open carry-forward finding Issues that the publish-evidence-and-tune-guardrails step links to the anchor Control.
workflow · Context
Security Control Assessment & POA&M Remediation
Run this assessment on the EXISTING Audit item for the engagement (audit_type: it_audit or compliance) — enrich that record, never create a duplicate: Audit.scope carries the authorization boundary and Audit.period_start/period_end the assessment window. Consumes, from the upstream SSP-development / system-categorization effort, the approved System Security Plan (SSP), the FIPS 199 system categorization, and the tailored NIST 800-53 baseline (existing Control items, framework: nist-800-53). Assess each in-scope control with 800-53A examine/interview/test methods, record satisfied / other-than-satisfied determinations, open a POA&M Issue for every gap, re-validate remediation, and issue the Security Assessment Report (SAR). In scope: control assessment, determinations, the POA&M lifecycle, and the SAR for the authorization boundary agreed at kickoff. Out of scope: the authorization (ATO) decision itself and the steady-state continuous-monitoring cadence. On completion, the frozen SAR and POA&M package are handed to the NIST RMF System Authorization (ATO) Cycle.
workflow · Context
ISO 27001 Stage 2 Annex A Controls Audit
Attach to the existing Audit engagement, owned by Internal Audit, using its approved Statement of Applicability, risk treatment plan, scope, review period and operating evidence; produce the Stage 2 Annex A Controls Audit report, four signed theme conclusions and finding register for the engagement and remediation owners. Apply the approved Statement of Applicability to ISO/IEC 27001:2022 Annex A.5.1–A.5.37, A.6.1–A.6.8, A.7.1–A.7.14 and A.8.1–A.8.34; document each exclusion and assess direct and inherited responsibilities. This Annex A assessment contributes to the engagement and does not independently establish full ISMS conformity or issue certification. Stage 1 and readiness remain separate workflows; any certification decision remains with the authorized certification body.
workflow · Context
ISO/IEC 42001 AI Management System Internal Audit
An independent internal-audit cycle for the AI management system: set scope and criteria, test governance, risk, documentation, operations, value-chain controls, and reporting, then issue findings and an evidence-backed conclusion. The audit supports management improvement and assurance; it does not certify conformity.
workflow · Context
AI Governance & Risk/Impact Assessment
Assess a single AI system end to end under ISO/IEC 42001 (AIMS) and the NIST AI RMF: govern and register the system, map context and risks, measure risks and impacts, manage treatment, produce transparency artifacts, and authorize deployment with monitoring. The instance attaches to an Audit item created for this assessment cycle (audit_type compliance, or advisory for a pre-deployment review); because Studio has no native AI System type, the system under assessment is named in that Audit's scope and its lifecycle/EU-AI-Act detail lives in the scoping memo and step documents. In scope: one named AI system or use case and its lifecycle risk posture. Out of scope: enterprise-wide AI policy authoring and detailed EU AI Act legal obligation mapping, which are consumed as an input handoff package from the EU AI Act Obligation Impact Analysis workflow. The named deliverable is the approved AI assessment package (executive summary plus recommended governance decision), backed by the AI risk register (Risk items, category ai_governance), the model/system card and AI impact-assessment record, and a deployment authorization with live drift/fairness monitoring; approved outputs hand off to Quarterly Board & Audit-Committee GRC Reporting.
workflow · Context
AI System Development, Data & Deployment Gate
Runs as a standalone per-release gate cycle — one workflow instance per new AI system or substantial modification. The schema has no native AI System item type, so the instance links (Item relationship) to the existing Control items it operates (UC-AI-04/05/06/07/09; framework eu-ai-act|iso-42001, domains ai_governance) and to the ai_governance Risk it mitigates, and all cycle evidence attaches to its steps. It consumes the AI system inventory profile and the enterprise responsible-AI objectives (Policy items) upstream; the release board then approves those objectives and the per-system requirements before build, governs training, validation, and test data with bias mitigation, executes the pre-deployment impact assessment and EU AI Act risk classification, applies the high-risk conformity obligations where they trigger, assembles Annex IV-grade technical documentation, and records verification-and-validation results and the deployment sign-off. Named deliverables: the risk-classification decision, the pre-deployment impact-assessment report, the Annex IV technical-documentation package, the verification-and-validation report, and the recorded sign-off. Retraining and substantial changes re-enter the same gate as a fresh instance (a self-loop, no downstream handoff).