SENAR Guide: AI Output Review Checklist

NOTE — Relationship to SENAR Core

SENAR Core includes an updated Verification Checklist (28 items, 3 tiers: Standard / High / Critical) that supersedes this checklist for Core adopters.

Tier mapping: Tier 1 (this guide) ≈ Standard (Core) | Tier 2 ≈ High | Tier 3 ≈ Critical

This Guide checklist is retained for teams adopting SENAR through the Standard (team-level entry point) who have not yet moved to Core. If your team uses SENAR Core, use the Core Verification Checklist instead.

NOTE — Status of this checklist under the Standard

This checklist is informative. It is not the normative requirement at QG-3, and no conformance claim rests on it.

The normative floor is Standard 8.4, AI Output Review Minimum Criteria — five checks that SHALL be performed at QG-3. Standard 8.4 states that those criteria replace dependency on this checklist: an organization is required to perform the five, not to adopt this document. That distinction was not stated anywhere the reader of this chapter would find it, and the chapter read as though ticking it were the obligation.

The five are covered here, and this is the mapping:

Standard 8.4 (SHALL at QG-3)Items in this checklist
a) Hallucination check — no non-existent APIs, methods, CLI flags24
b) Dependency validation — imports exist in the manifest3, 4
c) Pattern conformance — follows project architecture and conventions6, 7, 12, 20
d) Test validity — tests exercise the AC, not just coverage8, 9
e) Secrets check — no hardcoded credentials, keys, or secrets5, 10

Everything else here is beyond the floor, offered as an aid rather than as a requirement. Two consequences worth stating plainly: running all 24 items does not by itself satisfy 8.4 unless each of the five was actually performed and recorded; and a team that finds this list too long may reduce it to the five without falling below the Standard.

When reviewing AI-generated output, check for these AI-specific issues.

The Checklist

#CheckWhat to Look For
1ScopeDid AI modify files outside the task scope? (most common AI error)
2DeletionsDid AI silently remove or replace existing working code?
3Phantom importsAre all imported packages in the dependency file?
4Dependency versionsAre specified versions real and published?
5Hardcoded valuesMagic numbers, URLs, credentials, API keys in code?
6Over-engineeringUnnecessary abstractions, patterns, or generalization?
7DuplicationNew code that duplicates existing utilities?
8Test qualityDo tests verify behavior, or just mirror implementation?
9Test tamperingDid AI modify tests to pass instead of fixing the code?
10SecurityOpen CORS, hardcoded tokens, SQL without parameterization?
11Edge casesHappy path works — what about null, empty, boundary, concurrent?
12NamingDoes AI follow project naming conventions?
13Commit scopeIs the commit atomic and focused, or a kitchen sink?
14Null guard before comparisonNone == None is True in Python (JS: null === null is true; most languages have equivalent null-equality traps) — access checks bypassed when both sides null?
15Empty config bypassSecurity check skipped when config value is empty string? (if secret and ... fails open)
16Header trustX-Forwarded-For, X-Partner-ID, Content-Length used for security without proxy validation?
17IDORResource accessed by ID without verifying user’s access to that resource? Auth ≠ authorization.
18Return True shortcutAccess control function returns True/grants access without explicit ownership validation?
19Format string injectionstr.format(**untrusted_dict) — Python format supports attribute access, enables injection (JS: template literals with eval(); C/C++: printf format strings; any language: string interpolation with untrusted input)
20God functions/filesFunctions >50 lines, files >400 lines — strongest signal of unrefactored AI output
21Unreachable safety codereturn False after exhaustive exception handling — AI added “just in case” but can never execute
22Swallowed exceptionscatch/except blocks that discard errors (except Exception: pass, catch(e) {}, or logging at debug level) — hides real failures. Check for null/None/nil returns that silently mask error conditions
23Unsafe deserializationpickle.loads, yaml.load without SafeLoader, eval/exec on untrusted input, JSON prototype pollution? (Java: ObjectInputStream; JS: eval(JSON); any language: deserializing untrusted data without validation)
24Phantom APIsDo the called functions, methods, fields and CLI flags actually exist in the version of the library, framework or tool in use? Distinct from item 3: the package can be present in the manifest and the call into it still invented. Signatures are hallucinated far more readily than package names, and a plausible one survives review precisely by looking right — check it against the installed version’s documentation or by executing it, not against recollection

Priority Tiers

The tiers below say how much of this list to run and when. They do not govern Standard 8.4: the five criteria in the mapping above are performed at QG-3 on every Task, whichever tier that Task falls into. Several of the items those criteria rest on sit in Tier 2 and Tier 3 here — 8.4(b) on item 4, 8.4(c) on 6, 7, 12 and 20, 8.4(e) on 5 and 10 — and the tier placement says when this Guide suggests reading the item, not whether the Standard’s criterion applies.

Tier 1 — Always check (every task): Items 1–3 (scope, deletions, phantom imports), 24 (phantom APIs) and 8–9 (test quality, test tampering). These catch the most frequent and most dangerous AI defects. A 6-item check takes a couple of minutes.

Tier 2 — Security-sensitive tasks (auth, payment, data, API): Items 10 (security), 14–18 (null guard, empty config, header trust, IDOR, return True). These catch latent defect patterns — AI output that looks correct but fails under adversarial conditions. Identified through adversarial audits of production AI-generated code.

Tier 3 — Deep review (complex tasks, agent dispatch output): All remaining items (4–7, 11–13, 19–23). Apply when reviewing complex features, refactors, or any output from dispatched agents (Rule 15 L3).

Usage

Print this list. Use it at QG-2 (Implementation Gate) for every Task. Over time, some checks become automatic habits — but keep the list visible for new Supervisors.