Mapping AI Coding Tool Artifacts to Control Questions Auditors Ask
A practical method for inventorying AI coding tool artifacts, mapping each one to the control question it partially answers, and naming the owner who keeps it current.
The auditor's question was simple: "How was this change reviewed?" The change was a Copilot-assisted refactor of a token validation path, merged three weeks earlier. The team's answer was a dashboard tile showing that automated review had run on nearly every pull request that quarter. The auditor looked at it, nodded, and asked the question again.
That gap is the whole problem. Tool telemetry describes activity. Control questions ask about authority, qualification, and accountability. A dashboard that says review happened does not say who was authorized to approve, whether that person owned the path, or whether the check that passed tested the risk the control exists to manage.
Mapping AI coding tool artifacts to control questions takes four steps: inventory every artifact your tools emit, write the control question in the reviewer's words before you look at your evidence, map each question to a primary and a supporting artifact, then assign one named evidence owner and a review date to every row. This article gives you that method. It does not tell you whether your controls are adequate, and the exact question wording has to come from the compliance owner who signs your attestations.
The Screenshot That Did Not Answer the Question
Screenshots fail as evidence because they are a rendering, not a record. The underlying data may be correct, but the screenshot cannot be re-queried, cannot show what was excluded, and cannot survive a follow-up question about a specific pull request from four months ago.
The deeper failure is category confusion. Your coding tools emit events: a review comment was posted, a status check passed, a ruleset evaluated. Control frameworks ask about decisions: was this change authorized, was separation of duties preserved, can you trace what shipped back to what was reviewed. An event can be an input to answering a decision question. It is never the answer by itself.
Teams that get this right stop asking "what evidence do we have" and start asking "what will someone ask us." Those are different exercises and they produce different maps. The first produces a pile of exports. The second produces a document you can walk someone through in twenty minutes.
What Your Coding Tools Actually Emit
Start with an honest inventory, sorted into three classes. Configuration artifacts record what you told the tool. Event artifacts record what the tool did. Outcome artifacts record what actually shipped.
On the configuration side, GitHub documents repository custom instructions that apply to Copilot work in a repository, which means the constraints you give an agent live in version control and can be diffed [1]. Ownership configuration is similar: CODEOWNERS maps paths to accountable reviewers and can be required on pull requests [2]. Rulesets define which branch and tag rules apply to a given ref, with visibility into what was evaluated [3].
On the event side, Copilot code review produces pull request comments that can be steered with repository instructions [4]. Required status checks produce pass and fail results tied to a commit. Bypass and override actions produce their own records, and those are usually the most interesting rows in your inventory.
On the outcome side, GitHub artifact attestations establish provenance for builds, connecting a produced artifact back to the workflow and source that created it [5], and the SLSA provenance specification defines the format and expectations those attestations can be checked against [6].
Write the inventory down in a structure you can review like code:
artifacts:
- id: repo-custom-instructions
class: configuration
lives_in: repository, version controlled
proves: which constraints were in force at a given commit
does_not_prove: that the agent honored them
retention: git history, durable
- id: copilot-review-comments
class: event
lives_in: pull request timeline
proves: an automated pass ran and what it flagged
does_not_prove: that a qualified human evaluated the finding
retention: tied to PR; resolvable and collapsible
- id: ruleset-bypass-event
class: event
lives_in: org and repo audit surfaces
proves: a merge condition was overridden and by whom
does_not_prove: that the override was justified
retention: check your plan's audit log window
- id: build-provenance-attestation
class: outcome
lives_in: attestation store, verified in pipeline
proves: artifact-to-commit-to-workflow linkage
does_not_prove: that the commit was properly reviewed
retention: managed by your release processThree properties deserve a column of their own in any real inventory. Some artifacts expire on a retention window you do not control. Some are editable after the fact, including pull request descriptions and review comment threads. Some exist only inside a vendor interface, which means they are not exportable evidence until you make them exportable.
Writing the Control Question Before the Evidence
Question-first mapping prevents the most common form of self-deception, which is building a map out of whatever you happened to collect. If you start from artifacts, you will produce a map with excellent coverage of the things that were easy to instrument and silence on everything else.
Write the question the way a reviewer would ask it. Four examples that hold up in walkthroughs:
- Authorization: "For this merged change, which human approved it, and were they authorized to approve changes to these paths?"
- Separation of duties: "Did the identity that authored the change differ from the identity that approved and merged it?"
- Secure development practice: "Which of your secure development activities ran on this branch, and which were skipped because an agent authored it?"
- Traceability: "For this deployed artifact, show me the commit, the build, and the approvals without reconstructing anything by hand."
NIST SP 800-218 gives you practice-area vocabulary for phrasing the secure development questions [7], and SP 800-218A extends that framing to generative AI and dual-use foundation model development, which is useful when you need language for agent-specific activity [8]. Map your existing controls to practice areas first, then mark which practices an agent-authored change currently bypasses. That second step is where the real findings live.
Application-layer risk turns into questions too. The OWASP Top 10 for Large Language Model Applications catalogs prompt injection and insecure output handling [9], which in a coding workflow becomes: "When an agent reads an issue comment, a linked page, or tool output, what prevents that untrusted content from steering a code edit, and where is the record that a person reviewed the resulting diff rather than the agent's summary of it?" CISA's Secure by Design material pushes the same expectations toward defaults and accountable ownership, which is a useful frame when you decide whether a recurring finding gets fixed in a repository default instead of a checklist [10].
One caution worth repeating: framework wording changes between revisions, and your auditor's interpretation may be narrower than the published text. Have your compliance owner approve the question text before the map goes anywhere near an audit binder. If you are building this alongside a SOC 2 program, the same question-first discipline applies to SOC 2 engineering controls.
The Mapping Table That Survives an Audit Walkthrough
Five columns, no exceptions: Control Question, Primary Artifact, Supporting Artifact, Evidence Owner, Last Reviewed. The supporting column exists because a single artifact almost never answers a full control question. A merge record shows who clicked merge. It does not show whether that person owned the path, which is why the CODEOWNERS evaluation result belongs in the same row.
| Control Question | Primary Artifact | Supporting Artifact | Evidence Owner | Last Reviewed |
|---|---|---|---|---|
| Who authorized this agent-authored change to merge? | Merge record with approving reviewer identity | Agent app identity on authoring commits; ruleset evaluation result | Platform lead (named) | 2026-02-03 |
| Did a qualified reviewer see changes to high-risk paths? | CODEOWNERS review requirement satisfied on the PR | High-risk path list in repo; required check results | Security engineering manager | 2026-01-27 |
| Did dependency and secret scanning run on agent branches before merge? | Required status check results on the merge commit | Workflow definition in version control; ruleset requiring the checks | DevSecOps owner | 2026-02-10 |
| Does the deployed artifact correspond to the reviewed commit? | Verified build provenance attestation | Release workflow run; deployment record | Release engineer | 2026-02-10 |
| Were the agent constraints in force the ones we claim? | Repository instructions file at the merge commit | Decision record referenced in the PR description | Tech lead | 2026-01-20 (partial) |
That last row shows the marker that makes this document credible: record explicitly what part of the question no artifact currently answers. In this case, the repository instructions file proves what was configured, but nothing in the record captures the full runtime context the agent actually received. Write that down. An auditor who finds an unstated gap treats the whole map as unreliable. An auditor who reads your own gap note treats you as someone running a real program.
Keep the map in version control next to the code it describes. Changes to your evidence story then arrive as pull requests, get reviewed, and carry a history. A compliance matrix in a spreadsheet drifts silently. A compliance matrix in Git drifts visibly.
Why Tool Events Are Inputs, Not Proof
A passing check shows the check ran. It does not show the risk was addressed. That distinction is the single most common source of false confidence in AI-assisted pipelines, because green history accumulates faster than anyone reads it.
The bypass case makes it concrete. A ruleset that requires code owner review but is overridden every sprint by the same three accounts produces a clean-looking merge history and a real control gap. GitHub's ruleset model gives you the visibility to see which rules applied to a ref [3]. Whether the overrides were justified is a judgment nobody but a person can make, and if that judgment is never written down, the bypass log is just noise with timestamps.
Attribution gets ambiguous fast when an agent app authors a branch and a human clicks merge. Both identities belong in the record, and they answer different questions. The agent identity answers "what produced this code." The human identity answers "who accepted it." Give agent identities their own scoped account or app so the first question is answerable at all. Teams working through this at scale usually end up formalizing it as part of a broader AI code governance model rather than as a per-repository setting.
Outcome evidence is where the chain closes. Artifact attestations tie what you deployed back to the workflow and source that built it [5], checked against the SLSA provenance format [6]. Verify attestations before deployment and fail closed when verification does not succeed, otherwise the attestation is decoration.
Naming Evidence Owners Who Can Actually Answer
The owner test is simple: the named person can produce the artifact within one business day and explain its limits without calling someone else. If either half fails, you have a placeholder, not an owner.
Separate three roles explicitly. The artifact owner keeps the evidence producible. The control owner answers for whether the control design is adequate. The exception resolver decides what happens when a merge condition conflicts with a delivery deadline, and records that decision. On small teams one person may hold all three, but write all three names down anyway, because the moment you grow, the ambiguity costs you a quarter.
Audit CODEOWNERS before you audit anything else. Every directory an agent can touch needs a team that can actually review it [2]. Unowned paths in the repository become unowned rows in the evidence map, and they show up during a walkthrough as the awkward pause where nobody knows who to ask.
Departures are the quiet failure mode. When the only person who knew how to export a given artifact leaves, that row goes stale without changing appearance. Watching concentration risk on evidence ownership the way you watch it on code ownership is the fix, which is why bus-factor risk belongs in the same conversation as compliance mapping.
Review Dates, Drift, and the Quarterly Walkthrough
Make Last Reviewed a required field. An entry with no date is an assertion about the past that nobody has tested, and those accumulate faster in AI-assisted repositories because tooling changes underneath the map every few weeks.
Run a sample walkthrough quarterly. Pick one release. Trace the deployed artifact back to its provenance attestation, then to the commit, then to the approvals and check results on the pull request. Log every link that required manual reconstruction. Those logged links are your backlog, ranked by how badly they will hurt under time pressure. Nothing in this exercise requires an auditor to be present, and doing it unattended is exactly the point.
Track bypass and override events between reviews as a drift signal on the map itself. A rule that gets overridden repeatedly is either wrong or unenforced, and both conditions need a decision rather than a note. DORA's research program is the right frame for arguing about whether tightening a gate actually costs delivery speed, because it measures throughput and stability together rather than in isolation [11].
Superseding an entry works like superseding a decision record: mark the old row replaced, update the artifact reference in the same change, and let review catch drift. Two systems reduce the manual reconstruction work here. Typed decision records in Genie's institutional memory give you a stable, referenceable place for the constraint and its rationale, so a pull request can cite the decision it relied on instead of restating it. Guardian merge gate history retains the merge-time record, including which conditions were required, which passed, and which were bypassed and by whom. Both reduce reconstruction effort. Neither makes the compliance judgment: whether your controls are adequate remains a call for your compliance owner, working from evidence you can produce on demand.
Frequently Asked Questions
Is a Copilot code review comment acceptable audit evidence?
It is evidence that an automated review pass ran and what it flagged [4]. It is not evidence that a qualified person evaluated the finding. Pair it with the human approval record and the code owner requirement result.
What if our AI tool only exposes data in its own interface?
Treat that as a gap and write it in the map. Either find an export or API path, or mirror the relevant facts into a system you control at merge time. Evidence you cannot produce without logging into a vendor console will fail a walkthrough eventually.
How granular should control questions be?
Granular enough that one artifact pair can answer each one. If a question needs five artifacts and three explanations, split it. Overly broad questions hide partial coverage.
Should agent-authored changes have different controls than human ones?
Run the same dependency, secret, and static analysis gates on agent branches as on human ones, and block merges on failures rather than reporting after the fact. What changes is attribution: record which agent or app authored the branch and which human approved the merge.
Who should own the mapping document?
One named person in engineering owns keeping it current, and your compliance owner approves the question wording. Split those and the map either goes stale or stops matching what auditors ask.
What to Do in the Next 30 Minutes
Open a new file in your primary repository and write five control questions in an auditor's voice. Do not look at your tooling while you write them. Then, for each one, name the primary artifact, the supporting artifact, an owner, and a review date, leaving explicit blanks where you have nothing. The blanks are the output that matters.
This week, start tracking one signal: bypass and override events on your merge conditions, with the name attached. It is the cheapest early warning that your green history and your actual control posture have separated.
Then go back to the question that started this. If someone asked you today how a specific agent-assisted change to a sensitive path was reviewed, you should be able to open one file, find the row, and produce the record without a screenshot. That is the entire goal. Teams building this alongside an existing audit program can see how the pieces fit in the compliance solutions overview.
References
[1] GitHub Docs, "Adding repository custom instructions for GitHub Copilot," 2026. https://docs.github.com/en/copilot/customizing-copilot/adding-repository-custom-instructions-for-github-copilot
[2] GitHub Docs, "About code owners," 2026. https://docs.github.com/en/repositories/managing-your-repositorys-settings-and-features/customizing-your-repository/about-code-owners
[3] GitHub Docs, "About rulesets," 2026. https://docs.github.com/en/repositories/configuring-branches-and-merges-in-your-repository/managing-rulesets/about-rulesets
[4] GitHub Docs, "Using GitHub Copilot code review," 2026. https://docs.github.com/en/copilot/using-github-copilot/code-review/using-copilot-code-review
[5] GitHub Docs, "Using artifact attestations to establish provenance for builds," 2026. https://docs.github.com/en/actions/security-for-github-actions/using-artifact-attestations/using-artifact-attestations-to-establish-provenance-for-builds
[6] SLSA, "Provenance specification v1.0," 2025. https://slsa.dev/spec/v1.0/provenance
[7] NIST, "SP 800-218, Secure Software Development Framework (SSDF) Version 1.1," 2022. https://csrc.nist.gov/pubs/sp/800/218/final
[8] NIST, "SP 800-218A, Secure Software Development Practices for Generative AI and Dual-Use Foundation Models," 2024. https://csrc.nist.gov/pubs/sp/800/218/a/final
[9] OWASP, "Top 10 for Large Language Model Applications," 2025. https://owasp.org/www-project-top-10-for-large-language-model-applications/
[10] CISA, "Secure by Design," 2025. https://www.cisa.gov/securebydesign
[11] DORA, "Research," 2025. https://dora.dev/research/