This website uses cookies

Read our Privacy policy and Terms of use for more information.

A coding agent repairs a pull request after CI passes. The tests were green, the review was clean, and a human had approved it. Which of those facts still applies to the repaired code?

That question gets harder as more development happens through automated work and feedback cycles. By software factory, I mean a delivery system that coordinates implementation, verification, repair, review, and release. Its reliability depends on what each step is allowed to conclude.

My opinion: CI should produce evidence bound to the change it evaluated, and the workflow should explicitly decide what that evidence authorizes.

What does an attestation add to CI?

A LinkedIn article called “The Attestation Model: Continuous Integration for AI-Native SDLC” by Rafael Ramos, caught my attention and I thought I’d add to his awesome explanation with my own thoughts.

His proposal brings check results, AI assessments, repair history, and approvals into a structured record for downstream consumers.

CI already produces logs and artifacts. The value here is making their relationship to intent, policy, and the evaluated change explicit enough for another system to inspect.

Consider an illustrative change to an export service. The requirement says every matching record must appear exactly once, including when results span multiple pages. Tests pass on commit A, and an AI reviewer reports that the implementation satisfies that requirement.

For that assessment to mean anything, the requirement needs to come from an independently established specification. Asking the implementation agent to invent its own acceptance criteria invites circular verification.

I want to know which pagination cases ran, which specification the reviewer read, and what it could not establish. “Passed” can’t answer all of that.

Which layer owns the decision to proceed?

This is where I connect attestation to two projects I’m developing.

ThreadLoop manages the outer software-delivery lifecycle: task state, guarded transitions, evidence freshness, repair, and human completion. Governed Agent Autonomy Patterns, or GAAP, addresses decisions inside one bounded agent run, including planning, permission, tool trust, protected effects, and verification.

I think of the responsibilities this way:

Component

Responsibility

Attestation

Record what was evaluated, against which subject, with what evidence and uncertainty

ThreadLoop

Decide whether current evidence satisfies a lifecycle transition

GAAP

Govern the permitted actions and integrity of a bounded agent run

A protected effect is an operation subject to an authority check, such as a filesystem mutation or service call. Its execution history can contribute evidence to an attestation.

That relationship is an architectural proposal. The current projects do not implement this complete attestation handoff.

The separation matters because an agent’s account of its work cannot also grant every permission, validate every conclusion, and complete the delivery task.

What happens when a repair changes the subject?

Back to the export service. A reviewer finds a missing pagination case. An agent repairs it, producing commit B.

Evidence for A remains useful history. It cannot establish that B passed.

A simplified record might make the mismatch visible:

{
  "evaluated_commit": "A",
  "current_commit": "B",
  "checks": {"pagination_tests": "passed_on_A"},
  "review": {"subject": "A", "finding": "no_blockers"},
  "human_approval": {"subject": "A"}
}

Under the policy I’m describing, progression would stop until the required evidence covered B. The repair also needs its own authorization: approval to evaluate A does not automatically permit an arbitrary mutation.

The checkable behavior is specific. Present A’s evidence while B is current; the transition must be rejected. Supply valid evidence and required approval for B; the transition can become eligible. Completion still requires observing the merge required by the workflow.

That distinguishes producing evidence from accepting it.

How do loops and graphs use that evidence?

Loop engineering defines how work repeats and when repetition stops. For this example, the outer loop is implement, verify, diagnose, repair, and verify again.

A stop rule needs to distinguish a code defect from an unavailable test environment or an ambiguous requirement. Otherwise, automation can spend its repair budget changing code that was never meaningfully evaluated.

Graph engineering makes states, transitions, and authority requirements explicit. A guard is the condition that must hold before a transition is allowed.

Passing tests may satisfy part of the verification guard. A current review may permit progression toward human approval. Conflicting findings may require a different path.

ThreadLoop currently implements a fixed governed PR lifecycle. Configurable graph execution and GAAP receipt admission remain deferred. That scope matters when discussing how the architecture could compose.

The attestation supplies inputs to these guards. Each transition still needs a policy decision.

What does a signed record actually authorize?

The SLSA attestation model separates the authenticated statement from its consumer. A signature establishes who produced the statement; it does not make every assessment correct.

I would require the consumer to check the issuer, subject, evidence integrity, applicable policy, and unresolved uncertainty. A model’s confidence also needs empirical calibration before it can carry the weight assigned to it.

Rafael’s article made the evidence layer more concrete for me. The consequence for loop and graph engineering is equally concrete: every permitted transition needs an explainable relationship between current evidence and the authority accepting it.

Reply

Avatar

or to participate