Phase 2 · Guardrails
Build confidence to ship the work.
Use AI Software Toolkit to make quality checks and shared skills part of everyday delivery, with evidence the team can inspect.
The program at a glance
One shared practice. Five phases.
- Phase 0 · Suggested week 0
Prepare
Pod, backlog, access, baselines, tools, and structure.
Explore phase - Phase 1 · Suggested week 1
Execute
Agree the design together before generating a change.
Explore phase - Phase 2 · Suggested week 2
Guardrails
Keep changes reviewable and run the quality gates.
Current phase - Phase 3 · Suggested week 3
Trust
Use AI to understand, explain, and debug the code.
Explore phase - Phase 4 · Suggested week 4
Scale
Transfer the playbook with coaches and rollout waves.
Explore phase
Your destination
What success looks like at the end of this phase
The pod uses AI Software Toolkit on real work: shared skills guide repeatable practice, checks produce evidence for the correct change, and the team can explain its Proof scorecard and remaining gaps. A representative ticket has followed the approved production path. Confidence comes from a demonstrated process and accountable review.
The performance step
Turn faster generation into accepted delivery.
- What improves
- Repeatable checks catch problems earlier. Reviewers can inspect the result and its evidence together, helping generated work move through review.
- What to measure
- Track review turnaround, fixes after review, check failures, and escaped defects alongside accepted delivery. Use the Proof scorecard to locate missing or failing checks.
- Build on it
- Carry the checked change into Trust and demonstrate that the team can explain and investigate its behavior.
Where you are now
Useful changes are arriving, but validation varies between tickets or depends on one person remembering every step.
What you will learn: Connect real checks to the work, interpret a Proof scorecard, and demonstrate that a guardrail catches a failure.
What to do in this phase
Step 01
Agree the risks, owners, and change boundaries
With your quality partner, identify what could go wrong on the next ticket and which checks should catch it. Keep small logical commits and focused pull requests. The PPT’s ≤300 lines / ≤5 files is an example commit policy; the toolkit’s PR scope report has separately configured thresholds. Agree owners and exceptions, and map the repository’s architecture, testing, security, and release documents.
Step 02
Adopt the toolkit in one pilot repository
Choose a toolkit revision with your coach and record it. Follow the setup guide: discover the existing repository, preview the installation, and review the proposed files before applying it. Start with Proof and a small set of shared skills; functional QA is an optional addition. Review and commit the resulting configuration and lock file. Use the setup diagnostic to find gaps; installed or configured does not mean a check passed.
Step 03
Connect the repository’s real checks
Wire the actual build, lint, tests, and relevant security checks to their evidence providers. Start with the applicable Core checks, such as changed-code coverage, Semgrep, and Gitleaks; add supported providers where needed. Configure test commands that cover the repository’s actual suites. Keep commands in trusted configuration and credentials in approved secret stores. Assign a remediation owner to every gap. Preserve existing required safeguards while new toolkit capabilities begin in advisory mode.
Step 04
Practice skills and functional QA on a real ticket
Use a relevant shared skill to guide repeatable review, testing, or repair. Keep the pair and quality partner involved in acceptance criteria. If the cohort adopts functional QA, set it up explicitly and exercise representative behavior; record failures and missing regressions. Have another engineer repeat the workflow on a second ticket. AI review and agent-driven QA are advisory support; executable checks and accountable review remain essential.
Step 05
Prove a failure is caught, then review the evidence
In an isolated exercise, introduce a harmless regression, confirm the relevant check detects it, restore the implementation, and rerun. Inspect what ran, failed, or remains unverified. For GitHub delivery, inspect trusted evidence for the exact pull-request revision. Promote a new capability to enforced only after its provider, results, and repair ownership are reliable; required-check settings are a separate decision. Complete the existing approved release process before claiming shipping readiness.
Your guardrails kit
Put AI Software Toolkit to work.
Use AI Software Toolkit as this phase’s practical companion. Shared skills guide the work, tests and scanners produce results, and Proof checks those results against your policy. Your team owns the application, acceptance criteria, and release decision.
- Set up the toolkit — discover, preview, select components, and review the installation.
- Connect Proof to real repository checks — start with applicable checks and make gaps visible.
- Read the scorecard — distinguish a pass from an advisory gap, a block, or an inactive check.
- Choose shared skills and set up optional functional QA — practice them with your pair and quality partner.
Version note (source reviewed 4 October 2026): Proof is available in current source and is not yet released. The released v1.0.0 demo uses the earlier Guardrails name and different paths. Choose and record a revision, and follow its matching instructions.
Readiness rule: Installed is not passed. A scorecard’s ALLOW decision covers its enforced checks for that snapshot; it does not grant merge or release approval.
Practice · Read the evidence
Read a real guardrails scorecard.
Open AI Software Toolkit’s published PR scorecard to practice interpreting check results. It is a snapshot of a change in the toolkit repository. Use it as a reference, then review the scorecard for your own ticket with your coach.
- Check the snapshot. Look at the source run, subject revision, and publication time. Confirm which change the evidence describes; a published PR snapshot does not assess current main or your repository.
- Read results and policy together. Distinguish passed, failed, blocked, unverified, and inactive checks. Compare advisory and enforced modes. An ALLOW decision can coexist with advisory failures; when no controls are enforced, it does not establish readiness to ship.
- Trace a gap and assign the next action. Open a check’s details and source report. Explain what ran, what the evidence establishes, and what needs repair or verification. Record the owner and follow-up in your team’s working notes before the coach checkpoint.
The scorecard changes as new snapshots are published. Its results belong to the toolkit’s evaluated change; your phase evidence must come from your own repository and approved release process.
Your working session
Try this with your team
Create a guardrail adoption record for one ticket: toolkit revision, selected components, real commands, check owners, and the first scorecard. Run a controlled negative case, restore it, and compare the results for their respective revisions. Then repeat the workflow on a second ticket and review release evidence with the coach.
Make this: A toolkit adoption record, a scorecard with explained gaps, a caught-and-restored failure, and a shared skill another engineer has used.
See it on the page The authorization review →
Use a practice guide when you reach this step:
A worked example
The pair uses a test skill to cover the invitation’s authorization rule. In a local fixture, it introduces a harmless regression and sees the unit-test check fail. After restoration, that check passes for the new revision. A missing security-provider result remains a named gap. The pair then reviews independent PR checks where available and the approved release record; a local pass alone does not establish production or email-provider behavior.
How success feels
- We know which checks ran on this change and what they actually establish.
- Guardrails catch problems during the work, and every remaining gap has an owner.
- We can take a ticket through our approved production process using checks and shared routines we can demonstrate.
Use the signals below to check your progress with the team.
Signals of success
- The pod has a recorded toolkit revision, reviewed configuration, and a small set of shared skills in active use.
- The relevant checks run real repository commands and report results for the correct revision.
- The pod can explain passed, failed, missing, and inactive results without treating an absent check as a pass.
- A controlled failure is caught, the fix is checked again, and another engineer repeats the workflow.
- A representative ticket completes the approved release path with recorded evidence. A missing release check leaves shipping capability unverified and has an owner and follow-up.
If you are stuck
- The checklist grows without helping. Keep checks that address a stated risk and remove duplicated ceremony.
- Tests pass but behavior is wrong. Revisit whether assertions actually cover acceptance conditions.
- Review cannot keep up. Reduce batch size and address the review queue before increasing generation.
Coach checkpoint
Before you move on
Bring the toolkit adoption record, a representative ticket and scorecard, negative-case evidence, shared-skill feedback, and approved release evidence. Explain each important check’s owner, effective policy mode, and revision. Distinguish local feedback from PR evidence and release verification. Advance when the pod can use and explain the process; repair missing controls before claiming confidence to ship.
Your next action: Show the coach the first scorecard, the caught-and-restored failure, and how another engineer used the same workflow.
Reading independently? Use these questions for self-review. Discuss participation and artifact feedback with a coach through a cohort enquiry.