Client work Reflections

Design human approval into AI delivery

Legacy banking modernization

By Ted Cao

What to examine

Faster drafting creates more material to review. If the same experts must inspect everything without clear priorities, AI can move the bottleneck from production to approval.

Approach

Define who approves business rules, architecture, code, runtime behavior and business readiness. Give each reviewer the evidence and unresolved questions needed for that decision, with deeper review for uncertain or consequential changes.

Lesson

AI can prepare the work. Named people remain responsible for accepting it.

Approval is real work

An assistant can draft a rule document, a design, and a pull request faster than a person can understand all three. The resulting review is not a quick signature. A reviewer has to reconstruct the draft's assumptions, check them against source evidence, consider missing cases, and judge whether the work fits the broader system. In some cases that is harder than writing the first draft.

The delivery plan must reserve time for this judgment. If expert review is treated as leftover capacity, more generated output simply creates a queue. The aim is a pace at which people can actually challenge what AI produces, return it for correction, and still understand the system they own.

Put the right people at each decision

The business specialist validates the meaning of a recovered rule. An architect or technical lead approves the target and transition design. Developers review code and tests. QA examines actual behavior. Business owners decide whether a change is ready for use. These decisions operate at different levels; a clean code review cannot repair a misunderstood policy.

Each approval should have a defined artifact, named owner, and exit criteria. The reviewer needs the sources, the difference from the last accepted version, and any unresolved questions. Approval should state its scope. An accepted rule for one product or flow must not silently become authority for all related services.

Route attention by evidence and risk

A useful assistant can point to low-confidence findings, contradictory sources, missing tests, or a change with broad customer impact. That helps experts spend more time where their judgment matters. The signal needs calibration against reviewed examples; a model's self-reported confidence is not enough. Agreement among multiple AI answers also does not prove that a missing business rule has been found.

Risk should affect review depth. A documented, low-impact change may need less inspection than an uncertain transfer rule. Yet reduced depth is a governed decision made by accountable owners, not an automatic release permission. Uncertain or consequential findings should return to source artifacts and domain expertise.

Keep evidence attached to the handoff

The output of each stage should make the next stage reviewable. A rule document links to code, configuration, and a specialist's decision. A design links back to those rules and names its coexistence assumptions. A pull request links to its bounded acceptance criteria and tests. A validation report explains runtime discrepancies and their disposition.

This chain lets a reviewer ask why a choice was made without reconstructing an entire project history. It also makes rejection productive: a failed test can be routed to the rule, design, or implementation that introduced the error. Without that trail, more approvals can add ceremony while leaving the underlying uncertainty intact.

Measure where the work accumulates

Throughput alone cannot show whether AI helps delivery. Track how long drafts wait at each review point, how often they are rejected, how much correction they require, and where defects appear. Compare those measures with a manual baseline for the same kind of work. This reveals whether a faster draft moved effort into expert review or downstream rework.

Outcome measures matter too: regression differences, production incidents, and reconciliation exceptions test whether accepted artifacts held up in use. The source framework proposes these as measurements to collect; it does not report a validated improvement. A bounded pilot provides a way to learn before expanding the approach.

Release improvements deliberately

A correction at an approval gate is valuable feedback. Capture it as a proposed rule, test, or assistant instruction, then evaluate the change against known examples before release. Keep versions and owners for the knowledge base, retrieval settings, prompts, and other AI delivery assets. Silent updates can make an earlier decision difficult to reproduce and can introduce a new failure while appearing to fix an old one.

Human approval is therefore part of an operating system, not a final ceremony. The team decides what AI may draft, what evidence a reviewer needs, when to escalate, and how learning enters the next cycle. Accountability remains with the people who accept the business and operational risk.

Adapted from Human–AI Collaborative SDLC for Legacy Banking Modernization.

Where this work lives

Related capabilities and industries

All client work

Let’s talk about your next big initiative

Tell us about your project scope, modernization goals, or delivery team needs.

Let’s talk