An AI assistant can produce a convincing screen before the team has agreed what should happen behind it. A form submits, a success message appears, and the demonstration feels complete. The difficult questions begin when the request is repeated, a member of staff lacks permission, or an external service accepts an action without returning a response.
My review would start with those questions. The following is a proposed release checklist, using a hypothetical service business with customers in Saudi Arabia and an operations team in Egypt. It is not a claim that a particular client system passed these checks.
Write the expected behaviour before reviewing the code
Suppose a customer uploads a document and an employee approves it. I would first write down who can upload, who can approve, which order the document belongs to, and what the customer should see while it is waiting. Approval may also send an email or move the order to another stage. Those consequences belong in the requirement.
Then I would ask for a small change that implements one part of that behaviour. A narrow change is easier to understand, challenge, and remove. If the assistant introduces a new dependency or changes a shared permission rule, that deserves its own explanation. A shorter implementation is not automatically a safer implementation.
Trace permission through the whole request
Hiding an approval button does not establish that the approval endpoint is protected. I would follow the action from the browser into the server and check both the user’s role and access to this particular order. A valid employee account should not automatically gain access to every customer’s documents.
The test cases would include a different customer’s identifier, an expired session, and a staff account without approval rights. I would also check what appears in logs and notifications. The person debugging an issue may need a request identifier; they may not need the customer’s entire document or personal details copied into a log.
Test outcomes independently of the generated solution
Tests written alongside generated code can repeat the same mistaken assumption. I would compare them with the behaviour agreed earlier: a document is approved once, the correct customer sees the change, and a failed notification does not erase the approval record.
For a bilingual service, the examples should include Arabic names, long organisation names, mixed Arabic and Latin references, and dates displayed to users in different locations. A page that looks correct with short English sample data has not answered those questions. Arabic and English workflows need their own review.
Follow the failure path as carefully as the success path
I would deliberately test an unavailable email service and a repeated approval request. Can staff see that a notification is still pending? Can they retry it without approving the document again? Does the system keep enough information to explain what happened?
These are workflow decisions before they are coding decisions. My article on repeated integration events explains why a second event should not necessarily cause a second business action. For a release, I would also want a clear way to disable the new behaviour and restore the previous version if the change fails.
Leave the next person an understandable system
The handover should describe what changed, which assumptions remain, how to run the relevant checks, and what to do when a dependency fails. An assistant’s explanation can help draft that handover, but I would verify it against the actual code and observed behaviour.
GitHub’s guidance on responsible use of Copilot Chat also calls for reviewing and testing generated outputs. That is the useful boundary: assistance can accelerate implementation, while someone still has to understand and accept the result. For the workflow thinking behind my portfolio, see the OVZA case study.
Frequently asked questions
Is a working demonstration enough to approve AI-assisted code?
No. A demonstration usually covers a selected path. Approval should also consider permissions, invalid inputs, repeated actions, dependency failures, and recovery.
Should the same assistant write the code and its tests?
It can help with both, but the expected results should come from independently agreed requirements. Review the test cases for missing scenarios and shared assumptions.
Does every change need the same review effort?
No. A wording correction and a payment-state change have different consequences. Increase review depth when a change affects money, access, personal data, or several connected systems.