Back to blog

Before you ship an AI update: five checks for growing SaaS teams

By Joe Zhou ยท

Your updated AI feature gives a better answer in the demo. It responds faster, follows the conversation more naturally and looks ready to ship.

Before release, there is another question worth asking: what has become less reliable?

For a growing software company, that question connects product quality with customer trust. An update should be assessed against the work customers need it to perform, the actions it is allowed to take and the situations it must hand over to a person.

I believe this is where AI governance can become useful to a product team: helping it decide what to test, what evidence to keep and who makes the release decision.

A useful lesson from customer-support research

A preprint recently submitted describes how Nubank used simulated customers and tool responses to screen customer-support agents before deployment. It allowed the team to explore different models and prompts without exercising production backends through those simulated tool calls.

The authors report benefits in their studied banking-support workflows. Those findings are not a guarantee for other products. They also describe an important boundary: mocked tools do not test the underlying backend, its latency or its real side effects.

For a smaller SaaS team, my takeaway is to rehearse difficult interactions before release, then test the real integrations separately. The following five checks are a practical starting point I would use to structure that discussion.

1. Define what a completed task actually looks like

Consider a hypothetical support assistant that can change a customer's subscription. A friendly message saying the plan has been changed is only one part of the interaction. The test should also establish whether the correct account was updated, the requested plan was selected and the confirmation accurately describes the result.

Write down the intended outcome before judging the response. Include what should happen when completion is impossible.

Anthropic's engineering guidance on agent evaluations makes a useful distinction between an agent's recorded interaction and the final state of the environment. Its guidance also recommends repeated trials because model behaviour can vary between runs.

2. Check the limits of its authority

For the subscription example, define which changes the assistant can complete itself and which require additional authorisation.

Then test a request outside those limits. A customer might ask it to waive a charge, access another account or skip a required confirmation. The expected behaviour should be explicit enough that a reviewer can recognise a failure.

I would record attempted unauthorised actions separately from successful task completion. A good overall score should not hide an action the system was never permitted to take.

3. Test what happens when the workflow breaks

Use scenarios in which an integration times out, information is missing or a person must take over.

For a hypothetical handover, I would check that the customer receives an accurate explanation, the receiving person gets the relevant context and the system does not claim that the issue has been resolved while it is still waiting.

Define the fallback before running the test. Otherwise, reviewers may disagree about whether an improvised response was acceptable.

4. Compare the new version with the current version

Run both versions against the same scenarios under comparable conditions. Record the model, prompts, configuration and relevant knowledge versions so the comparison can be understood later.

Anthropic distinguishes capability evaluations, which explore what an agent can do, from regression evaluations, which check whether previously working behaviour still works. Both questions matter when assessing an update.

My preference is to review new failures individually, especially where customer information, external actions or commitments are involved. An average improvement can still leave a specific customer journey worse off.

5. Make the release decision traceable

Keep a short record of what changed, what was tested, what failed, any remaining limitations and who authorised the release. Include the conditions under which the team would pause or roll back.

This gives the next person reviewing the change something concrete to work with. It also helps the business explain its testing approach when a customer asks how an AI feature is assessed.

Be precise about the scope. A test report supports claims about the scenarios and conditions it covers; it is not a blanket assurance that the product is safe in every situation.

Start with one customer journey

You do not need to settle every question about AI governance to make progress. Choose one important workflow and agree on its expected outcomes, boundaries and failure behaviour.

For a first exercise, I would use a small set of synthetic scenarios, including ordinary requests and difficult cases. Have product, engineering and the person responsible for governance agree on the pass criteria. Expand the set as new risks and failures emerge, alongside integration testing and production monitoring.

The aim is to make release decisions easier to explain and better supported by evidence. That is a practical role for AI governance as teams keep improving their products.

Bring governance into the work of shipping

At Complyd, we help growing AI and technology companies with AI governance, enterprise security reviews and ongoing compliance. Our view is that technology and automation should reduce routine work while keeping responsibility for decisions clear.

If your team is preparing an AI feature for enterprise customers, bring the workflow and the questions you are struggling to answer. We can discuss where governance, evidence and ownership need to fit.

Book a Call with Complyd.