AI Agents

How should I test my AI agent before publishing?

Updated October 7, 2026

Before publishing an AI agent, test the complete experience from the customer's point of view.

A good test should confirm that:

  • Answers are accurate
  • The right workflow path is used
  • Documents and data are found correctly
  • Questions and forms collect the right information
  • AI Answer skills work as expected
  • Human Handoff works when needed
  • Missing or unclear information is handled correctly
  • The experience works well in the channels where customers will use it

Do not test only the easiest example.

Test normal requests, unclear requests, missing information, incorrect information, and situations where the AI agent should not continue on its own.

Start with the main customer journeys

List the most important jobs your AI agent is expected to handle.

For example:

  • Answer a policy question
  • Find a product
  • Check pricing
  • Collect lead information
  • Calculate an estimate
  • Book an appointment
  • Transfer a customer to a person

Test each journey from the first message to the final outcome.

Do not test nodes only in isolation.

A workflow can work correctly node by node but still feel confusing when customers move through the full conversation.

Test the AI agent as a customer would use it

Use natural language instead of only carefully written test prompts.

For example, do not test only:

“What is your cancellation policy?”

Also try:

“Can I cancel?”

“What happens if I need to cancel tomorrow?”

“Do I get my money back if I cancel?”

Customers will describe the same need in different ways.

The AI agent should still reach the right answer or workflow.

Test AI Answer responses

If your workflow uses AI Answer, check:

  • The answer is relevant to the question
  • Important instructions are followed
  • The response is not unnecessarily long
  • The AI agent does not invent information
  • The tone matches your intended experience
  • The answer stays within the AI agent's role
  • The AI agent asks follow-up questions only when needed

Also test questions that are outside the intended scope.

The agent should follow the boundaries you defined in AI Instructions.

If your workflow uses Document Search, ask questions that should be answered from your uploaded documents.

Test:

  • Direct questions
  • Questions phrased differently from the document
  • Questions where the answer appears deep inside a document
  • Similar topics that appear in more than one document
  • Questions where the answer is not present

Confirm that the AI agent finds the correct information and does not invent an answer when the source does not contain it.

If your workflow uses Data Search, test exact and realistic lookup requests.

For example:

  • A known product
  • A known customer record
  • A known price
  • A known location
  • A known inventory item

Also test:

  • Different capitalization
  • Partial names
  • Similar records
  • Missing records
  • Multiple matching records
  • Empty fields

If you use multiple datasets, test that related records are joined correctly.

Test Router paths

If your workflow uses Router, test every intent or route.

For each path, use several different customer messages that should lead to the same route.

For example, a booking route might be triggered by:

“I want to book.”

“Do you have anything Thursday?”

“Can I make an appointment?”

“I need a consultation.”

Then test messages that should not go to that route.

This helps identify overlapping or unclear routing logic.

Test AI Form

If your workflow uses AI Form, complete the form several times with different answers.

Check that:

  • Questions appear in the correct order
  • Each answer is saved correctly
  • AI Verification accepts valid answers
  • AI Verification rejects answers that do not meet the requirement
  • The AI agent re-asks when verification fails
  • AI Answer Cleanup stores the value in the format you expect
  • The confirmation step shows the correct information
  • Customers can correct an answer when needed

Also test what happens when the customer gives several answers in one message.

Test Question nodes

If your workflow uses Question, test:

  • A valid answer
  • An invalid answer
  • An unclear answer
  • A customer refusing to answer
  • A value written in a different format
  • A customer changing the answer later

Confirm that the saved value is suitable for later nodes that depend on it.

Test Message nodes

For Message nodes, check:

  • The wording is still accurate
  • The message appears at the right point
  • It does not repeat information the customer already received
  • The next step is clear

Because Message content is fixed, it is easy for old wording to remain after the surrounding workflow changes.

If Live Web Search is enabled, test both questions that should use the web and questions that should not.

Try:

  • Current weather
  • Current opening hours
  • A current public event
  • Information from an approved website
  • A stable question already covered by your own knowledge

If web access is restricted to approved websites, confirm that the answers come from the intended sources.

Test Calculations

If Calculations are enabled, test known inputs with known expected results.

Check:

  • Normal values
  • Minimum expected values
  • Maximum expected values
  • Decimal values
  • Missing inputs
  • Invalid inputs
  • Different categories or factors
  • Repeated calculations with the same inputs

The same inputs should produce the same result.

If the result is an estimate, confirm that the AI agent explains that clearly.

Test Scheduling

If Scheduling is enabled, test:

  • A specific date and time
  • A broad request such as “Thursday afternoon”
  • A time that is already booked
  • A request outside working hours
  • A request inside the minimum advance-notice period
  • A request beyond the booking window
  • Missing contact information
  • A customer changing the selected time
  • A successful booking

Confirm that the times offered by the AI agent match the actual calendar availability.

Test Product Cards, Image Answers, and Albums

If your AI Answer node can return visual or structured content, test when it should appear.

Check:

  • The correct product or service is shown
  • The correct image is shown
  • The correct album is shown
  • The visual result matches the customer's request
  • The AI agent does not show unrelated content
  • Text around the visual result is useful and not repetitive

Also test a request where no matching visual content exists.

Test email actions

If your AI agent can send email, test:

  • A valid email address
  • A missing email address
  • An incorrectly formatted email address
  • The correct content or material being sent
  • The correct trigger for sending
  • A customer changing the email address before sending

Do not assume the email action works just because the conversational response appears correct.

Test Human Handoff

If your workflow includes Human Handoff, test the complete transfer experience.

Try:

  • A customer explicitly asking for a person
  • A case that should require human review
  • A request the AI agent cannot resolve
  • A handoff after information has already been collected

Confirm that:

  • The handoff happens at the right time
  • The AI agent does not continue responding when it should stop
  • The human receives the useful conversation context
  • Collected information is available where expected
  • The customer understands what happens next

Test missing information

One of the most important tests is to leave out information the workflow requires.

For example:

“Can you calculate my premium?”

without providing the required vehicle information.

Or:

“I'd like to book.”

without giving a preferred day.

The AI agent should identify what is missing and ask for it.

It should not silently invent a value.

Test invalid information

Customers may enter values in unexpected ways.

For example:

  • “I don't know”
  • “maybe 150”
  • “next-ish Friday”
  • “none”
  • “N/A”
  • A sentence instead of a number
  • A typo in an email address
  • Several values in one message

Test the types of imperfect answers your real customers are likely to provide.

Test customers changing their mind

Real conversations are not always linear.

Try:

“Actually, make it Friday instead.”

“Use my work email, not that one.”

“I don't want that service anymore.”

“Forget the booking. I have another question.”

Check that the conversation can recover naturally.

Test multiple requests in one message

Customers may ask several things at once.

For example:

“How much does it cost, are you open Saturday, and can I book?”

Test whether the AI agent:

  • Understands the different requests
  • Handles them in a sensible order
  • Uses the correct sources or skills
  • Does not ignore an important part of the message

Test unsupported requests

Ask the AI agent to do something outside its intended role.

For example:

  • Ask a hotel agent for legal advice
  • Ask a support agent an unrelated trivia question
  • Ask the AI to invent a discount
  • Ask it to confirm inventory that does not exist in the data

Confirm that the response follows your boundaries.

Test conversations in supported languages

If your customers use more than one language, test the most important workflows in those languages.

Do not test only a greeting.

Complete the full journey.

For example:

  • Ask a question
  • Provide required information
  • Trigger a skill
  • Complete a booking or handoff

Make sure saved values and workflow behavior remain correct even when the conversation language changes.

Test the channels customers will actually use

A workflow that works in one channel should still be tested in every channel where it will be deployed.

For example:

  • Web messenger
  • WhatsApp
  • Chat Page

Check:

  • Message formatting
  • Buttons or interactive elements
  • Images
  • Product cards
  • Long responses
  • Handoffs
  • Booking flows

The customer experience can differ by channel even when the underlying workflow is the same.

Test on desktop and mobile

If customers will use the web experience, test it on both desktop and mobile.

Check that:

  • Messages are readable
  • Interactive content fits the screen
  • Cards and images display correctly
  • The conversation remains easy to follow
  • Important actions are easy to use

Do not rely only on the builder preview.

Test a fresh conversation

After making changes, start a new conversation and test again.

Old conversation context can hide problems.

A fresh conversation helps confirm that the workflow works correctly from the beginning.

Test after changing instructions, data, or workflow logic

Re-test the AI agent whenever you make meaningful changes.

Examples include:

  • Editing AI Instructions
  • Replacing a document
  • Updating a Data Search file
  • Changing Router logic
  • Adding a new AI Form field
  • Changing a calculation formula
  • Updating Scheduling rules
  • Adding or removing a skill
  • Changing Human Handoff behavior

A small configuration change can affect more than one part of the workflow.

Compare results against a known source

For business-critical information, test against something you already know is correct.

Examples:

  • Compare a calculation with your existing calculator
  • Compare Data Search results with the source spreadsheet
  • Compare a policy answer with the original document
  • Compare offered appointment times with the calendar

Do not validate the AI agent only by whether the answer “sounds right.”

Keep a small regression test set

For important AI agents, keep a short list of test conversations that you run after major changes.

For example:

  1. Main FAQ question
  2. Main product lookup
  3. Main calculation
  4. Main booking request
  5. Main lead form
  6. Human handoff request
  7. Out-of-scope question

This makes it easier to catch problems introduced by later edits.

Final pre-publish checklist

Before publishing, confirm:

  • The main customer journeys work from start to finish
  • AI Instructions are being followed
  • Document Search returns the correct knowledge
  • Data Search returns the correct records
  • Router sends requests to the correct path
  • AI Form and Question save the right values
  • Enabled AI Answer skills work correctly
  • Missing information is requested instead of guessed
  • Human Handoff works correctly
  • Out-of-scope requests are handled appropriately
  • Multilingual conversations work where needed
  • The experience works in each customer channel
  • The latest version of your files, formulas, and settings is being used

If any important path fails, fix it before publishing.

Tips for testing your AI agent

  • Test real customer wording, not only perfect prompts.
  • Test both expected and unexpected behavior.
  • Use known correct answers when checking data and calculations.
  • Test each route with more than one phrase.
  • Test missing and invalid information.
  • Test customers changing their answers or intent.
  • Test skills individually and together.
  • Start fresh conversations after major changes.
  • Re-test after updating data, instructions, or workflow logic.
  • Keep a reusable regression test set for important AI agents.

For more information, see:

  • How should I write AI Instructions?
  • Which AskHandle node should I use?
  • What are AI Answer skills?
  • How does Live Web Search work?
  • How do Calculations work?
  • How does Scheduling work?
  • What is a Document Search node?
  • What is a Data Search node?
  • What is an AI Form node?
  • How does Human Handoff work?