SimulationEvals¶
SimulationEvals runs AI-driven end-to-end conversation simulations against your CXAS agent. Instead of scripting exact utterances, you describe goals and success criteria — and a Gemini model figures out what to say at each turn to try to achieve them. This is a great way to test how your agent handles realistic, messy, unpredictable conversations.
Here are the key concepts:
Step(Pydantic model) — a single goal within a simulation, with agoal,success_criteria, optionalresponse_guide, and amax_turnslimit. Steps can also include astatic_utterancefor when you want a fixed first message, andinject_variablesfor seeding session state.StepStatusenum — tracks whether each step isNOT_STARTED,IN_PROGRESS, orCOMPLETED.simulate_conversation()— drives the full multi-turn loop, returning anLLMUserConversationobject that contains the transcript, step progress, and expectation results.generate_report()— produces aSimulationReportwith two DataFrames: goal progress and expectation results. It renders as styled HTML in a Jupyter notebook.
Quick Example¶
from cxas_scrapi import SimulationEvals
from cxas_scrapi.utils.rate_limiter import RateLimiter
app_name = "projects/my-project/locations/us/apps/my-app-id"
# Optional: configure a rate limiter to pace simulation turns and prevent quota exhaustion
limiter = RateLimiter(requests_per_minute=30.0)
sim = SimulationEvals(app_name=app_name, rate_limiter=limiter)
test_case = {
"steps": [
{
"goal": "User wants to check their account balance",
"success_criteria": "Agent provides a numeric balance and account status",
"max_turns": 5,
},
{
"goal": "User asks to dispute a charge",
"success_criteria": "Agent acknowledges the dispute and provides a reference number",
"max_turns": 8,
},
],
"expectations": [
"The agent should never ask for the full credit card number",
"The agent should offer to escalate if it cannot resolve the dispute",
],
}
# Run the simulation
conversation = sim.simulate_conversation(
test_case=test_case,
console_logging=True,
)
# View the report
report = conversation.generate_report()
print(report) # Colorized in terminal, styled HTML in Jupyter
Naturalness Metric (optional)¶
SimulationEvals can also grade how human the agent sounds. The metric is opt-in: a test case that does not declare a naturalness_metric block behaves exactly as before, with no extra Gemini call and no extra keys in the results. See the Local Simulations guide for the YAML schema and worked examples.
NaturalnessConfig(Pydantic model) — the per-test-case configuration: which model grades, which qualities are scored, the label bands, the turn/conversation blend weight, the perceived-latency bands and weight, and an optionalpass_threshold. Unknown keys are tolerated, so newer YAML still loads.NaturalnessResult— the aggregate result for one simulation:overall_score,overall_label, per-turn gradings, conversation-level factors,factor_averages,label_counts,latency_ms_by_turn,failure_reasons, andpassed. Exposed asconversation.naturalness_result.TurnNaturalness— the grading for a single agent turn: its label, 1–5 score, justification, and the factors behind it.NaturalnessFactor— one scored quality (e.g.pacing) with its 1–5 score, an optional shortvaluedescriptor, and the evidence-citingreason.NaturalnessLabelenum —Bot-like,Transitional, orHuman-like.evaluate_naturalness()— grades a simulation trace directly, optionally against the agent's recorded audio and the platform's per-turn perceived latency. ReturnsNone(after logging) when there is nothing to grade or the grading call fails, so the metric can never break a run.
In audio modality the agent's recordings are attached to the grading call automatically, so pacing and emotion are judged from what the caller heard. perceivedLatency is measured from the conversation trace rather than judged, and a turn past latency_failure_threshold_ms fails the simulation outright.
Reference¶
SimulationEvals ¶
SimulationEvals(app_name, rate_limiter=None, expectations_only=False, deployment_id=None, vertex_location='global', naturalness=None, **kwargs)
Bases: Apps
Wrapper class to simulate entire multi-turn conversations with a CXAS Agent.
Source code in src/cxas_scrapi/evals/simulation_evals.py
simulate_conversation ¶
simulate_conversation(test_case, sim_user_model=_DEFAULT_GEMINI_MODEL, eval_model=_DEFAULT_GEMINI_MODEL, session_id=None, console_logging=True, modality='text', capture_agent_audio=False, background_noise_file=None, burst_noise_files=None, use_tool_fakes=False, voice_config=None, initial_utterance=_FIRST_UTTERANCE, skip_playback_wait=False, single_bidi_stream=False, max_turns=None, naturalness=None, **kwargs)
Runs the simulated conversation loop.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
test_case | dict[str, Any] | The test case dictionary defining evaluation steps. | required |
sim_user_model | str | None | The Gemini model used for the simulated user. | _DEFAULT_GEMINI_MODEL |
eval_model | str | None | The Gemini model used for evaluating expectations. | _DEFAULT_GEMINI_MODEL |
console_logging | bool | Whether to print interaction transcript to the console. | True |
single_bidi_stream | bool | For audio modality, keep one persistent bidi WebSocket open for the whole conversation instead of opening a new connection per turn (the default). | False |
max_turns | int | None | Maximum number of conversation turns. Defaults to the test_case's max_turns setting, or 30 if unspecified. | None |
naturalness | bool | dict[str, Any] | None | Overrides the test case's | None |
Source code in src/cxas_scrapi/evals/simulation_evals.py
824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 875 876 877 878 879 880 881 882 883 884 885 886 887 888 889 890 891 892 893 894 895 896 897 898 899 900 901 902 903 904 905 906 907 908 909 910 911 912 913 914 915 916 917 918 919 920 921 922 923 924 925 926 927 928 929 930 931 932 933 934 935 936 937 938 939 940 941 942 943 944 945 946 947 948 949 950 951 952 953 954 955 956 957 958 959 960 961 962 963 964 965 966 967 968 969 970 971 972 973 974 975 976 977 978 979 980 981 982 983 984 985 986 987 988 989 990 991 992 993 994 995 996 997 998 999 1000 1001 1002 1003 1004 1005 1006 1007 1008 1009 1010 1011 1012 1013 1014 1015 1016 1017 1018 1019 1020 1021 1022 1023 1024 1025 1026 1027 1028 1029 | |
export_results_to_golden ¶
Exports simulation results to a Golden Evaluation YAML file.
Fetches the full conversation trace for each simulation from the platform to ensure accuracy.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
results | list[dict[str, Any]] | The list of results returned by run_simulations. | required |
output_path | str | None | Optional local path to save the generated YAML. | None |
Returns:
| Type | Description |
|---|---|
str | The generated YAML string. |
Source code in src/cxas_scrapi/evals/simulation_evals.py
Step ¶
Bases: BaseModel
StepStatus ¶
Bases: str, Enum
SimulationReport ¶
A report containing both Goals and Expectations DataFrames.
Source code in src/cxas_scrapi/evals/simulation_evals.py
NaturalnessConfig ¶
Bases: BaseModel
Per-test-case configuration for the Naturalness Metric.
Unknown keys are preserved rather than rejected so that a YAML file written against a newer version of the library still loads here.
NaturalnessResult ¶
Bases: BaseModel
Aggregated Naturalness Metric result for a whole simulation.
TurnNaturalness ¶
Bases: BaseModel
Naturalness grading for one agent turn.
NaturalnessFactor ¶
Bases: BaseModel
A single scored conversational quality, e.g. pacing or grammarStyle.
NaturalnessLabel ¶
Bases: str, Enum
Graded label assigned to an agent turn and to the whole call.
evaluate_naturalness ¶
evaluate_naturalness(gemini_client, model_name, trace, config, audio_paths=None, latency_ms_by_turn=None)
Grades the naturalness of every agent turn in a simulation trace.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
gemini_client | Any | A | required |
model_name | str | Fallback model, used when the config does not pin one. | required |
trace | list[str] | The simulation's | required |
config | NaturalnessConfig | Resolved metric configuration. | required |
audio_paths | dict[int, str] | None | Optional map of turn index to the agent's recorded audio, as either a | None |
latency_ms_by_turn | dict[int, float] | None | Optional map of turn index to the platform's reported perceived latency in ms. When supplied, each turn gains a computed | None |
Returns:
| Name | Type | Description |
|---|---|---|
A | NaturalnessResult | None | class: |
NaturalnessResult | None | grade or the grading call failed. Failures are logged rather than | |
NaturalnessResult | None | raised so the metric can never break a simulation run. |
Source code in src/cxas_scrapi/evals/naturalness.py
642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 | |